Session 10: Embedded & Usage-Based Insurance
Learning Objectives
- Explain the economics of embedded insurance and why conversion rates are 3–5× higher than standalone digital channels
- Map the Indian embedded insurance landscape across mobility, e-commerce, travel, and fintech platforms
- Distinguish between Usage-Based Insurance models — PAYD (Pay-As-You-Drive) vs. PHYD (Pay-How-You-Drive) — and understand their risk selection mechanics
- Calculate risk scores from telematics-style driving data using a weighted composite model in Python
- Analyse the privacy, ethics, and regulatory trade-offs of data-driven insurance pricing and articulate a defensible position
1. What is Embedded Insurance?
Embedded insurance is the integration of insurance products into the purchase journey of another product or service — at the moment the customer needs it, in the context where they need it, without requiring a separate insurance purchase process. The insurance is "embedded" into the customer experience of a non-insurance platform.
This is not simply "selling insurance through a partner." The distinction is fundamental: in a traditional partnership, the partner is a distribution channel — the customer must still actively choose to buy insurance. In embedded insurance, the insurance is part of the product experience. The customer buying a flight ticket on MakeMyTrip does not "buy travel insurance" — they check a box during checkout. The customer booking a ride on Ola does not "buy personal accident cover" — they accept a ₹5 add-on. The insurance is invisible until it is needed. This invisibility is the core innovation.
1.1 Embedded vs. Traditional Distribution
| Dimension | Traditional Digital Distribution | Embedded Insurance |
|---|---|---|
| Customer intent | Customer actively seeks insurance | Customer seeks another product; insurance is a contextual add-on |
| Purchase trigger | Renewal reminder, life event, need awareness | Purchase of primary product (flight, gadget, ride, loan) |
| Conversion rate | 1–5% (from site visitors to purchase) | 15–50% (of primary product buyers who see offer) |
| Customer acquisition cost | ₹1,500–5,000 (paid search, aggregator commission) | ₹0–500 (near-zero — the platform covers acquisition) |
| CAC savings source | Marketing spend to attract insurance intenders | The platform's existing traffic is the audience; no separate marketing needed |
| Customer context | Customer is comparing insurance options | Customer is in a buying mindset for the primary product — insurance is a natural addition |
| Policy complexity | Full product details, terms, comparisons | Simplified offer — one price, one coverage summary, one click to accept |
| Trust model | Insurer brand or aggregator platform trust | Trust transfers from the primary platform (Amazon, Ola, MakeMyTrip) |
🛒 Exercise 1.1 — Embedded or Not?
Classify each purchase journey into Embedded insurance or Traditional digital distribution using the §1.1 distinction — is insurance part of the primary product experience, or is the customer actively seeking insurance?
| # | Purchase Journey | Type |
|---|---|---|
| 1 | A customer adds gadget protection while buying a ₹25,000 smartphone on Amazon. | |
| 2 | A customer clicks a Google ad for "term insurance" and compares plans on PolicyBazaar. | |
| 3 | An Ola ride app adds a ₹5 personal-accident cover to the fare during ride booking. | |
| 4 | A bank's partner insurer calls a loan applicant to offer health cover — the customer must decide. | |
| 5 | MakeMyTrip shows a travel-insurance checkbox during flight checkout. | |
| 6 | A customer opens a renewal-reminder SMS and follows the link to the insurer's app. |
Reflection: Embedded insurance converts at 15–50% vs 1–5% for standalone digital. What makes the checkout decision psychologically easier for the customer?
Check Your Classifications
- Embedded — insurance rides the primary purchase (the phone); the customer is already in a buying mindset.
- Traditional — the customer actively seeks insurance; the ad is a distribution channel, not part of a product experience.
- Embedded — the cover is added to the fare during ride booking; enrolment is automatic (opt-out).
- Traditional — the bank is a distribution channel; the customer must still actively choose to buy insurance.
- Embedded — the offer appears at the moment of booking the primary product (the flight).
- Traditional — the customer must decide and go through a separate insurance purchase process.
Reflection note: the customer has already answered the hard questions ("should I spend money? which product? do I trust this?") when they buy the primary product. The insurance is a small, contextual add-on to a decision they have already made — "one price, one coverage summary, one click." The context does the selling that marketing budgets cannot. This is why conversion is 3–5× higher and CAC near zero.
2. The Economics of Embedded Insurance
The economic advantage of embedded insurance over traditional distribution is not incremental — it is structural. Embedded models change the fundamental unit economics of insurance distribution in four ways that standalone digital channels cannot replicate.
2.1 The Four-Part Economics Advantage
- Near-zero marginal CAC: The platform already pays to acquire its users (through marketing, product investment, brand). Adding insurance as an offer at checkout costs almost nothing — the marginal cost of displaying a checkbox on a checkout page that already exists. This compares to ₹2,000–₹5,000 CAC for a standalone digital insurance channel. The economic difference is enormous — and it is structural, not a matter of optimization.
- Contextual conversion: Embedded insurance converts at 15–50% because the customer is already in a purchase mindset. They have already decided to spend money — the incremental decision to add insurance is psychologically easy. In contrast, a standalone insurance website requires the customer to overcome multiple cognitive barriers: "Should I buy insurance? Which one? Is this a good price? Will the claim actually work?" By the time embedded insurance offers appear, the customer has already answered most of these questions implicitly.
- Adverse selection reversal: In standard insurance, the customers most likely to buy are those who believe they have the highest risk — a problem insurers manage through underwriting. In embedded insurance, the customer is buying the primary product (a flight, a phone, a ride) for reasons unrelated to the insurance. The insurance purchase is almost incidental. This dramatically reduces adverse selection — the insurance pool looks more like the general population, not a self-selected high-risk segment. This structural advantage can improve loss ratios by 5–15 points compared to standalone channels.
- Trust transfer: Customers trust Amazon with their payments, their address, and their product choices. When Amazon offers insurance on a phone purchase, some of that trust transfers to the insurance product. The customer does not need to research the insurer's claim settlement ratio — they trust Amazon to have vetted the partner. This trust transfer reduces the "trust barrier" that is one of the biggest friction points in standalone insurance purchase. In marketing terms, embedded insurance turns "pull" (I need to research and decide) into "push" (here is an offer from a platform I already trust).
2.2 The Revenue Split Model
In a typical embedded insurance arrangement, the premium is split three ways:
- Platform (distribution partner): 30–50% of premium — the platform's commission for providing access to its customer base and handling the checkout experience. This is their "take rate." For platforms with thin margins (e-commerce, mobility), insurance can be a high-margin incremental revenue stream.
- Insurer (risk carrier): 40–60% of premium — after paying claims (typically 50–75% of this share) and expenses (15–25%), the insurer earns an underwriting margin of 5–20% on its share.
- Enabler (API infrastructure provider): 5–15% of premium or a SaaS fee — the technology layer connecting the platform to multiple insurers. This is Riskcovry, Zopper, and similar API enabler business models.
For the insurer, the trade-off is simple: in embedded insurance it keeps a smaller slice of each premium. Of a ₹1,000 premium, the platform takes its commission and the enabler its fee, leaving the insurer roughly ₹400–600 — versus the full ₹1,000 it would keep on a standalone policy. The key question is whether the smaller slice still produces a better business.
It does, because the insurer's costs are also far smaller. Acquisition is nearly free — the platform's existing traffic is the audience, so there is no ₹1,500–5,000 CAC to recover and no aggregator commission. The pool is also better risk (the adverse-selection reversal from §2.1), so claims run lower. A standalone insurer must spend heavily to acquire each customer and then recover that spend across several years of renewals; an embedded insurer starts making money on the first policy.
The correct measure is therefore profit per policy, not premium per policy. Example, on a ₹1,000 premium: an embedded insurer keeps ₹500 after platform and enabler cuts, pays ₹300 in claims and ₹100 in expenses, and banks ₹100 profit with zero acquisition cost — immediately. A standalone insurer keeps the full ₹1,000, but after a 70% loss ratio (₹700), expenses (₹150) and a ₹1,500 acquisition cost, its first-year profit is often negative — it reaches ₹100 profit only after the customer renews for 2–3 years. And renewal is where embedded insurance compounds: embedded customers often auto-renew because the platform handles renewal as part of its ecosystem, while standalone renewal must be won again every year. Lower cost + better risk + higher retention together mean the embedded profit per policy can match or beat standalone — and it repeats every renewal.
🧮 Exercise 2.1 — The Embedded Economics Math
Work the §2.2 revenue split on a real product, then make the Pro Tip's "profit per policy" call.
Part A — calculate: A ₹600 gadget-protection premium is split Platform 40% / Enabler 10% / Insurer 50%. The insurer pays claims = 60% of its share and expenses = 18% of its share. Fill in the table:
| Player | Share of ₹600 | Your Answer (₹) |
|---|---|---|
| Platform (distribution partner) | 40% | |
| Enabler (API provider) | 10% | |
| Insurer (risk carrier) | 50% | |
| Insurer profit per policy = insurer share − claims − expenses | — |
Part B — decide: Two products compete for capital: Embedded — ₹500 premium, ₹200 profit per policy; Standalone — ₹1,000 premium, ₹100 profit per policy. Which creates more shareholder value, and why (use the Pro Tip)?
Part C — explain: Why can an embedded pool improve loss ratios by 5–15 points even before the insurer prices for it?
Check Your Math
Part A: Platform ₹240 · Enabler ₹60 · Insurer ₹300 · claims ₹180 (60% × 300) · expenses ₹54 (18% × 300) → insurer profit = ₹300 − ₹180 − ₹54 = ₹66 per policy. Note the insurer earns a healthy margin on a ₹600 product because its capital is deployed only on the risk it carries.
Part B: the embedded product — ₹200 profit per policy at ₹500 premium (40% margin) vs ₹100 on ₹1,000 (10% margin). It consumes less capital per rupee of profit, scales with the platform's growth, and typically auto-renews. Profit per policy, not premium per policy, is what builds shareholder value.
Part C: adverse selection reversal — the customer buys the primary product (the phone, the ride, the flight) for reasons unrelated to insurance, so the insurance pool resembles the general population instead of a self-selected high-risk segment. The self-selection that normally loads the insurance pool works in reverse here, improving the loss ratio structurally — before any behavioural pricing.
3. Embedded Insurance in India
India's embedded insurance market has grown rapidly, driven by the combination of digital public infrastructure (India Stack), large consumer platforms, and InsurTech enablers that provide the API plumbing to connect them. The ecosystem can be mapped across four primary distribution contexts.
3.1 Embedded Insurance by Platform Type
| Platform Type | Platform Example | Insurance Product(s) | Enabler / Insurer | How It Works |
|---|---|---|---|---|
| Mobility & Ride-Hailing | Ola, Uber, Rapido | Personal accident cover per ride, auto driver insurance | Acko (on Ola), ICICI Lombard (on Uber) | ₹0.5–₹5 per ride added to fare. Covers accidental death, permanent disability. Policy active for duration of ride. No separate purchase — enrolment is automatic (opt-out, not opt-in). |
| E-Commerce & Retail | Amazon, Flipkart, Myntra | Product protection plans (gadget damage + liquid + theft), extended warranty | Zopper (on Flipkart/Myntra), Acko (on Amazon) | Offered at checkout when purchasing electronics. 1–4 year plans. Customer enrolled instantly. Claims: WhatsApp photo → AI assessment → replacement or repair. 15–30% conversion of eligible purchases. |
| Travel & Hospitality | MakeMyTrip, IRCTC (via Acko), Yatra, EaseMyTrip | Travel insurance (medical abroad, trip cancellation, lost baggage, flight delay) | Acko, ICICI Lombard, Tata AIG | Offered during flight or hotel booking checkout. Premium: ₹99–₹999 depending on destination and trip duration. Highly contextual — destination with high medical costs = more likely to purchase. |
| Fintech & Payments | PhonePe, Paytm, Cred, Google Pay | Life insurance (sachet), accident cover, travel insurance, device protection | Multiple insurers through Riskcovry and other enablers | Offered within the payments/fintech app as a "sachet" product — very low premium (₹5–₹99), very short term (1 month to 1 year). High volume, low value. Low friction purchase — no forms, UPI-click to buy. |
3.2 The Enabler Model: Riskcovry and Zopper
Two companies have built the infrastructure that powers embedded insurance at scale in India. Understanding their model is essential to understanding how embedded insurance works operationally.
Riskcovry positions itself as "insurance middleware" — an API platform that connects insurers to any distribution platform. A mobility app, e-commerce site, or fintech app can integrate Riskcovry's API in 2–4 weeks and offer insurance products from multiple carriers without building insurance-specific capabilities. Riskcovry handles: product configuration, API connection to multiple insurers, quote engine, policy issuance, claims API, and compliance. It earns a per-transaction fee or a revenue share. Riskcovry does not underwrite risk — it provides the technology layer. This is a capital-light, high-margin business model.
Zopper started as an extended warranty company and evolved into an embedded insurance platform focused on e-commerce. Zopper partners with Flipkart, Myntra, Tata CLiQ, and others to offer product protection plans at checkout. Unlike Riskcovry (which is a pure technology enabler), Zopper works with insurers to design specific products for the e-commerce context — product protection plans that cover accidental damage, liquid damage, and theft, with AI-based claims assessment via photo upload.
The power imbalance is one-sided. When one platform supplies most of an InsurTech's volume, the platform can:
- Demand a higher commission — the InsurTech can hardly refuse without losing most of its book.
- Invite a second insurer to bid — competition squeezes the InsurTech's price and margin.
- Replace the partner at renewal — partnerships are exclusive or semi-exclusive only for a fixed period; then they are renegotiated or ended.
- Build its own insurance capability — the platform takes the margin itself.
That is why analysts watch one number: the share of volume from the single largest platform. Above roughly 40%, the business's survival depends on one counterparty's decisions — losing that partner means losing nearly half the business overnight, so investors treat it as a structural vulnerability. Below it, losing a partner is painful but survivable. When you evaluate any embedded InsurTech, ask: "What share of volume comes from platform #1 — and what would happen if that platform disappeared tomorrow?"
🌍 Exercise 3.1 — Match the Platform
For each embedded-insurance offer, pick the platform type it belongs to (from the §3.1 table) and the likely enabler/insurer.
| # | Offer | Platform Type | Enabler / Insurer |
|---|---|---|---|
| 1 | A ₹1-per-ride accident cover added to a ride-hailing fare. | ||
| 2 | A 2-year screen + liquid damage plan at smartphone checkout. | ||
| 3 | A ₹299 travel-medical cover checkbox at flight checkout. | ||
| 4 | A ₹50 sachet health plan inside a payments app. | ||
| 5 | An extended-warranty plan on a laptop bought from an e-commerce marketplace. |
Reflection: The §3.2 Warning says platform concentration risk is the biggest strategic risk for embedded insurers. What specific power does a platform have over an embedded InsurTech once the partnership exists?
Check Your Matches
- Mobility & Ride-Hailing · Acko — per-ride cover on Ola; enrolment is opt-out.
- E-Commerce & Retail · Acko (or Zopper) — checkout product protection on Amazon/Flipkart/Myntra.
- Travel & Hospitality · Acko — flight/hotel checkout on MakeMyTrip/IRCTC/Yatra.
- Fintech & Payments · Multiple via Riskcovry — sachet products inside PhonePe/Paytm/Cred/GPay.
- E-Commerce & Retail · Zopper — extended warranty / product protection on marketplaces.
Reflection note: the platform owns the customer relationship and the checkout — it can demand a higher commission, invite a second insurer tomorrow, or build its own insurance capability. If one platform is >40% of an InsurTech's volume, that single partner's decisions are existential. Diversification of distribution is the only durable defence.
4. Usage-Based Insurance (UBI)
Usage-Based Insurance uses data about actual behaviour — how much someone drives, how well they drive, what time of day they drive — to price insurance more accurately than traditional risk factors (age, gender, vehicle type, location). UBI is the most significant innovation in personal lines insurance pricing since the actuarial table, because it replaces inferred risk (what group you belong to) with observed risk (what you actually do).
4.1 PAYD vs. PHYD
| Dimension | PAYD — Pay-As-You-Drive | PHYD — Pay-How-You-Drive |
|---|---|---|
| Data collected | Distance driven (kilometres/miles) | Distance + speed + braking + cornering + time of day + phone usage |
| Data source | Odometer reading, GPS, smartphone sensor | Telematics device (OBD-II dongle), smartphone app, or vehicle API |
| Pricing model | Flat rate per km (may vary by road type) | Variable rate per km adjusted by behavioural risk score (higher risk = higher per-km rate) |
| Risk accuracy vs. traditional | Moderate improvement — mileage is a strong predictor of claim frequency | Significant improvement — driving behaviour predicts both claim frequency and severity |
| Premium savings for safe driver | 10–25% (if they drive less than average) | 20–60% (if they drive safely AND less than average) |
| Customer adoption | Higher — less intrusive, fewer privacy concerns | Lower — requires device installation or app permissions; privacy concerns higher |
| Regulatory acceptance | Higher — regulators view mileage-based pricing as fair | Mixed — some regulators concerned about behavioural scoring fairness and data privacy |
| Leading vendors | Allstate Milewise (US), Zego (UK) | Root Insurance (US), Progressive Snapshot (US), ICICI Lombard DriveTrack (India) |
4.2 Data Sources for UBI
UBI data comes from four sources, each with different cost, accuracy, and customer friction profiles:
- Smartphone app: Uses the phone's GPS, accelerometer, and gyroscope to track speed, braking, cornering, and phone use while driving. Lowest hardware cost (customers already have smartphones). Highest data quality challenge (phone position varies, cannot distinguish driver from passenger). Requires significant signal-processing investment to clean the data.
- OBD-II dongle: Plugs into the vehicle's diagnostic port. Reads real-time vehicle data: speed, RPM, engine load, fuel consumption, GPS. Most accurate driving data. High hardware cost (~₹2,000–₹5,000 per device) and installation friction (customer must plug it in). Widely used in early UBI programmes globally.
- Black-box / OEM integration: Some modern vehicles (Tesla, connected cars) natively capture telematics data and can share it via API. No additional hardware needed. Limited by the number of connected vehicles on the road — currently a small fraction of the Indian vehicle fleet.
- Dashcam with computer vision: Emerging — dashcam footage analysed by AI to detect driving events (hard braking, lane deviation, following distance, phone use). Combines the accuracy of OBD with the contextual richness of video. High data processing cost.
🚗 Exercise 4.1 — PAYD or PHYD?
Classify each pricing decision, then run the PAYD vs PHYD arithmetic.
| # | Pricing Decision | PAYD / PHYD |
|---|---|---|
| 1 | A driver pays a flat ₹1.50 per km regardless of how they drive. | |
| 2 | A driver's per-km rate rises 50% after a month of hard-braking alerts. | |
| 3 | A night-shift driver pays 1.3× the base rate for kilometres driven 10pm–5am. | |
| 4 | A low-mileage driver gets a 20% discount simply for driving 800 km/month. | |
| 5 | An app scores braking, cornering, and phone use and adjusts the renewal discount. |
Calculate: A driver does 1,800 km/month in the Moderate tier. Compute the monthly premium under PAYD (₹1.50/km) and PHYD (₹0.80/km × 1.10 multiplier).
| Model | Monthly Premium (₹) |
|---|---|
| PAYD | |
| PHYD |
Reflection: Which model risks underpricing an aggressive low-mileage driver — and why is that an adverse-selection risk for the insurer?
Check Your Answers
- PAYD — flat per-km, no behaviour adjustment.
- PHYD — behaviour (hard braking) changes the rate.
- PHYD — time-of-day behaviour adjusts the per-km rate.
- PAYD — the discount comes from distance alone, not behaviour.
- PHYD — a behavioural score drives the adjustment.
Calculation: PAYD = 1,800 × ₹1.50 = ₹2,700 · PHYD = 1,800 × ₹0.80 × 1.10 = ₹1,584. The safe moderate driver pays ₹1,116 less under PHYD than PAYD — behaviour is rewarded, not just mileage.
Reflection note: PAYD underprices an aggressive low-mileage driver — at 500 km/month they pay ₹750 (500 × ₹1.50) while driving dangerously, far below the ₹2,500 they would pay under a traditional flat premium. The insurer collects a premium that does not reflect the claim risk, and if such drivers self-select into PAYD, the pool's loss ratio worsens — adverse selection. PHYD fixes this with a behavioural multiplier on the per-km rate, so risky behaviour is priced even at low mileage. PHYD prices observed behaviour; PAYD prices only exposure.
5. UBI Data Analysis in Python
In this section, we work with a synthetic telematics dataset. The data represents a month of driving behaviour for 500 policyholders, including mileage, driving events, time-of-day patterns, and claims history. The objective is to build a risk score from behavioural data and compare PAYD (Pay-As-You-Drive) Vs PHYD (Pay-How-You-Drive).
data/telematics_data.csv is provided in the lab folder (InsuranceTech_code/data/) and on GitHub (link below) — save your notebook in InsuranceTech_code/ so the relative path works. A single complete runnable script is available at the end of §5 (and as session_10_ubi_analysis.py in InsuranceTech_code/): run venv/bin/python session_10_ubi_analysis.py to reproduce everything at once. GitHub: telematics_data.csv · complete script.
5.1 Loading the Telematics Data
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# Load synthetic telematics data
telematics = pd.read_csv('data/telematics_data.csv')
print(f"Drivers: {telematics.shape[0]:,}")
print(f"Columns: {telematics.columns.tolist()}")
print(f"\nData types:\n{telematics.dtypes}")
print(f"\nSummary statistics:\n{telematics.describe()}")
# Key columns:
# driver_id — Unique driver identifier
# total_km — Kilometres driven in the month
# avg_speed_kmph — Average speed
# hard_braking_events — Count of hard braking events
# hard_accel_events — Count of rapid acceleration events
# night_driving_pct — Percentage of driving at night (10pm-5am)
# phone_use_events — Count of phone usage events while driving
# previous_claims — Number of claims in the last 3 years
# claim_amount_last_3y — Total claim amount in the last 3 years
5.2 Building a Behavioural Risk Score
▶ Prerequisite: run the §5.1 block first in the same notebook — telematics is defined there. This block will not run standalone.
# Define risk factor weights (these would be calibrated from claims data)
RISK_WEIGHTS = {
'km_risk': 0.20, # Higher mileage = more exposure
'braking_risk': 0.25, # Hard braking = unsafe driving pattern
'accel_risk': 0.15, # Hard acceleration = aggressive driving
'night_risk': 0.20, # Night driving = higher accident risk
'phone_risk': 0.20 # Phone use = distraction risk
}
# Set thresholds for each risk factor
KM_THRESHOLD_HIGH = 2000 # km/month — above this is high mileage
BRAKING_THRESHOLD = 15 # events/month — above this is aggressive
ACCEL_THRESHOLD = 10 # events/month
NIGHT_THRESHOLD = 25 # % of driving at night
PHONE_THRESHOLD = 5 # events/month
# Calculate risk sub-scores (each normalised to 0–100)
telematics['km_risk'] = np.clip(
(telematics['total_km'] / KM_THRESHOLD_HIGH) * 100, 0, 100
)
telematics['braking_risk'] = np.clip(
(telematics['hard_braking_events'] / BRAKING_THRESHOLD) * 100, 0, 100
)
telematics['accel_risk'] = np.clip(
(telematics['hard_accel_events'] / ACCEL_THRESHOLD) * 100, 0, 100
)
telematics['night_risk'] = np.clip(
(telematics['night_driving_pct'] / NIGHT_THRESHOLD) * 100, 0, 100
)
telematics['phone_risk'] = np.clip(
(telematics['phone_use_events'] / PHONE_THRESHOLD) * 100, 0, 100
)
# Composite risk score (weighted average)
telematics['risk_score'] = (
telematics['km_risk'] * RISK_WEIGHTS['km_risk'] +
telematics['braking_risk'] * RISK_WEIGHTS['braking_risk'] +
telematics['accel_risk'] * RISK_WEIGHTS['accel_risk'] +
telematics['night_risk'] * RISK_WEIGHTS['night_risk'] +
telematics['phone_risk'] * RISK_WEIGHTS['phone_risk']
)
# Classify into risk tiers
telematics['risk_tier'] = pd.cut(
telematics['risk_score'],
bins=[0, 20, 40, 60, 100],
labels=['Very Low', 'Low', 'Moderate', 'High']
)
print("Risk Score Distribution:")
print(telematics['risk_tier'].value_counts().sort_index())
print(f"\nMean risk score: {telematics['risk_score'].mean():.1f}")
print(f"Median risk score: {telematics['risk_score'].median():.1f}")
print(f"Std dev: {telematics['risk_score'].std():.1f}")
5.3 PAYD vs. PHYD Pricing Comparison
▶ Prerequisite: continues from §5.1–§5.2 — telematics must include risk_score and risk_tier. pandas 3.x note: mapping a categorical with a dict can fail on arithmetic — write telematics['risk_tier'].astype(str).map(PHYD_RISK_MULTIPLIER).astype(float) instead of .map(...) directly.
# PAYD: flat rate per km (traditional UBI)
PAYD_RATE = 1.50 # ₹1.50 per km
# PHYD: base rate + behavioural adjustment
PHYD_BASE_RATE = 0.80 # ₹0.80 per km (lower base because behaviour-adjusted)
PHYD_RISK_MULTIPLIER = {
'Very Low': 0.60,
'Low': 0.85,
'Moderate': 1.10,
'High': 1.50
}
# Calculate monthly premium under each model
telematics['payd_premium'] = telematics['total_km'] * PAYD_RATE
telematics['phyd_rate'] = telematics['risk_tier'].map(PHYD_RISK_MULTIPLIER)
telematics['phyd_premium'] = telematics['total_km'] * PHYD_BASE_RATE * telematics['phyd_rate']
# Calculate savings vs. traditional flat premium (assume ₹2,500/month)
TRADITIONAL_FLAT = 2500
telematics['traditional_premium'] = TRADITIONAL_FLAT
telematics['payd_savings'] = TRADITIONAL_FLAT - telematics['payd_premium']
telematics['phyd_savings'] = TRADITIONAL_FLAT - telematics['phyd_premium']
# Aggregate by risk tier
comparison = telematics.groupby('risk_tier', observed=False).agg(
driver_count=('driver_id', 'count'),
avg_km=('total_km', 'mean'),
avg_risk_score=('risk_score', 'mean'),
avg_payd=('payd_premium', 'mean'),
avg_phyd=('phyd_premium', 'mean'),
avg_traditional=('traditional_premium', 'mean'),
avg_payd_savings=('payd_savings', 'mean'),
avg_phyd_savings=('phyd_savings', 'mean'),
).round(0)
print("=" * 110)
print(f"{'Risk Tier':12s} {'Drivers':>8s} {'Avg Km':>8s} {'Risk':>6s} {'PAYD ₹':>9s} {'PHYD ₹':>9s} {'Trad ₹':>9s} {'PAYD Save':>10s} {'PHYD Save':>10s}")
print("=" * 110)
for tier in ['Very Low', 'Low', 'Moderate', 'High']:
r = comparison.loc[tier]
print(f"{tier:12s} {r['driver_count']:>5.0f} {r['avg_km']:>5,.0f} {r['avg_risk_score']:>4.0f} ₹{r['avg_payd']:>6,.0f} ₹{r['avg_phyd']:>6,.0f} ₹{r['avg_traditional']:>5,.0f} {'+' if r['avg_payd_savings'] > 0 else ''}₹{r['avg_payd_savings']:>6,.0f} {'+' if r['avg_phyd_savings'] > 0 else ''}₹{r['avg_phyd_savings']:>6,.0f}")
# Visualize
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(14, 5))
# LEFT: Premium comparison by tier
tiers = ['Very Low', 'Low', 'Moderate', 'High']
x = np.arange(len(tiers))
width = 0.25
ax1.bar(x - width, comparison.loc[tiers, 'avg_traditional'], width, label='Traditional (Flat)', color='#a29bfe', alpha=0.7)
ax1.bar(x, comparison.loc[tiers, 'avg_payd'], width, label='PAYD', color='#6c5ce7')
ax1.bar(x + width, comparison.loc[tiers, 'avg_phyd'], width, label='PHYD', color='#00d2d3')
ax1.set_xlabel('Risk Tier')
ax1.set_ylabel('Monthly Premium (₹)')
ax1.set_title('Premium Comparison by Risk Tier', fontweight='bold')
ax1.set_xticks(x)
ax1.set_xticklabels(tiers)
ax1.legend()
ax1.spines['top'].set_visible(False)
ax1.spines['right'].set_visible(False)
# RIGHT: Scatter — risk score vs. claims
if 'claim_amount' in telematics.columns:
ax2.scatter(telematics['risk_score'], telematics['claim_amount'],
alpha=0.4, color='#6c5ce7', s=20)
# Add trend line
z = np.polyfit(telematics['risk_score'], telematics['claim_amount'], 1)
p = np.poly1d(z)
x_line = np.linspace(telematics['risk_score'].min(), telematics['risk_score'].max(), 100)
ax2.plot(x_line, p(x_line), 'r--', linewidth=2, label='Trend')
corr = telematics['risk_score'].corr(telematics['claim_amount'])
ax2.annotate(f'Correlation: {corr:.2f}', xy=(0.05, 0.95),
transform=ax2.transAxes, fontsize=11,
bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
ax2.set_title('Risk Score vs. Claim Amount', fontweight='bold')
ax2.set_xlabel('Behavioural Risk Score (0–100)')
ax2.set_ylabel('Claim Amount (₹)')
ax2.legend()
ax2.spines['top'].set_visible(False)
ax2.spines['right'].set_visible(False)
plt.tight_layout()
plt.show()
5.4 Interpreting the Results
The analysis reveals the core UBI insight: PAYD primarily rewards low-mileage drivers regardless of their driving quality, while PHYD rewards safe driving behaviour regardless of mileage.
- A driver who drives 500 km/month very aggressively (hard braking, phone use, night driving) gets a low premium from PAYD (₹750, saving ₹1,750 vs. traditional) — even though they are high-risk. This is an adverse selection risk for the insurer.
- A driver who drives 3,000 km/month very safely gets a high premium from PAYD (₹4,500, paying ₹2,000 more than traditional) — even though they are low-risk. This is a customer retention risk.
- PHYD solves both problems. The aggressive low-mileage driver's premium is adjusted upward by their behavioural multiplier (₹750 × 1.50 = ₹1,125). The safe high-mileage driver's premium is adjusted downward (₹4,500 × 0.60 = ₹2,700). Both are treated more fairly than under either traditional or PAYD pricing.
This is why PHYD is a better risk selection mechanism than PAYD — but it requires more data, more customer trust, and more regulatory acceptance to deploy at scale.
💻 Exercise 5.1 — Trace the Risk Score
Use the §5.2 thresholds and weights to hand-compute a driver's risk score, then price them under PHYD.
Given — Driver X: total_km 1,000 · hard_braking 6 · hard_accel 4 · night_driving 10% · phone_use 2. Compute each sub-score = value ÷ threshold × 100 (clip at 100):
| Factor | Threshold | Sub-score (0–100) |
|---|---|---|
| km_risk (weight 0.20) | 2,000 km | |
| braking_risk (0.25) | 15 events | |
| accel_risk (0.15) | 10 events | |
| night_risk (0.20) | 25% | |
| phone_risk (0.20) | 5 events | |
| Composite risk_score | weighted sum |
Then price: Which tier is Driver X in (§5.2 bins)? What PHYD multiplier applies (§5.3)? What is the monthly PHYD premium at ₹0.80/km?
Reflection: Why does the composite weight braking (0.25) more than acceleration (0.15)?
Check Your Score
- km_risk = 1000/2000 × 100 = 50
- braking_risk = 6/15 × 100 = 40
- accel_risk = 4/10 × 100 = 40
- night_risk = 10/25 × 100 = 40
- phone_risk = 2/5 × 100 = 40
- Composite = 0.20(50) + 0.25(40) + 0.15(40) + 0.20(40) + 0.20(40) = 10 + 10 + 6 + 8 + 8 = 42
Pricing: 42 falls in the Moderate tier (40–60) → multiplier 1.10 → PHYD premium = 1,000 × ₹0.80 × 1.10 = ₹880/month. Under PAYD the same driver pays ₹1,500 — PHYD's lower base rate rewards their moderate behaviour.
Reflection note: hard braking is the strongest single predictor of claim severity (and, with the §5.3 scatter, of claim amount), so it carries the largest weight. Acceleration correlates with aggression but is a weaker predictor on Indian roads than night driving or phone use — the weights mirror real calibration, where braking and night driving typically dominate.
Complete Runnable Script — Copy & Run (sections 5.1–5.4)
Everything from §5.1–§5.4 in one script, so nothing is left undefined. Save as session_10_ubi_analysis.py inside InsuranceTech_code/ and run venv/bin/python session_10_ubi_analysis.py, or paste it into one Jupyter cell. Requires data/telematics_data.csv (provided in the lab folder and on GitHub — link in the §5 run note). Outputs: the console tables below plus ubi_premium_comparison.png.
# ============================================================================
# Session 10 — Usage-Based Insurance: UBI Data Analysis (COMPLETE RUNNABLE SCRIPT)
# InsurTech & Digital Risk Solutions (MBA) — Woxsen University
#
# HOW TO RUN:
# venv/bin/python session_10_ubi_analysis.py
# (or run top-to-bottom in one Jupyter notebook — the page's §5.1–§5.4 blocks
# are merged here in order so nothing is left undefined)
#
# DATA FILE: data/telematics_data.csv (provided) — 500 drivers, one month of
# synthetic driving behaviour. Falls back to telematics_data.csv
# in the current folder.
#
# OUTPUTS: console tables + ubi_premium_comparison.png + ubi_risk_scatter.png
# ============================================================================
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# ----------------------------------------------------------------------------
# SECTION 5.1 — Loading the Telematics Data
# ----------------------------------------------------------------------------
try:
telematics = pd.read_csv('data/telematics_data.csv')
except FileNotFoundError:
telematics = pd.read_csv('telematics_data.csv')
print(f"Drivers: {telematics.shape[0]:,}")
print(f"Columns: {telematics.columns.tolist()}")
print(f"\nData types:\n{telematics.dtypes}")
print(f"\nSummary statistics:\n{telematics.describe().round(1)}")
# Key columns:
# driver_id — Unique driver identifier
# total_km — Kilometres driven in the month
# avg_speed_kmph — Average speed
# hard_braking_events — Count of hard braking events
# hard_accel_events — Count of rapid acceleration events
# night_driving_pct — Percentage of driving at night (10pm-5am)
# phone_use_events — Count of phone usage events while driving
# previous_claims — Number of claims in the last 3 years
# claim_amount_last_3y — Total claim amount in the last 3 years
# ----------------------------------------------------------------------------
# SECTION 5.2 — Building a Behavioural Risk Score
# ----------------------------------------------------------------------------
RISK_WEIGHTS = {
'km_risk': 0.20, # Higher mileage = more exposure
'braking_risk': 0.25, # Hard braking = unsafe driving pattern
'accel_risk': 0.15, # Hard acceleration = aggressive driving
'night_risk': 0.20, # Night driving = higher accident risk
'phone_risk': 0.20 # Phone use = distraction risk
}
KM_THRESHOLD_HIGH = 2000 # km/month — above this is high mileage
BRAKING_THRESHOLD = 15 # events/month — above this is aggressive
ACCEL_THRESHOLD = 10 # events/month
NIGHT_THRESHOLD = 25 # % of driving at night
PHONE_THRESHOLD = 5 # events/month
telematics['km_risk'] = np.clip((telematics['total_km'] / KM_THRESHOLD_HIGH) * 100, 0, 100)
telematics['braking_risk'] = np.clip((telematics['hard_braking_events'] / BRAKING_THRESHOLD) * 100, 0, 100)
telematics['accel_risk'] = np.clip((telematics['hard_accel_events'] / ACCEL_THRESHOLD) * 100, 0, 100)
telematics['night_risk'] = np.clip((telematics['night_driving_pct'] / NIGHT_THRESHOLD) * 100, 0, 100)
telematics['phone_risk'] = np.clip((telematics['phone_use_events'] / PHONE_THRESHOLD) * 100, 0, 100)
telematics['risk_score'] = (
telematics['km_risk'] * RISK_WEIGHTS['km_risk'] +
telematics['braking_risk'] * RISK_WEIGHTS['braking_risk'] +
telematics['accel_risk'] * RISK_WEIGHTS['accel_risk'] +
telematics['night_risk'] * RISK_WEIGHTS['night_risk'] +
telematics['phone_risk'] * RISK_WEIGHTS['phone_risk']
)
telematics['risk_tier'] = pd.cut(
telematics['risk_score'],
bins=[0, 20, 40, 60, 100],
labels=['Very Low', 'Low', 'Moderate', 'High'],
include_lowest=True
)
print("Risk Score Distribution:")
print(telematics['risk_tier'].value_counts().sort_index())
print(f"\nMean risk score: {telematics['risk_score'].mean():.1f}")
print(f"Median risk score: {telematics['risk_score'].median():.1f}")
print(f"Std dev: {telematics['risk_score'].std():.1f}")
# ----------------------------------------------------------------------------
# SECTION 5.3 — PAYD vs. PHYD Pricing Comparison
# ----------------------------------------------------------------------------
PAYD_RATE = 1.50 # ₹1.50 per km
PHYD_BASE_RATE = 0.80 # ₹0.80 per km (lower base because behaviour-adjusted)
PHYD_RISK_MULTIPLIER = {
'Very Low': 0.60,
'Low': 0.85,
'Moderate': 1.10,
'High': 1.50
}
telematics['payd_premium'] = telematics['total_km'] * PAYD_RATE
# pandas 3.x: map a categorical with a dict can return a non-arithmetic dtype,
# so cast to str first, then to float.
telematics['phyd_rate'] = telematics['risk_tier'].astype(str).map(PHYD_RISK_MULTIPLIER).astype(float)
telematics['phyd_premium'] = telematics['total_km'] * PHYD_BASE_RATE * telematics['phyd_rate']
TRADITIONAL_FLAT = 2500
telematics['traditional_premium'] = TRADITIONAL_FLAT
telematics['payd_savings'] = TRADITIONAL_FLAT - telematics['payd_premium']
telematics['phyd_savings'] = TRADITIONAL_FLAT - telematics['phyd_premium']
comparison = telematics.groupby('risk_tier', observed=False).agg(
driver_count=('driver_id', 'count'),
avg_km=('total_km', 'mean'),
avg_risk_score=('risk_score', 'mean'),
avg_payd=('payd_premium', 'mean'),
avg_phyd=('phyd_premium', 'mean'),
avg_traditional=('traditional_premium', 'mean'),
avg_payd_savings=('payd_savings', 'mean'),
avg_phyd_savings=('phyd_savings', 'mean'),
).round(0)
print("=" * 110)
print(f"{'Risk Tier':12s} {'Drivers':>8s} {'Avg Km':>8s} {'Risk':>6s} {'PAYD ₹':>9s} {'PHYD ₹':>9s} {'Trad ₹':>9s} {'PAYD Save':>10s} {'PHYD Save':>10s}")
print("=" * 110)
for tier in ['Very Low', 'Low', 'Moderate', 'High']:
r = comparison.loc[tier]
print(f"{tier:12s} {r['driver_count']:>5.0f} {r['avg_km']:>5,.0f} {r['avg_risk_score']:>4.0f} ₹{r['avg_payd']:>6,.0f} ₹{r['avg_phyd']:>6,.0f} ₹{r['avg_traditional']:>5,.0f} {'+' if r['avg_payd_savings'] > 0 else ''}₹{r['avg_payd_savings']:>6,.0f} {'+' if r['avg_phyd_savings'] > 0 else ''}₹{r['avg_phyd_savings']:>6,.0f}")
# Visualize
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(14, 5))
tiers = ['Very Low', 'Low', 'Moderate', 'High']
x = np.arange(len(tiers))
width = 0.25
ax1.bar(x - width, comparison.loc[tiers, 'avg_traditional'], width, label='Traditional (Flat)', color='#a29bfe', alpha=0.7)
ax1.bar(x, comparison.loc[tiers, 'avg_payd'], width, label='PAYD', color='#6c5ce7')
ax1.bar(x + width, comparison.loc[tiers, 'avg_phyd'], width, label='PHYD', color='#00d2d3')
ax1.set_xlabel('Risk Tier')
ax1.set_ylabel('Monthly Premium (₹)')
ax1.set_title('Premium Comparison by Risk Tier', fontweight='bold')
ax1.set_xticks(x)
ax1.set_xticklabels(tiers)
ax1.legend()
ax1.spines['top'].set_visible(False)
ax1.spines['right'].set_visible(False)
if 'claim_amount' in telematics.columns:
ax2.scatter(telematics['risk_score'], telematics['claim_amount'],
alpha=0.4, color='#6c5ce7', s=20)
z = np.polyfit(telematics['risk_score'], telematics['claim_amount'], 1)
p = np.poly1d(z)
x_line = np.linspace(telematics['risk_score'].min(), telematics['risk_score'].max(), 100)
ax2.plot(x_line, p(x_line), 'r--', linewidth=2, label='Trend')
corr = telematics['risk_score'].corr(telematics['claim_amount'])
ax2.annotate(f'Correlation: {corr:.2f}', xy=(0.05, 0.95),
transform=ax2.transAxes, fontsize=11,
bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
ax2.set_title('Risk Score vs. Claim Amount', fontweight='bold')
ax2.set_xlabel('Behavioural Risk Score (0–100)')
ax2.set_ylabel('Claim Amount (₹)')
ax2.legend()
ax2.spines['top'].set_visible(False)
ax2.spines['right'].set_visible(False)
plt.tight_layout()
plt.savefig('ubi_premium_comparison.png', dpi=120) # artifact for submission
plt.show()
# ----------------------------------------------------------------------------
# SECTION 5.4 — Interpreting the Results (worked examples from the page)
# ----------------------------------------------------------------------------
print("\n" + "=" * 70)
print("SECTION 5.4 — Interpretation examples")
print("=" * 70)
aggressive_low_km = 500 # km/month, aggressive driving (High tier)
payd = aggressive_low_km * PAYD_RATE
phyd = aggressive_low_km * PHYD_BASE_RATE * 1.50
print(f"Aggressive low-mileage driver (500 km, High tier):")
print(f" PAYD = ₹{payd:.0f} | PHYD = ₹{phyd:.0f} | Traditional = ₹{TRADITIONAL_FLAT}")
print(f" -> PAYD underprices the risk; PHYD adjusts it up by the behavioural multiplier.")
safe_high_km = 3000 # km/month, safe driving (Very Low tier)
payd = safe_high_km * PAYD_RATE
phyd = safe_high_km * PHYD_BASE_RATE * 0.60
print(f"\nSafe high-mileage driver (3000 km, Very Low tier):")
print(f" PAYD = ₹{payd:.0f} | PHYD = ₹{phyd:.0f} | Traditional = ₹{TRADITIONAL_FLAT}")
print(f" -> PAYD overcharges the safe driver; PHYD rewards them with a discount.")
print("\nCharts saved: ubi_premium_comparison.png")
6. Global UBI Leaders
Usage-Based Insurance has a 15+ year history globally, with several companies reaching significant scale. Their experiences provide lessons for any market — including India — where UBI is still nascent.
6.1 Root Insurance (US)
Root Insurance (founded 2015, IPO 2020 at $6B+) was the most ambitious UBI InsurTech globally. Root's model was PHYD-only: customers install a smartphone app that tracks driving for 2–4 weeks, Root calculates a driving score, and if the score is good enough, Root offers a premium. If the score is poor, Root declines to offer coverage — the customer is "uninsurable by Root's standards." This is the opposite of traditional insurance, where the insurer must accept all applicants (subject to regulatory constraints) and price accordingly.
What went right: Root's risk selection was genuinely superior — its loss ratio on policies that went through the full telematics scoring process was approximately 15–20 points better than industry average. The self-selection effect (safer drivers are more willing to be tracked) was powerful.
What went wrong: Root's growth stalled because nearly 70% of potential customers did not complete the app-based driving test — they either declined to install the app, installed but did not drive enough to generate a score, or scored poorly and were rejected. The customer acquisition cost for the remaining 30% was extremely high. By 2023, Root's market cap had fallen to ~$100M (a 95%+ decline from its peak) as growth stalled and underwriting losses from its early expansion years caught up.
Key lesson: A superior risk selection mechanism is not enough if the customer acquisition funnel is so restrictive that it limits volume and drives up CAC. UBI programmes must balance risk accuracy with customer accessibility. The best UBI model may be one that does not require a pre-purchase driving test but instead adjusts premiums dynamically as driving data accumulates — "start with traditional pricing, earn a discount as you prove you are a safe driver."
6.2 Progressive Snapshot (US)
Progressive's Snapshot programme is the largest UBI programme in the US, with millions of active participants. It uses an OBD-II dongle (or smartphone app) that monitors driving behaviour and offers discounts of up to 30% at renewal based on safe driving. Progressive's approach differs from Root in two critical ways: (a) Snapshot is optional and available to all Progressive customers — it does not gate access to coverage on the driving score, and (b) the programme rewards safe driving with discounts rather than penalizing risky driving with surcharges. This "carrot, not stick" design makes Snapshot far more acceptable to customers and regulators.
6.3 ICICI Lombard DriveTrack (India)
India's most established UBI programme is ICICI Lombard's DriveTrack, which gives customers a free OBD-II dongle that tracks driving behaviour. The programme structure is closer to Progressive's model: customers enrol voluntarily, receive a driving score after each policy period, and safe drivers earn premium discounts at renewal. DriveTrack has been active since approximately 2018 and has accumulated significant data on Indian driving patterns. The programme has not been transformative for ICICI Lombard's motor insurance business (it covers a small fraction of the overall portfolio), but it has generated valuable data that the company uses to refine its pricing models for all motor customers.
🏁 Exercise 6.1 — UBI Strategy Clinic
Three UBI launches need a diagnosis and a strategy call, using the §6 lessons from Root, Progressive, and DriveTrack.
| # | Situation | Diagnosis | Strategy Call |
|---|---|---|---|
| 1 | A startup requires a 3-week app driving test before it will quote anyone (Root-style). | ||
| 2 | An insurer offers a discount at renewal to safe drivers, never a surcharge (Progressive-style). | ||
| 3 | A new Indian UBI launch must pick a data-collection device for mass adoption. |
Design: Write the 3-sentence product memo for the Indian launch — discount structure, observation period, and the one risk you mitigate first.
Check Your Diagnoses
- Restrictive funnel / CAC problem — Root's pre-purchase test saw ~70% of customers drop out, crushing volume and CAC. The fix: stop gating access — "start with traditional pricing, earn a discount as you prove you drive safely" (the §6.1 key lesson).
- Discount-only model — the "carrot, not stick" design is why Snapshot has run at scale for a decade: no customer ever receives a higher bill, so trust and regulatory goodwill are preserved. It leaves the surcharge on the table but keeps enrolment and retention high.
- Mass adoption → smartphone app — zero hardware cost (vs ₹2,000–5,000 OBD dongles) and the widest reach, per the hands-on memo's logic; OBD/OEM can follow for precision later. (Any of app-first answers with a defensible reason is acceptable.)
Model memo: "Base premium ₹2,500 (same as traditional — no customer pays MORE by enrolling). Safe drivers earn up to 40% off at renewal after a 3-month observation period; moderate and high-risk tiers simply pay the traditional premium. The one risk we mitigate first is premium shock — a customer discovering a higher bill — because it is the single biggest reputational risk for UBI (§7 Warning)."
7. Privacy, Ethics & UBI
Usage-Based Insurance sits at the intersection of two powerful trends — data-driven personalisation and consumer privacy — that are increasingly in tension. As UBI expands, the privacy and ethics questions it raises become not peripheral concerns but central strategic issues that will determine how quickly and how broadly UBI can be adopted.
7.1 The Data Collection Question
A PHYD programme may collect: GPS location every few seconds, speed, acceleration on three axes, braking force, cornering G-force, time of day, phone usage events, and — if the smartphone app is used — background location data even when not driving (to distinguish driving from other travel modes). For a customer, the question is: what does the insurer know about me, when, and for how long do they keep it? For the insurer, the question is: can we build a commercially viable UBI programme without collecting data that customers and regulators find intrusive?
The tension is real. Studies consistently show that 40–60% of consumers say they are "very concerned" about sharing driving data with their insurer — but 60–80% say they would share data for a meaningful premium discount. The "privacy premium" — how much discount a customer demands in exchange for sharing data — varies by demographics, product type, and cultural context. Indian customers have generally shown less privacy sensitivity than European customers (where GDPR imposes strict limits), but sensitivity is rising with awareness.
7.2 The Fairness Question
Usage-based pricing raises three fairness questions that regulators and insurers are actively debating:
- Segregation of risk pools: If the lowest-risk drivers all migrate to UBI programmes (where they get lower premiums), the traditional pool retains a higher proportion of higher-risk drivers. Their premiums must rise. Is society comfortable with a two-tier insurance system where data-rich customers pay less and data-poor customers pay more — even if the data-poor customers are not poor by choice but by privacy preference? This is the "privacy tax" problem.
- Behavioural discrimination by price: If an insurer charges a 50% higher per-km rate for night driving, it is effectively saying: "if you cannot avoid driving at night (because you work night shifts, have a family emergency, or live in an area with poor public transport), you will pay more." Is this fair? The customer may not have a choice about when they drive. Pricing based on behaviour that customers cannot easily change functions more like a penalty than a risk-reflective price.
- Data ownership and portability: Driving data belongs to the customer, but only the insurer that collected it can use it to price the risk. If a customer wants to switch insurers, their accumulated safe-driving data stays with the old insurer. The new insurer starts from scratch — either requiring a new tracking period or reverting to traditional pricing. Data portability — the ability for customers to take their driving data to any insurer — would create a more competitive UBI market but raises significant technical, privacy, and competitive challenges.
7.3 The DPDP Act and UBI
India's Digital Personal Data Protection Act (DPDP Act, 2023) has direct implications for UBI programmes. Key requirements: explicit consent for data collection (separate from the insurance contract), purpose limitation (UBI data cannot be used for other purposes without fresh consent), data minimisation (collect only what is necessary), and the right to withdraw consent (the customer can stop sharing data at any time — which would require the insurer to revert to a non-UBI pricing mechanism).
For UBI providers in India, the DPDP Act means: (a) the consent form for telematics data must be written in clear, simple language — not buried in the policy terms, (b) data collection should be limited to what is actually used for pricing — collecting GPS location continuously when the purpose is mileage calculation may violate the data minimisation principle, and (c) the option to opt out of data sharing must be as easy as opting in. UBI programmes designed with the DPDP Act in mind from the start will have a significant regulatory advantage over those that treat privacy compliance as an afterthought.
⚖ Exercise 7.1 — The Privacy Tax Debate
Identify the principle behind each UBI concern, then play regulator on a filing.
| # | Concern | Principle |
|---|---|---|
| 1 | A night-shift driver pays 50% more per km because they cannot avoid driving at night. | |
| 2 | Safe drivers migrate to UBI; the remaining traditional pool's premiums must rise. | |
| 3 | A customer switches insurers but their safe-driving data stays with the old insurer. | |
| 4 | An insurer collects continuous GPS when only mileage is used for pricing. |
Be the regulator: A UBI filing collects app telematics for 6 months, then adjusts renewal premiums up to ±40%. Approve or reject — and if you approve, write the THREE DPDP conditions you impose.
Check Your Analysis
- Behavioural discrimination by price — pricing based on behaviour the customer cannot easily change functions more like a penalty than a risk-reflective price (§7.2).
- Segregation / privacy tax — data-poor customers (by privacy preference, not by choice) pay more as the low-risk migrate out (§7.2).
- Data ownership & portability — driving data belongs to the customer but stays with the collecting insurer (§7.2).
- Data minimisation (DPDP) — collecting more than the pricing purpose needs violates the DPDP data-minimisation principle (§7.3).
Regulator's verdict (model): Approve with conditions — (1) explicit, plain-language consent separate from the policy contract; (2) purpose limitation: collect only the driving events used for pricing, no continuous GPS; (3) one-click opt-out as easy as opt-in, reverting the customer to traditional pricing. Also consider a cap: given the §7 Warning, many regulators would reject a ±40% surcharge and demand discount-only (±0 to −40%) to protect trust.
Hands-On Project: Build a UBI Risk Scoring Model
You are a data scientist at "SafeDrive Insurance," a hypothetical InsurTech launching a Usage-Based Insurance product for private car owners in India. Your task is to build a behavioural risk scoring model from telematics data, compare PAYD vs. PHYD pricing for 100 policyholders, and recommend a product strategy.
Steps
- Load the telematics data from `data/telematics_data.csv`. Explore the columns and distributions. Print summary statistics for all numeric variables. Identify any data quality issues (missing values, extreme outliers).
- Build a behavioural risk score using the following factors with equal weights (20% each):
- Hard braking events per 100 km (normalize to 0–100)
- Hard acceleration events per 100 km (normalize to 0–100)
- Night driving percentage (0–100 scale)
- Phone use events per 100 km (0–100 scale)
- Average speed deviation from the speed limit (0–100 scale, where 0 = no deviation)
- Segment drivers into risk tiers using the 25th, 50th, and 75th percentiles of the risk score as thresholds. Label as Very Low (≤25th percentile), Low (25–50th), Moderate (50–75th), and High (>75th).
- Calculate premium under three models for each driver:
- Traditional: flat ₹3,000/month for all drivers
- PAYD: ₹1.75 per km driven
- PHYD: ₹1.00 per km × risk multiplier (Very Low: 0.55, Low: 0.80, Moderate: 1.15, High: 1.60)
- Analyse the results: For each risk tier, calculate the average premium under each model and the savings vs. traditional. Identify: (a) Which drivers benefit most from PAYD? (b) Which benefit most from PHYD? (c) Are there drivers who would pay more under BOTH UBI models than traditional?
- Create three visualizations: (a) Histogram of risk scores with quartile lines. (b) Premium comparison (grouped bar chart by risk tier for all 3 models). (c) Scatter plot of risk score vs. premium paid under PHYD, coloured by risk tier.
- Write a 300-word product recommendation to the CEO of SafeDrive Insurance: would you launch PAYD, PHYD, or both? What pricing structure do you recommend? What are the key risks and how would you mitigate them? What customer segments would you target first?
View Solution / Walkthrough
Solution Approach
The solution follows the analysis framework from Sections 5.2–5.4. Key code components are identical to the code blocks in this session. The main deliverable is the product recommendation memo. A sample memo structure is provided below.
Sample Product Recommendation Memo
To: CEO, SafeDrive Insurance
From: Data Science — UBI Product Team
Subject: UBI Product Launch Recommendation
Executive Summary:
Our analysis of the telematics pilot data (n=100 drivers) confirms that behavioural risk scoring predicts claim experience significantly better than traditional mileage-only metrics. The composite risk score shows a correlation of approximately 0.35 with historical claim amounts, compared to 0.18 for mileage alone. A UBI programme is strategically viable. We recommend launching a PHYD-only product with a "discount-only" pricing structure — no surcharges above the traditional premium — to minimise customer friction and regulatory risk.
Recommended Product Structure:
- Base monthly premium: ₹2,500 (same as traditional, so no customer pays MORE by enrolling)
- Maximum discount: 40% (₹1,500/month) for Very Low risk tier
- Typical discount: 15–20% for Low risk tier
- No discount (but no surcharge) for Moderate and High risk tiers — they simply pay the traditional premium
- Discount is applied at renewal, not immediately — customers earn their discount over a 3-month observation period
Rationale: A discount-only PHYD model addresses the three key risks. First, it eliminates the "premium shock" risk — no customer receives a higher bill because of UBI. Second, it reduces regulatory risk — IRDAI is more likely to approve a discount-only programme than a surcharge programme. Third, it improves customer acquisition — marketing "save up to 40% on your car insurance" will drive significantly higher enrolment than a message about "fair pricing based on your driving."
Launch recommendation: Target a controlled pilot of 5,000 policyholders in two cities (Delhi NCR and Bengaluru — chosen for their demographic diversity and telematics coverage). Use a smartphone app (not OBD dongle) for data collection — it has zero hardware cost and reaches the widest customer base. Validate loss ratios for 6 months before expanding to 50,000 policyholders. After 12 months and statistical validation of the pricing model, consider introducing a "premium adjustment" for the High risk tier — but only with clear customer communication and regulatory approval.
Key risk and mitigation: The most significant risk is low enrolment due to privacy concerns. Mitigation: (1) a clear, simple privacy notice written in plain language (not legalese), (2) data collection limited to driving events (no continuous GPS tracking), (3) a monthly "privacy report" showing each customer exactly what data was collected, and (4) a one-click option to pause data collection and revert to traditional pricing for that month. We believe this combination of transparency and control will achieve an enrolment rate of 25–35% of willing customers — sufficient for a statistically viable UBI programme.
Key Takeaways
Embedded insurance converts at 3–5× the rate of standalone digital channels because it benefits from contextual conversion, near-zero marginal CAC, trust transfer from the platform, and the "reverse adverse selection" effect.
India's embedded insurance ecosystem spans mobility, e-commerce, travel, and fintech — enabled by API platforms (Riskcovry, Zopper) that connect insurers to distribution partners. Platform concentration risk is the primary strategic vulnerability.
PAYD (flat per-km pricing) rewards low mileage regardless of driving quality. PHYD (behaviour-adjusted per-km pricing) rewards safe driving. PHYD is a better risk selection mechanism but requires more data and customer trust.
The global UBI experience shows that a superior risk model is not enough — if the enrolment funnel is too restrictive (Root's 30% completion rate), the CAC makes the economics unviable. The most successful UBI programmes use a "discount-only" structure and do not gate access to coverage.
UBI raises unresolved privacy and fairness questions — the "privacy tax" on data-poor customers, behavioural discrimination (night-driving pricing), and data portability. DPDP Act compliance in India requires explicit consent, purpose limitation, and data minimisation from day one.
3-2-1 Reflection — Before You Move On
Insurers read behaviour — practise translating today's frameworks into decisions. Write from memory, don't scroll back.
3 Things I Learned Today
2 Embedded / UBI Concepts I Can Now Explain
1 Question I Still Have About Embedded & Usage-Based Insurance
Test Your Understanding
1. An embedded insurance offer during e-commerce checkout consistently converts at 25% while the same insurer's standalone website converts at 3%. The MAIN reason for this difference is:
2. Under the PHYD pricing model, a driver who drives 1,500 km/month with a "High" risk tier (multiplier = 1.50) and a base rate of ₹1.00/km would pay:
3. An InsurTech that derives 60% of its embedded insurance premium volume from a single e-commerce platform partnership faces what strategic risk?
4. Root Insurance's UBI programme failed to achieve sustainable scale primarily because:
5. A "discount-only" UBI programme (where telematics data can only LOWER a customer's premium, never increase it) is recommended as the safest launch strategy because: