Module 3 · Session 10 · 90 min · Python Lab

Session 10: Embedded & Usage-Based Insurance

CILO-1, CILO-2 · Domain Knowledge & Analytics · Lecture & Data Analysis (Python) · Jupyter Notebook required

Learning Objectives

1. What is Embedded Insurance?

Embedded insurance is the integration of insurance products into the purchase journey of another product or service — at the moment the customer needs it, in the context where they need it, without requiring a separate insurance purchase process. The insurance is "embedded" into the customer experience of a non-insurance platform.

This is not simply "selling insurance through a partner." The distinction is fundamental: in a traditional partnership, the partner is a distribution channel — the customer must still actively choose to buy insurance. In embedded insurance, the insurance is part of the product experience. The customer buying a flight ticket on MakeMyTrip does not "buy travel insurance" — they check a box during checkout. The customer booking a ride on Ola does not "buy personal accident cover" — they accept a ₹5 add-on. The insurance is invisible until it is needed. This invisibility is the core innovation.

1.1 Embedded vs. Traditional Distribution

DimensionTraditional Digital DistributionEmbedded Insurance
Customer intentCustomer actively seeks insuranceCustomer seeks another product; insurance is a contextual add-on
Purchase triggerRenewal reminder, life event, need awarenessPurchase of primary product (flight, gadget, ride, loan)
Conversion rate1–5% (from site visitors to purchase)15–50% (of primary product buyers who see offer)
Customer acquisition cost₹1,500–5,000 (paid search, aggregator commission)₹0–500 (near-zero — the platform covers acquisition)
CAC savings sourceMarketing spend to attract insurance intendersThe platform's existing traffic is the audience; no separate marketing needed
Customer contextCustomer is comparing insurance optionsCustomer is in a buying mindset for the primary product — insurance is a natural addition
Policy complexityFull product details, terms, comparisonsSimplified offer — one price, one coverage summary, one click to accept
Trust modelInsurer brand or aggregator platform trustTrust transfers from the primary platform (Amazon, Ola, MakeMyTrip)
🌎
Real World: When Amazon India partnered with Acko to offer gadget protection during checkout for electronic devices, the conversion rate — customers who saw the offer and accepted it — exceeded 25% in its first month. Compare this to a standalone insurance website where 2–5% of visitors convert to purchase. The customer who just added a ₹25,000 smartphone to their cart is in a buying mindset. They understand the value of the device. The insurance offer (typically ₹499–₹999 for 1–2 years of accidental damage + liquid damage cover) is a trivial addition to the purchase decision. The context does the selling that insurance marketing budgets can never achieve.

🛒 Exercise 1.1 — Embedded or Not?

Classify each purchase journey into Embedded insurance or Traditional digital distribution using the §1.1 distinction — is insurance part of the primary product experience, or is the customer actively seeking insurance?

#Purchase JourneyType
1A customer adds gadget protection while buying a ₹25,000 smartphone on Amazon.
2A customer clicks a Google ad for "term insurance" and compares plans on PolicyBazaar.
3An Ola ride app adds a ₹5 personal-accident cover to the fare during ride booking.
4A bank's partner insurer calls a loan applicant to offer health cover — the customer must decide.
5MakeMyTrip shows a travel-insurance checkbox during flight checkout.
6A customer opens a renewal-reminder SMS and follows the link to the insurer's app.

Reflection: Embedded insurance converts at 15–50% vs 1–5% for standalone digital. What makes the checkout decision psychologically easier for the customer?

Check Your Classifications
  1. Embedded — insurance rides the primary purchase (the phone); the customer is already in a buying mindset.
  2. Traditional — the customer actively seeks insurance; the ad is a distribution channel, not part of a product experience.
  3. Embedded — the cover is added to the fare during ride booking; enrolment is automatic (opt-out).
  4. Traditional — the bank is a distribution channel; the customer must still actively choose to buy insurance.
  5. Embedded — the offer appears at the moment of booking the primary product (the flight).
  6. Traditional — the customer must decide and go through a separate insurance purchase process.

Reflection note: the customer has already answered the hard questions ("should I spend money? which product? do I trust this?") when they buy the primary product. The insurance is a small, contextual add-on to a decision they have already made — "one price, one coverage summary, one click." The context does the selling that marketing budgets cannot. This is why conversion is 3–5× higher and CAC near zero.

2. The Economics of Embedded Insurance

The economic advantage of embedded insurance over traditional distribution is not incremental — it is structural. Embedded models change the fundamental unit economics of insurance distribution in four ways that standalone digital channels cannot replicate.

2.1 The Four-Part Economics Advantage

  1. Near-zero marginal CAC: The platform already pays to acquire its users (through marketing, product investment, brand). Adding insurance as an offer at checkout costs almost nothing — the marginal cost of displaying a checkbox on a checkout page that already exists. This compares to ₹2,000–₹5,000 CAC for a standalone digital insurance channel. The economic difference is enormous — and it is structural, not a matter of optimization.
  2. Contextual conversion: Embedded insurance converts at 15–50% because the customer is already in a purchase mindset. They have already decided to spend money — the incremental decision to add insurance is psychologically easy. In contrast, a standalone insurance website requires the customer to overcome multiple cognitive barriers: "Should I buy insurance? Which one? Is this a good price? Will the claim actually work?" By the time embedded insurance offers appear, the customer has already answered most of these questions implicitly.
  3. Adverse selection reversal: In standard insurance, the customers most likely to buy are those who believe they have the highest risk — a problem insurers manage through underwriting. In embedded insurance, the customer is buying the primary product (a flight, a phone, a ride) for reasons unrelated to the insurance. The insurance purchase is almost incidental. This dramatically reduces adverse selection — the insurance pool looks more like the general population, not a self-selected high-risk segment. This structural advantage can improve loss ratios by 5–15 points compared to standalone channels.
  4. Trust transfer: Customers trust Amazon with their payments, their address, and their product choices. When Amazon offers insurance on a phone purchase, some of that trust transfers to the insurance product. The customer does not need to research the insurer's claim settlement ratio — they trust Amazon to have vetted the partner. This trust transfer reduces the "trust barrier" that is one of the biggest friction points in standalone insurance purchase. In marketing terms, embedded insurance turns "pull" (I need to research and decide) into "push" (here is an offer from a platform I already trust).

2.2 The Revenue Split Model

In a typical embedded insurance arrangement, the premium is split three ways:

For the insurer, the trade-off is simple: in embedded insurance it keeps a smaller slice of each premium. Of a ₹1,000 premium, the platform takes its commission and the enabler its fee, leaving the insurer roughly ₹400–600 — versus the full ₹1,000 it would keep on a standalone policy. The key question is whether the smaller slice still produces a better business.

It does, because the insurer's costs are also far smaller. Acquisition is nearly free — the platform's existing traffic is the audience, so there is no ₹1,500–5,000 CAC to recover and no aggregator commission. The pool is also better risk (the adverse-selection reversal from §2.1), so claims run lower. A standalone insurer must spend heavily to acquire each customer and then recover that spend across several years of renewals; an embedded insurer starts making money on the first policy.

The correct measure is therefore profit per policy, not premium per policy. Example, on a ₹1,000 premium: an embedded insurer keeps ₹500 after platform and enabler cuts, pays ₹300 in claims and ₹100 in expenses, and banks ₹100 profit with zero acquisition cost — immediately. A standalone insurer keeps the full ₹1,000, but after a 70% loss ratio (₹700), expenses (₹150) and a ₹1,500 acquisition cost, its first-year profit is often negative — it reaches ₹100 profit only after the customer renews for 2–3 years. And renewal is where embedded insurance compounds: embedded customers often auto-renew because the platform handles renewal as part of its ecosystem, while standalone renewal must be won again every year. Lower cost + better risk + higher retention together mean the embedded profit per policy can match or beat standalone — and it repeats every renewal.

💡
Pro Tip: When evaluating an embedded insurance opportunity, do not focus on premium volume alone. The key metric is profit per policy, not premium per policy. An embedded product with ₹500 premium but ₹200 profit may be more valuable than a standalone product with ₹1,000 premium and ₹100 profit — because the embedded product consumes less capital per rupee of profit and scales faster through the platform's growth. This is the metric that matters for shareholder value.

🧮 Exercise 2.1 — The Embedded Economics Math

Work the §2.2 revenue split on a real product, then make the Pro Tip's "profit per policy" call.

Part A — calculate: A ₹600 gadget-protection premium is split Platform 40% / Enabler 10% / Insurer 50%. The insurer pays claims = 60% of its share and expenses = 18% of its share. Fill in the table:

PlayerShare of ₹600Your Answer (₹)
Platform (distribution partner)40%
Enabler (API provider)10%
Insurer (risk carrier)50%
Insurer profit per policy = insurer share − claims − expenses

Part B — decide: Two products compete for capital: Embedded — ₹500 premium, ₹200 profit per policy; Standalone — ₹1,000 premium, ₹100 profit per policy. Which creates more shareholder value, and why (use the Pro Tip)?

Part C — explain: Why can an embedded pool improve loss ratios by 5–15 points even before the insurer prices for it?

Check Your Math

Part A: Platform ₹240 · Enabler ₹60 · Insurer ₹300 · claims ₹180 (60% × 300) · expenses ₹54 (18% × 300) → insurer profit = ₹300 − ₹180 − ₹54 = ₹66 per policy. Note the insurer earns a healthy margin on a ₹600 product because its capital is deployed only on the risk it carries.

Part B: the embedded product — ₹200 profit per policy at ₹500 premium (40% margin) vs ₹100 on ₹1,000 (10% margin). It consumes less capital per rupee of profit, scales with the platform's growth, and typically auto-renews. Profit per policy, not premium per policy, is what builds shareholder value.

Part C: adverse selection reversal — the customer buys the primary product (the phone, the ride, the flight) for reasons unrelated to insurance, so the insurance pool resembles the general population instead of a self-selected high-risk segment. The self-selection that normally loads the insurance pool works in reverse here, improving the loss ratio structurally — before any behavioural pricing.

⚠ Volatile content — Reviewed: July 2026 · Next review: October 2026

3. Embedded Insurance in India

India's embedded insurance market has grown rapidly, driven by the combination of digital public infrastructure (India Stack), large consumer platforms, and InsurTech enablers that provide the API plumbing to connect them. The ecosystem can be mapped across four primary distribution contexts.

3.1 Embedded Insurance by Platform Type

Platform TypePlatform ExampleInsurance Product(s)Enabler / InsurerHow It Works
Mobility & Ride-Hailing Ola, Uber, Rapido Personal accident cover per ride, auto driver insurance Acko (on Ola), ICICI Lombard (on Uber) ₹0.5–₹5 per ride added to fare. Covers accidental death, permanent disability. Policy active for duration of ride. No separate purchase — enrolment is automatic (opt-out, not opt-in).
E-Commerce & Retail Amazon, Flipkart, Myntra Product protection plans (gadget damage + liquid + theft), extended warranty Zopper (on Flipkart/Myntra), Acko (on Amazon) Offered at checkout when purchasing electronics. 1–4 year plans. Customer enrolled instantly. Claims: WhatsApp photo → AI assessment → replacement or repair. 15–30% conversion of eligible purchases.
Travel & Hospitality MakeMyTrip, IRCTC (via Acko), Yatra, EaseMyTrip Travel insurance (medical abroad, trip cancellation, lost baggage, flight delay) Acko, ICICI Lombard, Tata AIG Offered during flight or hotel booking checkout. Premium: ₹99–₹999 depending on destination and trip duration. Highly contextual — destination with high medical costs = more likely to purchase.
Fintech & Payments PhonePe, Paytm, Cred, Google Pay Life insurance (sachet), accident cover, travel insurance, device protection Multiple insurers through Riskcovry and other enablers Offered within the payments/fintech app as a "sachet" product — very low premium (₹5–₹99), very short term (1 month to 1 year). High volume, low value. Low friction purchase — no forms, UPI-click to buy.

3.2 The Enabler Model: Riskcovry and Zopper

Two companies have built the infrastructure that powers embedded insurance at scale in India. Understanding their model is essential to understanding how embedded insurance works operationally.

Riskcovry positions itself as "insurance middleware" — an API platform that connects insurers to any distribution platform. A mobility app, e-commerce site, or fintech app can integrate Riskcovry's API in 2–4 weeks and offer insurance products from multiple carriers without building insurance-specific capabilities. Riskcovry handles: product configuration, API connection to multiple insurers, quote engine, policy issuance, claims API, and compliance. It earns a per-transaction fee or a revenue share. Riskcovry does not underwrite risk — it provides the technology layer. This is a capital-light, high-margin business model.

Zopper started as an extended warranty company and evolved into an embedded insurance platform focused on e-commerce. Zopper partners with Flipkart, Myntra, Tata CLiQ, and others to offer product protection plans at checkout. Unlike Riskcovry (which is a pure technology enabler), Zopper works with insurers to design specific products for the e-commerce context — product protection plans that cover accidental damage, liquid damage, and theft, with AI-based claims assessment via photo upload.

Warning — Platform concentration risk: embedded insurance's strength and its risk are the same thing: the InsurTech borrows another company's customers. That is what makes acquisition nearly free — but it also means the platform owns the relationship, and the InsurTech's access is rented, not owned.

The power imbalance is one-sided. When one platform supplies most of an InsurTech's volume, the platform can:

  • Demand a higher commission — the InsurTech can hardly refuse without losing most of its book.
  • Invite a second insurer to bid — competition squeezes the InsurTech's price and margin.
  • Replace the partner at renewal — partnerships are exclusive or semi-exclusive only for a fixed period; then they are renegotiated or ended.
  • Build its own insurance capability — the platform takes the margin itself.

That is why analysts watch one number: the share of volume from the single largest platform. Above roughly 40%, the business's survival depends on one counterparty's decisions — losing that partner means losing nearly half the business overnight, so investors treat it as a structural vulnerability. Below it, losing a partner is painful but survivable. When you evaluate any embedded InsurTech, ask: "What share of volume comes from platform #1 — and what would happen if that platform disappeared tomorrow?"

🌍 Exercise 3.1 — Match the Platform

For each embedded-insurance offer, pick the platform type it belongs to (from the §3.1 table) and the likely enabler/insurer.

#OfferPlatform TypeEnabler / Insurer
1A ₹1-per-ride accident cover added to a ride-hailing fare.
2A 2-year screen + liquid damage plan at smartphone checkout.
3A ₹299 travel-medical cover checkbox at flight checkout.
4A ₹50 sachet health plan inside a payments app.
5An extended-warranty plan on a laptop bought from an e-commerce marketplace.

Reflection: The §3.2 Warning says platform concentration risk is the biggest strategic risk for embedded insurers. What specific power does a platform have over an embedded InsurTech once the partnership exists?

Check Your Matches
  1. Mobility & Ride-Hailing · Acko — per-ride cover on Ola; enrolment is opt-out.
  2. E-Commerce & Retail · Acko (or Zopper) — checkout product protection on Amazon/Flipkart/Myntra.
  3. Travel & Hospitality · Acko — flight/hotel checkout on MakeMyTrip/IRCTC/Yatra.
  4. Fintech & Payments · Multiple via Riskcovry — sachet products inside PhonePe/Paytm/Cred/GPay.
  5. E-Commerce & Retail · Zopper — extended warranty / product protection on marketplaces.

Reflection note: the platform owns the customer relationship and the checkout — it can demand a higher commission, invite a second insurer tomorrow, or build its own insurance capability. If one platform is >40% of an InsurTech's volume, that single partner's decisions are existential. Diversification of distribution is the only durable defence.

4. Usage-Based Insurance (UBI)

Usage-Based Insurance uses data about actual behaviour — how much someone drives, how well they drive, what time of day they drive — to price insurance more accurately than traditional risk factors (age, gender, vehicle type, location). UBI is the most significant innovation in personal lines insurance pricing since the actuarial table, because it replaces inferred risk (what group you belong to) with observed risk (what you actually do).

4.1 PAYD vs. PHYD

DimensionPAYD — Pay-As-You-DrivePHYD — Pay-How-You-Drive
Data collectedDistance driven (kilometres/miles)Distance + speed + braking + cornering + time of day + phone usage
Data sourceOdometer reading, GPS, smartphone sensorTelematics device (OBD-II dongle), smartphone app, or vehicle API
Pricing modelFlat rate per km (may vary by road type)Variable rate per km adjusted by behavioural risk score (higher risk = higher per-km rate)
Risk accuracy vs. traditionalModerate improvement — mileage is a strong predictor of claim frequencySignificant improvement — driving behaviour predicts both claim frequency and severity
Premium savings for safe driver10–25% (if they drive less than average)20–60% (if they drive safely AND less than average)
Customer adoptionHigher — less intrusive, fewer privacy concernsLower — requires device installation or app permissions; privacy concerns higher
Regulatory acceptanceHigher — regulators view mileage-based pricing as fairMixed — some regulators concerned about behavioural scoring fairness and data privacy
Leading vendorsAllstate Milewise (US), Zego (UK)Root Insurance (US), Progressive Snapshot (US), ICICI Lombard DriveTrack (India)

4.2 Data Sources for UBI

UBI data comes from four sources, each with different cost, accuracy, and customer friction profiles:

📝
Note: The key UBI insight that many observers miss is: UBI customers are self-selected safer drivers. A driver who is willing to install a telematics device or an app that tracks their driving is, on average, a more conscientious driver than one who refuses. This creates a "reverse adverse selection" effect — the UBI pool has a better risk profile than the general pool BEFORE the behavioural data is even used to adjust prices. This self-selection effect may account for 40–60% of the observed loss ratio improvement in UBI programmes, with behavioural pricing adjustments accounting for the rest. When evaluating UBI results, separate the selection effect from the pricing effect.

🚗 Exercise 4.1 — PAYD or PHYD?

Classify each pricing decision, then run the PAYD vs PHYD arithmetic.

#Pricing DecisionPAYD / PHYD
1A driver pays a flat ₹1.50 per km regardless of how they drive.
2A driver's per-km rate rises 50% after a month of hard-braking alerts.
3A night-shift driver pays 1.3× the base rate for kilometres driven 10pm–5am.
4A low-mileage driver gets a 20% discount simply for driving 800 km/month.
5An app scores braking, cornering, and phone use and adjusts the renewal discount.

Calculate: A driver does 1,800 km/month in the Moderate tier. Compute the monthly premium under PAYD (₹1.50/km) and PHYD (₹0.80/km × 1.10 multiplier).

ModelMonthly Premium (₹)
PAYD
PHYD

Reflection: Which model risks underpricing an aggressive low-mileage driver — and why is that an adverse-selection risk for the insurer?

Check Your Answers
  1. PAYD — flat per-km, no behaviour adjustment.
  2. PHYD — behaviour (hard braking) changes the rate.
  3. PHYD — time-of-day behaviour adjusts the per-km rate.
  4. PAYD — the discount comes from distance alone, not behaviour.
  5. PHYD — a behavioural score drives the adjustment.

Calculation: PAYD = 1,800 × ₹1.50 = ₹2,700 · PHYD = 1,800 × ₹0.80 × 1.10 = ₹1,584. The safe moderate driver pays ₹1,116 less under PHYD than PAYD — behaviour is rewarded, not just mileage.

Reflection note: PAYD underprices an aggressive low-mileage driver — at 500 km/month they pay ₹750 (500 × ₹1.50) while driving dangerously, far below the ₹2,500 they would pay under a traditional flat premium. The insurer collects a premium that does not reflect the claim risk, and if such drivers self-select into PAYD, the pool's loss ratio worsens — adverse selection. PHYD fixes this with a behavioural multiplier on the per-km rate, so risky behaviour is priced even at low mileage. PHYD prices observed behaviour; PAYD prices only exposure.

5. UBI Data Analysis in Python

In this section, we work with a synthetic telematics dataset. The data represents a month of driving behaviour for 500 policyholders, including mileage, driving events, time-of-day patterns, and claims history. The objective is to build a risk score from behavioural data and compare PAYD (Pay-As-You-Drive) Vs PHYD (Pay-How-You-Drive).

🔗
How to run this section: the §5.1–§5.3 blocks below are one continuous script split across the page — each continues from the previous one, so copy them in order into the same Jupyter notebook cell sequence. The data file data/telematics_data.csv is provided in the lab folder (InsuranceTech_code/data/) and on GitHub (link below) — save your notebook in InsuranceTech_code/ so the relative path works. A single complete runnable script is available at the end of §5 (and as session_10_ubi_analysis.py in InsuranceTech_code/): run venv/bin/python session_10_ubi_analysis.py to reproduce everything at once. GitHub: telematics_data.csv · complete script.

5.1 Loading the Telematics Data

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# Load synthetic telematics data
telematics = pd.read_csv('data/telematics_data.csv')

print(f"Drivers: {telematics.shape[0]:,}")
print(f"Columns: {telematics.columns.tolist()}")
print(f"\nData types:\n{telematics.dtypes}")
print(f"\nSummary statistics:\n{telematics.describe()}")

# Key columns:
# driver_id — Unique driver identifier
# total_km — Kilometres driven in the month
# avg_speed_kmph — Average speed
# hard_braking_events — Count of hard braking events
# hard_accel_events — Count of rapid acceleration events
# night_driving_pct — Percentage of driving at night (10pm-5am)
# phone_use_events — Count of phone usage events while driving
# previous_claims — Number of claims in the last 3 years
# claim_amount_last_3y — Total claim amount in the last 3 years

5.2 Building a Behavioural Risk Score

▶ Prerequisite: run the §5.1 block first in the same notebook — telematics is defined there. This block will not run standalone.

# Define risk factor weights (these would be calibrated from claims data)
RISK_WEIGHTS = {
    'km_risk': 0.20,              # Higher mileage = more exposure
    'braking_risk': 0.25,         # Hard braking = unsafe driving pattern
    'accel_risk': 0.15,           # Hard acceleration = aggressive driving
    'night_risk': 0.20,           # Night driving = higher accident risk
    'phone_risk': 0.20            # Phone use = distraction risk
}

# Set thresholds for each risk factor
KM_THRESHOLD_HIGH = 2000       # km/month — above this is high mileage
BRAKING_THRESHOLD = 15         # events/month — above this is aggressive
ACCEL_THRESHOLD = 10           # events/month
NIGHT_THRESHOLD = 25           # % of driving at night
PHONE_THRESHOLD = 5            # events/month

# Calculate risk sub-scores (each normalised to 0–100)
telematics['km_risk'] = np.clip(
    (telematics['total_km'] / KM_THRESHOLD_HIGH) * 100, 0, 100
)

telematics['braking_risk'] = np.clip(
    (telematics['hard_braking_events'] / BRAKING_THRESHOLD) * 100, 0, 100
)

telematics['accel_risk'] = np.clip(
    (telematics['hard_accel_events'] / ACCEL_THRESHOLD) * 100, 0, 100
)

telematics['night_risk'] = np.clip(
    (telematics['night_driving_pct'] / NIGHT_THRESHOLD) * 100, 0, 100
)

telematics['phone_risk'] = np.clip(
    (telematics['phone_use_events'] / PHONE_THRESHOLD) * 100, 0, 100
)

# Composite risk score (weighted average)
telematics['risk_score'] = (
    telematics['km_risk'] * RISK_WEIGHTS['km_risk'] +
    telematics['braking_risk'] * RISK_WEIGHTS['braking_risk'] +
    telematics['accel_risk'] * RISK_WEIGHTS['accel_risk'] +
    telematics['night_risk'] * RISK_WEIGHTS['night_risk'] +
    telematics['phone_risk'] * RISK_WEIGHTS['phone_risk']
)

# Classify into risk tiers
telematics['risk_tier'] = pd.cut(
    telematics['risk_score'],
    bins=[0, 20, 40, 60, 100],
    labels=['Very Low', 'Low', 'Moderate', 'High']
)

print("Risk Score Distribution:")
print(telematics['risk_tier'].value_counts().sort_index())
print(f"\nMean risk score: {telematics['risk_score'].mean():.1f}")
print(f"Median risk score: {telematics['risk_score'].median():.1f}")
print(f"Std dev: {telematics['risk_score'].std():.1f}")

5.3 PAYD vs. PHYD Pricing Comparison

▶ Prerequisite: continues from §5.1–§5.2 — telematics must include risk_score and risk_tier. pandas 3.x note: mapping a categorical with a dict can fail on arithmetic — write telematics['risk_tier'].astype(str).map(PHYD_RISK_MULTIPLIER).astype(float) instead of .map(...) directly.

# PAYD: flat rate per km (traditional UBI)
PAYD_RATE = 1.50  # ₹1.50 per km

# PHYD: base rate + behavioural adjustment
PHYD_BASE_RATE = 0.80        # ₹0.80 per km (lower base because behaviour-adjusted)
PHYD_RISK_MULTIPLIER = {
    'Very Low': 0.60,
    'Low': 0.85,
    'Moderate': 1.10,
    'High': 1.50
}

# Calculate monthly premium under each model
telematics['payd_premium'] = telematics['total_km'] * PAYD_RATE

telematics['phyd_rate'] = telematics['risk_tier'].map(PHYD_RISK_MULTIPLIER)
telematics['phyd_premium'] = telematics['total_km'] * PHYD_BASE_RATE * telematics['phyd_rate']

# Calculate savings vs. traditional flat premium (assume ₹2,500/month)
TRADITIONAL_FLAT = 2500
telematics['traditional_premium'] = TRADITIONAL_FLAT

telematics['payd_savings'] = TRADITIONAL_FLAT - telematics['payd_premium']
telematics['phyd_savings'] = TRADITIONAL_FLAT - telematics['phyd_premium']

# Aggregate by risk tier
comparison = telematics.groupby('risk_tier', observed=False).agg(
    driver_count=('driver_id', 'count'),
    avg_km=('total_km', 'mean'),
    avg_risk_score=('risk_score', 'mean'),
    avg_payd=('payd_premium', 'mean'),
    avg_phyd=('phyd_premium', 'mean'),
    avg_traditional=('traditional_premium', 'mean'),
    avg_payd_savings=('payd_savings', 'mean'),
    avg_phyd_savings=('phyd_savings', 'mean'),
).round(0)

print("=" * 110)
print(f"{'Risk Tier':12s} {'Drivers':>8s} {'Avg Km':>8s} {'Risk':>6s} {'PAYD ₹':>9s} {'PHYD ₹':>9s} {'Trad ₹':>9s} {'PAYD Save':>10s} {'PHYD Save':>10s}")
print("=" * 110)
for tier in ['Very Low', 'Low', 'Moderate', 'High']:
    r = comparison.loc[tier]
    print(f"{tier:12s} {r['driver_count']:>5.0f}   {r['avg_km']:>5,.0f}  {r['avg_risk_score']:>4.0f}  ₹{r['avg_payd']:>6,.0f}  ₹{r['avg_phyd']:>6,.0f}  ₹{r['avg_traditional']:>5,.0f}  {'+' if r['avg_payd_savings'] > 0 else ''}₹{r['avg_payd_savings']:>6,.0f}  {'+' if r['avg_phyd_savings'] > 0 else ''}₹{r['avg_phyd_savings']:>6,.0f}")

# Visualize
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(14, 5))

# LEFT: Premium comparison by tier
tiers = ['Very Low', 'Low', 'Moderate', 'High']
x = np.arange(len(tiers))
width = 0.25

ax1.bar(x - width, comparison.loc[tiers, 'avg_traditional'], width, label='Traditional (Flat)', color='#a29bfe', alpha=0.7)
ax1.bar(x, comparison.loc[tiers, 'avg_payd'], width, label='PAYD', color='#6c5ce7')
ax1.bar(x + width, comparison.loc[tiers, 'avg_phyd'], width, label='PHYD', color='#00d2d3')
ax1.set_xlabel('Risk Tier')
ax1.set_ylabel('Monthly Premium (₹)')
ax1.set_title('Premium Comparison by Risk Tier', fontweight='bold')
ax1.set_xticks(x)
ax1.set_xticklabels(tiers)
ax1.legend()
ax1.spines['top'].set_visible(False)
ax1.spines['right'].set_visible(False)

# RIGHT: Scatter — risk score vs. claims
if 'claim_amount' in telematics.columns:
    ax2.scatter(telematics['risk_score'], telematics['claim_amount'],
               alpha=0.4, color='#6c5ce7', s=20)
    # Add trend line
    z = np.polyfit(telematics['risk_score'], telematics['claim_amount'], 1)
    p = np.poly1d(z)
    x_line = np.linspace(telematics['risk_score'].min(), telematics['risk_score'].max(), 100)
    ax2.plot(x_line, p(x_line), 'r--', linewidth=2, label='Trend')
    corr = telematics['risk_score'].corr(telematics['claim_amount'])
    ax2.annotate(f'Correlation: {corr:.2f}', xy=(0.05, 0.95),
                transform=ax2.transAxes, fontsize=11,
                bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
    ax2.set_title('Risk Score vs. Claim Amount', fontweight='bold')
    ax2.set_xlabel('Behavioural Risk Score (0–100)')
    ax2.set_ylabel('Claim Amount (₹)')
    ax2.legend()
    ax2.spines['top'].set_visible(False)
    ax2.spines['right'].set_visible(False)

plt.tight_layout()
plt.show()

5.4 Interpreting the Results

The analysis reveals the core UBI insight: PAYD primarily rewards low-mileage drivers regardless of their driving quality, while PHYD rewards safe driving behaviour regardless of mileage.

This is why PHYD is a better risk selection mechanism than PAYD — but it requires more data, more customer trust, and more regulatory acceptance to deploy at scale.

🌎
Real World: ICICI Lombard's DriveTrack programme in India uses an OBD-II dongle plugged into the car's diagnostic port to collect driving data. The device tracks speed, braking, acceleration, cornering, and time of day. At each renewal, the customer receives a driving score and a premium adjustment — safe drivers can save up to 25%. DriveTrack was one of the first telematics programmes in India and has provided ICICI Lombard with a rich dataset that it uses to calibrate its traditional pricing models as well. The data from DriveTrack revealed, for example, that the relationship between speeding and claims was significantly weaker in India than in the US (probably because Indian roads rarely permit sustained high-speed driving), while the relationship between night driving and claims was stronger than expected — insights that improved ICICI Lombard's pricing accuracy for all motor customers, not just DriveTrack participants.

💻 Exercise 5.1 — Trace the Risk Score

Use the §5.2 thresholds and weights to hand-compute a driver's risk score, then price them under PHYD.

Given — Driver X: total_km 1,000 · hard_braking 6 · hard_accel 4 · night_driving 10% · phone_use 2. Compute each sub-score = value ÷ threshold × 100 (clip at 100):

FactorThresholdSub-score (0–100)
km_risk (weight 0.20)2,000 km
braking_risk (0.25)15 events
accel_risk (0.15)10 events
night_risk (0.20)25%
phone_risk (0.20)5 events
Composite risk_scoreweighted sum

Then price: Which tier is Driver X in (§5.2 bins)? What PHYD multiplier applies (§5.3)? What is the monthly PHYD premium at ₹0.80/km?

Reflection: Why does the composite weight braking (0.25) more than acceleration (0.15)?

Check Your Score
  1. km_risk = 1000/2000 × 100 = 50
  2. braking_risk = 6/15 × 100 = 40
  3. accel_risk = 4/10 × 100 = 40
  4. night_risk = 10/25 × 100 = 40
  5. phone_risk = 2/5 × 100 = 40
  6. Composite = 0.20(50) + 0.25(40) + 0.15(40) + 0.20(40) + 0.20(40) = 10 + 10 + 6 + 8 + 8 = 42

Pricing: 42 falls in the Moderate tier (40–60) → multiplier 1.10 → PHYD premium = 1,000 × ₹0.80 × 1.10 = ₹880/month. Under PAYD the same driver pays ₹1,500 — PHYD's lower base rate rewards their moderate behaviour.

Reflection note: hard braking is the strongest single predictor of claim severity (and, with the §5.3 scatter, of claim amount), so it carries the largest weight. Acceleration correlates with aggression but is a weaker predictor on Indian roads than night driving or phone use — the weights mirror real calibration, where braking and night driving typically dominate.

Complete Runnable Script — Copy & Run (sections 5.1–5.4)

Everything from §5.1–§5.4 in one script, so nothing is left undefined. Save as session_10_ubi_analysis.py inside InsuranceTech_code/ and run venv/bin/python session_10_ubi_analysis.py, or paste it into one Jupyter cell. Requires data/telematics_data.csv (provided in the lab folder and on GitHub — link in the §5 run note). Outputs: the console tables below plus ubi_premium_comparison.png.

# ============================================================================
# Session 10 — Usage-Based Insurance: UBI Data Analysis (COMPLETE RUNNABLE SCRIPT)
# InsurTech & Digital Risk Solutions (MBA) — Woxsen University
#
# HOW TO RUN:
#   venv/bin/python session_10_ubi_analysis.py
#   (or run top-to-bottom in one Jupyter notebook — the page's §5.1–§5.4 blocks
#    are merged here in order so nothing is left undefined)
#
# DATA FILE: data/telematics_data.csv (provided) — 500 drivers, one month of
#            synthetic driving behaviour. Falls back to telematics_data.csv
#            in the current folder.
#
# OUTPUTS:   console tables + ubi_premium_comparison.png + ubi_risk_scatter.png
# ============================================================================
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# ----------------------------------------------------------------------------
# SECTION 5.1 — Loading the Telematics Data
# ----------------------------------------------------------------------------
try:
    telematics = pd.read_csv('data/telematics_data.csv')
except FileNotFoundError:
    telematics = pd.read_csv('telematics_data.csv')

print(f"Drivers: {telematics.shape[0]:,}")
print(f"Columns: {telematics.columns.tolist()}")
print(f"\nData types:\n{telematics.dtypes}")
print(f"\nSummary statistics:\n{telematics.describe().round(1)}")

# Key columns:
# driver_id — Unique driver identifier
# total_km — Kilometres driven in the month
# avg_speed_kmph — Average speed
# hard_braking_events — Count of hard braking events
# hard_accel_events — Count of rapid acceleration events
# night_driving_pct — Percentage of driving at night (10pm-5am)
# phone_use_events — Count of phone usage events while driving
# previous_claims — Number of claims in the last 3 years
# claim_amount_last_3y — Total claim amount in the last 3 years

# ----------------------------------------------------------------------------
# SECTION 5.2 — Building a Behavioural Risk Score
# ----------------------------------------------------------------------------
RISK_WEIGHTS = {
    'km_risk': 0.20,              # Higher mileage = more exposure
    'braking_risk': 0.25,         # Hard braking = unsafe driving pattern
    'accel_risk': 0.15,           # Hard acceleration = aggressive driving
    'night_risk': 0.20,           # Night driving = higher accident risk
    'phone_risk': 0.20            # Phone use = distraction risk
}

KM_THRESHOLD_HIGH = 2000       # km/month — above this is high mileage
BRAKING_THRESHOLD = 15         # events/month — above this is aggressive
ACCEL_THRESHOLD = 10           # events/month
NIGHT_THRESHOLD = 25           # % of driving at night
PHONE_THRESHOLD = 5            # events/month

telematics['km_risk'] = np.clip((telematics['total_km'] / KM_THRESHOLD_HIGH) * 100, 0, 100)
telematics['braking_risk'] = np.clip((telematics['hard_braking_events'] / BRAKING_THRESHOLD) * 100, 0, 100)
telematics['accel_risk'] = np.clip((telematics['hard_accel_events'] / ACCEL_THRESHOLD) * 100, 0, 100)
telematics['night_risk'] = np.clip((telematics['night_driving_pct'] / NIGHT_THRESHOLD) * 100, 0, 100)
telematics['phone_risk'] = np.clip((telematics['phone_use_events'] / PHONE_THRESHOLD) * 100, 0, 100)

telematics['risk_score'] = (
    telematics['km_risk'] * RISK_WEIGHTS['km_risk'] +
    telematics['braking_risk'] * RISK_WEIGHTS['braking_risk'] +
    telematics['accel_risk'] * RISK_WEIGHTS['accel_risk'] +
    telematics['night_risk'] * RISK_WEIGHTS['night_risk'] +
    telematics['phone_risk'] * RISK_WEIGHTS['phone_risk']
)

telematics['risk_tier'] = pd.cut(
    telematics['risk_score'],
    bins=[0, 20, 40, 60, 100],
    labels=['Very Low', 'Low', 'Moderate', 'High'],
    include_lowest=True
)

print("Risk Score Distribution:")
print(telematics['risk_tier'].value_counts().sort_index())
print(f"\nMean risk score: {telematics['risk_score'].mean():.1f}")
print(f"Median risk score: {telematics['risk_score'].median():.1f}")
print(f"Std dev: {telematics['risk_score'].std():.1f}")

# ----------------------------------------------------------------------------
# SECTION 5.3 — PAYD vs. PHYD Pricing Comparison
# ----------------------------------------------------------------------------
PAYD_RATE = 1.50  # ₹1.50 per km
PHYD_BASE_RATE = 0.80        # ₹0.80 per km (lower base because behaviour-adjusted)
PHYD_RISK_MULTIPLIER = {
    'Very Low': 0.60,
    'Low': 0.85,
    'Moderate': 1.10,
    'High': 1.50
}

telematics['payd_premium'] = telematics['total_km'] * PAYD_RATE
# pandas 3.x: map a categorical with a dict can return a non-arithmetic dtype,
# so cast to str first, then to float.
telematics['phyd_rate'] = telematics['risk_tier'].astype(str).map(PHYD_RISK_MULTIPLIER).astype(float)
telematics['phyd_premium'] = telematics['total_km'] * PHYD_BASE_RATE * telematics['phyd_rate']

TRADITIONAL_FLAT = 2500
telematics['traditional_premium'] = TRADITIONAL_FLAT
telematics['payd_savings'] = TRADITIONAL_FLAT - telematics['payd_premium']
telematics['phyd_savings'] = TRADITIONAL_FLAT - telematics['phyd_premium']

comparison = telematics.groupby('risk_tier', observed=False).agg(
    driver_count=('driver_id', 'count'),
    avg_km=('total_km', 'mean'),
    avg_risk_score=('risk_score', 'mean'),
    avg_payd=('payd_premium', 'mean'),
    avg_phyd=('phyd_premium', 'mean'),
    avg_traditional=('traditional_premium', 'mean'),
    avg_payd_savings=('payd_savings', 'mean'),
    avg_phyd_savings=('phyd_savings', 'mean'),
).round(0)

print("=" * 110)
print(f"{'Risk Tier':12s} {'Drivers':>8s} {'Avg Km':>8s} {'Risk':>6s} {'PAYD ₹':>9s} {'PHYD ₹':>9s} {'Trad ₹':>9s} {'PAYD Save':>10s} {'PHYD Save':>10s}")
print("=" * 110)
for tier in ['Very Low', 'Low', 'Moderate', 'High']:
    r = comparison.loc[tier]
    print(f"{tier:12s} {r['driver_count']:>5.0f}   {r['avg_km']:>5,.0f}  {r['avg_risk_score']:>4.0f}  ₹{r['avg_payd']:>6,.0f}  ₹{r['avg_phyd']:>6,.0f}  ₹{r['avg_traditional']:>5,.0f}  {'+' if r['avg_payd_savings'] > 0 else ''}₹{r['avg_payd_savings']:>6,.0f}  {'+' if r['avg_phyd_savings'] > 0 else ''}₹{r['avg_phyd_savings']:>6,.0f}")

# Visualize
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(14, 5))

tiers = ['Very Low', 'Low', 'Moderate', 'High']
x = np.arange(len(tiers))
width = 0.25

ax1.bar(x - width, comparison.loc[tiers, 'avg_traditional'], width, label='Traditional (Flat)', color='#a29bfe', alpha=0.7)
ax1.bar(x, comparison.loc[tiers, 'avg_payd'], width, label='PAYD', color='#6c5ce7')
ax1.bar(x + width, comparison.loc[tiers, 'avg_phyd'], width, label='PHYD', color='#00d2d3')
ax1.set_xlabel('Risk Tier')
ax1.set_ylabel('Monthly Premium (₹)')
ax1.set_title('Premium Comparison by Risk Tier', fontweight='bold')
ax1.set_xticks(x)
ax1.set_xticklabels(tiers)
ax1.legend()
ax1.spines['top'].set_visible(False)
ax1.spines['right'].set_visible(False)

if 'claim_amount' in telematics.columns:
    ax2.scatter(telematics['risk_score'], telematics['claim_amount'],
               alpha=0.4, color='#6c5ce7', s=20)
    z = np.polyfit(telematics['risk_score'], telematics['claim_amount'], 1)
    p = np.poly1d(z)
    x_line = np.linspace(telematics['risk_score'].min(), telematics['risk_score'].max(), 100)
    ax2.plot(x_line, p(x_line), 'r--', linewidth=2, label='Trend')
    corr = telematics['risk_score'].corr(telematics['claim_amount'])
    ax2.annotate(f'Correlation: {corr:.2f}', xy=(0.05, 0.95),
                transform=ax2.transAxes, fontsize=11,
                bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
    ax2.set_title('Risk Score vs. Claim Amount', fontweight='bold')
    ax2.set_xlabel('Behavioural Risk Score (0–100)')
    ax2.set_ylabel('Claim Amount (₹)')
    ax2.legend()
    ax2.spines['top'].set_visible(False)
    ax2.spines['right'].set_visible(False)

plt.tight_layout()
plt.savefig('ubi_premium_comparison.png', dpi=120)   # artifact for submission
plt.show()

# ----------------------------------------------------------------------------
# SECTION 5.4 — Interpreting the Results (worked examples from the page)
# ----------------------------------------------------------------------------
print("\n" + "=" * 70)
print("SECTION 5.4 — Interpretation examples")
print("=" * 70)

aggressive_low_km = 500     # km/month, aggressive driving (High tier)
payd = aggressive_low_km * PAYD_RATE
phyd = aggressive_low_km * PHYD_BASE_RATE * 1.50
print(f"Aggressive low-mileage driver (500 km, High tier):")
print(f"  PAYD = ₹{payd:.0f}  |  PHYD = ₹{phyd:.0f}  |  Traditional = ₹{TRADITIONAL_FLAT}")
print(f"  -> PAYD underprices the risk; PHYD adjusts it up by the behavioural multiplier.")

safe_high_km = 3000        # km/month, safe driving (Very Low tier)
payd = safe_high_km * PAYD_RATE
phyd = safe_high_km * PHYD_BASE_RATE * 0.60
print(f"\nSafe high-mileage driver (3000 km, Very Low tier):")
print(f"  PAYD = ₹{payd:.0f}  |  PHYD = ₹{phyd:.0f}  |  Traditional = ₹{TRADITIONAL_FLAT}")
print(f"  -> PAYD overcharges the safe driver; PHYD rewards them with a discount.")

print("\nCharts saved: ubi_premium_comparison.png")
📋 Stable content — Reviewed: July 2026

6. Global UBI Leaders

Usage-Based Insurance has a 15+ year history globally, with several companies reaching significant scale. Their experiences provide lessons for any market — including India — where UBI is still nascent.

6.1 Root Insurance (US)

Root Insurance (founded 2015, IPO 2020 at $6B+) was the most ambitious UBI InsurTech globally. Root's model was PHYD-only: customers install a smartphone app that tracks driving for 2–4 weeks, Root calculates a driving score, and if the score is good enough, Root offers a premium. If the score is poor, Root declines to offer coverage — the customer is "uninsurable by Root's standards." This is the opposite of traditional insurance, where the insurer must accept all applicants (subject to regulatory constraints) and price accordingly.

What went right: Root's risk selection was genuinely superior — its loss ratio on policies that went through the full telematics scoring process was approximately 15–20 points better than industry average. The self-selection effect (safer drivers are more willing to be tracked) was powerful.

What went wrong: Root's growth stalled because nearly 70% of potential customers did not complete the app-based driving test — they either declined to install the app, installed but did not drive enough to generate a score, or scored poorly and were rejected. The customer acquisition cost for the remaining 30% was extremely high. By 2023, Root's market cap had fallen to ~$100M (a 95%+ decline from its peak) as growth stalled and underwriting losses from its early expansion years caught up.

Key lesson: A superior risk selection mechanism is not enough if the customer acquisition funnel is so restrictive that it limits volume and drives up CAC. UBI programmes must balance risk accuracy with customer accessibility. The best UBI model may be one that does not require a pre-purchase driving test but instead adjusts premiums dynamically as driving data accumulates — "start with traditional pricing, earn a discount as you prove you are a safe driver."

6.2 Progressive Snapshot (US)

Progressive's Snapshot programme is the largest UBI programme in the US, with millions of active participants. It uses an OBD-II dongle (or smartphone app) that monitors driving behaviour and offers discounts of up to 30% at renewal based on safe driving. Progressive's approach differs from Root in two critical ways: (a) Snapshot is optional and available to all Progressive customers — it does not gate access to coverage on the driving score, and (b) the programme rewards safe driving with discounts rather than penalizing risky driving with surcharges. This "carrot, not stick" design makes Snapshot far more acceptable to customers and regulators.

6.3 ICICI Lombard DriveTrack (India)

India's most established UBI programme is ICICI Lombard's DriveTrack, which gives customers a free OBD-II dongle that tracks driving behaviour. The programme structure is closer to Progressive's model: customers enrol voluntarily, receive a driving score after each policy period, and safe drivers earn premium discounts at renewal. DriveTrack has been active since approximately 2018 and has accumulated significant data on Indian driving patterns. The programme has not been transformative for ICICI Lombard's motor insurance business (it covers a small fraction of the overall portfolio), but it has generated valuable data that the company uses to refine its pricing models for all motor customers.

🌎
Real World: The Root Insurance case is one of the most instructive failures in InsurTech. Root's investors lost billions. Root's customers who completed the driving test got excellent insurance at fair prices. The paradox at the heart of UBI is: the customers who benefit most from usage-based pricing (high-risk drivers who would be overcharged in a traditional pool, and low-risk drivers who want to prove their safety) are the ones least likely to enrol — the first group because they know their score will be bad, the second because they do not need to prove anything to an incumbent insurer that already gives them a good rate. The customers most willing to enrol are moderate-risk drivers who believe they are better than average — and they are often right. But this self-selection creates a pool that UBI must price correctly, which requires data that only the UBI programme itself can generate. This circular dependency — "we need data to price correctly, but we need the right pricing to attract the data" — is the core UBI business model challenge.

🏁 Exercise 6.1 — UBI Strategy Clinic

Three UBI launches need a diagnosis and a strategy call, using the §6 lessons from Root, Progressive, and DriveTrack.

#SituationDiagnosisStrategy Call
1A startup requires a 3-week app driving test before it will quote anyone (Root-style).
2An insurer offers a discount at renewal to safe drivers, never a surcharge (Progressive-style).
3A new Indian UBI launch must pick a data-collection device for mass adoption.

Design: Write the 3-sentence product memo for the Indian launch — discount structure, observation period, and the one risk you mitigate first.

Check Your Diagnoses
  1. Restrictive funnel / CAC problem — Root's pre-purchase test saw ~70% of customers drop out, crushing volume and CAC. The fix: stop gating access — "start with traditional pricing, earn a discount as you prove you drive safely" (the §6.1 key lesson).
  2. Discount-only model — the "carrot, not stick" design is why Snapshot has run at scale for a decade: no customer ever receives a higher bill, so trust and regulatory goodwill are preserved. It leaves the surcharge on the table but keeps enrolment and retention high.
  3. Mass adoption → smartphone app — zero hardware cost (vs ₹2,000–5,000 OBD dongles) and the widest reach, per the hands-on memo's logic; OBD/OEM can follow for precision later. (Any of app-first answers with a defensible reason is acceptable.)

Model memo: "Base premium ₹2,500 (same as traditional — no customer pays MORE by enrolling). Safe drivers earn up to 40% off at renewal after a 3-month observation period; moderate and high-risk tiers simply pay the traditional premium. The one risk we mitigate first is premium shock — a customer discovering a higher bill — because it is the single biggest reputational risk for UBI (§7 Warning)."

7. Privacy, Ethics & UBI

Usage-Based Insurance sits at the intersection of two powerful trends — data-driven personalisation and consumer privacy — that are increasingly in tension. As UBI expands, the privacy and ethics questions it raises become not peripheral concerns but central strategic issues that will determine how quickly and how broadly UBI can be adopted.

7.1 The Data Collection Question

A PHYD programme may collect: GPS location every few seconds, speed, acceleration on three axes, braking force, cornering G-force, time of day, phone usage events, and — if the smartphone app is used — background location data even when not driving (to distinguish driving from other travel modes). For a customer, the question is: what does the insurer know about me, when, and for how long do they keep it? For the insurer, the question is: can we build a commercially viable UBI programme without collecting data that customers and regulators find intrusive?

The tension is real. Studies consistently show that 40–60% of consumers say they are "very concerned" about sharing driving data with their insurer — but 60–80% say they would share data for a meaningful premium discount. The "privacy premium" — how much discount a customer demands in exchange for sharing data — varies by demographics, product type, and cultural context. Indian customers have generally shown less privacy sensitivity than European customers (where GDPR imposes strict limits), but sensitivity is rising with awareness.

7.2 The Fairness Question

Usage-based pricing raises three fairness questions that regulators and insurers are actively debating:

7.3 The DPDP Act and UBI

India's Digital Personal Data Protection Act (DPDP Act, 2023) has direct implications for UBI programmes. Key requirements: explicit consent for data collection (separate from the insurance contract), purpose limitation (UBI data cannot be used for other purposes without fresh consent), data minimisation (collect only what is necessary), and the right to withdraw consent (the customer can stop sharing data at any time — which would require the insurer to revert to a non-UBI pricing mechanism).

For UBI providers in India, the DPDP Act means: (a) the consent form for telematics data must be written in clear, simple language — not buried in the policy terms, (b) data collection should be limited to what is actually used for pricing — collecting GPS location continuously when the purpose is mileage calculation may violate the data minimisation principle, and (c) the option to opt out of data sharing must be as easy as opting in. UBI programmes designed with the DPDP Act in mind from the start will have a significant regulatory advantage over those that treat privacy compliance as an afterthought.

Warning: The single biggest reputational risk for UBI is not a data breach — it is a customer discovering that their premium increased because of their driving data. A driver who enrols in a UBI programme expecting a discount and instead receives a surcharge is a customer who will: (a) never enrol in UBI again, (b) tell everyone they know that the programme is a "trap," and (c) potentially contact the regulator. The safest UBI programme design is "discount only" — use telematics data only to offer discounts to safe drivers, never to surcharge risky ones. This leaves money on the table (the insurer cannot capture the extra premium that risky drivers should pay), but it preserves customer trust and regulatory goodwill. Progressive's Snapshot has operated on this model for over a decade, and it is the primary reason the programme has been sustainable at scale.

⚖ Exercise 7.1 — The Privacy Tax Debate

Identify the principle behind each UBI concern, then play regulator on a filing.

#ConcernPrinciple
1A night-shift driver pays 50% more per km because they cannot avoid driving at night.
2Safe drivers migrate to UBI; the remaining traditional pool's premiums must rise.
3A customer switches insurers but their safe-driving data stays with the old insurer.
4An insurer collects continuous GPS when only mileage is used for pricing.

Be the regulator: A UBI filing collects app telematics for 6 months, then adjusts renewal premiums up to ±40%. Approve or reject — and if you approve, write the THREE DPDP conditions you impose.

Check Your Analysis
  1. Behavioural discrimination by price — pricing based on behaviour the customer cannot easily change functions more like a penalty than a risk-reflective price (§7.2).
  2. Segregation / privacy tax — data-poor customers (by privacy preference, not by choice) pay more as the low-risk migrate out (§7.2).
  3. Data ownership & portability — driving data belongs to the customer but stays with the collecting insurer (§7.2).
  4. Data minimisation (DPDP) — collecting more than the pricing purpose needs violates the DPDP data-minimisation principle (§7.3).

Regulator's verdict (model): Approve with conditions — (1) explicit, plain-language consent separate from the policy contract; (2) purpose limitation: collect only the driving events used for pricing, no continuous GPS; (3) one-click opt-out as easy as opt-in, reverting the customer to traditional pricing. Also consider a cap: given the §7 Warning, many regulators would reject a ±40% surcharge and demand discount-only (±0 to −40%) to protect trust.

Hands-On Project: Build a UBI Risk Scoring Model

You are a data scientist at "SafeDrive Insurance," a hypothetical InsurTech launching a Usage-Based Insurance product for private car owners in India. Your task is to build a behavioural risk scoring model from telematics data, compare PAYD vs. PHYD pricing for 100 policyholders, and recommend a product strategy.

Steps

  1. Load the telematics data from `data/telematics_data.csv`. Explore the columns and distributions. Print summary statistics for all numeric variables. Identify any data quality issues (missing values, extreme outliers).
  2. Build a behavioural risk score using the following factors with equal weights (20% each):
    • Hard braking events per 100 km (normalize to 0–100)
    • Hard acceleration events per 100 km (normalize to 0–100)
    • Night driving percentage (0–100 scale)
    • Phone use events per 100 km (0–100 scale)
    • Average speed deviation from the speed limit (0–100 scale, where 0 = no deviation)
  3. Segment drivers into risk tiers using the 25th, 50th, and 75th percentiles of the risk score as thresholds. Label as Very Low (≤25th percentile), Low (25–50th), Moderate (50–75th), and High (>75th).
  4. Calculate premium under three models for each driver:
    • Traditional: flat ₹3,000/month for all drivers
    • PAYD: ₹1.75 per km driven
    • PHYD: ₹1.00 per km × risk multiplier (Very Low: 0.55, Low: 0.80, Moderate: 1.15, High: 1.60)
  5. Analyse the results: For each risk tier, calculate the average premium under each model and the savings vs. traditional. Identify: (a) Which drivers benefit most from PAYD? (b) Which benefit most from PHYD? (c) Are there drivers who would pay more under BOTH UBI models than traditional?
  6. Create three visualizations: (a) Histogram of risk scores with quartile lines. (b) Premium comparison (grouped bar chart by risk tier for all 3 models). (c) Scatter plot of risk score vs. premium paid under PHYD, coloured by risk tier.
  7. Write a 300-word product recommendation to the CEO of SafeDrive Insurance: would you launch PAYD, PHYD, or both? What pricing structure do you recommend? What are the key risks and how would you mitigate them? What customer segments would you target first?
View Solution / Walkthrough

Solution Approach

The solution follows the analysis framework from Sections 5.2–5.4. Key code components are identical to the code blocks in this session. The main deliverable is the product recommendation memo. A sample memo structure is provided below.

Sample Product Recommendation Memo

To: CEO, SafeDrive Insurance
From: Data Science — UBI Product Team
Subject: UBI Product Launch Recommendation

Executive Summary:
Our analysis of the telematics pilot data (n=100 drivers) confirms that behavioural risk scoring predicts claim experience significantly better than traditional mileage-only metrics. The composite risk score shows a correlation of approximately 0.35 with historical claim amounts, compared to 0.18 for mileage alone. A UBI programme is strategically viable. We recommend launching a PHYD-only product with a "discount-only" pricing structure — no surcharges above the traditional premium — to minimise customer friction and regulatory risk.

Recommended Product Structure:

  • Base monthly premium: ₹2,500 (same as traditional, so no customer pays MORE by enrolling)
  • Maximum discount: 40% (₹1,500/month) for Very Low risk tier
  • Typical discount: 15–20% for Low risk tier
  • No discount (but no surcharge) for Moderate and High risk tiers — they simply pay the traditional premium
  • Discount is applied at renewal, not immediately — customers earn their discount over a 3-month observation period

Rationale: A discount-only PHYD model addresses the three key risks. First, it eliminates the "premium shock" risk — no customer receives a higher bill because of UBI. Second, it reduces regulatory risk — IRDAI is more likely to approve a discount-only programme than a surcharge programme. Third, it improves customer acquisition — marketing "save up to 40% on your car insurance" will drive significantly higher enrolment than a message about "fair pricing based on your driving."

Launch recommendation: Target a controlled pilot of 5,000 policyholders in two cities (Delhi NCR and Bengaluru — chosen for their demographic diversity and telematics coverage). Use a smartphone app (not OBD dongle) for data collection — it has zero hardware cost and reaches the widest customer base. Validate loss ratios for 6 months before expanding to 50,000 policyholders. After 12 months and statistical validation of the pricing model, consider introducing a "premium adjustment" for the High risk tier — but only with clear customer communication and regulatory approval.

Key risk and mitigation: The most significant risk is low enrolment due to privacy concerns. Mitigation: (1) a clear, simple privacy notice written in plain language (not legalese), (2) data collection limited to driving events (no continuous GPS tracking), (3) a monthly "privacy report" showing each customer exactly what data was collected, and (4) a one-click option to pause data collection and revert to traditional pricing for that month. We believe this combination of transparency and control will achieve an enrolment rate of 25–35% of willing customers — sufficient for a statistically viable UBI programme.

Key Takeaways

1

Embedded insurance converts at 3–5× the rate of standalone digital channels because it benefits from contextual conversion, near-zero marginal CAC, trust transfer from the platform, and the "reverse adverse selection" effect.

2

India's embedded insurance ecosystem spans mobility, e-commerce, travel, and fintech — enabled by API platforms (Riskcovry, Zopper) that connect insurers to distribution partners. Platform concentration risk is the primary strategic vulnerability.

3

PAYD (flat per-km pricing) rewards low mileage regardless of driving quality. PHYD (behaviour-adjusted per-km pricing) rewards safe driving. PHYD is a better risk selection mechanism but requires more data and customer trust.

4

The global UBI experience shows that a superior risk model is not enough — if the enrolment funnel is too restrictive (Root's 30% completion rate), the CAC makes the economics unviable. The most successful UBI programmes use a "discount-only" structure and do not gate access to coverage.

5

UBI raises unresolved privacy and fairness questions — the "privacy tax" on data-poor customers, behavioural discrimination (night-driving pricing), and data portability. DPDP Act compliance in India requires explicit consent, purpose limitation, and data minimisation from day one.

3-2-1 Reflection — Before You Move On

Insurers read behaviour — practise translating today's frameworks into decisions. Write from memory, don't scroll back.

3 Things I Learned Today

2 Embedded / UBI Concepts I Can Now Explain

1 Question I Still Have About Embedded & Usage-Based Insurance

Test Your Understanding

1. An embedded insurance offer during e-commerce checkout consistently converts at 25% while the same insurer's standalone website converts at 3%. The MAIN reason for this difference is:

2. Under the PHYD pricing model, a driver who drives 1,500 km/month with a "High" risk tier (multiplier = 1.50) and a base rate of ₹1.00/km would pay:

3. An InsurTech that derives 60% of its embedded insurance premium volume from a single e-commerce platform partnership faces what strategic risk?

4. Root Insurance's UBI programme failed to achieve sustainable scale primarily because:

5. A "discount-only" UBI programme (where telematics data can only LOWER a customer's premium, never increase it) is recommended as the safest launch strategy because: