Adoption is near-universal. Deployment is not.
The November 2024 joint survey by the Bank of England and the FCA covered 118 UK financial-services firms. Three-quarters reported active use of AI, with a further 10% planning adoption inside three years. Insurance was the highest-adopting sector at 95% — ahead of international banks at 94%.
The second set of numbers in the same survey is the one that should shape your buying decisions. Foundation models account for 17% of all AI use cases. Only 2% of use cases are fully autonomous; 24% are semi-autonomous; 55% involve some form of automated decision-making. The dominant pattern is augmentation, not replacement — AI extracts data, ranks priorities, drafts responses, predicts likelihoods, and flags anomalies, while a person approves, rejects, or overrides.
The picture repeats elsewhere. Conning’s 2025 survey shows LLM adoption among US insurers rising from 18% to 63% in a single year, while only 7% have scaled AI across the enterprise. The Lloyd’s Market Association found that 14% of London-market firms have deployed agentic or generative AI in underwriting, and 65% have deployed it in no part of underwriting or claims at all.
Three claims the evidence supports
Lead scoring produces measurable conversion lift
Supervised-learning lead models produce a consistent two-to-three-times conversion uplift on the top-scored segment, plus an efficiency saving from deprioritising the bottom. One documented mid-market insurance deployment showed top-scoring leads converting at 3.5 times the average, with the bottom fifth excluded from outreach entirely, and a 1.5% profit-line improvement inside the first months. Progressive has publicly attributed roughly $2bn in new premium volume to a single ML-driven app-feature improvement.
The technique is not exotic — gradient-boosted decision trees, increasingly with transformer architectures for the unstructured inputs. What it needs is roughly 18 to 36 months of clean attribution data, a defined conversion event, and producers who will act on the score. The third condition is the one most often underestimated. A score producers do not trust changes nothing.
Document AI cuts submission intake time by 60–90%
On submission intake the evidence is unusually consistent across vendors and operators. Time-per-submission falls 60–90% where intelligent document processing is deployed on the dominant ACORD forms, with first-pass extraction accuracy of 97–98% on printed fields and 93–96% on handwritten. The mechanism is well understood: a commercial-lines underwriter spends 30–45% of the working day on data entry and document handling before any risk evaluation begins.
Algorithmic follow placement works at Lloyd’s scale
Ki, launched by Brit in 2020 and separated into Fairfax Group in January 2025, wrote over $1bn of gross written premium in 2024 and $1.11bn in 2025, including more than $200m of partner capacity from Beazley, QBE, Aspen, Travelers, and Tokio Marine Kiln. A broker submits a risk and receives an algorithmically priced follow line in as little as ten seconds.
These three share a shape. Narrow, repeatable tasks. Published operating numbers from named firms. Augmentation rather than replacement. And the regulatory perimeter left intact.
Three claims it does not
AI will replace agents and brokers
Lemonade is the most architecturally committed AI-first insurer in the market, and its own filings are the clearest counter-evidence. At the end of 2025 it reported 96% of first notices of loss taken without human intervention — and 55% of claims automated end to end. The second number is the meaningful one: 45% of claims still need a person. Lemonade also posted a $1,464.3m net loss in 2025. Klarna’s 2024 replacement of 700 customer-service staff, followed within months by rehiring after service quality collapsed, is the cautionary case outside insurance.
Generative AI will replace underwriters
Peer-reviewed work on agentic AI for commercial underwriting finds LLMs hallucinating on 11.3% of cases without adversarial-critique architectures, and 3.8% with them. Frontier-model benchmarking in late 2025 found some small model variants hallucinating on up to 19% of completed underwriting traces. A 3.8% rate is workable for augmentation. It is not a basis for binding commercial coverage unsupervised.
This is a description of the current state, not a permanent verdict. But procurement decisions made in 2026 commit you to a governance architecture for three to five years, and that architecture should be calibrated to what the technology does now.
AI compliance is principally an IT problem
The EU AI Act, the NAIC Model Bulletin, Colorado’s SB 21-169, and the UK’s Consumer Duty all treat AI governance as an actuarial, legal, and senior-management responsibility. The most common procurement failure we see is an AI programme run as a CTO project, where the documentation, bias-testing, and human-oversight architecture needed to defend the system was never written into the contract. Retrofitting that costs considerably more than building it in.
Submission intake: the largest measured gain in distribution
A commercial-lines underwriter receives 40 to 60 submissions a week. Each is a bundle: an ACORD 125, line-specific supplements, three years of loss runs, financial statements, supplemental questionnaires, and a broker cover narrative — arriving in different formats, scanned at varying quality, often with handwritten margin notes. Before any risk evaluation begins, someone reads every page and keys the data into a workbench.
Carriers quote only about half the submissions they receive. The constraint is not underwriting expertise; it is the speed at which submissions can be ingested, normalised, and routed to the underwriter best equipped to write them.
Modern intake combines three capabilities: classification (this attachment is an ACORD 125, that one is a loss run), extraction (pull the structured fields, including from handwriting), and triage (score against appetite and remaining capacity, then route). Reported results cluster tightly — a 25-minute-to-2-minute reduction in triage time in one deployment; a 1,200-submission renewal queue processed without the six staff on overtime it previously needed.
In the April 2025 LMA survey of 81 firms, including 45 managing agents representing the bulk of the market’s £56.2bn stamp capacity, data extraction from unstructured documents was the single most prevalent use case at 74%, with submission preparation at 54%.
This is the use case that most directly addresses the distribution capacity constraint. A senior commercial underwriter is the most expensive seat in distribution and the supply is tight. AI-augmented intake creates effective underwriting capacity out of the staff you already have. In a hardening market that is worth more than any customer-facing application.
Algorithmic follow: capacity, not efficiency
In the London market a lead underwriter takes the first share and sets terms; follow underwriters take subsequent shares at those terms. Filling out a placement has historically meant a broker walking from box to box, or its electronic equivalent. Algorithmic follow replaces that with a model that prices the submission, decides how much capacity to offer against carrier-defined appetite and portfolio composition, and commits the line — in seconds.
The governance architecture around it is now well developed, and it is the part worth copying:
- Carrier-defined appetite constraints expressed as algorithmic parameters.
- Continuous portfolio monitoring against accumulation, rate adequacy, and class mix.
- Pre-defined human-escalation thresholds for risks outside the algorithm’s authority.
- An audit trail of every algorithmic decision, traceable to the input data.
- Periodic actuarial review and recalibration with documented sign-off.
The pattern generalises beyond Lloyd’s — algorithmic capacity allocation in property-catastrophe reinsurance, AI-augmented binding authorities in US E&S, API-based capacity provisioning for embedded distribution. What it creates is not speed. It is capacity: the carrier writes a class at a volume and granularity that skilled-underwriter headcount alone could not cover, while the appetite parameters hold underwriting discipline in place. It is the most strategically important AI use case in distribution and gets a fraction of the attention that chatbots do.
Conversational interfaces: where they hold and where they break
This is the most visible category and the one where the distance between demo and production is widest. It works, and is now standard practice, in three places: quote intake for low-complexity personal lines (Lemonade’s onboarding agent asks 13 questions while collecting more than 1,600 underlying data points, binding inside 90 seconds); first notice of loss in personal lines; and policy self-service, where one US carrier’s assistant handles 60% of routine enquiries and escalates the rest.
It underperforms in three places that matter to distribution:
- Complex claims handling. Past simple FNOL capture, the handoff to a human handler is fragile, and the regulatory exposure on automated denial decisions is significant.
- Producer-facing assistants. Despite years of vendor claims, the assistant that helps an experienced producer write better business remains a complement to judgement, not a substitute for it.
- Cross-channel handoffs. A conversation that starts in web chat, escalates to voice, and returns by email is where most deployments lose context — and where record-keeping expectations make the engineering hardest.
One procurement date to diarise: under the May 2026 Digital Omnibus agreement, watermarking and provenance-labelling obligations for AI-generated content apply from 2 December 2026. Any customer-facing chatbot or voice agent producing content for EU customers needs technical traceability. Treat it as a selection criterion now, not a planning item for later.
Producer analytics and the politics of prediction
Distribution generates performance data that should be a goldmine for ML. In practice this is the slowest-adopted application area, because the politics of producer compensation make any predictive intervention sensitive.
Four applications change behaviour in the published evidence:
- Account-rounding propensity. Identify the 5–10% of single-line policyholders most likely to accept a bundle and route them first. Low implementation cost, large uplift on the targeted segment, minimal producer resistance because the work is additive.
- At-risk renewal prediction. Identify the 10–15% of renewals most likely to lapse without intervention, so retention effort is allocated by risk-adjusted expected value rather than by renewal date. Among the highest-return applications measured in retention pounds per producer hour.
- Carrier-appetite matching. On the broker side, identify the carriers most likely to bind a given risk at competitive terms, cutting time spent on placements that were never going to bind.
- New-producer ramp prediction. Predict which new producers will hit retention thresholds in their first 12 months, so coaching starts early.
A related and separately deployable application is commission reconciliation — matching incoming carrier statements against expected schedules, flagging discrepancies, routing corrections. High volume, low judgement, immediate return, and the audit-trail benefit outlasts the operating saving. We set out the numbers in The Hidden Cost of Commission Leakage.
Hallucination, bias, and what human-in-the-loop actually means
Three risks cut across every use case above.
Hallucination
Any LLM application touching a binding, pricing, or denial decision needs explicit mitigation in its architecture: adversarial self-critique, retrieval-augmented generation with verifiable sources, structured-output enforcement, human verification. None of these are free. A procurement that does not specify them is accepting a risk tolerance the regulators will not.
Bias
Bias in distribution AI is actively regulated in at least three jurisdictions. Colorado is the most quantitative, the EU the most procedural, the UK the most outcomes-focused — but the expectation is the same everywhere: systems used in pricing or denial decisions must be tested for disparate impact, the testing documented, and the documentation produced on request. The common failure is treating bias-testing as a vendor certification. Both the NAIC Model Bulletin and Colorado’s regulations are explicit that responsibility stays with the insurer whoever built the model. A vendor model card is an input, not a substitute.
Human-in-the-loop
Most deployments claim a human in the loop. Three things separate the ones that hold from the decorative ones:
- Trigger sensitivity. The reviewer is triggered by informative signals — confidence thresholds, anomaly detection, exception flags — not a uniform sampling rate. A 5% random review is theatre; a 100% review of low-confidence outputs is governance.
- Reviewer authority. The reviewer can override the model in practice, not only in policy. An architecture that nominally includes a reviewer but defaults to the model output is automation with documentation.
- Audit trail integrity. Every override and every approval logged with timestamp, reviewer identity, and rationale. This is the line most often underestimated at procurement.
Doing this properly is not cheap. Doing it badly costs more in expectation, once regulatory exposure and litigation risk are priced in.
The regulatory perimeter is hardening
The EU AI Act classifies AI used for risk assessment and pricing in life and health insurance as high-risk under Annex III. Motor, home, commercial property, and casualty pricing are not explicitly listed, but the Act’s provisions on automated decision-making affecting essential services may reach them — the prudent posture is to govern substantial P&C pricing AI as if it were high-risk.
The May 2026 Digital Omnibus agreement moved the deadlines: stand-alone high-risk systems from 2 August 2026 to 2 December 2027, product-embedded systems to 2 August 2028, and content watermarking forward to 2 December 2026. The amendments are provisional pending formal adoption, so plan against the new dates and keep the ability to fall back to the originals.
| Jurisdiction | Framework | Key obligation |
|---|---|---|
| EU | EU AI Act (Regulation 2024/1689) | Conformity assessment, data governance, logging, human oversight, and post-market monitoring for high-risk AI in life and health pricing |
| UK | FCA principles, Consumer Duty, SM&CR | Fair value across the AI lifecycle; senior-manager accountability that cannot be delegated to a model or a vendor |
| US — Colorado | SB 21-169 / Reg 10-1-1; SB 24-205 | Quantitative disparate-impact testing and annual attestation; scope widened in October 2025 to motor and health |
| US — 24+ states | NAIC Model Bulletin | A written AI Systems Program covering governance, risk management, and third-party vendor oversight |
The UK has deliberately not written sector-specific AI rules. The FCA’s position is that Consumer Duty, SM&CR, SYSC governance, and the outsourcing and operational-resilience regime already cover the substance. In practice, for a firm selling into both markets, the EU Act is the binding constraint on system architecture and the Consumer Duty is the binding constraint on outcomes monitoring. A single programme can satisfy both — but only if it was designed for both. Retrofitting either onto a system built for the other is expensive.
From January 2026, twelve US states began running the NAIC AI Systems Evaluation Tool, with a deep dive on high-risk systems and a data-source review covering proxy-discrimination screening. If you write business in a pilot state, expect substantive examination of AI governance now. If you do not, expect it within 12 to 24 months.
Build, buy, and the integration tax
The answer to build-versus-buy is nearly always both. The mix is what matters:
- Buy foundation models. Building your own is uneconomic for any insurer or distributor short of the very largest.
- Buy specialist vendors for well-characterised workflows where the vendor has a real reference book — intake, pricing, fraud, claims. Building these in-house typically produces a worse result at higher cost.
- Build integration, orchestration, and data infrastructure. This is where proprietary advantage is created and where off-the-shelf is least viable.
- Build governance, audit trail, and bias-testing. Vendor documentation is supporting evidence, not a programme.
Three drivers explain most of it. Core-system interface maturity varies module by module. Data normalisation is usually the dominant line, because AI vendors need structured, normalised inputs and most core systems do not produce them natively — and it is the work you cannot skip, because data quality dominates model performance. And governance integration takes both technical and process work, of which the process half decides whether the governance holds in production.
Buyers who succeed treat this as a five-year capability decision rather than a twelve-month software purchase. They specify the integration architecture before selecting the vendor, write audit rights and bias-testing cooperation into the contract, and sequence low-regulatory-risk workflows — intake, internal triage — ahead of pricing and denial decisions.
The data flywheel is the part that lasts
AI is deployed against a workflow. The deployment generates structured data about customer behaviour, risk characteristics, and outcomes that previously existed only as unstructured information or did not exist at all. That data improves the model, the model improves outcomes, better outcomes attract more business, and more business produces more data.
Three conditions have to hold for the loop to close. The deployment has to touch the moment of interaction, not just the back office around it — submission triage produces data about broker behaviour, quote intake produces data about customer intent, while pure paperwork extraction produces none of it. The data has to be operationally consequential: logging it is not enough, the model has to retrain on it. And the volume has to clear the threshold where improvements are measurable rather than noise.
MGAs are the segment best positioned to capture this. They combine direct distribution relationships, proprietary specialty data, and short procurement cycles — a combination rarely co-located elsewhere. US MGA direct premium nearly doubled from $47bn in 2020 to $97bn in 2024, roughly 14% a year, well above the broader market. The structural advantage is real. Converting it into a durable position is the work.
What to do next
- Inventory every AI system in use, including vendor-supplied ones, and classify each against the Annex III taxonomy and the NAIC Bulletin’s three pillars.
- Name an accountable owner for each — not a technology owner, an outcome owner who can answer to a senior manager.
- Start with intake and internal triage. Lowest regulatory risk, largest measured return, and it builds the structured data the later use cases need.
- Fix the data substrate before buying the model. Data normalisation dominates model performance, and paying that cost once as infrastructure beats paying it per AI capability. See From Spreadsheets to a Unified Distribution Platform.
- Budget two to four times the licence for first-year integration, and write audit rights, model documentation, and bias-testing cooperation into the contract before signing.
- Test your human-in-the-loop architecture against the three questions above: is the trigger informative, does the reviewer have real authority, and is every decision logged?
