When anyone thinks of AI in insurance, Claims is the oft-quoted use case. And, as you may have guessed, it’s the one that disappoints most reliably in production. The pattern is so consistent it’s become part of the folklore.
A carrier announces a multi-year transformation. A press release names a partner. Eighteen months later, the cycle-time improvements are real but smaller than promised, the straight-through rates have nudged up by a few points, and the executive who sponsored the program has moved on. The next carrier announces a similar program.
This is not a story about AI maturity. It’s a story about how claims has been scoped, sequenced, and rebuilt, and about why a different sequence, anchored in how the work is actually performed rather than how it’s described, produces a different outcome.
Claims is also the place I’d point an operations leader who wants to see what agentic AI in regulated work will look like in practice. The structure of the work (sub-process density, variant explosion, regulatory overlay, and judgment calls that should remain human) is a near-perfect test bed. If you can do agents well here, you can do them well almost anywhere.
What the claims journey actually contains
A homeowners or auto claim from first notice of loss through settlement looks, on the map, like five or six steps. FNOL. Coverage verification. Investigation and assessment. Reserving. Negotiation and settlement. Closure and subrogation.
The map is correct in the same way that a map of the United States is correct. It tells you the shape, not the territory. The actual journey, traced through a month of real claims at a top-twenty US carrier, contains something closer to twelve to fifteen sub-processes for that one line of business, depending on how you slice them. Each sub-process has variants. Each variant has its own exception patterns. The work moves across three to seven systems, depending on the line of business. It crosses functional boundaries (adjusting, special investigations, salvage, vendor management, customer service), and the handoffs are where most of the cycle-time pain lives.
If you compare the documented process to the observed one for a typical auto bodily injury claim, the documented version has about thirty steps. The observed version has closer to two hundred discrete interactions, distributed across roughly forty variant paths. Most of the variation is invisible to the people who drew the map, because most of it is in the second-by-second behavior of the adjuster, not in the workflow states the system records.
That gap, between thirty documented steps and two hundred observed ones, is where every claims automation program of the last decade has gotten into trouble.
Why straight-through processing has been the wrong frame
The dominant frame for claims automation has been straight-through processing. The goal, stated and restated, is to move some percentage of claims from FNOL to settlement without human intervention. The percentage gets revised upward in every roadmap and downward in every retrospective.
The frame is structurally wrong. Here’s why.
STP treats the claim as the unit of automation. The question is: can this whole claim go through without a person touching it? But a claim isn’t a process. A claim is a case, a particular configuration of policy, loss, claimant, and evidence, that moves through a sequence of sub-processes, each with its own automation potential.
A clean comprehensive auto claim with photo-documented damage on a low-deductible policy might have eight sub-processes that are completely automatable, two that are partially automatable, and one (say, the conversation with the insured about a rental car timeline) that should remain human because the customer experience of having an agent handle it is worse. The right question isn’t whether the claim goes through automatically. It’s which sub-processes go through automatically, which get agent assistance, and which stay with humans by design.
This sounds like a small reframing. It changes everything downstream. The KPI changes from “STP rate” to “sub-process automation coverage.” The architecture changes from one big workflow to a set of composable agents. The scope changes from “all claims” to “the variants of each sub-process where the work is stable enough to entrust.” The governance changes from one go/no-go decision to a portfolio of them, each with its own evidence base.
This reframing is what separates the carriers making real progress from the ones still relaunching the same program.
A sub-process view of the work
Let me sketch what this looks like with a concrete sub-process: coverage verification on a homeowner’s claim.
Coverage verification, on the map, is one step. Confirm the loss is covered under the policy. In the work, it’s a sequence of decisions that varies substantially by what kind of loss is being claimed.
For a wind damage claim in a non-named-storm region, coverage verification is mostly checking policy effective dates, deductibles, and a short list of exclusions. The work is stable. The variants are limited. The judgment content is low. This sub-process is a strong candidate for agentic execution. An agent equipped with the policy document, the loss details, and the carrier’s coverage interpretation guidelines can perform it reliably, with clear escalation paths for the cases that don’t fit.
For a water damage claim, coverage verification is a different exercise. The carrier has to determine whether the damage is from a covered cause (sudden discharge from a plumbing failure) versus an excluded one (long-term seepage, flood). The determination depends on the inspection report, the claimant’s narrative, sometimes the public adjuster’s framing, and a body of carrier-specific interpretation that has accreted over years of litigation and settlements. The work is less stable. The variants are wider. The judgment content is high. This sub-process is a candidate for agent assistance: the agent prepares the case, flags the relevant interpretation precedents, and presents a recommended position to a human adjuster who makes the call.
For a smoke damage claim from a wildfire that’s still being investigated, coverage verification is largely a holding action: the carrier is waiting on the cause-and-origin report from the fire marshal, the regulatory environment may be shifting, and the claim is being managed against a backdrop of catastrophe response that affects everything from reserving to communication tone. This sub-process probably stays human, end to end, until the cat event resolves.
Three variants of the same sub-process. Three different deployment strategies. The agent doesn’t handle “coverage verification.” It handles the wind-damage variant of coverage verification in non-named-storm regions, under stable policy conditions. The scope is precise because it has to be.
This is the level of resolution at which agentic claims processing actually works.
What a pragmatic rebuild looks like
If a carrier asked me what a process-first rebuild of claims should look like, the sequence would be something like this.
First, observe the work for 30-60 days. Not interview people about the work. Observe it — across all the systems, all the lines of business, all the variants. The output is a process record, not a process map. It captures what people do, in what order, in which systems, with what exceptions, producing what outcomes. The record covers documented steps and undocumented ones, including workarounds, shortcuts, and cases that don’t fit any of the named variants.
Second, segment the record into sub-processes and variants. This is the work consultants used to do with stickies on a wall. It’s now done by analyzing the observed data, which gives you frequencies and variation the stickies never had. The output is a sub-process inventory. For a mid-size carrier, this typically contains forty to sixty sub-processes across all lines of business, each with three to twelve variants. The sub-process inventory is the unit of analysis for everything downstream.
Third, score each variant on stability, judgment content, exception density, and customer-experience sensitivity. Stability is how consistent the variant is across cases. High-stability variants are good automation candidates; low-stability ones are not. Judgment content is how much of the work depends on calls that should remain human. Exception density is how often the variant produces unexpected paths. Customer-experience sensitivity is how much the customer cares whether a person handled it.
Those four scores are what point a variant at its deployment. Stable, low-judgment, low-exception, and low-CX-sensitivity, and the variant can run autonomously. Add judgment to that mix and it becomes agent-assisted, with the agent preparing and a human deciding. Low stability or high judgment, and the work stays with a person, perhaps with a copilot alongside. And anything customers care deeply about stays human regardless of what the other three scores say.
Fourth, deploy in waves, with feedback loops. Each variant moves to its target deployment in a controlled rollout. The observation layer keeps running. The agent’s behavior is monitored against the human baseline that the record established. When the agent’s confidence drops, or when the variant drifts, or when an exception pattern emerges that wasn’t in the original record, the system surfaces it before it becomes a problem. The carrier learns continuously.
Fifth, retire the project model. This is the step most programs skip, and skipping it is why the cycle keeps repeating. The work I’ve just described isn’t a project. It’s the new operating model. The observation layer keeps running. The variant catalog keeps evolving. The agent estate keeps being tuned. There’s no completion event, no transition from “transformation” back to “business as usual.” Claims processing, from here, looks more like a continuously tuned engineering system than a periodically transformed business process.
Which steps stay human, by design
A serious claims rebuild is also serious about what doesn’t change. Some sub-processes should not move to agents, and the discipline of naming them up front is part of what makes the rest of the program credible.
Catastrophic loss conversations. When a customer has lost their home, the first conversation isn’t an efficiency problem to optimize. It’s a moment that defines the customer’s relationship with the carrier for the rest of the claim and often beyond. An agent can handle the paperwork, the documentation requests, and the scheduling. The conversation itself stays with a person.
Coverage disputes. When a coverage decision is going to result in a partial or full denial, the communication is consequential: for the customer, for the regulator, and for the carrier’s reputation. An agent can prepare the case. A human owns the communication.
Suspected fraud, on the human side. Special Investigations Unit work is judgment-dense, exception-heavy, and bound by investigative practices that change slowly and carry high stakes. The agent’s role here is data preparation and pattern surfacing. The investigative judgment stays with the SIU.
Bad-faith adjacent moments. Any point in the claim where the carrier’s behavior, if mishandled, could expose it to extra-contractual liability. The threshold for “could mishandle” is set conservatively. This is not where a program should be trying to push the frontier.
The point of this list isn’t the specific items. It’s that a credible claims program names them, and names them early. A program that says it’s going to automate everything is a program that hasn’t done the scoping seriously. The scoping is part of the work.
What this requires that most carriers don’t have
The precondition for any of this is a faithful record of how the work is currently performed. Without it, the sub-process inventory is guesswork, the variant scoring is guesswork, and the deployment strategy is guesswork. The agents will be deployed against an idealized version of the work, and they will fail against the real one.
The carriers that have moved fastest in the last two years have, almost without exception, invested in continuous process observation before investing in agent deployment. The ones that did it the other way around have spent the same money and gotten less for it.
This is not a coincidence. Agentic systems work in proportion to the quality of the context they are given. In claims, the context is the work. If the carrier knows how its claims are actually being handled, across every variant, every system, every exception, the agents work. If the carrier knows only how its claims are supposed to be handled, the agents underperform, the program loses credibility, and the next executive starts the cycle over.
The path from here
Insurance has spent a decade announcing the future of claims and then producing modest, incremental improvements to the present of claims. The reason isn’t that the technology wasn’t ready. It’s that the work hadn’t been mapped at the resolution the technology required.
The good news: this is now a solvable problem. Continuous observation of claims work at scale is technically feasible and operationally affordable in 2026 in a way it wasn’t in 2020. The sub-process view is well-established in carriers that have done the work. The deployment patterns (autonomous on the stable variants, assistive on the judgment-rich ones, human-only where it matters) are no longer theoretical.
Carriers that take this seriously over the next eighteen months will end the decade with a claims operation that’s meaningfully faster, more consistent, and more defensible. The ones that keep buying agents without rebuilding the substrate will end the decade roughly where they are. The technology won’t be the differentiator. The discipline of doing the process work first will be.
That’s the real frontier in claims. Not the model. The work.
