Technology

The most dangerous automation is the one underwriters quietly repair

When experts routinely correct AI or workflow output without recording the intervention, the organization hides both model risk and its best source of product learning.

The most dangerous automation is the one underwriters quietly repair

The submission arrives with the wrong occupancy, a missing attachment and a risk classification that does not fit the wording in the broker’s email. The intake system has produced a usable-looking record, but the underwriter knows better. She changes three fields, finds the document and proceeds to quote.

The case is saved. The automation appears to have succeeded.

That small act of repair creates a serious management problem for managing general agents. As AI-assisted intake moves into production, an expert can prevent a bad output from becoming a bad underwriting decision while simultaneously concealing the weakness that produced it. If the correction is not recorded, management sees neither the model risk nor the expert knowledge needed to improve the product.

This is not an argument for recording every click. It is an argument for treating material repair as operating evidence: capturing interventions that change the interpretation, eligibility, pricing path or handling of a risk; distinguishing their causes; and assigning repeated failures to an accountable owner. The strongest objection is legitimate. Excessive capture will slow underwriters and generate telemetry that nobody uses. The answer is not comprehensive surveillance. It is selective, decision-relevant evidence.

What the evidence supports

The supplied external evidence is limited to the National Association of Insurance Commissioners’ artificial intelligence resource, which sets the relevant context as governance expectations for insurance AI. That source supports a narrow proposition: insurers using AI should regard its operation as a governance matter, rather than treating it solely as a software feature.

It does not, on the evidence supplied here, establish how frequently underwriters correct automated output, which kinds of repair are most common or how much financial loss goes undetected. No claims should be made about prevalence, industry benchmarks or return on investment. The case instead rests on an observable control gap: when a system’s output is corrected before it reaches the final record, the final record alone cannot reveal that the system was wrong.

The distinction matters because a clean outcome is not the same as a sound process. A file may be accurate because the automation worked, or because an experienced employee repaired it. Those paths carry different implications for scaling, staffing, oversight and product design. An MGA that cannot distinguish them may overestimate automation quality precisely because its strongest underwriters are compensating for it.

The invisible control layer

Quiet repair is often rational at the individual level. The underwriter wants to move the submission forward, not diagnose a workflow. Correcting a field may take seconds; reporting the failure may require a ticket, a screenshot and an uncertain wait for resolution. Production targets reward the completed case. They rarely reward evidence that the path to completion was defective.

The result is an informal control layer built from judgement and memory. Experienced underwriters learn which outputs to distrust, which brokers’ documents require manual reconciliation and which classifications need contextual interpretation. Newer staff may not know the same exceptions. Managers may see acceptable turnaround times without seeing the expertise consumed to achieve them.

This creates two distinct risks.

The first is conventional model or workflow risk. A wrong output may eventually pass through uncorrected, particularly when volumes rise, staffing changes or a case looks routine. Human review is only a reliable control if management knows what reviewers are expected to catch and whether they are catching it.

The second is product-learning risk. Each expert correction can contain a small piece of specification: this document takes precedence over that field; this phrase changes the occupancy; this exposure belongs outside the apparent class; this case reveals an unresolved appetite boundary. If those interventions disappear into the final record, the organization loses examples that could refine rules, prompts, data mappings, training or underwriting guidance.

Not every edit contains such value. That is why the unit of capture should be material repair, not activity.

A practical threshold for materiality

A repair is material when leaving the output unchanged could alter the underwriting path. That includes interventions that affect eligibility, referral, classification, exposure interpretation, pricing inputs, required evidence or the decision to quote. Cosmetic formatting, navigation choices and harmless wording changes do not meet the threshold.

This approach requires a comparison between proposed output and accepted output, but not necessarily a narrative from the underwriter each time. In many workflows, the system can retain changed values automatically and ask for a short reason only when the change crosses a defined threshold. The operating principle is that capture should be proportionate to consequence.

The threshold also exposes a design choice. If every correction is labelled an override, employees may infer that disagreement is discouraged or that their performance is being monitored. They may then accept dubious outputs, make changes outside the system or choose the least controversial reason code. A repair log designed as surveillance will corrupt its own evidence.

Management should therefore separate correction rates from individual performance assessment unless there is a clear control reason not to do so. A high repair rate may indicate careful underwriting, poor automation, difficult business or ambiguous appetite. It is not, by itself, evidence of low productivity or resistance to technology.

Three kinds of failure

Captured repairs become useful only when they are classified. A single “AI error” category would obscure the different operating responses required.

A data error occurs when the workflow has the wrong, incomplete or conflicting input. The remedy may sit in document handling, broker submission requirements, field mapping or data precedence. Retraining a model will not fix a missing endorsement or an incorrect source field.

A model error occurs when adequate inputs are available but the system extracts, classifies, summarises or recommends incorrectly. That points towards model evaluation, rules, prompts, confidence thresholds or human-review design. Even here, one corrected case proves little. Repetition across comparable cases is the stronger signal.

Appetite ambiguity is different again. The system may have interpreted the material reasonably, yet the MGA’s own rules do not resolve the case. Underwriters then repair the output by applying tacit judgement. Calling that a model failure would send the problem to the wrong team. The necessary decision belongs to underwriting leadership: clarify the appetite, preserve discretion or specify a referral.

These categories are not always clean. A poor source document can induce an erroneous model output, while vague appetite can make the “correct” classification contestable. The purpose is not perfect taxonomy. It is to prevent all exceptions from disappearing into a generic technology backlog.

Ownership after capture

Telemetry without ownership merely makes neglect measurable. Recurring repairs need a workflow owner with authority to change the relevant process and an obligation to close the loop.

That owner should not automatically be the technology team. Document defects may belong with distribution operations; mapping failures with data owners; model behaviour with the product or engineering function; appetite ambiguity with underwriting leadership. The accountable party is the one able to change the source of recurrence.

This has a second-order consequence. Once repair evidence is visible, departments may dispute classification because ownership carries cost. Technology teams may call a failure bad data. Operations may call it model behaviour. Underwriting may describe an unclear rule as an exceptional case. Governance must therefore examine patterns using the underlying examples, not rely only on summary labels.

The strongest counterposition

The case against systematic capture deserves weight. Underwriters already face administrative burdens. Additional prompts can interrupt judgement, delay response and encourage mechanical coding. Large volumes of trivial changes may overwhelm review teams, while dashboards create the appearance of control without improving decisions. In a fast-moving MGA, the cure could cost more than the defect.

That argument defeats indiscriminate logging, not selective capture. A defensible system should collect most evidence passively, trigger explanation only for consequential changes and review patterns rather than investigate every event. It should also retire fields and reason codes that do not lead to decisions. If telemetry does not change a rule, workflow, model test or control, its value should be questioned.

There is also a legitimate case for preserving expert discretion without forcing every judgement into a rigid taxonomy. Some underwriting decisions are contextual. The objective is not to eliminate judgement but to identify where the business depends on unrecorded judgement at scale.

Predictions that can be tested

If quiet repair is masking automation weakness, three results should become visible after material-repair capture is introduced.

First, apparent straight-through success should decline before it improves, because previously hidden intervention will be reclassified. If reported automation performance remains unchanged while repair capture rises, either the repairs are immaterial or the original performance measure was not sensitive to intervention.

Second, repairs should cluster around a limited number of fields, document types, rules or appetite boundaries. If they remain evenly dispersed with no recurrence, the business case for dedicated remediation will be weaker.

Third, targeted fixes should reduce repeat repairs in the affected category without increasing downstream referrals or corrections. A lower override rate alone is insufficient: employees may simply stop recording interventions. The prediction is falsified if recorded repairs fall while later defects or off-system work rise.

Questions for executives

Executives should ask where accepted output can differ from proposed output without leaving evidence. Which changes can alter eligibility, pricing or referral? Who owns each recurring repair category? Are underwriters rewarded for exposing defects, or only for working around them? Can management tell whether lower correction rates reflect better automation or weaker reporting? And if the most experienced underwriter left tomorrow, which invisible repairs would leave with her?

The dangerous system is not necessarily the one that makes obvious mistakes. It is the one whose mistakes are made invisible by competent people.