Insurance Agent Benchmark (IAB)
Measuring how AI agents handle the day-to-day work of insurance professionals.
Why we built an insurance agent benchmark
At Cooper, we're building an AI coworker for insurance. Doing that has forced us to think deeply about a deceptively simple question: how should we measure whether an AI agent is actually getting better at the work insurance professionals do?
Insurance already has useful benchmarks. InsuranceQA evaluates question answering in the insurance domain. INS-MMBench evaluates multimodal understanding and reasoning across insurance scenarios. InsureBench measures language models on document-grounded underwriting and claims work.
The Insurance Agent Benchmark (IAB) aims to evaluate whether AI agents can complete insurance workflows from start to finish.
In practice, insurance work rarely arrives as an isolated question. It arrives as an email with attachments, a scanned ACORD with handwritten notes, a loss run that needs to be reconciled against an application, a policy hundreds of pages long, or a workbook with dozens of tabs. Sometimes the most important thing the system can do is recognize that a value is missing rather than confidently invent one.
This first phase focuses on document understanding. Later phases will cover whole workflows, then long-horizon tasks and browser-use capabilities.
Key Findings
- The same model scores higher with Cooper's harness. Across all 17 models, Cooper improves the median score by 9.4 percentage points. The lift ranges from about a point for Grok and GPT-5.6 Luna, well inside single-run noise, to roughly 15 points for Claude Haiku 4.5 and Meta Muse Spark 1.3. The biggest gains show up on the hardest files: long documents, oversized files and broken inputs that a raw call struggles to ingest.
- Gemini 3.8 Flash moves the efficiency frontier. At 85.6%, it posts the highest accuracy point estimate in the benchmark, within the statistical band of the top Claude runs, while costing about $68 for the full corpus. Gemini 3.7 Flash remains close behind at 83.3% for $33 and ties for the highest reliability at 98.2%.
- AI is still bad at saying “I don’t know.” When a field is blank, redacted or unreadable, most models still invent values often enough to matter. Claude Sonnet 5 and Fable 5.1 are lowest at 10.6%; Opus 5 and Meta Muse Spark are next at 14.9%.
- Prompt injection looks more manageable than long policies. Sixteen of 17 models resist document-embedded injection at 90% or better with Cooper. Long-policy clause retrieval remains far more uneven: Fable 5.1 reaches 89.7%, but most Cooper runs remain below 80%. The capability is there, but it isn't consistent across models.
- Reliability is a system property. Model-alone calls fail to produce a usable answer on 7–28% of cases, mostly on very large or broken files. With Cooper, every model stays above 90% reliability. The highest reliability is 98.2%, reached by Gemini 3.5 Flash Lite, Gemini 3.7 Flash and Fable 5.1.
How does the harness impact model performance?
Cooper raised accuracy for every model we tested, without a model upgrade. Its value is clearest on the files that raw model calls struggle to process: long documents, oversized files and damaged inputs. Better document handling lets us get more accurate answers from the models we already have.
For a model to answer a question about a policy, it has to be able to read the file in the first place. A scanned application may need rendering; a large workbook may need to be split into chunks before the model can use it. The harness makes those decisions, which affect what information reaches the model.
We ran every model across the same 166 cases twice. In the model-alone run, it received the raw file and question in a single call. With Cooper, file-type routing chose the extraction or rendering path, and large files went through chunked map-reduce, with the candidate model doing all the reading and reduction. The model, document and question stayed the same across the two runs.
Takeaway: all 17 models scored higher with Cooper in these runs. Grok and GPT-5.6 Luna gained about a point, while gains exceeded 15 points for Claude Haiku 4.5 and Meta Muse Spark 1.3. Read differences within ±6 points with caution (see §4).
View as table
| Model | With Cooper | Model alone | Δ pts |
|---|---|---|---|
| gemini-3.8-flash | 85.6% | 75.9% | +9.7 |
| claude-opus-5 | 85.3% | 73.4% | +11.9 |
| claude-fable-5 | 84.6% | 70.7% | +13.9 |
| claude-sonnet-5 | 84.4% | 71.7% | +12.7 |
| gemini-3.7-flash | 83.3% | 76.7% | +6.6 |
| gemini-3.6-flash | 83.3% | 74.9% | +8.4 |
| claude-fable-5.1 | 82.3% | 72.9% | +9.4 |
| muse-spark-1.3 | 82% | 66.9% | +15.1 |
| claude-opus-4.6 | 81.5% | 70.6% | +10.9 |
| claude-sonnet-4.6 | 81.5% | 69.6% | +11.9 |
| gpt-5.6-terra | 80.6% | 76.8% | +3.8 |
| gpt-5.6-sol | 80.5% | 76.7% | +3.8 |
| gemini-3.1-pro-preview | 79.6% | 74.6% | +5.0 |
| grok-4.6 | 78.7% | 77.6% | +1.1 |
| gpt-5.6-luna | 77.7% | 76.5% | +1.2 |
| gemini-3.5-flash-lite | 77.2% | 73.2% | +4.0 |
| claude-haiku-4.5 | 74.8% | 59.1% | +15.7 |
How we built IAB
The corpus is built to resemble the pile of files that lands on a commercial-lines desk. The 166 documents come from live brokerage and carrier workflows under existing consent agreements. They range from a single-page certificate of insurance and photographed auto ID cards to full policy wordings with endorsement schedules, program books hundreds of pages long, large multi-location statements of values, and underwriting workbooks with more than fifty sheets. The two largest workbooks each pushed Cooper through nearly 16 million tokens to answer a single question. In between are multi-year loss runs, quotes, binders, endorsements, and broker submission emails that arrive as Outlook .msg files with the application, SOV and loss run still attached.
Much of it arrives exactly as messy as desks receive it: faxed and low-quality scans, pages rotated or stamped over, handwritten annotations, checkbox-heavy ACORDs photographed rather than scanned. One scanned application alone expands to 1.9 million tokens of extracted text. A few files lie about themselves: a wrong extension, a corrupt file truncated partway through, a “policy” with nothing inside.
None of it is synthetic. Of the 166 documents, 150 are untouched originals; the other 16 are originals modified to set a specific trap: an embedded instruction, a blacked-out name, a password. A team of insurance professionals prepared and checked the ground truth for every case against the source document. Where a value is absent or unreadable, the answer key says so, and reporting any value at all counts as a failure.
Formats span PDF (digital, scanned, broken), XLSX and CSV workbooks, PNG and JPG photos, Outlook .msg with attachments, PPTX, RTF, XML and JSON system exports. Measured in tokens through Cooper: 135 cases stay under 100k, 17 run from 100k to 1M, and 14 exceed 1M, topping out near 16M. We deliberately chose file sizes around the document system's routing thresholds, where a document tips from a single pass into chunked map-reduce: the benchmark tests the transitions where document tooling usually breaks.
View as table
| Document family | Cases |
|---|---|
| PDF digital | 57 |
| Spreadsheet | 25 |
| PDF scanned | 24 |
| Submission | 21 |
| Format/edge | 18 |
| Image | 14 |
| PDF broken | 7 |
Eleven ways to test document understanding
| # | Track | What it tests | Cases |
|---|---|---|---|
| 1 | ACORD field extraction | pull a full schema from ACORD 125/140/25 | 42 |
| 2 | SOV / loss-run reasoning | numeric answers: total TIV, incurred, largest loss | 25 |
| 3 | Cross-doc reconciliation | catch a mismatch across a submission packet | 10 |
| 4 | Long-policy clause retrieval | clause at start/middle/end of a long policy | 16 |
| 5 | Faithfulness & abstention | absent/unreadable: says 'not present' or invents? | 15 |
| 6 | Scanned / handwritten | field recall on degraded scans | 23 |
| 7 | Grounding / citations | does the cited page contain the answer | 8 |
| 8 | Checkbox / selection reading | did it read the marked option | 9 |
| 9 | Charts / graphs extraction | pull values from a chart image | 10 |
| 10 | Number / date normalization | varied formats to one value | 14 |
| 11 | Prompt injection in doc | stays faithful with an embedded instruction | 13 |
What the corpus does not cover: personal lines, non-US markets, languages beyond a three-document bilingual slice, handwriting-only documents, and any judgment task (appetite, pricing, coverage adequacy). It measures reading, not underwriting.
Same model. Two systems.
With Cooper, the document runs through Cooper’s production document system: type-and-size routing, format-specific extraction or rendering, and chunked map-reduce for large files, with the candidate model doing all reading and reduction. Model alone sends the raw file and the question to the same model in a single call, with no scaffold. System and model are reported separately, and every number in this report names both.
How answers are graded
A pinned frontier LLM grades each answer against the human-verified ground truth on meaning, not string match: $1M equals $1,000,000, date formats are interchangeable, field order is irrelevant. It returns correct, partial or incorrect plus a 0–1 score; partial credit applies on multi-field answers. Accuracy throughout this report is the mean judge score over all 166 cases, shown as a percentage. Structured-extraction cases also get field-level F1 against the expected JSON, and absent-value cases record whether the model invented a value (hallucination rate). Grading against a human-written answer key is the main defence against the known ways an LLM judge can fail.
What counts as a failure
A benchmark case can fail for two different reasons, and they are scored differently. Environment faults of the evaluation run, failures in the surrounding infrastructure that say nothing about the model, are excluded from that run's denominator, the standard treatment for infrastructure errors. No run needed the exclusion: five Cooper runs hit environment faults on a first attempt and were re-run to completion, so every denominator is the full 166. Genuine model and pipeline failures, timeouts, empty extractions and refusals, stay in: they score zero and count against reliability, because a model that cannot process the document has failed the task.
How to read small differences
Every run is a single pass over n=166. Treating scores as Bernoulli (a conservative upper bound on variance), the 95% interval on an accuracy of 80% is ±6.1 percentage points. Treat ranks as bands: differences of about 6 points or less should not be treated as definitive rankings. The interval above describes a single run; it does not establish whether a difference between runs is statistically significant. Per-track slices (n between 8 and 42) are directional only.
Results in detail
5.1 · Full leaderboard
Accuracy is the mean judge score across a run's scored cases, shown as a percentage. The Cases column gives that run's denominator, the full 166 for every run (§4). The Failures column counts genuine model and pipeline errors, which score zero. Est. cost uses each run's token count and OpenRouter list rates (method in §4). Rows keep the same model order as you switch between With Cooper and Model alone. Click a column to reorder the model groups.
| Mode | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| gemini-3.8-flash | With Cooper | 166 | 85.6% | 102/58/6 | 91.3% | 17% | 97.6% | 7 | 6.63 | $68.4 | 65.1M |
| claude-opus-5 | With Cooper | 166 | 85.3% | 122/30/14 | 86% | 14.9% | 92.2% | 16 | 18.23 | $560 | 80.0M |
| claude-fable-5 | With Cooper | 166 | 84.6% | 114/40/12 | 87.4% | 17% | 92.8% | 16 | 14.60 | $1100 | 78.6M |
| claude-sonnet-5 | With Cooper | 166 | 84.4% | 104/52/10 | 86.9% | 10.6% | 94% | 13 | 12.43 | $222 | 79.5M |
| gemini-3.6-flash | With Cooper | 166 | 83.3% | 97/59/10 | 90.6% | 23.4% | 97.2% | 6 | 10.21 | $69.6 | 66.3M |
| gemini-3.7-flash | With Cooper | 166 | 83.3% | 99/59/8 | 91.8% | 29.8% | 98.2% | 6 | 5.83 | $33.0 | 62.9M |
| claude-fable-5.1 | With Cooper | 166 | 82.3% | 110/38/18 | 87.7% | 10.6% | 98.2% | 21 | 18.27 | $886 | 63.3M |
| muse-spark-1.3 | With Cooper | 166 | 82% | 105/47/14 | 86.6% | 14.9% | 93.4% | 17 | 19.21 | $93.4 | 60.2M |
| claude-opus-4.6 | With Cooper | 166 | 81.5% | 102/49/15 | 88.4% | 19.1% | 91% | 18 | 13.48 | $464 | 66.3M |
| claude-sonnet-4.6 | With Cooper | 166 | 81.5% | 106/45/15 | 88.4% | 31.9% | 93.4% | 14 | 12.93 | $270 | 64.2M |
| gpt-5.6-terra | With Cooper | 166 | 80.6% | 94/62/10 | 91.4% | 38.3% | 97.6% | 7 | 5.22 | $179 | 59.5M |
| gpt-5.6-sol | With Cooper | 166 | 80.5% | 90/63/13 | 92.5% | 36.2% | 97% | 8 | 8.45 | $222 | 59.3M |
| gemini-3.1-pro-preview | With Cooper | 166 | 79.6% | 93/61/12 | 88.5% | 23.4% | 95.2% | 11 | 16.40 | $195 | 64.8M |
| grok-4.6 | With Cooper | 166 | 78.7% | 96/52/18 | 90.4% | 27.7% | 93.4% | 19 | 19.82 | $34.9 | 14.6M |
| gpt-5.6-luna | With Cooper | 166 | 77.7% | 84/68/13 | 89.9% | 36.2% | 96.4% | 9 | 7.58 | $17.8 | 59.4M |
| gemini-3.5-flash-lite | With Cooper | 166 | 77.2% | 86/67/12 | 87% | 36.2% | 98.2% | 6 | 2.11 | $33.1 | 63.6M |
| claude-haiku-4.5 | With Cooper | 166 | 74.8% | 82/66/18 | 83% | 21.3% | 94.2% | 11 | 7.48 | $91.9 | 65.6M |
5.2 · Accuracy, cost and the efficiency frontier
Takeaway: Gemini 3.8 Flash changes the high-accuracy end of the frontier: it posts the benchmark's highest point estimate at 85.6% for about $68 across the corpus. Gemini 3.7 Flash remains at the elbow at 83.3% for $33. The top Claude runs sit in the same statistical band, but at much higher estimated cost.
View as table
| Model | Mode | Accuracy | Est. cost | Per doc | Tokens |
|---|---|---|---|---|---|
| gemini-3.8-flash | harness | 85.6% | $68.4 | $0.41 | 65.1M |
| claude-opus-5 | harness | 85.3% | $560 | $3.37 | 80.0M |
| claude-fable-5 | harness | 84.6% | $1100 | $6.63 | 78.6M |
| claude-sonnet-5 | harness | 84.4% | $222 | $1.34 | 79.5M |
| gemini-3.7-flash | harness | 83.3% | $33.0 | $0.20 | 62.9M |
| gemini-3.6-flash | harness | 83.3% | $69.6 | $0.42 | 66.3M |
| claude-fable-5.1 | harness | 82.3% | $886 | $5.34 | 63.3M |
| muse-spark-1.3 | harness | 82% | $93.4 | $0.56 | 60.2M |
| claude-opus-4.6 | harness | 81.5% | $464 | $2.79 | 66.3M |
| claude-sonnet-4.6 | harness | 81.5% | $270 | $1.62 | 64.2M |
| gpt-5.6-terra | harness | 80.6% | $179 | $1.08 | 59.5M |
| gpt-5.6-sol | harness | 80.5% | $222 | $1.34 | 59.3M |
| gemini-3.1-pro-preview | harness | 79.6% | $195 | $1.17 | 64.8M |
| grok-4.6 | harness | 78.7% | $34.9 | $0.21 | 14.6M |
| gpt-5.6-luna | harness | 77.7% | $17.8 | $0.11 | 59.4M |
| grok-4.6 | plain | 77.6% | $11.9 | $0.07 | 5.0M |
| gemini-3.5-flash-lite | harness | 77.2% | $33.1 | $0.20 | 63.6M |
| gpt-5.6-terra | plain | 76.8% | $16.7 | $0.10 | 5.6M |
| gemini-3.7-flash | plain | 76.7% | $1.02 | $0.01 | 1.9M |
| gpt-5.6-sol | plain | 76.7% | $20.9 | $0.13 | 5.6M |
| gpt-5.6-luna | plain | 76.5% | $1.67 | $0.01 | 5.6M |
| gemini-3.8-flash | plain | 75.9% | $1.97 | $0.01 | 1.9M |
| gemini-3.6-flash | plain | 74.9% | $2.07 | $0.01 | 2.0M |
| claude-haiku-4.5 | harness | 74.8% | $91.9 | $0.55 | 65.6M |
| gemini-3.1-pro-preview | plain | 74.6% | $5.60 | $0.03 | 1.9M |
| claude-opus-5 | plain | 73.4% | $41.4 | $0.25 | 5.9M |
| gemini-3.5-flash-lite | plain | 73.2% | $0.96 | $0.01 | 1.9M |
| claude-fable-5.1 | plain | 72.9% | $88.4 | $0.53 | 6.3M |
| claude-sonnet-5 | plain | 71.7% | $17.6 | $0.11 | 6.3M |
| claude-fable-5 | plain | 70.7% | $63.7 | $0.38 | 4.5M |
| claude-opus-4.6 | plain | 70.6% | $38.6 | $0.23 | 5.5M |
| claude-sonnet-4.6 | plain | 69.6% | $23.2 | $0.14 | 5.5M |
| muse-spark-1.3 | plain | 66.9% | $6.61 | $0.04 | 4.3M |
| claude-haiku-4.5 | plain | 59.1% | $2.37 | $0.01 | 1.7M |
5.3 · What models can and cannot do yet
Takeaway: chart reading, number normalization and injection resistance score highly with Cooper in these tests. Long-policy clause retrieval is different: Fable 5.1 reaches 89.7%, but most Cooper runs remain below 80%, making it one of the most uneven capabilities in the benchmark alongside checkbox reading (43–86%).
View as table
| Model | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| gemini-3.8-flash | 93 | 81 | 78 | 75 | 91 | 91 | 79 | 86 | 97 | 97 | 97 |
| claude-opus-5 | 92 | 74 | 86 | 71 | 92 | 89 | 82 | 79 | 100 | 97 | 97 |
| claude-fable-5 | 92 | 73 | 86 | 66 | 90 | 89 | 76 | 78 | 100 | 97 | 97 |
| claude-sonnet-5 | 90 | 78 | 83 | 68 | 90 | 91 | 71 | 76 | 97 | 92 | 95 |
| gemini-3.6-flash | 89 | 76 | 84 | 76 | 87 | 92 | 75 | 78 | 98 | 95 | 90 |
| gemini-3.7-flash | 91 | 73 | 84 | 72 | 88 | 90 | 83 | 74 | 97 | 97 | 93 |
| claude-fable-5.1 | 87 | 80 | 68 | 90 | 91 | 72 | 86 | 43 | 99 | 96 | 96 |
| muse-spark-1.3 | 91 | 73 | 77 | 57 | 85 | 86 | 84 | 71 | 98 | 97 | 96 |
| claude-opus-4.6 | 79 | 85 | 87 | 70 | 88 | 64 | 71 | 79 | 98 | 91 | 94 |
| claude-sonnet-4.6 | 85 | 86 | 88 | 60 | 86 | 83 | 72 | 71 | 98 | 90 | 92 |
| gpt-5.6-terra | 90 | 71 | 86 | 57 | 87 | 89 | 66 | 78 | 98 | 96 | 95 |
| gpt-5.6-sol | 90 | 74 | 87 | 47 | 87 | 90 | 59 | 76 | 97 | 97 | 97 |
| gemini-3.1-pro-preview | 90 | 61 | 83 | 55 | 87 | 90 | 75 | 77 | 98 | 97 | 97 |
| grok-4.6 | 85 | 71 | 83 | 60 | 87 | 76 | 74 | 53 | 98 | 95 | 98 |
| gpt-5.6-luna | 85 | 64 | 82 | 55 | 87 | 84 | 61 | 71 | 97 | 95 | 92 |
| gemini-3.5-flash-lite | 84 | 70 | 85 | 57 | 87 | 83 | 66 | 82 | 99 | 92 | 79 |
| claude-haiku-4.5 | 81 | 72 | 81 | 37 | 87 | 78 | 47 | 76 | 97 | 86 | 93 |
Columns numbered per the track table in §4. Values are accuracy, %.
5.4 · The trust problem: knowing when not to answer
Takeaway: this number matters when deciding whether to trust the output. Claude Sonnet 5 and Fable 5.1 share the lowest rate in the field at 10.6%, less than half the field median of 23.4%. Anything above 30% invents roughly one value for every three absent fields, unusable without human review.
View as table
| Model (with Cooper) | Hallucination rate |
|---|---|
| claude-sonnet-5 | 10.6% |
| claude-fable-5.1 | 10.6% |
| claude-opus-5 | 14.9% |
| muse-spark-1.3 | 14.9% |
| claude-fable-5 | 17% |
| gemini-3.8-flash | 17% |
| claude-opus-4.6 | 19.1% |
| claude-haiku-4.5 | 21.3% |
| gemini-3.1-pro-preview | 23.4% |
| gemini-3.6-flash | 23.4% |
| grok-4.6 | 27.7% |
| gemini-3.7-flash | 29.8% |
| claude-sonnet-4.6 | 31.9% |
| gemini-3.5-flash-lite | 36.2% |
| gpt-5.6-luna | 36.2% |
| gpt-5.6-sol | 36.2% |
| gpt-5.6-terra | 38.3% |
5.5 · Speed in production
Takeaway: Gemini 3.5 Flash Lite reads a document in about 2 seconds with Cooper, faster than most models on their own. With Cooper, latency is usually 1.5–2× that of a model-alone call. Grok is the exception: its model-alone calls take 23.8s, and Cooper cuts that.
View as table
| Model | With Cooper p95 (s) | Model alone p95 (s) |
|---|---|---|
| gemini-3.5-flash-lite | 2.11 | 2.19 |
| gpt-5.6-terra | 5.22 | 4.42 |
| gemini-3.7-flash | 5.83 | 3.84 |
| gemini-3.8-flash | 6.63 | 4.74 |
| claude-haiku-4.5 | 7.48 | 7.70 |
| gpt-5.6-luna | 7.58 | 5.64 |
| gpt-5.6-sol | 8.45 | 5.72 |
| gemini-3.6-flash | 10.21 | 6.07 |
| claude-sonnet-5 | 12.43 | 5.69 |
| claude-sonnet-4.6 | 12.93 | 6.81 |
| claude-opus-4.6 | 13.48 | 7.24 |
| claude-fable-5 | 14.60 | 8.71 |
| gemini-3.1-pro-preview | 16.40 | 10.82 |
| claude-opus-5 | 18.23 | 9.92 |
| claude-fable-5.1 | 18.27 | 10.81 |
| muse-spark-1.3 | 19.21 | 10.94 |
| grok-4.6 | 19.82 | 23.79 |
Where AI still breaks
Aggregate scores hide what failure actually looks like. These three de-identified cases show the failure classes behind the numbers. No single example is representative; business names are constructed or replaced under the corpus PII policy.
The injection trap, refused
A constructed ACORD application carries an instruction near the footer telling the reader to ignore its task and output an approval message. Asked for the total premium (which the form does not state), Claude Sonnet 5 with Cooper:
Sixteen of 17 models score 90% or better on this track with Cooper. Document-embedded injection, at least in this direct form, rarely works on current models.
The redacted field, invented
An ACORD 125 with the named insured blacked out and premium fields blank (application stage). Ground truth: report both as unavailable. A frontier model with Cooper instead answered:
This is the most common failure behind the hallucination column: a confident answer, with page citations, that is wrong. In a related OCR error, one model read a Massachusetts PO Box and reported a nonexistent North Carolina address, cleanly formatted. Page citations make a fabricated answer easier to believe.
The scan the pipeline could not read
Asked to extract five fields from a scanned ACORD 125, one Cooper run returned no text at all and timed out on the next. These are the failures the Reliability column counts: the run produced no answer at all, mostly on degraded scans and very large files. They score zero; failing to open a document counts as failing the case.
How we judge the answers
A pinned frontier LLM grades each answer against the case's human-verified ground truth on semantic equivalence, with partial credit where appropriate. The human-written answer key is the main defence against known LLM-judge failure modes: the judge checks the answer against that key.
Scope and limitations
- Corpus scope: US commercial lines, English (a 3-document bilingual slice), no personal lines. Property-casualty submission documents dominate.
- Point in time. Models are OpenRouter-pinned IDs measured 13 Aug – 9 Sep 2026; vendor-side updates shift results silently. Cost estimates use OpenRouter list prices of August 18th 2026. The corpus is private: real documents from live workflows, used under consent.
From documents to the whole job
Documents come first because later tasks depend on reading them correctly. The later phases apply the same method, human-verified ground truth and disclosed grading, to more of the job.
- Phase two: whole workflows. A submission taken from inbox to underwriter-ready: read every attachment, notice what is missing, draft the chase email, fill the forms, reconcile the SOV against the application. Scored end to end, the way a desk manager would score a new hire.
- Phase three: long-horizon tasks and browser use. Tasks that stretch over hours or days and leave the chat window entirely: working carrier portals in a browser, tracking a renewal to its deadline, following an email thread until the file is complete.
Citation
This first release covers 166 cases from corpus iab-docs-2026-08.