Part three of three on how we're building Cooper, the AI coworker for insurance.
When you hire someone, part of what you're paying for is what they'll learn. They'll figure out which report the client actually means, why one insurer needs an extra explanation and which of the three templates in the shared folder is the one people really use. You'll correct them a few times. Eventually they'll need less explaining, and at some point they'll suggest a better way to do something you've been doing for years. Getting better is part of the job.
When you hire someone, part of what you're paying for is what they'll learn.
We want the same progression from Cooper. If you correct how it prepares a report, it should apply that correction to the next report. If Cooper spends an hour figuring out how to use an online rater, it should remember the steps and finish faster next time. Finishing the task matters, of course, but so does what Cooper learns along the way.
In the first post, I argued that context is the hard part of building an AI coworker. In the second, I described why Cooper's core agent contains no insurance knowledge. It gets the procedures, memories and reference material it needs as context for the task. We use what Cooper learns from doing the work to update its instructions and memories. Every task should make Cooper better at the next one.
Learning from production
We've given Cooper a job on our engineering team. It reviews production conversations, finds problems and proposes fixes. The same general-purpose agent can do both jobs because reading evidence, using tools and working through an unfamiliar problem are useful in engineering too. What changes is the context. For insurance work, Cooper gets customer documents and office procedures; on the engineering team, it gets production traces and our instructions for investigating problems and testing fixes.
We've given Cooper a job on our engineering team.
A correct report doesn't tell you how much work it took to produce it. Cooper might have needed several attempts, spent most of its time on a detour or only got the answer right after the user corrected it. A production trace records those steps: what Cooper tried, what each tool returned and where a person had to step in. We review successful tasks as well as failures, because a task can succeed and still reveal a problem worth fixing.
Cooper reviews these traces in a process we call reflection. It works out what caused the problem and what should change. A missing step may need to be added to a standard operating procedure, while a user's preference belongs in memory. A broken tool needs an engineering fix; more instructions won't reliably solve it.
Browser scripts are a good example. When a website changes, Cooper can repair the script it uses to interact with the site and continue the task. The repair helps with that task, but the shared script still needs updating or the next task will hit the same problem. During reflection, Cooper checks what the website showed and whether the repair worked, then proposes the change to the shared script. Successful tasks help too: steps Cooper worked out manually can become a script for future tasks. The next task should benefit from what the last one taught us.
The next task should benefit from what the last one taught us.
Testing whether it got better
A script can get through a website and still enter the wrong information. It might select the wrong option and carry on anyway. We have to check the result, however convincing Cooper's explanation of the repair sounds. An improvement needs a test it can fail.
An improvement needs a test it can fail.
So we turn real customer tasks into evals: repeatable tests with clear criteria. We get the expected facts from the source documents and define what Cooper should produce and which questions it needs to ask. We don't use Cooper's original answer as the answer key. It's useful to study, but it may contain the very mistake we're trying to catch.
An eval can check whether Cooper asks for a missing document before finishing a report. It can also catch unnecessary questions, such as asking the user for information that's already in the files. The final report might look perfectly reasonable even when Cooper made someone do work it could have done itself. We test the decisions along the way as well as the final result.
When a procedure changes, we rerun these evals to check that the fix helped without breaking something that used to work. Cooper is not allowed to weaken a test just to make its fix pass. Sometimes the eval itself is wrong, and changing it goes through review too. A test that measures the wrong thing can make a bad change look like an improvement.
Review is part of the learning
Cooper submits each proposed change as a pull request for our engineers to review. It includes the evidence behind the change and the eval results. The engineers approve some, ask for changes to others and reject the rest, just as they do with each other's work. Cooper gets its pull requests rejected too. It does not merge its own. We've merged hundreds of these changes so far, and reviewing them has become an ordinary part of maintaining the product.
Cooper gets its pull requests rejected too.
When an engineer asks for changes, Cooper revises its proposal using those comments and feedback on earlier attempts. We also look across reviews for patterns and update the instructions it follows when proposing improvements. Reviewers have caught duplicate proposals and advice for one organization presented as a general rule. They've also caught attempts to fix a software bug by adding more instructions. That feedback helps Cooper avoid repeating the same mistakes. The improvement process needs to improve too.
The improvement process needs to improve too.
We use hill climbing to improve procedures, browser scripts and even the evals themselves. We look at real work, propose a small change, test it and keep it if it helps. Then we repeat. Evals tell us whether a change improved the behavior we tested. Engineers judge whether the tests and evidence are enough to trust the change. We can inspect each change and roll it back if we need to.
Learning how your office works
Most of those changes improve Cooper for everyone. It also has to learn how a particular person or office works. You might explain how to lay out a report, correct a client detail or tell Cooper when to ask you to review its work. After each exchange, our memory layer picks out facts worth keeping and compares them with what Cooper already knows, saving only the parts of the conversation that will help with future work.
You can also teach Cooper a procedure directly. Explain when to use it, give it a template or an example and describe what a good result looks like, much as you would when onboarding a colleague. A team's reporting procedure might specify the layout, comparison rules and who needs to review the result. We keep this knowledge at three levels: personal preferences, procedures shared within an organization and a general knowledge base available to every customer. Getting the level right matters as much as getting the content right. One person's formatting preference shouldn't become everybody else's standard.
Giving memory time to settle
Memory needs maintenance too. We think of it a little like sleep, or taking a proper break from work. You let the day's experiences settle and come back with a clearer sense of what mattered, how it fits with what you already knew and which details you can let go. Reconciliation is Cooper's process for organizing what it learns after a task.
After a task, Cooper compares new information with its existing memories. It can add a fact, update an older memory or remove one that's no longer true. A fact already in memory may not need to be saved again. If someone explains the same preference three times in different words, Cooper merges those memories into one, so it can use the preference without comparing three versions of it again.
Conflicting instructions need more care. A request to use a different layout for one report is probably an exception, but an office switching to a new template changes the default for future reports. Cooper needs to keep that distinction and remember where each correction came from. We keep exceptions with the procedures they apply to, so a remembered instruction doesn't get used in the wrong situation.
Old information loses relevance as the work changes. Yesterday's deadline becomes irrelevant quickly, while a review requirement might matter for years. Cooper gives stale memories less weight over time and retires them when something replaces them. Age is only one signal, so Cooper keeps older facts and instructions that still apply. Learning includes deciding what to stop relying on.
Learning includes deciding what to stop relying on.
Cooper applies corrections during the task. The memory layer saves and organizes lasting facts afterward, and scheduled reviews look for improvements that apply more broadly. You shouldn't have to wait for memory cleanup before getting a report. On the next task, Cooper needs to remember the right information and find it when it needs it.
I think this is some of the most interesting work in building agents right now. It gets a little meta. Cooper proposes improvements to the procedures and software it uses, we test those proposals against real tasks, and feedback from reviewers helps it make better proposals next time. The person is still our blueprint. A good coworker learns how the office works, needs less explaining as time goes on and eventually helps the whole office work better. That's the progression we want for Cooper.
If you'd like to work on problems like these, we're hiring.
Get new posts in your inbox
Notes on building an AI coworker for insurance, from the people building it. No more than we'd want to receive ourselves.