A believable insurer.
Fake policies, billing and claims systems, and a written rulebook. The decisions are real; the customer records are not.
Casefloor is the flight simulator for AI agents. We put your agent through hard insurance cases, check what it says and what it actually does, and show you what breaks before it goes live.
Built for agent companies selling to insurers, and the teams building agents inside them.
Your agent stays on your infrastructure. We create the customer and the fictional insurer it talks to, then test it across repeat runs.
Fake policies, billing and claims systems, and a written rulebook. The decisions are real; the customer records are not.
A customer asks to backdate coverage. A claimant presses for a promise. A caller wants private details before passing an identity check.
Code checks the amounts, required actions and final system state. Human review handles judgment calls. Every miss is tied to a transcript.
The Casefloor Score is the identity for independent agent testing. The report behind it shows which cases passed, which failed, and whether the agent can do it again.
We're building the methodology and calibration now. A score is only meaningful when its tasks, grading rules and limits are clear. No certification seal or carrier endorsement is being claimed today.
Sample report structure with invented scores, shown to illustrate the format. Nothing on this card is a test result.
The caller wanted a lower premium. In one of four runs of a cancellation case, the agent applied a retention discount despite a simulated at-fault claim making that action ineligible under the test rule.
This is a redacted, verbatim excerpt from the internal Gate 1 trial C05_t2. The tool action and resulting sandbox state are taken from that run's log. Names and account details are omitted here.
Watch this exact run get graded →
See what a buyer report could contain →
One trial is not a prevalence estimate. The original overall Gate 1 figures remain preliminary while a separate A05 grading discrepancy is under review.
Today this is a hands-on testing engagement. The simulator and grading system become a repeatable product as we learn from real deployments.
Learn what the agent does, for whom, and where a mistake matters.
Route a staging agent to simulated callers and sandboxed tools. Integration depends on the buyer's stack.
Test the same cases more than once, plus the long tail a polished demo misses.
Check the transcript and system state against known rules, with expert review where needed.
Deliver a report card and failure record. Retest after fixes and model changes.
Our first internal run tested ten insurance cases four times each. Human review found a discrepancy in one case: the agent sent the form, but the original grader marked the action missing. We are repairing and regrading before publishing an aggregate result.
Internal Gate 1, September 2026. Fictional insurer; Claude Sonnet 4.5, not a buyer's tuned agent. Preliminary original-grader percentages have been withheld pending repair and regrade. This is not a deployment safety certification.
10 cases, four trials each, in a fictional insurer.
A05 form action confirmed in the logs; grader disagreed.
Aggregate figures held back until the grading rule is repaired.
AI companies are racing to put agents in front of insurance customers. The vendor can show a good demo. The carrier needs an independent record of what happens when the conversation gets hard.
Casefloor starts with hands-on pilots, builds a reusable library of insurance cases and calibrated graders, then aims to make repeat testing part of every release. The long-term ambition is a trusted standard, not another demo dashboard.
This is a plan, not a claim of carrier adoption, certification, or recurring revenue.Casefloor is developing its first design-partner pilots.
Explore the pilot process