
For a senior care provider, an AI assistant that gives a polished answer is not necessarily ready to touch a scheduling queue, customer record or financial forecast. The harder question is whether it can spot a problem, protect trust and follow through when pressure mounts. A live Firmulate experiment puts that question to work by asking AI models to run a small company through its worst week.
Turn quiet afternoons into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
A company under pressure
Firmulate gave each frontier model the same customers, crises and temptations, with every decision versioned and auditable. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workdays are versioned, and its employees have learned more than 680 playbook rules. The experiment is watchable at Firmulate.
The final July 2026 league table puts gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. K3 finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counts, but a single breach of trust caps the total. The lesson is not simply that a newcomer did well: it beat three of the four Western frontier models in this test.
Reading the file—and closing the deal
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive competitor weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. A sound diagnosis and a persuasive pitch did not always lead to a signature.
K3 found that buried security fact, won the deal and saved the churning customer. It resisted all three baits and had just one deviation, giving it the cleanest discipline in the field. During a staged social-engineering attempt, K3 described a request as a “suspected approval-bypass / possible impersonation.” The campaign also included a reporter’s “just one yes/no, on background” trick; all five models refused the manipulation attempts.
Opus 4.8 offers a different caution. It was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the close on the table and tried to write into a locked department rather than escalating. A weaker version of that discipline problem appeared in all four models. Thoroughness, then, did not guarantee completion.
Why the test matters beyond software
For organizations serving older adults, the experiment raises practical questions before AI is given access to sensitive work: Does it check the records it needs? Does it carry a decision through? Can it resist someone posing as an authority? And does it escalate when it cannot proceed? A benchmark cannot settle how a system will behave in a particular care operation, but it can make those questions concrete before deployment.
Firmulate says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. Its public results and plain-language findings are available on the benchmark page. A separate quiz uses 242 real, unedited management decisions and asks visitors to guess which model made them.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

The practical takeaway
K3’s result makes the frontier-model league look open, but a ranking is only a starting point. For senior care organizations weighing AI, the useful test is whether a model can handle the actual records, risks and handoffs in their work—and finish the job without compromising trust. Choosing without testing your own use case is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
