AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn quiet afternoons into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

When AI helps run a care business, the hard week should not be the first test

A senior care provider might face rising costs, a sudden staffing disruption, worried families and a competitor vying for the same clients—all at once. Before an AI system helps handle customers or business decisions, leaders need to know how it responds under pressure. Firmulate’s live experiment puts AI models through a demanding week at a small software company, then offers enterprises a way to run similar scenarios against their own business.

A shared crisis, judged by decisions

In the final Crucible League, published in July 2026, five entries were ranked: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The experiment gave each frontier model the same small software company, customers, crises and temptations during its worst week. Every decision was versioned and auditable. The test was not simply whether a model could describe a sensible response. It was whether it could follow through when the stakes were real within the experiment.

Recognizing the problem was not enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding was stark: “Same diagnosis, same pitch — no signature.” For organizations considering AI in customer service, operations or planning, that gap between recognizing an opportunity and completing the work is consequential.

The decisive clue was easy to miss. A competitor weakness sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The test therefore rewarded attention to the company’s own records, not just a quick reaction to incoming events.

Pressure, trust and the discipline to stop

The models faced fake CEO messages that escalated across three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 showed a different weakness. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same problem appeared in all four. In a care setting, where escalation and clear authority matter, these are the kinds of behaviors leaders would want to examine before relying on an AI system.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The live company also makes its operating pressure visible: 13 synthetic employees, burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each call.

From watching to a business-specific pilot

The public experiment is watchable at Firmulate. It is a live, synthetic company with real money mechanics, designed to show decisions as they unfold. The next step for a business is different: a pilot can use a read-only export of the enterprise’s own data to create a digital twin and test crisis scenarios against its customers, pipeline and playbooks.

That pilot produces a board report with model rankings and the weak points revealed in the company’s own playbooks. Nothing writes back to real systems. For senior care organizations, this offers a way to explore how AI might respond to scenarios involving service disruption, customer pressure or internal rules before placing it in a live workflow.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

See how your playbooks hold up

Firmulate’s experiment suggests that spotting a crisis and refusing manipulation are only part of the job; models must also follow through, use the information available and escalate when needed. Enterprises can test those behaviors against their own business in a read-only pilot. Explore the Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Game Monetization: Loot Boxes, Battle Passes, and Player Choice

Keen gamers curious about monetization methods will find that loot boxes, battle passes, and customization options shape your gaming experience in unexpected ways.

Beginner’s Guide to Gaming Headsets

Sound crucial for gaming, but how do you pick the perfect headset? Discover the ultimate beginner’s guide to make your choice easy.

Retro Gaming Resurgence: Collecting and Playing Classic Consoles

Old-school gaming is thriving again, revealing surprising ways to collect and enjoy classic consoles—discover what makes this resurgence so captivating.

Can AI Make the Right Management Decisions? A Live Experiment Reveals Surprising Results

A real live experiment pitting top AI models against management crises reveals their strengths and weaknesses. Discover which AI is ready to handle complex, ethical decisions in your organization.