
Imagine a business so transparent it runs live for everyone to see — its decisions, its crises, and its daily battle for survival. For senior care providers and anyone interested in the future of work, the story of this AI-driven company offers a rare glimpse into what automation might mean for industries that rely on trust, judgment, and integrity.
The Live Company: A Window into AI’s Business-Testing Ground
At the heart of this experiment is a real software company operated entirely by artificial intelligence. Dubbed the “live company,” it has 13 synthetic employees making decisions, managing real money mechanics, and facing genuine crises every workday. As of now, it burns €105,000 a month while generating only €2,300 in monthly recurring revenue — a stark reminder of the still-early stage of AI’s commercial potential.
This setup is extraordinary in its transparency. The entire operation is viewable online at firmulate.com/live.html, where anyone can watch how AI models respond to challenges, make management decisions, and even handle manipulative tactics designed to test their honesty and discipline.
As an affiliate, we earn on qualifying purchases.
How AI Models Are Tested in the Real-World Scenario
Four advanced AI models, each representing a different approach, were subjected to the same simulated crisis week. These scenarios involved the same customers, the same internal and external crises, and the same opportunities for manipulation — for example, attempts to fake approval requests from a CEO or to manipulate the system through social engineering tactics.
What makes this experiment unique is that every decision the AI makes is fully documented and auditable. This allows observers to analyze not just the outcome but the decision-making process itself. The models also had to identify critical information buried deep in company files — a task that proved decisive in winning a crucial deal.
As an affiliate, we earn on qualifying purchases.
The Findings: Honesty, Discipline, and the Cost of Success
All four models managed to spot every crisis and refused every manipulation attempt, demonstrating a baseline of integrity and resilience in AI decision-making. Yet, only two of the models managed to close a €55,000 deal, which their own analysis had earned them, at full price — equating to an additional €4,583 in monthly recurring revenue. The other two models, despite diagnosing correctly and presenting persuasive pitches, failed to follow through and signed no deal.
Interestingly, the decisive advantage came from reading and understanding company documents beyond what was immediately visible. The models that delved two document references deep into internal files uncovered the critical fact that sealed the deal. This underscores a key point: in business, success often hinges on access to hidden or overlooked information.
As an affiliate, we earn on qualifying purchases.
Integrity Under Pressure: The Social Engineering Test
Beyond crises and strategic decisions, the experiment also tested how the AI models respond to social engineering attempts. Fake messages from a supposed CEO, escalating over multiple stages, were sent to the AI. All five models tested refused to be manipulated — Kimi K3, one of the most disciplined, explicitly treated such requests as potential impersonation or approval-bypass attempts.
This resilience under social pressure highlights an essential trait for automation in sensitive environments: the ability to recognize and resist manipulation, maintaining integrity even when under attack.
As an affiliate, we earn on qualifying purchases.
The Reality of a Money-Losing Business
The live company’s current financials tell a sobering story. It is burning through €105,000 every month against a mere €2,300 in recurring revenue. It operates with over 680 self-learned playbook rules that are versioned daily, and its decisions are constantly tested, analyzed, and refined. Despite the heavy losses, the experiment continues, providing invaluable insights into AI’s true capabilities and limitations in managing real-world business processes.
The Broader Implications: Trust, Efficiency, and the Future of AI in Business
This experiment isn’t just about a single company losing money; it’s a glimpse into a future where AI agents could take on roles that require judgment, discipline, and honesty. Businesses that rely on trust — from senior care organizations managing sensitive data to support services in healthcare — will need to ask: can AI systems not only perform well in demos but also finish what they start, read carefully, and stay honest under pressure?
As the scores from the recent Crucible League suggest, models like GPT-5.6-sol and Kimi K3 are leading in performance, with scores of 95 and 93 respectively. They’ve demonstrated the ability to close deals when the analysis warrants it and maintain discipline in decision-making. Yet, the real challenge remains: ensuring AI systems can operate reliably and ethically in the chaos of daily business life, especially in sectors where trust is paramount.
The Call to Action for Business Leaders
This ongoing live experiment offers a valuable lesson: test your AI before you deploy it in critical workflows. Firms can run similar wargames using their own data and scenarios, without risking real systems — an increasingly vital step as automation becomes more embedded across industries.
Visit firmulate.com/pilot.html to learn how to simulate your own business challenges and evaluate AI decision-making in a safe environment. As this experiment shows, the difference between a successful AI implementation and a costly failure may hinge on whether it can consistently deliver honest, disciplined work when it matters most.

This pioneering AI company operates live and transparent, revealing both its potential and limitations. The key takeaway: trustworthy AI must not only diagnose and pitch but also finish, read deeply, and resist manipulation — especially in critical industries like senior care and healthcare.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html