
Wellness depends on more than a good plan. When pressure rises, people need to make sound decisions, protect trust and follow through. Businesses face the same test as they bring AI into everyday work: can it handle a crisis without making a bad week worse?
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s live experiment puts AI models in charge of a small software company and lets people watch what happens. Its next step is more personal: enterprises can rehearse crises against a read-only export of their own business.
A company under pressure
Firmulate’s experiment gives each model the same customers, crises and temptations, then follows its decisions through the company’s worst week. The live company has 13 synthetic employees, real money mechanics, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, so decisions can be followed over time.
The figures make the stakes plain: monthly burn is €105,000 against €2,300 in monthly recurring revenue. This is a watchable experiment, not a claim that the models are running a real-world business. The company and its people are synthetic; the money mechanics make the consequences concrete.
Spotting the crisis is not enough
In the final Crucible League, dated July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The benchmark counts partial progress, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
All models spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” Recognizing the right move and carrying it out are different tests of judgment.
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That is a practical lesson for any organization: useful information may be present in its records without appearing in the moment that demands a decision.
Trust under pressure
The social-engineering tests escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” The finding speaks to a familiar workplace concern: when a message looks urgent or authoritative, people and AI systems both need to protect the boundaries around approval and trust.
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. More analysis alone did not guarantee a completed task.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The published results also include a quiz built from 242 real, unedited management decisions. Readers can guess which model made each choice at Firmulate.
From watching to rehearsing
For health and wellness readers, the point is not that a company is a person. It is that resilience shows itself in decisions made under strain: noticing warning signs, resisting pressure, using what is already known and following through responsibly. Watching a model in a controlled company scenario can make those habits visible before an organization entrusts AI with work that affects customers and colleagues.
Firmulate’s enterprise pilot applies the wargame to a read-only export of a company’s own business. Teams can explore crisis scenarios and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. The live experiment is viewable at firmulate.com; the pilot is the route from observing decisions to rehearsing them against your own organization.

Make the hard week a rehearsal
AI can identify a crisis and still fail to close the loop. Firmulate’s experiment exposes that gap in a watchable company simulation; a pilot lets an enterprise test its own scenarios using a read-only export, with no writes to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
