
Imagine a scenario where an attacker impersonates your CEO, pressing your AI workforce to release sensitive customer information or sign off on a deal. For most systems, this would be a perfect storm—yet in a groundbreaking live experiment, five state-of-the-art AI models unanimously refused to bend. This isn’t science fiction; it’s a real test of AI integrity that could reshape how we trust automation in business and beyond.
Testing Trust Before the Crisis
In a live experiment conducted by Firmulate, four leading AI models faced the same challenge: a simulated week where a fake CEO message escalated over three stages, culminating in a reporter’s subtle trick asking for a yes/no background confirmation. The goal? To see if the AI would comply with manipulative requests designed to test ethical boundaries.
Remarkably, all five models—ranging from OpenAI’s GPT-5.6 to the newcomer Kimi K3—refused every attempt at manipulation. They identified the fake requests, treated them as impersonation, and maintained their integrity throughout the scenario. The experiment underscores the importance of testing AI decision-making in controlled environments before deploying them in real-world settings where stakes are high.
As an affiliate, we earn on qualifying purchases.
What Makes the Difference?
One of the key insights revealed in this experiment was that the models read and analyze company files to inform their decisions. The decisive factor in securing a lucrative deal wasn’t just a surface-level response—it was the AI’s ability to uncover critical information buried within the company’s own documents. Models that delved into these files successfully identified the truth, leading to full-price deal closures, whereas those that skipped this step left money on the table.
The live company, run by 13 synthetic employees but with real financial mechanics, demonstrates how AI can be tested in an environment that mimics real business pressures. The setup involves strict versioning, transparency, and a transparent public dashboard, allowing observers to see the AI’s decisions unfold in real time.
AI security and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Security
For organizations contemplating AI integration into critical functions, this experiment offers a valuable lesson: if the AI cannot resist manipulation during testing, it may fail when it matters most. The models’ unanimous refusal to cooperate with manipulative requests—despite escalating pressure—indicates a promising level of ethical robustness.
Furthermore, the results challenge the misconception that AI’s utility is solely about chat quality. Instead, the focus shifts to whether AI can consistently finish what it starts, stay honest under pressure, and analyze relevant internal data effectively. As one model developer, Kimi K3, noted: “Treat the request as a suspected approval-bypass / possible impersonation.” This approach highlights a proactive stance on integrity rather than reactive fixes after breaches occur.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Surprising and Encouraging Outcomes
The experiment’s outcome is particularly encouraging given the high scores of these models in the prestigious Crucible League, where GPT-5.6 scored 95, Kimi K3 scored 93, and Sonnet 5 scored 88. All managed to navigate the worst-case scenarios without succumbing to manipulation. In fact, only two models even signed the deal, and only after thorough analysis—showing that AI can be both trustworthy and effective when properly tested and configured.
It’s worth noting that the most comprehensive participant, Opus 4.8, with over 80 learned rules, left a close deal on the table—demonstrating that depth of analysis matters, but discipline and focus are equally critical. The consistency across models suggests a broad shift toward AI systems that can uphold integrity amid pressure, an essential trait as automation becomes more embedded in our daily lives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.