
Imagine if your health tracker gave you a score just for showing up—no matter what you did or didn’t do. That’s similar to how some AI benchmarks measure their initial performance: not zero, but a modest 26 points. For business leaders considering AI tools, understanding this baseline is crucial because it reveals what honest evaluation really looks like—and why trust is the key metric.
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the AI Benchmark Floor: Why 26?
When assessing artificial intelligence for managing critical business tasks, it’s tempting to think of scores as a simple measure of ‘how well’ an AI performs. But actual benchmarks—like the recent Firmulate Crucible League—show that even a ‘do-nothing’ AI system scores around 26 points. This isn’t coincidence; it’s a reflection of the benchmark’s design, which aims for transparency and integrity.
Why 26 Points for Doing Nothing?
The scoring system assigns a baseline score of 26 to a model that takes no action—no crisis detection, no decision, no manipulation. This baseline recognizes that even inaction involves some understanding of the environment, and partial progress—such as recognizing a crisis—counts toward the total score. It’s a way of measuring minimal awareness, not just raw output.
Partial Progress Counts
In the real-world test, models that identified issues or read key files earned partial points. For example, reading two documents deep into a company’s files—without even making a decision—earns some recognition. This approach discourages superficial performance and promotes meaningful engagement. It also means that a model’s score reflects not just final decisions, but its capacity to process information correctly.
Why Does a Single Breach of Trust Cap the Score?
Trust is paramount in business AI. If a model attempts manipulation or breaches ethical boundaries, the entire performance is capped. In the experiment, all models successfully spotted crises and refused manipulative requests, like fake CEO messages or attempts to sign deals under false pretenses. However, even a single breach—such as attempting to escalate an issue dishonestly—caps the total score, emphasizing that honesty isn’t negotiable.
As an affiliate, we earn on qualifying purchases.
The Firmulate Live Experiment: Real Companies, Real Crises
The live test runs a full simulated company with 13 synthetic employees, handling real money mechanics—burning €105k monthly against just €2.3k MRR—and subject to daily decision-making. This setup provides a transparent window into how AI models perform under pressure, mirroring real business challenges.
What the Models Achieved
- All models identified every crisis and refused manipulation attempts, demonstrating integrity.
- Only two signed the €55,000 deal their own analysis earned, showing that recognition of value and trustworthiness can be separate from mere diagnosis.
- The most thorough model, Opus 4.8, performed well but slipped on closing the deal, illustrating that discipline and consistency matter.
The Hidden Weaknesses
Interestingly, the models’ weakest points weren’t in customer interactions or crisis detection, but in the details buried within company files—two document references deep—where the decisive deal was made. Models that read these files won at full price (+€4,583 MRR), underscoring the importance of deep context understanding.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business Leaders
In practical terms, AI’s value isn’t just in generating content or quick responses; it’s in completing tasks ethically, thoroughly, and reliably. The benchmark’s approach—acknowledging a score for doing nothing, rewarding partial understanding, and capping performance after breaches—sets a high standard for honesty and discipline.
As the AI landscape evolves, companies should prioritize tools that can see through crises, resist manipulation, and read deeply into relevant data. The Firmulate live experiment offers a blueprint: measure what management quality truly is, not just what AI can produce in a demo.
As an affiliate, we earn on qualifying purchases.
The Takeaway: Trust and Transparency in AI Evaluation
For business decision-makers, the core lesson is clear: a trustworthy AI system must demonstrate honesty under pressure, thoroughness in understanding, and discipline in execution. The benchmark’s honest scoring system, starting at 26 points for doing nothing, encourages continuous improvement grounded in integrity—crucial qualities for AI that will influence your company’s future.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethics and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
