
Imagine your fitness routine — every workout, every decision, tested under pressure. Now, picture AI models running a company through its toughest week, with real consequences and high stakes. That’s exactly what the latest experiment by Firmulate has accomplished, revealing how advanced AI can handle real-world business crises just as well as, or better than, human managers.
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Experiment: Putting AI to the Test in a Live Business Environment
In a groundbreaking live experiment, four leading AI models were tasked with managing a small software company during its most challenging week. Every detail was real: the same customers, identical crises, and the same temptations to cheat or cut corners. The companies’ responses were fully transparent, with decisions recorded and auditable, ensuring the results are both reliable and replicable.
Key Findings Show AI’s Potential for Enterprise Management
All four models demonstrated impressive competence: each spotted every crisis and refused every manipulation attempt. This alone underscores the growing trustworthiness of AI in sensitive decision-making. However, the standout was Moonshot’s Kimi K3, which not only identified the buried security flaw but also sealed a €55,000 deal — translating to over €4,500 in monthly recurring revenue. Its performance was just two points shy of the top scorer, gpt-5.6-sol, which achieved a score of 95 out of 100.
The Hidden Weaknesses and the Power of Deep Data Reading
Interestingly, the decisive edge came from reading company files deeply. While all models detected crises from external cues, the ones that delved into internal documents uncovered critical, buried information that clinched the deal. This demonstrates that success hinges on sophisticated data comprehension, not merely surface-level analysis.
The Testing of Social Engineering Tactics
In a simulated social engineering scenario, fake CEO messages escalated over three stages, plus a journalist trick involving background approval. All models correctly refused to be manipulated, with K3 reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows AI’s resilience against common corporate scams.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real Business Mechanics and the Cost of AI Performance
The experiment isn’t just a game. The company involved operates with 13 synthetic employees, manages real money, and faces a burn rate of €105,000 monthly against a revenue of €2,300. Every workday, the models make decisions based on over 680 self-learned rules, which are versioned and visible live at firmulate.com/live.
Discipline and Deep Analysis Matter
The least successful participant, Opus 4.8, with over 80 learned rules, displayed the deepest analyses but faltered at closing the deal — illustrating that thoroughness isn’t enough if discipline slips or decisions aren’t escalated properly. This highlights a vital lesson for enterprise AI: strategic discipline and decision protocols are crucial, especially under pressure.
As an affiliate, we earn on qualifying purchases.
The League Table and Its Implications for Enterprise AI Choice
Here are the final scores:
- gpt-5.6-sol — 95, found the buried fact, closed the deal, showing complete performance
- Kimi K3 — 93, the new contender, also closed the deal using the cleanest discipline
- Sonnet 5 — 88, closed with minor slips
- Fable 5 — 77, closed with more process slips
- Opus 4.8 — 73, left the close on the table due to weaker discipline
Notably, K3 ran without an effort parameter (the API default), while the others were set at xhigh, making K3’s high performance even more impressive.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
If AI models will soon touch your customer management, support, or forecasting systems, the real question is not just “Can it write well?” but: Can it finish what it starts, read critical documents, maintain honesty under pressure, and deliver real, measurable work? This experiment shows that the best AI models are capable of these qualities, provided they are tested and validated in real-world scenarios.

The latest live experiment from Firmulate proves that advanced AI models can outperform traditional governance in company management, especially when deep data reading and discipline are involved. For enterprises, choosing the right AI isn’t just about chat quality — it’s about reliability, honesty, and the ability to deliver real results under pressure. The league is open, and the stakes are high. Test your AI workforce before you hire it, and ensure it can handle your toughest crises.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI cybersecurity and fraud detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
