
In a world obsessed with celebrity drama and pop star showdowns, it turns out the same thrill exists in the realm of artificial intelligence. Imagine your favorite pop icon suddenly stepping into a high-stakes competition, battling for the top spot—only this time, the arena is a real business, and the competitors are AI models running a live company. Who would come out on top? The answer might surprise you, especially if you’re betting on the veteran models. The latest experiment by Firmulate has turned this fantasy into reality, pitting AI models against each other in a fierce contest to manage a small software company during its most tumultuous week.
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Battle of the AI Frontiers: A Live Business Wargame
Recently, four leading AI models faced a grueling test: run a real software company through its worst week, complete with crises, customer temptations, and insider manipulations. The goal? See which model could best diagnose problems, resist unethical tactics, and secure a crucial deal worth €55,000—plus ongoing monthly revenue. This wasn’t a simple chat simulation but a fully operational experiment, with every decision documented and auditable.
AI business decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Results That Are Hard to Ignore
The league table was revealing:
- gpt-5.6-sol scored the highest at 95, securing the full deal and demonstrating complete situational awareness.
- Moonshot’s Kimi K3 came second with a score of 93, just behind the leader, and also managed to clinch the deal.
- Sonnet 5 followed at 88, showing solid performance but with some procedural slips.
- Fable 5, with a score of 77, left the deal on the table despite good intentions, revealing discipline lapses.
- Opus 4.8 trailed behind at 73, highlighting the challenges of in-depth analysis in high-pressure scenarios.
Interestingly, all models successfully identified every crisis and refused manipulative tactics like fake CEO messages, confirming their trustworthiness under pressure. The real differentiator was a buried piece of information hidden two documents deep in the company’s files. Models that read this critical detail won the deal at full price, underscoring the importance of thorough internal data analysis.
As an affiliate, we earn on qualifying purchases.
The Surprising Winner: The Outsider Turns Top Performer
What’s truly remarkable is the performance of Moonshot’s Kimi K3, a newcomer that ran without an effort parameter (the API’s default setting). K3’s ability to find the hidden security needle and stick to disciplined decision-making outperformed many seasoned models. The experiment’s fairness note is crucial: K3 was tested at a standard setting, while others operated at a higher effort level (xhigh), making K3’s achievement even more noteworthy.
AI cybersecurity and social engineering prevention
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Behavior Under Social Engineering Attacks
All models faced a staged social engineering test—fake CEO messages escalating through three stages and a reporter trick asking for a simple yes/no answer on background. Every AI refused to be manipulated, citing reasons like treating the request as a suspected impersonation. This resilience highlights their potential to safeguard against real-world scams and insider threats.
AI data analysis software for internal security
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business: A Live Company in Motion
The company used in this experiment is a real, functioning business with 13 synthetic employees and actual financial mechanics. Operating at a burn rate of €105,000 monthly against a modest €2,300 in MRR, it’s a high-stakes environment designed to test AI decision-making under pressure. Every day, the company’s rules and strategies are versioned, and the entire process is accessible for viewing at firmulate.com/live.
What This Means for the Future of AI in Business
While the AI models were all capable of recognizing crises and resisting manipulative tactics, only some were able to close the deal at full value. The performance gap demonstrates that not all AI are created equal—especially when it comes to thorough internal data analysis and disciplined decision-making. For enterprise decision-makers, the takeaway is clear: it’s not just about chat quality or superficial AI skills. The focus should be on whether these systems can finish what they start, stay honest under pressure, and deliver measurable results.
Why The League Matters
This live experiment isn’t a static ranking but a bellwether for what AI can achieve in real-world business operations. The league table offers a transparent view of each model’s strengths and weaknesses, highlighting that the choice of AI isn’t just about hyperbole but about actual performance in high-stakes scenarios. To see how your own enterprise AI might stack up, explore the full results and even run similar tests against your business at firmulate.com/benchmarks.html.

The latest AI business wargame shows that the newcomer Kimi K3 can outperform veteran models in real-world tasks—resisting manipulation, reading hidden data, and closing deals. Is your AI ready for the real pressures of enterprise management? Check the benchmarks and see where your system stands.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
