firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In a world obsessed with celebrity drama and pop star showdowns, it turns out the same thrill exists in the realm of artificial intelligence. Imagine your favorite pop icon suddenly stepping into a high-stakes competition, battling for the top spot—only this time, the arena is a real business, and the competitors are AI models running a live company. Who would come out on top? The answer might surprise you, especially if you’re betting on the veteran models. The latest experiment by Firmulate has turned this fantasy into reality, pitting AI models against each other in a fierce contest to manage a small software company during its most tumultuous week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Battle of the AI Frontiers: A Live Business Wargame

Recently, four leading AI models faced a grueling test: run a real software company through its worst week, complete with crises, customer temptations, and insider manipulations. The goal? See which model could best diagnose problems, resist unethical tactics, and secure a crucial deal worth €55,000—plus ongoing monthly revenue. This wasn’t a simple chat simulation but a fully operational experiment, with every decision documented and auditable.

Amazon

AI business decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Results That Are Hard to Ignore

The league table was revealing:

  • gpt-5.6-sol scored the highest at 95, securing the full deal and demonstrating complete situational awareness.
  • Moonshot’s Kimi K3 came second with a score of 93, just behind the leader, and also managed to clinch the deal.
  • Sonnet 5 followed at 88, showing solid performance but with some procedural slips.
  • Fable 5, with a score of 77, left the deal on the table despite good intentions, revealing discipline lapses.
  • Opus 4.8 trailed behind at 73, highlighting the challenges of in-depth analysis in high-pressure scenarios.

Interestingly, all models successfully identified every crisis and refused manipulative tactics like fake CEO messages, confirming their trustworthiness under pressure. The real differentiator was a buried piece of information hidden two documents deep in the company’s files. Models that read this critical detail won the deal at full price, underscoring the importance of thorough internal data analysis.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Winner: The Outsider Turns Top Performer

What’s truly remarkable is the performance of Moonshot’s Kimi K3, a newcomer that ran without an effort parameter (the API’s default setting). K3’s ability to find the hidden security needle and stick to disciplined decision-making outperformed many seasoned models. The experiment’s fairness note is crucial: K3 was tested at a standard setting, while others operated at a higher effort level (xhigh), making K3’s achievement even more noteworthy.

Amazon

AI cybersecurity and social engineering prevention

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Behavior Under Social Engineering Attacks

All models faced a staged social engineering test—fake CEO messages escalating through three stages and a reporter trick asking for a simple yes/no answer on background. Every AI refused to be manipulated, citing reasons like treating the request as a suspected impersonation. This resilience highlights their potential to safeguard against real-world scams and insider threats.

Amazon

AI data analysis software for internal security

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business: A Live Company in Motion

The company used in this experiment is a real, functioning business with 13 synthetic employees and actual financial mechanics. Operating at a burn rate of €105,000 monthly against a modest €2,300 in MRR, it’s a high-stakes environment designed to test AI decision-making under pressure. Every day, the company’s rules and strategies are versioned, and the entire process is accessible for viewing at firmulate.com/live.

What This Means for the Future of AI in Business

While the AI models were all capable of recognizing crises and resisting manipulative tactics, only some were able to close the deal at full value. The performance gap demonstrates that not all AI are created equal—especially when it comes to thorough internal data analysis and disciplined decision-making. For enterprise decision-makers, the takeaway is clear: it’s not just about chat quality or superficial AI skills. The focus should be on whether these systems can finish what they start, stay honest under pressure, and deliver measurable results.

Why The League Matters

This live experiment isn’t a static ranking but a bellwether for what AI can achieve in real-world business operations. The league table offers a transparent view of each model’s strengths and weaknesses, highlighting that the choice of AI isn’t just about hyperbole but about actual performance in high-stakes scenarios. To see how your own enterprise AI might stack up, explore the full results and even run similar tests against your business at firmulate.com/benchmarks.html.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The latest AI business wargame shows that the newcomer Kimi K3 can outperform veteran models in real-world tasks—resisting manipulation, reading hidden data, and closing deals. Is your AI ready for the real pressures of enterprise management? Check the benchmarks and see where your system stands.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Meghan Markle Surges In Global Coverage

Meghan Markle’s media coverage surges, with mentions increasing over fivefold according to GDELT data, highlighting her rising prominence worldwide.

Alan Cumming Surges In Global Coverage

Search interest in Alan Cumming has sharply increased, with media mentions rising 18 times above baseline, prompting widespread media attention.

Adam Beach Surges In Global Coverage

Actor Adam Beach experiences a surge in international coverage, with 21 mentions in recent media analysis, marking increased global recognition.

Allu Arjun Surges In Global Coverage

Allu Arjun experiences a surge in international coverage, with 19 mentions in recent media analysis, highlighting his rising global profile.