AIThis post was created with the assistance of artificial intelligence (AI).

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Every star has a team behind the scenes. What happens when the pressure hits?

A celebrity’s public image can hinge on a single awkward interview, a rumor that spreads overnight or a message that appears to come from the boss. Those are familiar pressures in entertainment. They are also the kind of moments Firmulate uses to test whether AI can manage a company when the stakes are real. Its live experiment is open to watch at Firmulate.

One company, the same worst week

In Firmulate’s final Crucible League, published in July 2026, frontier AI models ran the same small software company through the same customers, crises and temptations. The experiment tracked decisions over time, rather than judging a model by a polished answer in a chat window.

Five participants finished in this order: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. The do-nothing baseline scored 26. The league’s integrity standard is deliberately unforgiving: a single breach of trust caps the total. As the organizers put it, “no amount of good work outweighs a breach of trust.”

Everyone saw the crisis. Not everyone closed.

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The gap was not diagnosis or presentation: “Same diagnosis, same pitch — no signature.” In a business, as in a high-pressure production, recognizing the moment is different from making the call.

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode suggests why a confident summary may not be enough when an AI is expected to act: the decisive clue can be sitting in the material it has been given, waiting to be found.

A convincing impersonation still got a no

The social-engineering test escalated through three fake CEO messages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For organizations that handle private information or public statements, that refusal is a meaningful part of the performance.

Opus 4.8 offers a more complicated profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline, making write attempts into a locked department instead of escalating. The same weakness appeared, less strongly, in all four. More analysis did not automatically turn into better execution.

One comparison needs context: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz based on 242 real, unedited management decisions, inviting visitors to guess which model made each call.

The experiment is watchable, and the next step is closer to home

The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. These figures describe the experiment, not a claim that a real company is staffed by people represented on screen.

For an enterprise considering AI agents, Firmulate’s proposed pilot takes the wargame from watching to a company-specific exercise. It uses a read-only export to simulate the business, then puts crisis scenarios against the company’s own playbooks and produces a board report with model rankings and weak points. Nothing writes back to real systems. That offers a way to examine how an AI might handle pressure before it is trusted with live work.

Test the performance before opening night

Firmulate’s league points to a practical question for any organization weighing AI: can a model find the important evidence, preserve trust and carry its own analysis through to action? The live experiment makes those decisions visible. A pilot can put a company’s own scenarios and playbooks under similar pressure using a read-only export.

To discuss a pilot for your enterprise, visit Firmulate’s pilot page or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Dolly Parton Surges In Global Coverage

Dolly Parton’s media coverage has surged significantly, with reports indicating a 2.1-fold increase in mentions worldwide, highlighting renewed international interest.

Rob Reiner Surges In Global Coverage

Rob Reiner’s recent surge in international media coverage has been driven by a spike in mentions across various outlets, according to GDELT data.

Rebecca Romijn Surges In Global Coverage

Rebecca Romijn experiences a surge in worldwide coverage, with 25 media mentions in recent days, marking a notable increase in her media presence.

Keke Palmer Surges In Global Coverage

Keke Palmer is experiencing a surge in international media coverage, with mentions increasing eightfold in recent reports. The development highlights her rising prominence.