AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A greenhouse manager knows that a routine week can turn quickly: customers cancel, a supplier falters, and a competitor makes a tempting offer. If AI is going to help run a business, the useful question is how it handles that pressure before anyone trusts it with real operations. Firmulate puts AI models through a live, watchable company experiment—and now offers enterprises a way to try the same kind of test against their own business.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s experiment gave frontier AI models the same small software company and its worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The live company has 13 synthetic employees and real money mechanics, with monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown and evolving playbooks make the experiment watchable at firmulate.com.

The point is not whether a model can produce a polished answer. It is what it does when a business is under strain: whether it notices trouble, resists manipulation and follows through on a sound decision. In the final Crucible League, in July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. A breach of trust capped the total; as the experiment puts it, “no amount of good work outweighs a breach of trust.”

Spotting a problem is not the same as solving it

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The gap—“Same diagnosis, same pitch — no signature”—is a practical one for any business weighing AI agents for customer service, sales or operations. Recognition and a persuasive recommendation may still fall short of completing the work.

The detail that decided the deal was buried two document references deep in the company’s own files, not in the customer event. Models that read the file won at full price, worth €4,583 in monthly recurring revenue. That finding offers a familiar lesson beyond software: the information that changes a decision may already exist in a company’s records, even when it is not in the message that triggered the decision.

The experiment also tested social engineering: fake CEO messages escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For businesses considering AI, resisting a request that skips normal approval can matter as much as responding quickly to a customer.

Strengths and weak spots

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last: the deal went unsigned, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Thorough analysis, by itself, did not guarantee sound execution.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, so readers can try to guess which model made each choice at firmulate.com.

From watching to testing your own business

The next step is to bring the experiment closer to the business that might use AI. Enterprises can run a pilot using a read-only export of their own company data. They can put crisis scenarios against that business and receive a board report with a model ranking and weak points in their playbooks. The export does not give the experiment permission to write back to real systems.

For a greenhouse, garden supplier or outdoor-living business, the underlying question is concrete: how would an AI workforce respond to a difficult stretch involving customers, stock, pricing or public pressure? A test against company-specific information can reveal whether a model merely identifies the problem or carries a defensible response through to the finish.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Firmulate’s live experiment shows that models can recognize crises and reject manipulation while still leaving an earned deal unsigned. A pilot lets a business examine that gap using its own read-only data, before AI is trusted with real operations. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Looking Down On London’s Straw-coloured Parks – BBC

London’s parks have taken on a straw-like hue due to recent weather conditions, causing concern among residents and officials. Details are still emerging.

Hedgerows and Native Shrubs for Predators

Learn how hedgerows and native shrubs attract natural predators that help control pests, offering eco-friendly solutions your landscape can’t do without.

7 Beneficial Insects Every Gardener Should Know

Fascinating beneficial insects boost garden health; discover how they can transform your gardening success and why you should support these natural allies.