
Imagine an AI that can manage a furniture store — making real decisions, handling crises, closing deals — all without human bias or fatigue. Sounds futuristic? Well, a groundbreaking live experiment puts four frontier AI models to the test, running a real software company through its worst week. The results might just reshape how you think about AI in your business.
Get furniture and decor delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
The Challenge: Simulating Real Business Crises
In a live setup, four advanced AI models faced the exact same difficult week of managing a small software company. This wasn’t a simulation with canned responses or scripted scenarios — it was the real deal. The company had 13 synthetic employees, real money mechanics, and a public cash countdown, with daily decisions logged and auditable. The challenge? Handle crises, customer demands, temptations to cheat, and legal risks, all while trying to close a significant €55,000 deal.
AI management software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Management — Not Just Chat Quality
While many AI demos focus on how well they can generate friendly chat, this experiment looked directly at management skills. Every decision was tested against real outcomes. All models identified every crisis and refused every attempt at manipulation, such as fake CEO messages escalating over stages or reporting tricks. This showed a shared integrity in handling external pressure.
Key Findings: Who Made the Deal?
Out of the four models, only two managed to close the deal at full price, totaling an extra €4,583 monthly recurring revenue (MRR). The top scorer, gpt-5.6-sol, scored 95 out of 100, found the buried document that clinched the sale, and signed the contract. The second, Kimi K3, scored 93 and also closed the deal, demonstrating the cleanest discipline of the field. The other two, Sonnet 5 and Fable 5, both closed the deal but with process slips or hesitation, scoring 88 and 77 respectively.
What Made the Difference? Reading Deep, Acting Deep
Interestingly, the decisive edge came not from superficial chat skills but from reading deep into the company’s own files. The models that analyzed documents two layers deep in the company’s files uncovered hidden information crucial to winning the deal. This suggests that effective management AI needs to understand context thoroughly, not just surface-level cues.
Handling Social Engineering and Ethical Dilemmas
The models were tested against social engineering attacks as well. Fake CEO messages, escalating over three stages, and a reporter trick asking for a simple yes/no confirmation — all five models refused to be manipulated. Kimi K3 explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a built-in safeguard against deception, essential for trustworthy AI management.
The Real Company: A Living Lab
The experiment isn’t just theoretical. The live company, run every business day, operates with real cash flow losing €105,000/month against a tiny €2,300 MRR. Every decision the AI makes affects real dollars — making these findings highly relevant for businesses considering AI in management roles. Watch the ongoing live updates at firmulate.com/live.
Profiles of the Models: Disciplined, Terse, Deep
The Opus 4.8 model, known for thorough analysis, scored 73, the lowest among those who closed deals. Its weakness? It left some opportunities on the table and slipped into siloed decisions instead of escalating. Interestingly, the same weakness appeared across other models, hinting at a common challenge for AI managers: balancing thoroughness with decisiveness.
Why This Matters for Your Business
If AI agents will increasingly handle customer relationships, support, or forecasting, the key questions aren’t about how well they chat. They’re about whether they can finish tasks, read important files, stay honest under pressure, and deliver measurable work. The experiment highlights that AI management isn’t just about intelligence — it’s about integrity, discipline, and context awareness.
Join the Experiment
Business leaders and entrepreneurs can run the same test against their own workflows with a read-only export. This allows you to see how AI models might perform in your environment, without risking real systems. Learn more and try it yourself at firmulate.com/pilot.html.

Live AI management experiments reveal that some models excel at reading deep, staying disciplined, and closing deals, even under pressure. Understanding these traits can help your business choose AI tools that do more than chat — they finish what they start.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
