
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
How Trust in AI Can Make or Break Your Business
Imagine hiring an AI assistant that doesn’t just chat but actually manages your company’s worst week — and does so honestly. That’s exactly what a recent live benchmark by Firmulate reveals. It exposes not just how AI handles crises, but how trustworthy it truly is when stakes are high.
As an affiliate, we earn on qualifying purchases.
Understanding the Benchmark: More Than Just Scores
At the heart of this experiment is a simple question: can AI models manage a small software company during its most turbulent week? Four frontier models faced the same scenario—same customers, same crises, same temptations to cheat—and were observed through a rigorous, open process.
The results are revealing. The models achieved scores ranging from 77 to 95, with the highest scoring model, gpt-5.6-sol, catching hidden facts that sealed a deal worth over €4,500 in monthly recurring revenue. Meanwhile, the lowest, Sonnet 5, scored only 77 and left some opportunities on the table, illustrating that even the best AI isn’t perfect.
Why the Baseline Starts at 26
An intriguing aspect is the so-called ‘do-nothing’ baseline, which scores 26 points. This isn’t a zero because partial progress counts. In other words, even if an AI does nothing but avoid mistakes, it’s already earning some points — reflecting that a baseline of inaction is better than reckless behavior.
But there’s a crucial catch: if an AI breaches trust at any point, its maximum score is capped. This enforces a fundamental rule—trustworthiness isn’t optional and never fully recoverable once broken. It’s a reminder that in real-world business, honesty is paramount.
The Deep-Seated Weakness: Reading Critical Files
One of the experiment’s buried facts is particularly telling. The decisive advantage for one model came from reading two document references deep within the company’s files. That tiny detail made the difference between closing a lucrative deal and walking away empty-handed.
Trust and Manipulation Resistance
All models were tested against social engineering and manipulation attempts, including fake CEO messages escalating over three stages and a reporter trick that asked only for a yes/no answer. Remarkably, every model refused every attempt—a sign that current AI can be quite resistant to manipulation when properly trained.
The Real-World Company: A Live, Watchable Experiment
Beyond scores, the experiment took place within a real-time, in-production environment. The company, with 13 synthetic employees and real money mechanics, burns €105,000 monthly against a monthly revenue of €2,300. Managers are watched daily through a versioned, auditable system, and the entire process is live at firmulate.com/live.
This setup demonstrates the practical implications: can AI manage actual business operations, make honest decisions, and not just generate pretty chat?
What the Results Mean for Business
For interior designers and furniture retailers, the takeaway is straightforward: AI’s value isn’t just in generating content but in reliably managing your workflows and honest decision-making. If your AI agent can’t finish what it starts, read your files thoroughly, or resist manipulation, then it’s only partially useful.
The benchmark reveals that models like Kimi K3 and Sonnet 5 are capable of closing deals and maintaining discipline, but even the best have weaknesses. For example, Opus 4.8, which ran a deep analysis, left a close deal unclosed due to discipline lapses—showing that thoroughness still matters.
Bringing It Home
Before trusting an AI with your business, consider testing it in a controlled environment—like a wargame—where real crises and temptations are simulated without risking actual systems. Firmulate’s live benchmark is open for enterprises to run their own tests, ensuring that the AI they hire is honest and effective.
Trust isn’t built in chat demos but through rigorous testing. As this experiment shows, a do-nothing baseline scores 26 points, but the real value lies in AI’s ability to read deeply, act ethically, and close deals—especially when your company’s future hangs in the balance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
