
Imagine a fitness class where the trainer is a virtual assistant, constantly tested under pressure, and every decision is recorded for the world to see. Now, transpose that idea into the business world: a real company run entirely by AI models, fighting to survive, with no human employees in sight. Welcome to the world of Firmulate, where a small software company is being put through its paces in a public, transparent experiment that reveals the true strengths—and weaknesses—of AI management.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Watching AI in Action
At the core of this experiment are four cutting-edge AI models, each tasked with running a small, real business through its worst week. The same challenges, the same clients, and the same crises are faced by each model, making it a unique test of AI decision-making. Every move, every decision, is versioned and auditable, providing a clear view of how these models handle complex management tasks in real-time.
As an affiliate, we earn on qualifying purchases.
Real Business, Real Money, Real Stakes
This is not some simulation or demo: the business operates with a monthly recurring revenue (MRR) of just €2,300, yet it burns through €105,000 each month. The site, firmulate.com/live.html, offers a live window into this relentless struggle for survival. Each workday, the company’s decision logs are updated, showing how the AI models navigate crises, strategic choices, and ethical dilemmas.
AI decision-making tools for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Models Did — and Didn’t Do
All four models demonstrated impressive awareness of crises: they identified every problem and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. Kimi K3, for example, explicitly cited concerns over impersonation when faced with a suspicious request. Despite their vigilance, only two models managed to close the €55,000 deal that their own analysis justified. The other two, despite correct diagnoses, left the opportunity unclaimed, illustrating a critical gap between recognizing value and executing on it.
internal document analysis AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses: Reading Between the Files
The most revealing discovery was in the details buried in internal files—information not visible in customer interactions but crucial for closing deals. The models that read and interpret these internal documents at depth won the deal at full price (+€4,583 MRR). This highlights a vital point: effective AI decision-making depends on access to and understanding of internal context, not just surface-level customer interactions.
As an affiliate, we earn on qualifying purchases.
Ethics and Trust in AI Management
Beyond strategic decisions, the experiment tested the models’ resistance to unethical pressure. When presented with staged social engineering requests—like escalating fake CEO messages—the AI models uniformly refused. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This underscores the importance of built-in safeguards in AI systems managing sensitive business operations.
The Reality of a Zero-Employee Company
This experiment is hosted by Firmulate, a platform that runs AI models as complete companies, complete with everyday crises, real money mechanics, and a public cash countdown. The company’s setup makes clear that AI can simulate management at a granular, decision-by-decision level, but it also exposes the fragility of such systems when it comes to closing deals or making profitable moves.
Lessons for Business and AI Developers
The results are stark: while all models detected crises and refused manipulation, only the most disciplined and thorough models succeeded in closing deals—a key measure of management effectiveness. The experiment’s leaderboard reveals that GPT-5.6-sol scored highest at 95, followed by Kimi K3 at 93, then Sonnet 5 at 88 and 77 respectively. The close calls and slips highlight how easily discipline can slip when models are less comprehensive or when they don’t run at higher effort levels.
What This Means for Your Business
For companies considering AI for customer service, sales, or operational management, the take-home is clear: the real question isn’t whether AI can generate decent chat, but whether it can finish what it starts, read important internal information, and stay honest under pressure. The live experiment offers a rare, unfiltered look at AI’s true capabilities and limitations in a real-world, high-stakes environment.
See It for Yourself
If you’re curious, you can watch the company in action, read the actual decision logs, or even run your own business wargame against a read-only export. This transparency is what sets Firmulate apart—providing a glimpse into the future of AI-managed enterprises and the challenges they face.

Watching a real, money-losing AI-run company in live time reveals vital truths: AI can detect crises and refuse manipulation, but closing deals and reading internal documents remains a challenge—an essential insight for future business automation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.