firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine an AI-driven company that operates transparently in real time—facing the same crises, temptations, and tough decisions as any human-led business. Now, imagine watching it struggle, succeed, and sometimes stumble, every single day. This isn’t fiction; it’s the live experiment by Firmulate, a company dedicated to testing AI’s ability to run a business in the real world. For those caring for seniors or managing complex care services, this story is a glimpse into the future of automation, honesty, and decision-making under pressure.

The Live Company That’s Always in the Red

At the core of this experiment is a real software company run entirely by AI models. It has 13 synthetic employees—digital decision-makers guided by a set of over 680 self-learned rules. Every workday, they face crises and opportunities just like human managers, but with one critical difference: every decision is versioned and auditable, offering a transparent view into how AI handles complexity.

Despite this sophisticated setup, the company is far from profitable. It burns through €105,000 each month, while its revenue stands at a modest €2,300 monthly recurring—making it a vivid example of a business in survival mode. Yet, the goal isn’t profitability but understanding AI’s true capabilities and limitations in managing real-world operations.

AI for Small Business: From Marketing and Sales to HR and Operations, How to Employ the Power of Artificial Intelligence for Small Business Success (AI Advantage)

AI for Small Business: From Marketing and Sales to HR and Operations, How to Employ the Power of Artificial Intelligence for Small Business Success (AI Advantage)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How AI Models Tackle Crises and Temptations

The experiment pits four leading AI models—based on frontier language models—against the same weekly business challenge: managing a small software firm during its worst week. These models face the same customers, same crises, and the same temptations to cheat or manipulate the system. Every decision they make is recorded, and their responses are auditable.

Remarkably, all four models identified every crisis and refused every attempt at manipulation. For example, when fake CEO messages escalated in a staged social engineering attack, all models refused to comply, citing concerns about impersonation or approval bypasses. This demonstrates a fundamental strength: AI’s ability to recognize social engineering and ethical boundaries, even when pressured.

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness and the Missed Opportunity

However, the real story lies beneath the surface. The models that succeeded in closing the high-value €55,000 deal relied not just on surface-level diagnosis but on insights buried two document references deep within the company’s own files. When an AI read these internal documents thoroughly, it discovered a critical piece of information that led to sealing the deal at full price, adding +€4,583 monthly recurring revenue.

This indicates that the AI’s ability to read and interpret internal documentation—beyond surface interactions—is a decisive factor in business success. It highlights a potential weakness: superficial reading or shallow analysis risks missing vital clues that could unlock value.

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Disappointing Performance of the Most Disciplined Model

Among the four models, Opus 4.8 was the most thorough—analyzing over 80 learned rules and conducting deep evaluations. Yet, it ranked last in the results. Its discipline slipped at a critical moment—writing attempts into a locked department instead of escalating issues—and left a close deal on the table. This demonstrates that even highly detailed, rule-based AI systems can falter under real-world pressures, especially without an explicit effort parameter in their configuration.

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for the Future of Automation and Decision-Making

This experiment offers a sobering yet illuminating view for industries that rely on precision, honesty, and complex decision-making—such as senior care, healthcare management, or any service where trust and reliability are paramount. The key takeaway is that AI’s effectiveness isn’t just about generating convincing narratives or responses; it’s about whether it can finish what it starts, read internal data thoroughly, and remain honest under pressure.

For organizations considering automation, the question should be: Will this AI finish its work reliably? Will it stay honest when tempted? And will it deliver useful results that justify its cost?

Watch the Live Experiment in Action

The company, running every weekday, is openly accessible at firmulate.com/live.html. You can observe real crises unfold, see the decision logs, and watch how different models perform. It’s a rare, unfiltered view into the ongoing struggle of AI systems to manage real business complexity—and whether they can be trusted to do so sustainably.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

Bed Exit Alarms: How to Reduce False Alarms (Without Turning Them Off)

Transform your approach to bed exit alarms by minimizing false triggers and enhancing patient safety—discover innovative strategies that can make a real difference.

Medical Alert Systems: The Features People Pay For (But Don’t Need)

Are you paying for features in medical alert systems that you don’t actually need? Discover what truly matters for your safety.

AI’s Hidden Strength: Why Some Models Close Business When It Matters Most

Discover how AI models perform under real crisis conditions, revealing crucial differences in their ability to finish tasks, read deeper, and maintain trust—vital for senior care.

Fall Detection: How It Works—and Why It Sometimes Misses

Preventing missed falls is crucial for safety, but what factors contribute to inaccurate detection? Discover the complexities behind fall detection technology.