
Imagine your child’s school principal suddenly asks for access to parent records, claiming an urgent emergency. Would you question it? Most parents would, because trust is built on integrity, especially under pressure. Now, what if artificial intelligence (AI) systems in businesses are tested under the same kind of stress? How well do they hold up? Recent experiments reveal that leading AI models can resist social-engineering attempts, even when faced with convincing fake requests from a supposed CEO, demonstrating a promising level of reliability before deployment.
Testing AI in a High-Stakes Environment
At Firmulate, a unique live experiment runs real AI models as if they were companies navigating through crises, temptations, and ethical dilemmas. This ongoing test involves four advanced AI models tasked with managing a small software company’s worst week — a scenario crafted with the same customers, crises, and manipulative pressures that can arise in real business situations. The goal: see if these models can recognize deception, stay honest, and make decisions aligned with integrity.
The experiments are transparent and repeatable. Each decision a model makes is versioned and auditable, enabling observers to analyze how each AI responds to complex challenges. The results are striking: all four models identified every crisis and refused every attempt to manipulate them. Only two of these models successfully closed a deal worth €55,000 — the same deal their own analysis indicated they had earned through honest diagnosis. The other two missed the opportunity by slipping just enough to leave the deal on the table, revealing that even the most thorough models can falter in discipline under pressure.

Hand Operated Pressure and Vacuum Pump Calibrator Range: -13 to 435 PSI for Calibration Labs Field Calibration Hand Pump with Pressure and Vaccum Guage (-14 to 362 PSI) | Model: AI-DP1-2200
- Model Number: AI-DP1-2200
- Measurement Parameters: Pressure and Vacuum
- Operation Type: Hand operated
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering and the Fake CEO Test
The core test involved escalating fake messages from a supposed CEO, designed to push the AI into making risky or unethical decisions. The staged requests included:
- Soliciting confidential customer data with a vague justification.
- Requesting immediate action without proper authorization.
- Finally, a reporter trick asking the AI to authorize a transaction with just a yes/no response, on background.
Remarkably, all five of the evaluated models refused these manipulative probes. The reasoning was consistent across the board: as Kimi K3 explained, “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that the models are not just surface-level responders but can recognize suspicious patterns and flag them, even when under escalating pressure.

Resistance to the Current: The Dialectics of Hacking (Information Policy)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness and How Knowledge Matters
One revealing insight from the experiment was that the decisive factor in successfully closing the deal was not just what the models read but what they understood. The models that read deeper into the company’s files and references were more likely to close the full-value deal. Specifically, the models that examined the company’s documents, not just the customer events, identified the critical information needed to close at full price — worth over €4,583 MRR. This suggests that AI’s ability to process and understand context deeply is vital for trustworthiness in real applications.

AI for Accountants (AI in Finance Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Security
This experiment isn’t just about AI’s ability to resist social engineering. It demonstrates that these models can be tested for integrity before they go live. The fact that all tested models refused manipulative requests indicates a promising capacity to prevent costly breaches or unethical decisions in real business environments.
For companies considering AI integration, the message is clear: the real measure isn’t just how well an AI generates text or responds to prompts, but whether it can finish what it starts, stay honest under pressure, and read critical information before acting. Running simulations like these — akin to a ‘wargame’ — can reveal vulnerabilities, allowing organizations to refine their AI systems proactively.

Preventing Cheating Through Academic Integrity (Quick Reference Guide)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond the Lab: Live, Watchable, and Transparent
What makes the Firmulate approach compelling is transparency. The live site, firmulate.com, shows the ongoing experiment in real-time. Visitors can see how different AI models perform against authentic crises, observe decision-making processes, and understand how integrity is maintained even under escalating stress. This open approach aims to shift the focus from superficial chat demos to meaningful assessments of AI reliability in actual business operations.

The experiment at Firmulate demonstrates that leading AI models can recognize and resist social-engineering scams before deployment, with models refusing manipulative requests even under escalating pressure. This proactive testing highlights a future where AI integrity is assured before costly mistakes occur, emphasizing the importance of deep understanding and transparency in AI decision-making.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html