85% of Companies Burned by AI Mistakes Cut Human Deployment Decisions, Survey Finds
Carl Franzen
August 18, 2026, 8:35 am PT
(Image credit: VentureBeat made with Midjourney v8.2)
Enterprises that suffered from AI agents passing evaluations but failing in production are swiftly removing humans from deployment decisions, contrary to decreasing trust in automated evaluation, according to new VB Pulse research.
In July:
- 13% of 108 surveyed enterprises trusted automated evaluation, up from 5% the previous month.
- Concern over poor alignment between tests and real-world results dropped from 29% to 19%.
However, 49% of respondents reported that an AI agent or LLM-powered feature that passed company testing subsequently caused customer-visible problems, with 24% experiencing this issue more than once.
The latest findings from VentureBeat Intelligence highlight a more concerning trend: the gap between confidence in evaluation and evidence of improved failure prevention.
Key Insights:
- 4% of enterprises with AI feature failures placed complete faith in automated checks, while 24% with no comparable incidents expressed full confidence.
- Companies like Raindrop.ai, offering automated agent error monitoring and mitigation, are witnessing significant market growth.
VB Pulse surveyed 108 individuals from companies with at least 100 employees, with 69% considering themselves the final AI-buying authorities. The sample primarily consisted of midsize organizations (63% with 100-2,499 employees).
Notable Findings:
- Increased confidence in automated evaluations, but unchanged outcome rates.
- More respondents expressing complete confidence in automated evaluations.
- Fewer citing poor real-world alignment as a concern.
- Integration ease becoming the primary buying factor.
- Braintrust gaining primary-platform share.