AI News

Enterprise AI teams appear to be making a strategic decision about where their agent systems will run: increasingly, on the platforms of the major model providers. But a new cluster of VentureBeat Pulse Research surveys argues that the harder problem is no longer choosing a platform. It is getting real agent deployments, controls and operating discipline to catch up with the ambition.

Across four June 2026 survey waves published by VentureBeat AI, respondents described a market that is standardizing around provider-native stacks from companies such as Anthropic, OpenAI, Microsoft and Google, while still struggling with the foundations that make autonomous software usable in production. Those weaknesses show up across orchestration, retrieval, evaluation and security.

The most pointed finding comes from the orchestration survey: 71% of enterprises said a quarter or fewer of their deployed “agents” are actually multi-step orchestrated workflows rather than simple chatbot wrappers. That matters because the same respondents said they judge orchestration success by reliable task completion and multi-step workflow management, not by chatbot UX. In other words, many teams are buying and planning for agentic systems before most of what they have in production fully qualifies as agents.

Taken together, the four surveys suggest a broad enterprise pattern. Teams are comfortable adopting the bundled controls and infrastructure that come with frontier model platforms, but they remain far less mature in identity, context governance, real-world evaluation and cost enforcement. For builders and buyers, the implication is sharp: the next competitive edge may come less from another orchestration framework and more from proving that agents can be trusted, measured and contained.

One story across four layers of the stack

The orchestration report covered 101 enterprises with more than 100 employees and found that Anthropic’s Claude was the primary orchestration platform for 40% of respondents, ahead of Microsoft at 18% and OpenAI at 13%. VentureBeat explicitly cautioned that these figures are not market share and should be read as a directional snapshot of a self-selected sample, not a spend-weighted industry census.

Even with that caveat, the pattern aligns with the other three surveys. In retrieval and context infrastructure, OpenAI file search and Google Vertex AI Search led production usage ahead of specialist vector databases. In security, OpenAI guardrails and controls from Google and Microsoft dominated current deployments. In evaluation, OpenAI’s native evals and traces led the field, tied with respondents that said they used no dedicated evaluation tooling at all.

That consistency matters. It suggests that enterprise teams are defaulting to the tools closest to the underlying model and cloud stack they already use. VentureBeat framed this as “model gravity” in orchestration and a provider-native default in the other layers. The practical reading is simpler: buyers want fewer moving pieces, faster deployment and tighter integration, at least for first-wave production systems.

But those same teams are not signaling deep long-term loyalty. The surveys found 68% plan to adopt, add or replace an orchestration platform within a year, 57% plan to switch or add retrieval providers, 64% expect changes in evaluation tooling and 59% plan agent-security tooling changes. Standardization is happening, but the stack is still fluid.

The chatbot label is outrunning the software reality

The orchestration data provides the clearest example of the gap between marketing language and operational reality. According to the survey, only 10% of enterprises said more than half of their deployed “agents” are true multi-step orchestrated workflows. Most admitted that the portfolio is still dominated by single-prompt assistants.

That finding lands at a moment when “AI agents” has become the default label for a wide range of products. The VentureBeat data implies many enterprise deployments are better described as advanced chatbots with enterprise connectors than as autonomous systems coordinating tools, memory and policy over multiple steps.

This distinction is not semantic nitpicking. It affects architecture, testing, security and budgeting. A chatbot wrapper may need good prompts, grounding and UX polish. A real multi-step agent needs runtime permissions, rollback paths, traceability, spend controls, durable state, evaluation aligned to task outcomes and often a hybrid control plane that sits above any one model vendor.

Respondents appear to understand that gap. By the end of 2026, 51% said they expect a hybrid control plane, combining provider-native systems with external orchestration. Only 6% expected to rely fully on a provider-managed service. The main reason, according to the survey, was vendor lock-in, which 35% named as the biggest risk if control lives inside a model-provider platform.

This is why the story is not simply “providers are winning.” They are winning the current deployment default, but enterprises are simultaneously building escape hatches.

Security, context and evaluation are still behind the rollout

The three companion surveys suggest the orchestration gap is only one piece of a broader deployment maturity problem.

In security, VentureBeat’s survey of 107 enterprises found that 54% reported either a confirmed AI agent security incident or a near-miss. Only 32% said every agent has its own scoped identity, while 69% reported credential sharing somewhere in the fleet. Just 30% isolate their highest-risk agents in sandboxes. The publication noted that the association between scoped identities and lower incident rates is suggestive, but not proof of causation, especially given the modest, self-selected sample.

In context infrastructure, a separate 101-enterprise survey found 57% had seen agents produce confident but wrong answers traced to missing or inconsistent business context. Retrieval-augmented generation was the primary context source for 38% of respondents, and provider-native tools led actual usage. At the same time, 58% said they either run or are building a governed semantic layer, implying that many organizations now see raw retrieval as insufficient without a stronger business-definition layer.

In evaluation, the warning signs were equally direct. Half of the 157 surveyed enterprises said they had shipped an agent or LLM feature that passed internal evaluations and then failed in front of a customer. Only 5% said they fully trust automated evaluation today. Yet 66% either already allow or are building toward zero-human-in-the-loop deployment for low-risk agents.

That combination is probably the most important signal in the cluster. Enterprises are not waiting for perfect assurance before automating more decisions. They are increasing autonomy while admitting that the controls meant to certify that autonomy are incomplete.

What the evidence does and does not prove

All four source items come from VentureBeat AI’s own Pulse Research program rather than company filings, product documentation or broad independent market datasets. That makes the cluster useful as a directional indicator of current buyer behavior, but not definitive market measurement.

VentureBeat itself repeatedly states the limitations: the samples were self-selected, skewed toward mid-market organizations in several surveys, and drawn from single June 2026 waves rather than pooled longitudinal panels. In the orchestration survey, respondents were also asked for a single primary platform, which produces a snapshot of current preference rather than a full view of multilayer usage or spending.

That means platform-specific numbers for Claude, OpenAI, Microsoft, Google, Vertex AI Search or LangChain/LangGraph should be handled carefully. They show where this cohort says it has placed its primary bet, not the full industry leaderboard.

Still, the broader patterns are harder to dismiss because they recur across the four surveys. Provider-native tooling appears to lead in orchestration, retrieval, security and evals. Enterprises report substantial incident and failure rates. And large majorities plan to revisit their stack within the year. The individual percentages may move with a different sample; the structural story looks more robust.

What this means for builders and enterprise buyers

For product teams building agent systems, the cluster suggests that orchestration is becoming a control-plane problem, not just a framework choice. Buying into Claude, OpenAI or Microsoft may accelerate initial deployment, but it does not remove the need for governed context, runtime identity, production evals and cost cutoffs.

For infrastructure startups, the opening may be narrower than “replace the platform” but larger than it first appears. The surveys show limited current penetration for specialist vendors, yet strong intent to add layers that the native stacks do not fully solve. That is especially relevant for non-human identity, semantic governance, live quality monitoring and spend management.

For enterprise AI buyers, the message is to examine whether internal “agent” portfolios are really autonomous workflows or mostly assisted chat interfaces. If they are the latter, the right next investment may not be another orchestration abstraction. It may be the operational disciplines that make a workflow trustworthy when it becomes truly autonomous.

The surveys also suggest that many organizations still buy for convenience and then monitor for correctness later. That ordering works during experimentation. It becomes riskier once agents can act across internal systems or ship changes without human approval.

What to watch next

The clearest follow-up signal is whether provider-native dominance hardens or plateaus. If OpenAI, Anthropic, Microsoft and Google continue to absorb orchestration, retrieval, evals and security into one stack, enterprises may accept more lock-in than they currently say they want.

A second signal is whether hybrid control planes move from stated intent to deployed architecture. If external layers around LangChain/LangGraph, custom orchestration and identity controls gain real production share, that would confirm that buyers want providers for model access but not for complete control.

Third, watch whether enterprise budgets shift from workflow tooling toward the less visible layers implicated by the surveys: semantic governance, scoped identities, real-time output monitoring and programmatic cost enforcement. Those are not headline features, but they are the capabilities most directly linked to the failures respondents reported.

Finally, the biggest operational test is definitional. As enterprises move more systems into production, the market will need clearer separation between chatbot interfaces, tool-using assistants and fully orchestrated AI agents. Right now, the label is ahead of the implementation.

Creati.ai perspective

The most revealing part of this story is that enterprises do not seem confused about what good agents should do. They say they want reliable multi-step execution, controlled permissions, accurate context and evals that predict real-world outcomes. The gap is that their current deployments often do not meet that bar, even as procurement and platform choices race ahead.

That is why this looks less like a platform war than an operations reckoning. The model providers are winning the right to be the default starting point, but the harder commercial problem is still unsolved: turning LLM-powered workflows into dependable software systems. The teams that close that gap first will likely do it with better controls and clearer definitions, not just better demos.

Featured

Enterprise ‘agent’ stacks are consolidating on model providers, but new data suggests deployment discipline is the real bottleneck

VentureBeat survey data suggests enterprise AI teams are standardizing agent stacks fast, but security, context, evals and cost controls lag badly.