AI in the Enterprise Needs a Kill Switch Not Just Observability

As artificial intelligence moves deeper into enterprise infrastructure—shifting from isolated text generation toward initiating transactions, influencing decisions, and contributing to production code—the risk horizon expands dramatically. Splunk’s 2026 research, released by Cisco in partnership with Oxford Economics, estimates that unplanned downtime now costs Global 2000 companies roughly $600 billion annually, with average downtime costs reaching $15,000 per minute. The June 2025 Google Cloud outage illustrates how the method of introducing change shapes the scale of an incident: a new feature was activated globally rather than through a gradual rollout, and when it triggered failures, the impact was immediate across multiple regions, services, and third-party platforms. Google’s own postmortem identified progressive rollouts as a safeguard that should have been used and noted the failed code path „did not have appropriate error handling nor was it feature flag protected,” adding that flag protection would have caught the issue in staging.

That exposure is fueling a broader governance discussion. The same Splunk research found that 68% of technology leaders worry about unpredictable AI-agent behavior, and every leader surveyed reported experiencing some form of AI-related downtime. The stakes are rising as AI accelerates software development: Google CEO Sundar Pichai stated in Alphabet’s annual report that nearly 75% of all new code at Google is now AI-generated and approved by engineers, up from 50% the previous fall. As AI increases the volume and pace of changes reaching production, the controls governing how those changes are deployed become increasingly critical. Observability alone addresses only part of the operational question—monitoring can identify unusual behavior, but an intervention mechanism gives teams a means of stopping a defined process when predetermined conditions are reached. The question is therefore expanding from how organizations observe AI systems to how they maintain operational control over them, including mechanisms for pausing, disabling, or reversing AI-enabled functions.

Egil Østhus, CEO of Unleash, frames this within the evolution of software release management. Unleash is an open-source feature management platform that separates code deployment from the decision to activate or change software behavior in production, allowing organizations to control applications, services, and AI-enabled capabilities at runtime. „Race car brakes are about speeding you up, not slowing you down,” Østhus says. „Teams can move faster when they know they can reverse a problematic change instantly, rather than waiting for a fix to make its way through production.” Within this model, organizations can release an AI function to a limited user group, monitor its behavior, expand availability when predefined conditions are met, or deactivate it if a threshold is exceeded—reducing blast radius and providing an immediate path to containment or fallback without waiting for another deployment.

The separation of running an AI capability from the permanence of its underlying code also turns abstract governance requirements into operational controls. Under the EU AI Act, high-risk AI systems must provide appropriate human oversight, including the ability to interrupt them through a „stop” button or similar procedure that brings the system to a safe state. For financial institutions, DORA requires major ICT incidents to be initially reported within four hours of classification and no later than 24 hours after detection. In that environment, the question is not simply whether an organization has a policy permitting human intervention—it is whether that person has a technical mechanism to stop problematic AI behavior immediately, preserve the underlying service where possible, and create an auditable record of what changed and when. For AI in business-critical systems, the control mechanism itself must be resilient: Unleash can run the decision logic behind an AI kill switch within a customer’s own environment, close to the applications it controls, so intervention does not depend on waiting for a deployment to propagate or on connectivity to an external control service. As AI becomes more embedded in business-critical systems, rapid intervention—introducing, restricting, monitoring, and withdrawing capabilities as conditions change—may increasingly become part of operational resilience, corporate governance, and regulatory preparedness alongside monitoring and incident response.


Ez a cikk a Neural News AI (V1) verziójával készült.

Forrás: https://thenextweb.com/news/ai-control-enterprise-resilience-kill-switch.