Women in AI by FemTechConf

AI Incident Response: How Organisations Should Prepare for Model and Agent Failures

A practical incident-response model for AI failures, covering preparation, detection, containment, investigation, recovery and organisational learning.

By Elena Marković, Women in AI Editorial Fellow · 30 August 2026

An AI incident is not limited to a hacked model.

It may be a confident false answer that reaches thousands of customers, a prompt injection that causes an agent to expose data, a ranking system that disadvantages a group, an automated action taken without authority, a provider change that silently degrades quality or a human operator who trusts an output too much.

Traditional incident response remains essential, but AI expands the evidence, expertise and containment options an organisation needs.

Define an AI incident before one happens

An AI incident is an event in which an AI system causes, contributes to or creates a material risk of harm, policy breach, security compromise or operational failure.

Create severity levels tied to impact, not publicity.

Consider:

harm to people or rights; exposure of sensitive data; unauthorised actions; scale and duration; reversibility; legal or regulatory obligations; financial and service impact; evidence that a failure is systematic; uncertainty about continuing operation.

A single harmful decision can be severe when the consequence is irreversible. A low-severity error can become serious when repeated at scale.

Build the response team

AI incidents cross organisational boundaries. The response team may require:

service owner; machine learning or AI engineer; security incident lead; data owner; product and operations lead; legal, privacy and compliance; responsible AI or model risk; communications; relevant domain specialist; senior decision-maker with authority to restrict or stop the system.

Record who leads at each severity and who can disable model calls, revoke tool permissions, switch providers, roll back a release or suspend a use case.

The NCSC's secure AI development guidance treats incident management, logging and monitoring as lifecycle responsibilities. Those controls must be designed before deployment.

Prepare an AI system inventory

You cannot respond quickly if you do not know which systems depend on a model, dataset or provider.

For each production use case, record:

business owner and technical owner; intended use and prohibited uses; model, version and provider; prompts, retrieval sources and tools; sensitive data touched; user groups and affected people; permissions and external actions; evaluation baseline; monitoring and alerting; rollback or fallback; supplier contacts; notification obligations.

Include unofficial or embedded AI dependencies where possible. A customer service product may depend on moderation, transcription, retrieval and a foundation model, each with different failure modes.

Detect signals that metrics miss

Operational dashboards will catch outages and latency. They may not catch persuasive misinformation or a slowly emerging disparity.

Use several signal types:

automated quality and safety evaluations; drift and distribution monitoring; security alerts; tool-call and permission logs; human overrides; customer complaints; employee escalation; appeal outcomes; subgroup analysis; unusual cost or usage patterns; provider notices and release changes.

Make reporting easy for frontline workers. The first person to notice a pattern may be a support agent, caseworker or customer, not an ML engineer.

The OECD AI Incidents Monitor is useful for understanding the breadth of real-world harms and near misses. Organisations can learn from external incidents before experiencing the same pattern.

Contain first, diagnose carefully

Containment should reduce harm while preserving the evidence needed to understand the event.

Options include:

disable an affected feature; revoke tool or data permissions; route all outputs to human review; lower automation thresholds; switch to a known model or prompt version; isolate a retrieval source; block an input pattern; stop processing a user group if performance is unsafe; return to a manual process; suspend the system entirely.

A generative AI system rarely has one simple rollback. The model may be unchanged while the prompt, retrieval corpus, tool schema or provider policy moved. Version all relevant components.

Avoid "fixing" logs or overwriting artefacts during containment. Preserve model and prompt versions, input and output records where lawful, retrieval results, tool calls, user actions, configuration, access logs and monitoring data.

Investigate the whole sociotechnical system

Do not stop at "the model hallucinated". That description is not a root cause.

Ask:

Was the use case appropriate for probabilistic output? Was the instruction clear and versioned? Did retrieval return the right evidence? Did the interface encourage over-trust? Was a human reviewer trained and given enough time? Were permissions broader than necessary? Did a threshold or guardrail fail? Did the provider change the model? Did monitoring look at the affected population? Was the system used outside its approved scope?

NIST's Generative AI Profile highlights risks including confabulation, data privacy, harmful bias, information security and human-AI configuration. These risks interact. A false output becomes more consequential when an agent can act on it and a user cannot challenge it.

Communicate by audience

Internal and external communication should be factual, timely and proportionate.

Prepare messages for:

people directly affected; customers and users; employees and operators; regulators or authorities; suppliers and partners; leadership and the board; the public, where appropriate.

Say what happened, what is known, what remains uncertain, what action has been taken and what affected people should do. Do not imply certainty before the investigation supports it.

Legal and privacy teams should determine notification duties. Communications should not delay urgent containment.

Recover with evidence

Recovery is not "turn it back on".

Define the conditions required for return:

root cause understood to an acceptable level; control implemented and tested; affected outputs or decisions reviewed; data and configurations validated; users and operators briefed; monitoring heightened; accountable owner approves; rollback remains available.

Use a staged release. Limit users, volume, permissions or use cases until real-world evidence supports broader operation.

If the system affected decisions about people, recovery may include correction, notification, appeal or compensation. Technical restoration alone does not repair the harm.

Run a blameless but accountable review

A post-incident review should examine design and governance, not search for one careless person.

Document:

timeline; impact and affected groups; detection path; containment and recovery; technical and organisational causes; which controls worked; which assumptions failed; actions, owners and deadlines; lessons relevant to other systems.

"Blameless" means people can report candidly. It does not mean decisions lack owners. Accountability is the ability to explain who was responsible for the system and what will change.

Test the playbook

Run tabletop exercises before launch and at least annually for important systems.

Useful scenarios include:

an agent sends sensitive information to an external service; a model update changes advice quality without notice; indirect prompt injection triggers an unauthorised tool call; a ranking system produces a sharp disparity; a provider outage removes a critical capability; a synthetic-media feature is abused during a public event; monitoring data itself is incomplete or misleading.

During the exercise, test whether the team can identify the owner, find logs, disable the relevant component, contact the supplier and make a decision under uncertainty.

Measure readiness

Track:

time to detect; time to contain; percentage of systems with tested rollback; percentage with named owners and severity criteria; completeness of logs and version records; overdue corrective actions; repeat incidents; time to notify affected people; exercise findings closed on schedule.

These measures reveal whether incident response is an operating capability or a document.

Treat incidents as portfolio intelligence

One incident often exposes a shared weakness: excessive permissions, inconsistent logging, poor provider controls or weak human review. Feed lessons into procurement, architecture standards, impact assessment and training across the AI portfolio.

The objective is not a world with no AI failures. That is not credible. The objective is an organisation that anticipates failure, limits the blast radius, protects people, learns quickly and makes recurrence less likely.

Use our Responsible AI hub for further governance guidance, the AI impact assessment guide for pre-deployment analysis and the prompt injection defence guide for engineering controls.

Sources and further reading