AI Impact Assessments: A Practical Guide to Identifying Risk Before Deployment
How to run an AI impact assessment that identifies affected people, tests evidence, assigns controls and informs a real deployment decision.
By Elena Marković, Women in AI Editorial Fellow · 26 August 2026
An AI impact assessment should answer a decision, not complete a template.
The decision may be whether to proceed, restrict the use case, change the design, add human review, collect better evidence or stop deployment. If the assessment cannot influence any of those outcomes, it is documentation after the fact.
The most useful assessments begin before procurement or development is locked in, involve people beyond the project team and return to the system after launch. They translate broad concerns such as fairness, privacy and safety into specific claims, tests, owners and controls.
What an AI impact assessment is
An AI impact assessment is a structured examination of how an AI system could affect individuals, groups, organisations and society. It records the intended use, context, affected people, potential benefits, possible harms, evidence, mitigations and residual risk.
It is related to, but not identical with:
a data protection impact assessment; a security threat model; a model validation report; a fundamental rights impact assessment; a safety case; a procurement due-diligence review.
A mature organisation connects these processes so teams do not answer the same question in separate documents while leaving gaps between them.
NIST's AI Risk Management Framework organises risk activity around four functions: govern, map, measure and manage. An impact assessment is strongest when it covers all four. It clarifies authority, maps the context, measures relevant risk and records how risk will be managed.
When to conduct one
Begin during discovery, before a supplier is selected or a model is trained. Update the assessment when important facts change.
The Government of Canada's Algorithmic Impact Assessment guidance recommends completing an assessment early in design, validating it before production and reviewing it when system functionality or scope changes. That lifecycle approach is more useful than a one-time approval.
Trigger a review when:
the purpose or user population changes; a new data source is added; the model or provider changes; the system gains access to tools or sensitive records; automation replaces a human decision; monitoring reveals a new failure pattern; a legal or regulatory obligation changes.
Step 1: define the decision and boundaries
Write the intended use in operational terms.
Weak: "Use AI to improve recruitment."
Stronger: "Rank external applicants for UK customer-support roles and present the top 20% to a recruiter, without automatically rejecting any applicant."
The stronger statement identifies the population, action and role of the human. It makes meaningful questions possible.
Record:
the owner and accountable executive; the users and people subject to outputs; the decision the system informs or makes; where and when it operates; the data it receives; the systems and suppliers it depends on; what is explicitly outside scope.
Scope boundaries are risk controls. If the system is not approved for disciplinary, medical or credit decisions, say so and enforce it technically where possible.
Step 2: identify affected people
The project team is not the complete stakeholder map.
Consider:
direct users; people evaluated or represented in data; people affected by a decision; workers whose responsibilities change; customers who cannot use the digital channel; communities exposed to aggregated or indirect effects; operators who must intervene when the system fails.
Include groups who may experience the system differently because of disability, language, age, gender, race, income, immigration status or another relevant characteristic. Intersectional effects matter because aggregate performance can hide concentrated harm.
Consultation should occur while the design can still change. Record disagreements rather than forcing artificial consensus.
Step 3: describe benefits as testable claims
Benefits deserve the same scrutiny as risks.
"Improve efficiency" is not measurable. Specify the expected outcome, baseline, beneficiary and timeframe.
For example:
reduce average handling time by 15% without lowering resolution quality; increase access to a service for people using assistive technology; reduce inconsistent decisions while preserving a route to human review.
This prevents a project from accepting concrete risks in exchange for a vague promise.
Step 4: map harms and failure modes
Consider harm across several dimensions:
rights and fairness: discrimination, exclusion, loss of recourse; privacy: unnecessary collection, inference, leakage or repurposing; safety: physical, psychological or financial injury; security: manipulation, unauthorised access or data exfiltration; reliability: false outputs, drift, outages or inconsistent behaviour; economic impact: lost work, denied opportunity or concentrated cost; agency: people unable to understand, challenge or avoid a decision; organisational impact: legal exposure, operational dependence or reputational damage.
Describe the causal path. "Bias risk" is too broad. "Historical promotion data may encode earlier unequal access to management assignments, causing the model to rank women with equivalent performance lower" is a testable risk statement.
Step 5: assess severity, likelihood and reversibility
A simple score can support prioritisation, but it should not replace judgement.
For each risk, record:
who is affected; how severe the outcome could be; how likely it is in the intended context; how many people may be exposed; how long the effect lasts; whether it can be reversed; how quickly it would be detected; the quality of available evidence.
Canada's assessment tool explicitly considers rights, wellbeing, economic interests, duration and reversibility. These dimensions help teams avoid treating a frequent inconvenience and a rare irreversible harm as equivalent.
Step 6: define evidence and tests
Every important risk should connect to evidence.
Evidence may include:
dataset analysis and provenance; subgroup performance; scenario and adversarial testing; usability research; accessibility testing; security testing; calibration and error analysis; human-factors evaluation; supplier documentation; post-deployment monitoring.
The ICO's AI and data protection risk toolkit offers practical prompts around fairness, lawfulness, transparency, security, data minimisation and individual rights. Where personal data is involved, connect the AI assessment to the formal privacy process.
Do not claim a system is fair because protected attributes were removed. Proxy variables, historical patterns and unequal error costs can remain.
Step 7: select controls
Controls can change the model, the surrounding system or the use case.
Examples include:
improving or restricting training data; setting minimum performance thresholds; routing uncertain cases to a human; preventing automated adverse decisions; limiting available tools and permissions; giving users notice, explanation and recourse; logging decisions and overrides; monitoring by subgroup; narrowing the approved population; pausing deployment.
Name an owner and completion date for each control. State how effectiveness will be tested.
Step 8: make the approval explicit
The final record should support one of four outcomes:
approve with stated conditions; run a limited pilot with heightened monitoring; redesign and reassess; do not proceed.
Record residual risk, dissenting views and who accepted the decision. High-impact systems should not be approved solely by the delivery team.
The EU AI Act creates specific assessment and risk-management obligations in defined circumstances, including a fundamental rights impact assessment for certain deployers of high-risk systems. Organisations should obtain legal advice on whether those provisions apply. A general impact assessment is useful, but it does not automatically satisfy every legal requirement.
Step 9: monitor the real system
Pre-deployment testing cannot reproduce every condition of use. Define what will be watched after launch:
error and override rates; outcomes for relevant groups; complaints and appeals; unexpected uses; model, prompt or data changes; security events; user dependence or automation bias; benefit measures promised at approval.
Set thresholds that trigger investigation, restriction or shutdown. Monitoring without action criteria is observation, not control.
A one-page decision summary
Executives rarely need the entire working file. Give them a concise decision record:
use case and accountable owner; affected populations; expected benefit and baseline; three most material risks; key evidence and its limits; required controls; residual risk; approval decision; review date and stop conditions.
The supporting assessment can remain detailed and auditable.
The standard is better decisions
A polished template can create false comfort. The quality of an impact assessment depends on whether it discovers something, changes something and assigns responsibility.
The process should make uncertainty visible. It should also make it safe to stop. Teams under delivery pressure often interpret governance as a route to approval. Responsible governance preserves the possibility that the right outcome is a narrower system or no system at all.
For related guidance, explore our Responsible AI hub, AI governance guide and human oversight guide. Our enterprise AI procurement guide applies the same evidence standard before a supplier is chosen.