Skip to content
AIRAS Cloud

Resource

AI risk assessment: a practical method

Most AI risk assessments fail for the same reason: they are questionnaires, not decisions. This guide sets out the method that produces an outcome an accountable person can sign, and a regulator can follow.

Start with scope, not with a form

An assessment is only meaningful for a defined use. The same model can be low consequence in one context and high consequence in another, so the assessable unit is the AI system in its deployment context: this model, for this purpose, on this data, affecting these people, with this level of autonomy.

Scope drift is the most common source of invalid assessments. If the purpose, population or autonomy changes after approval, the original outcome no longer describes the system in production and the assessment must be reopened.

Capture the context that changes the answer

  • Purpose, business process and the decision the AI influences
  • The organisation's role: provider, deployer, importer or distributor
  • Data categories, lawful basis, special-category and minor involvement
  • Affected populations and whether they can contest an outcome
  • Autonomy: recommend, act with approval, or act without approval
  • Model provenance, version and vendor dependency
  • Failure modes and the realistic severity if they occur
  • Existing controls and the evidence that they operate

Screen for prohibited and restricted practice first

Screening precedes scoring. A practice that is prohibited or restricted cannot be brought into range by strong controls, and a scored band would imply otherwise. Screening runs against the published criteria and returns a hard outcome with the reason recorded.

Where screening returns a possible match, the assessment is stopped and routed for legal review rather than continuing to a numeric result. That routing is itself part of the record.

Score deterministically, not probabilistically

Scoring must be reproducible. AIRAS Cloud evaluates each assessment against a versioned ruleset: the same inputs always produce the same score, the same band and the same explanation trace, and the ruleset version is recorded against the outcome.

Mandatory floors matter more than the arithmetic. Certain contexts — safety-relevant use, vulnerable populations, autonomous action without approval, special-category data — set a minimum band regardless of mitigating answers, so an otherwise favourable questionnaire cannot pull a serious system into a low band.

Produce a readable explanation trace

The output of an assessment is not a number. It is a statement of which criteria applied, which inputs triggered them, which floors were imposed, which obligations follow, and what evidence is still outstanding — written so that a reviewer who was not present can follow the reasoning.

If the reasoning cannot be reconstructed without the original assessor, the assessment is not defensible.

Separate assessment from approval

Independent review is the control that makes an assessment credible. The reviewer must be someone other than the assessor, with authority to approve, approve with conditions, or refuse, and with the ability to send the assessment back for evidence.

Conditions should be captured as owned, dated obligations rather than free text. A condition without an owner and a date is a note, not a control.

Close the loop with monitoring and reassessment

Approval is a point-in-time statement. Governance holds only if the system is monitored for incidents, drift, changed usage and vendor changes, and if material change forces reassessment automatically rather than relying on someone remembering.

Frequently asked questions

What should an AI risk assessment cover?
Purpose and deployment context, the organisation's role, data categories and lawful basis, affected populations, autonomy and human oversight, model provenance, failure modes and consequence severity, existing controls, and the evidence available to support each answer.
Should a language model score AI risk?
No. Scoring must be reproducible and reconstructable months later. AIRAS Cloud uses deterministic, version-controlled rules for scoring and classification, so identical inputs always produce the same band and the same explanation trace. AI may assist with drafting or summarising, but it never holds the decision.
How often should an AI system be reassessed?
On a fixed periodic cycle set by risk band, and immediately on material change: a new purpose, a new population, a model or vendor change, an increase in autonomy, a new data category, or an incident that reveals an unassessed failure mode.
What makes an AI risk assessment defensible to a regulator?
A published ruleset version, a complete intake record, a readable explanation of how the outcome was reached, named owners, an independent reviewer separate from the assessor, an attributable decision, and an append-only audit record with the supporting evidence attached.

See a deterministic assessment run end to end

A short executive briefing covering your AI estate, the criteria that would apply and the evidence a reviewer would expect.

No pricing commitment. No confidential information required.