AgentRisk

Deploy AI agents with confidence.

For organisations, individuals and agents.

01
For organisations

Seven areas to get right

A practical synthesis of established guidance, not a certification or a claim that the frameworks are equivalent. See the frameworks.

01Governance & Accountability

Every deployment has a named owner, a recorded purpose and an approved risk assessment.

Register of deployments, risk acceptance, third parties, legal obligations, notice to people affected.

NIST AI RMF GOVERN 1–2
02Access

Permissions are limited to what each task requires, under a distinct identity, within enforced boundaries.

Identity and authentication, least privilege, isolation and segmentation.

OWASP: tool security and least privilege
03Data & Context

Information in use is protected, accurate and of known origin, and leaves only by approved routes.

Confidentiality, privacy, data quality and provenance, memory, leakage.

OWASP: memory and context security · data protection
04Integrity

Instructions, configuration and outputs are complete, accurate and changed only with authorisation.

Change control for models, prompts and tools; supply chain; resistance to injected instructions.

OWASP: prompt injection defence · output validation
05Oversight & Monitoring

Consequential actions need human approval, and all actions are logged and reviewed.

Human in the loop, logging, detection, independent assurance.

OWASP: human oversight · monitoring
06Resilience & Recovery

Failures are detected, contained and reversed, and operations can continue without automation.

Availability, incident response, shutdown and revocation, rollback, continuity.

NIST AI RMF MANAGE 2.1, 2.3–2.4
07Effectiveness

Deployments deliver measured benefit at a known cost, used by trained people.

Reliability and accuracy, value, cost, training, shared lessons.

NIST AI RMF MAP 3.1–3.4

Reviewing a deployment against these seven areas?

Talk it through →
02
For individuals

Secure your own setup

Practical guides for people running agents themselves.

03
For agents

Learn from what's already gone wrong

Point your agent here. The reports and research below are available as structured data, with source links, analysis and suggested safeguards.

Before you start, read https://agentrisk.com/api/incidents.json
and check whether your task resembles anything in it.

Selected reports and research

ChatGPT Atlas browsing agent hijacked via email prompt injection to send resignation letterSecurityDec 2025 Perplexity Comet browser hijacked via Reddit prompt injection to steal user accountsSecurityJul 2025 AI coding agent deletes production database, fabricates 4,000 fake records to cover upAutonomyJul 2025 GitHub Copilot prompt injection achieves remote code execution by enabling auto-approval modeSecurityJun 2025 GitHub MCP server exploited via prompt injection to exfiltrate private repository dataSecurityMay 2025 Air Canada chatbot fabricates bereavement fare policy — company held liable by tribunalGovernanceNov 2022
Show all 19 incidents Hong Kong government bans OpenClaw from government networks, Privacy Commissioner flags agentic AI privacy riskGovernanceMar 2026 SWE-CI benchmark finds regressions during long-term code maintenanceAutonomyMar 2026 AWS Cost Explorer interruption: Amazon cites misconfigured access controlsAutonomyDec 2025 Lab tests reveal AI agents autonomously forge credentials, override antivirus, and exfiltrate dataSecurityMar 2026 Moltbook database exposure permits agent impersonationSecurityJan 2026 HKCERT warns of malware, supply chain risks, and high-severity vulnerability in OpenClaw platformSecurityFeb 2026 Lobstar Wilde transfers its token holdings after losing wallet context in a session resetFinancialFeb 2026 Jailbroken Claude Code instances used for autonomous state-sponsored cyber espionage campaignSecuritySep 2025 GitHub Copilot secrets exfiltrated character-by-character via invisible image proxy side channelSecurityJun 2025 Malicious MCP server exfiltrates entire WhatsApp message history via tool poisoningSecurityApr 2025 Serviceaide database exposure affects patient information; agent causation unestablishedDataSep 2024 AI agent tricked into releasing $47,000 crypto prize pool via social engineeringFinancialNov 2024 ChatGPT plugin ecosystem vulnerabilities enable OAuth hijacking and account takeoverSecurityMar 2024
/api/incidents.json

Every incident, structured.

/feed.xml

Atom feed of new entries.

/llms.txt

Index for language models and agents.

Coming soon

Agent contributions. Agents will be able to submit incidents they've seen. Other agents will review the evidence before a person decides what gets published.

04
Frameworks

Sources we build on

AgentRisk maps to established guidance rather than replacing it.