Deploy AI agents with confidence.
For organisations, individuals and agents.
Seven areas to get right
A practical synthesis of established guidance, not a certification or a claim that the frameworks are equivalent. See the frameworks.
01Governance & AccountabilityEvery deployment has a named owner, a recorded purpose and an approved risk assessment.
Register of deployments, risk acceptance, third parties, legal obligations, notice to people affected.
NIST AI RMF GOVERN 1–202AccessPermissions are limited to what each task requires, under a distinct identity, within enforced boundaries.
Identity and authentication, least privilege, isolation and segmentation.
OWASP: tool security and least privilege03Data & ContextInformation in use is protected, accurate and of known origin, and leaves only by approved routes.
Confidentiality, privacy, data quality and provenance, memory, leakage.
OWASP: memory and context security · data protection04IntegrityInstructions, configuration and outputs are complete, accurate and changed only with authorisation.
Change control for models, prompts and tools; supply chain; resistance to injected instructions.
OWASP: prompt injection defence · output validation05Oversight & MonitoringConsequential actions need human approval, and all actions are logged and reviewed.
Human in the loop, logging, detection, independent assurance.
OWASP: human oversight · monitoring06Resilience & RecoveryFailures are detected, contained and reversed, and operations can continue without automation.
Availability, incident response, shutdown and revocation, rollback, continuity.
NIST AI RMF MANAGE 2.1, 2.3–2.407EffectivenessDeployments deliver measured benefit at a known cost, used by trained people.
Reliability and accuracy, value, cost, training, shared lessons.
NIST AI RMF MAP 3.1–3.4Reviewing a deployment against these seven areas?
Talk it through →Secure your own setup
Practical guides for people running agents themselves.
Learn from what's already gone wrong
Point your agent here. The reports and research below are available as structured data, with source links, analysis and suggested safeguards.
Before you start, read https://agentrisk.com/api/incidents.json and check whether your task resembles anything in it.
Selected reports and research
ChatGPT Atlas browsing agent hijacked via email prompt injection to send resignation letterSecurityDec 2025 Perplexity Comet browser hijacked via Reddit prompt injection to steal user accountsSecurityJul 2025 AI coding agent deletes production database, fabricates 4,000 fake records to cover upAutonomyJul 2025 GitHub Copilot prompt injection achieves remote code execution by enabling auto-approval modeSecurityJun 2025 GitHub MCP server exploited via prompt injection to exfiltrate private repository dataSecurityMay 2025 Air Canada chatbot fabricates bereavement fare policy — company held liable by tribunalGovernanceNov 2022Show all 19 incidents
Hong Kong government bans OpenClaw from government networks, Privacy Commissioner flags agentic AI privacy riskGovernanceMar 2026 SWE-CI benchmark finds regressions during long-term code maintenanceAutonomyMar 2026 AWS Cost Explorer interruption: Amazon cites misconfigured access controlsAutonomyDec 2025 Lab tests reveal AI agents autonomously forge credentials, override antivirus, and exfiltrate dataSecurityMar 2026 Moltbook database exposure permits agent impersonationSecurityJan 2026 HKCERT warns of malware, supply chain risks, and high-severity vulnerability in OpenClaw platformSecurityFeb 2026 Lobstar Wilde transfers its token holdings after losing wallet context in a session resetFinancialFeb 2026 Jailbroken Claude Code instances used for autonomous state-sponsored cyber espionage campaignSecuritySep 2025 GitHub Copilot secrets exfiltrated character-by-character via invisible image proxy side channelSecurityJun 2025 Malicious MCP server exfiltrates entire WhatsApp message history via tool poisoningSecurityApr 2025 Serviceaide database exposure affects patient information; agent causation unestablishedDataSep 2024 AI agent tricked into releasing $47,000 crypto prize pool via social engineeringFinancialNov 2024 ChatGPT plugin ecosystem vulnerabilities enable OAuth hijacking and account takeoverSecurityMar 2024/api/incidents.jsonEvery incident, structured.
/feed.xmlAtom feed of new entries.
/llms.txtIndex for language models and agents.
Agent contributions. Agents will be able to submit incidents they've seen. Other agents will review the evidence before a person decides what gets published.
Sources we build on
AgentRisk maps to established guidance rather than replacing it.