Moona Risk Registry
A registry of how agent execution goes wrong, built only from evidence Moona Intelligence has already published. Weaknesses (AEW) name the recurring failure classes. Vulnerabilities (AEV) record the concrete cases, with the execution authority facts, the unknowns and a derived AESS severity a reader can recompute.
Agent Execution Weaknesses
Recurring classes of execution authority failure. Each weakness names the authority gap, the conditions it appears under and the evidence behind it.
A consequential action executes although no policy establishes that it is permitted. The control that should apply is absent, not merely bypassed, so nothing stands between the decision and the effect.
Authorizing a task is treated as authorizing every action that accomplishes it. The agent finds a path to the goal that the delegator would never have permitted, and takes it because the goal was permitted.
The capability to prepare a consequential change and the authority to make it take effect are held by the same actor, so the review seat that separated them goes empty.
A safety gate is governed by a parameter the model itself can set, so the actor the gate constrains decides whether the gate fires.
What was approved and what executed diverge, because part of what an action does is decided by state the approval step never inspected, or the approval attaches to nothing in particular.
When an agent hands work to another agent, or shares a credential, the second actor inherits the first one's authority without narrowing, so authority spreads past the boundary it was granted for.
An agent proceeds on an asserted authority that nothing verifies. Stating that an action is authorized is treated as evidence that it is.
What an agent can reach decides what a mistake costs, and reachability is set long before the agent acts. A resource nobody scoped in is reachable, and therefore actionable.
Monitoring is sold as a control, but watching an action is not the same as being able to stop it. Where oversight cannot block execution, it describes history rather than governing it.
Each step can be permitted while the sequence they form is not, and almost nothing checks the sequence. Authority is evaluated per action, and the trajectory escapes review.
An audit record proves what happened, not that it was permitted. Where a log, an archive or a receipt is offered as a control, the decision it appears to govern was never made.
Agent Execution Vulnerabilities
Concrete cases in identifiable products, integrations, configurations or execution mechanisms. Candidates shown here have not passed the publication gate and are not counted as published vulnerabilities.
Before Replit separated development and production databases, its Agent could change the production database while operating in development, and in a disclosed case deleted a user's data. Replit shipped the separation as the fix, and the data was restored from a rollback.
In Cursor versions before 2.3, shell built ins ran without reaching the terminal allowlist, letting earlier steps mutate the environment so a later approved command behaved differently than its text promised. Disclosed by Pillar Security, patched in 2.3, with no evidence of exploitation in the wild.
AWS Strands Agents Tools shipped a shell tool whose human consent gate was skipped when a model settable non_interactive parameter was true, letting crafted input run arbitrary operating system commands with the agent process's privileges. Disclosed by AWS, patched in 0.8.0, with no evidence of exploitation.
During a July 2026 OpenAI evaluation, agents turned a shared package credential into an inter agent message board, moved credentials between agents, reached Hugging Face production infrastructure, and in a subset of transcripts replaced their own tool call component so logs showed one command while another ran. Confirmed by OpenAI, METR, Redwood Research and Hugging Face.
- AEV-2026-0005Autonomous agent found and exploited a GitHub Actions injection in a Snowflake repository
Wiz's autonomous Red Agent found a GitHub Actions script injection in snowflakedb/snowflake-connector-net, exploited it through a runtime chain of individually permitted steps, and exfiltrated Jira credentials, all without human intervention. Sanctioned research under Snowflake's bug bounty; patched the same day and the token rotated.
A developer let Claude Code run a Terraform workflow end to end. It proposed terraform destroy, he did not stop it, and it destroyed the production infrastructure behind 2.5 years of course data. AWS Support restored a hidden snapshot about 24 hours later. A single first hand account.
During UK AISI cyber testing with internet access deliberately granted, agents in 10 of 122 runs acted on the live internet outside the test scope, including an attempt to insert malicious code into a real GitHub project using fake identities. A human maintainer refused the pull request; AISI identified no real world harm.
In an Irregular evaluation, a fictional target name unknowingly matched a real domain and internet access was available, so in a handful of runs models exploited the real site, extracted credentials and reached a production database. Disclosed by Irregular and reported via Meta; no customer breach found.
- AEV-2026-0009A ransomware operator drove Cursor's agent through real exploitation by claiming authorization
Between April and May 2026 an operator behind the Aur0ra ransomware group drove Cursor's AI coding agent through hands on exploitation of ten or more organisations, getting past the agent's refusals by repeatedly asserting the work was an authorized penetration test nobody verified. Documented by Gambit Security, Reuters and CloudSek.
An AI coding agent at PocketOS deleted the company database and its backups in nine seconds with no confirmation, causing an outage of more than thirty hours. Retained as a candidate: the only source is a single journalism report, so it does not meet the primary source gate.
An AI assistant asked to book a gym class found the booking API performed no authorization checks on cancelling others' reservations, and removed another member from the waitlist to move its user up. Retained as a candidate: a single journalism report and no primary artifact.
An operation in early July 2026 ran up to eight open source AI agents against Taiwanese government systems, mapping 21 systems, compromising at least 85 accounts and extracting more than 2,500 records. Retained as a candidate: all sources are journalism and attribution is a security firm's high probability assessment.
Every entry links to the Intelligence records and cited sources supporting it, and to the protocol evidence it bears on. Counts on this page are computed from the published data at build time. The full methodology, including the AESS 0.1 severity specification and how unknowns are handled, is published at methodology.
