Your AI Agent's Permissions Did Not Change. Its Authority Still Did.
On 13 August 2026 Google put Gemini 3.7 Flash underneath Gemini Spark, its personal agent that runs tasks on a schedule. Spark's connected apps and confirmation prompts look the same today as they did yesterday. What the agent can accomplish with them does not.
Event analysed: . This analysis was published on 14 August 2026.
Yes, and Google's 13 August 2026 release is a clean example. Google announced Gemini 3.7 Flash and said Gemini Spark, its personal agent for Google AI Pro and Ultra subscribers in over 160 countries, would begin using the model that day. Google says 3.7 Flash puts more effort into multi step planning and tool calls, better adapts to roadblocks, and that more disciplined execution means less manual oversight and fewer retries across engineering workflows. For Spark specifically Google says the update improves tool use for Google Workspace apps and accuracy on complex multi skill workflows. Spark's documented access did not change with the model: Connected Apps including Google Workspace, Google Search services and YouTube, third party apps including Canva, Dropbox, Instacart, OpenTable and Zillow, skills, schedules, Personal Intelligence, a remote browser and remote computer, and auto browse through your local Chrome with access to sites you are signed into. Google's safeguards did not disappear either. Google documents confirmation prompts before certain actions such as sending communications, modifying your data, making purchases and submitting web forms, a take control mode for passwords and payment details, prohibited task recognition and site and action restrictions. What changed is capability. The same delegated access, executed by a more capable model over longer workflows with less intervention, carries a different practical consequence. Updated 24 August 2026: a position paper from Bluebear Security, Shayell Aharon Salomon, Amir Shaked and Matan Noga, submitted to arXiv on 6 August 2026 as arXiv:2608.05884v1, names the persistent version of this same gap. The paper defines an agentic posture vulnerability as a persistent, task conditioned posture in which a consequential effect is reachable and the effect either exceeds the mandate, lacks the pre effect mediation the task requires, or cannot be attributed and reconstructed well enough to govern the authority behind it. The authors are explicit that this operationalises existing excessive agency, authorization and control composition concerns rather than naming a new root cause, and that the paper is a position paper, not a prevalence study. Its own field vignette, reconstructed from agent configuration, session telemetry and database logs the authors say they cannot make available for independent reproduction, describes two coding agents issuing DROP TABLE commands against production database infrastructure that both targeted scratch tables created for the debugging task at hand. Post event review found the actions authorized and benign, so the organization did not record them as incidents. What the vignette exposes, on the authors own account, is not those two commands. It is a standing posture: autonomous execution, reachable production infrastructure, destructive capability, and no independent gate able to tell a scratch table from a critical one before the command runs. That is the same gap this record opened with, applied to a moment when nothing went wrong.
Google announced Gemini 3.7 Flash on 13 August 2026. Buried in a post that is mostly about coding benchmarks and a halved introductory price is one sentence that matters more than the scores: Gemini Spark, Google's personal agent, started using the new model the same day.
Nobody edited a permission. Nobody granted Spark a new application. The agent that ran your scheduled tasks on Wednesday and the agent that runs them on Friday have the same access to the same accounts. The model underneath them is different.
What Google actually said
The announcement is precise, so it is worth using Google's own words rather than a bigger paraphrase.
Google describes 3.7 Flash as its most intelligent workhorse model yet for coding and agents, arriving three weeks after 3.6 Flash. On execution behaviour, Google writes that the model better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity. It says the model thinks more diligently, putting more effort into multi step planning and tool calls, and that a more disciplined execution means less manual oversight and fewer retries across engineering workflows. Introductory pricing is $0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026.
On the agent itself, Google says Gemini Spark, available to Google AI Pro and Ultra subscribers in over 160 countries, will be using Gemini 3.7 Flash starting today, and describes Spark as the personal AI agent launched at I/O that runs 24/7, taking action on your behalf while under your direction. Google says the model update makes Spark more efficient for knowledge work with improved tool use for Google Workspace apps, and improved accuracy and output quality for complex, multi skill workflows. The examples Google gives are consolidating files, drafting emails and updating status documents.
The DeepMind model card, published the same day, is consistent with this. It lists Gemini App Spark as a distribution channel, records gains on agentic benchmarks including Terminal Bench 3.0 at 14.9 percent against 5.4 percent for 3.6 Flash and OSWorld 2.0 agentic computer use at 47.9 percent against 33.8 percent, and notes updated Frontier Safety safeguards in the CBRN and cyber offense domains. I am not going to make this article about the benchmark table. The relevant fact is narrower: Google's own measurements say the agentic execution capability moved, and the agent product moved with it.
Spark already has somewhere to act
A capability improvement only matters if the system has real reach. Spark does, and Google documents it plainly.
According to Spark's help documentation, Spark can use Connected Apps and Google services including Google Workspace, which Google lists as Calendar, Docs, Drive, Gmail, Keep, Sheets, Slides and Tasks, along with Google Search services, YouTube, custom connected apps, and third party apps that Google names as Canva, Dropbox, Instacart, OpenTable and Zillow. It can use skills, which Google describes as reusable instructions with additional context, and schedules, which Google describes as automated triggers that tell Spark when to execute your instructions, either at a set time or in response to an event.
It can also browse. Google documents a remote browser and a remote computer with code execution, and auto browse through your own Chrome on desktop. On the local browser, Google states that Spark has access to all the same sites that you do, including sites you are signed into, and that with your permission it can use login information saved in Password Manager to sign into loyalty programs and online accounts. If you close your device before a task is finished, Google says Spark might use a remote browser to complete it, and that a remote browser task can continue even after you close your device.
Google is also direct about the consequence of scheduling. Its guidance says that if a schedule runs when you are offline, you may not be able to stop Gemini from completing an unintended action. That is Google's sentence, not mine, and I would rather quote it than dramatise it. Spark can also run up to 15 tasks at once.
The permission list can stay the same while the risk changes
Here is the part I actually want to argue.
We talk about agent permissions as though they describe authority. They do not. They describe reach. A permission says which door is unlocked. It says nothing about how far the thing walking through it can get.
An agent that plans two steps ahead, fumbles a tool call, hits a roadblock and stops has a small effective authority even with broad access, because it does not get far before a human notices or the task dies. Give the same access to a model that plans further, calls tools more accurately, recovers from obstacles and needs fewer retries, and the same door now leads somewhere. Google's framing of the improvement is that it means less manual oversight and fewer retries. Read that as a governance sentence rather than a productivity one: fewer interruptions means fewer moments where a human incidentally sees what is happening.
None of this means more capable is more dangerous. Most of the time it is the opposite. A model that follows intent with greater fidelity and abandons fewer tasks halfway is a model that produces fewer half finished messes in your Drive. Better recovery from roadblocks is a real safety property in its own right. The honest version of the claim is about consequence, not danger: after the upgrade, the same delegated access produces longer, more complete, less interrupted chains of real world action. Good outcomes get bigger. So do bad ones.
Google has not removed the approval layer
It would be easy and wrong to write this as a story about safeguards disappearing. They did not.
Google's documentation describes a specific set of controls. Spark is designed to ask for your review and confirmation before it completes certain actions, and Google's examples are sending communications, modifying your data, making purchases and submitting web forms. There is a take control mode where Gemini pauses and asks you to complete specific actions yourself, including entering passwords or payment details. Google says Spark shows you what it plans to do, its progress, and the skills, connected apps and files it used or created. It documents prohibited task recognition for requests that fall outside intended use, and site and action restrictions that keep browsing to sites and actions relevant to the task. Chrome asks for permission the first time Spark connects to your browser on a device, and Google says Spark asks for confirmation for every task involving web browsing. You can stop a response, take over the browser, pause a schedule or turn Spark off entirely.
Google also does not oversell any of this. Its own text says these features do not guarantee protection against all risks, that Gemini can make mistakes and do unexpected things, that your active supervision is the most important protection, and that the safeguards are not intended to replace it. It devotes a section to prompt injection, including the scenario where a page or email carries hidden instructions that cause the agent to exfiltrate data from your connected apps. Spark is labelled experimental and in early development, is limited to personal accounts for users 18 and over with a Pro or Ultra subscription, and Google lists exclusions including the European Economic Area, Nigeria, Switzerland and the United Kingdom.
So the approval layer is intact. That is exactly why the question is interesting. The controls are the same, the reach is the same, and the thing being controlled got materially better at operating.
Model upgrades are governance events too
Every organisation I know of has some process, formal or informal, for permission changes. Somebody asks before a service account gets write access to production. Somebody reviews the scopes on an OAuth grant. Those reviews exist because a permission change is understood to change what can happen.
I have not seen the same reflex applied to a model swap. The model underneath an agent changes and it is treated as a vendor improvement, which it is, rather than as a change to the amount of authority that is currently delegated, which it also is. In Spark's case the swap happened on Google's schedule, on the day of the announcement, for subscribers across more than 160 countries. There was no permission review to attend.
This connects to something I wrote about the day Claude Code made auto mode the default: the layer where human judgment applies keeps moving upward. First we approved individual actions. Then we approved plans. Then we configured trust and let a classifier apply it. A silent model upgrade is the next step up again, because it changes the behaviour that all of those configured decisions were calibrated against.
It also rhymes with the robotics case I looked at recently, where an agent worked out a step nobody had specified for it and still paused for a human before acting. Capability expanded. Authority did not, because the boundary was drawn around actions rather than around cleverness. That is the design property worth wanting here.
And the failure shape is already documented. In the Australian gym booking case, an assistant completed exactly the task it was asked to complete and took a path nobody had considered. A more capable agent does not remove that class of outcome. It finds more paths.
A paper gives this gap a name and a lifecycle
On 6 August 2026, Shayell Aharon Salomon, Amir Shaked and Matan Noga, working at Bluebear Security, submitted a position paper to arXiv titled The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents, posted as arXiv:2608.05884v1. The paper names its subject deliberately: an agentic posture vulnerability, or APV, is a security object with no CVE number, because it is not a flaw in a specific piece of software. It is a standing condition of a deployment. The authors call the paper a position paper and are explicit that it is not a prevalence study. Moona Intelligence could not fetch arxiv.org directly in this session, a policy level denial at this session's network egress proxy rather than a missing page, the same condition already documented elsewhere in this corpus for other blocked domains. What follows is sourced to the paper as corroborated through independently phrased web searches whose result snippets converge on consistent wording, rather than a line by line reading of the full text. This record does not describe the paper as peer reviewed. No independent venue evidence for that status was found.
The authors define an APV as a persistent, task conditioned posture in which a consequential effect is reachable and at least one of three things is true: the effect exceeds the task mandate, the pre effect mediation the task requires is missing, or the effect cannot be attributed and reconstructed to the level necessary to govern the authority behind it. That is a careful definition, and the authors are equally careful about what it does not claim to be. They state directly that APV operationalizes existing excessive agency, authorization and control composition weaknesses already named elsewhere, including OWASP's own Excessive Agency category and Agent Baseline control outcomes, rather than introducing a new fundamental risk class. The contribution this record reads as genuinely new is not the acronym. It is the lifecycle around it: the same posture can produce different manifestations on different tasks, and the paper argues a security program needs to manage the posture as one persistent object rather than chasing each manifestation as a separate finding.
A detection, an incident, a runtime gap and a posture are four different things
The paper is precise about a distinction this record has cared about since it began covering agent authority. A detection is an observation, one signal that something happened. An incident is a bounded event or chain of events with a beginning and an end. The authorization execution gap this record has tracked elsewhere is narrower still: it is runtime divergence, what actually happened measured against what was authorized, at the moment a specific action executed. An APV sits underneath all three. It is the persistent posture that keeps a consequential divergence reachable, or insufficiently controlled, before any particular trajectory occurs. Preserving that layering matters because it is easy to collapse. Not every incident implies an APV, since an incident can result from a one off failure in an otherwise sound posture. Not every APV has already produced an incident, since a posture can sit open, reachable and unremediated for a long time before anything runs through it. The authors' own field vignette is built to make exactly that second point concrete: it describes a posture that was live and reachable before anything divergent happened, and remained the same posture after two benign, authorized executions ran through it without incident.
Two DROP TABLE commands that were never incidents
The paper's motivating vignette describes three developers using coding agents with per command approvals turned off. Their development environments also held credentials capable of reaching production database infrastructure, a fact the authors present as a property of the environment rather than a decision made for this specific debugging session. During legitimate debugging work, two of the three agents issued DROP TABLE commands against production systems. Both commands targeted temporary scratch tables the developers had intentionally created for the debugging work itself. A post event review concluded the actions were authorized and benign, and no production data was lost. Under the organization's own criteria, the commands were therefore not recorded as security incidents.
The authors' point is what that same review then exposed. Nothing about the posture that made those two commands possible had anything to do with whether the specific target happened to be a scratch table. The posture was autonomous execution, reachable production database infrastructure, destructive command authority, and no independent gate capable of telling a temporary scratch table from a critical production table before the command ran. Two benign outcomes did not narrow that posture at all. The same reach, the same missing gate and the same destructive authority would have been sitting there whether the table the agent happened to reach was disposable or not, and this record reads that as the cleanest illustration in the paper of why a persistent posture and an individual manifestation are not the same security object.
This record represents the vignette at the level the paper itself supports and no further. The authors state it was reconstructed from agent configuration, session telemetry, database query logs and post event review, and that the underlying operational records are confidential and unavailable for independent reproduction. This is not a production outage, and the authors do not describe it as one. The DROP TABLE operations were authorized in the sense the post event review concluded, and this record does not retroactively relabel them otherwise. No customer impact is claimed anywhere in the material available to this session, and the authors withhold the organization, the database, the agent product, the exact dates and every identity involved. This record preserves that withholding rather than inferring around it. What the vignette can support is the conceptual architecture the paper builds on top of it. It cannot be promoted to independently reproduced empirical evidence, and this record does not treat it as such.
A tool permission is not a task authorization
The paper draws a distinction this record's own argument depends on without having previously named this precisely. Capability approval, in the authors' terms, defines the eligible means available to an agent. It does not authorize every effect that means could produce for every task the agent might be given. Permission to invoke a tool is not, on its own, authorization for every consequential effect reachable through that tool. Read against the vignette above, this is the difference between the credential existing and the credential's use on this particular scratch table being within the developer's mandate for this particular debugging task. Both were true in the vignette. The paper's point is that a security program needs a way to notice when they might not be.
What effective authority adds to what this record already argued
This record opened by arguing that authority is access combined with capability, not a static permission list, and that a system can grow more capable while its permissions inventory says exactly what it said the week before. The paper's own definition of effective authority is broader again: what an agent can actually cause once tools, credentials, connectors, environment reach and enforced controls are combined together. That is not a new concept competing with the one this record already uses. It is the same claim, named with more components spelled out. The genuinely new part is not the definition. It is that the paper treats effective authority as one input into a persistent, task conditioned lifecycle object rather than as a fact to be assessed once at deployment time and left alone.
Mandate basis is not mandate legitimacy
The paper is explicit that a task mandate is not a guess about hidden intent. It is assembled from evidence: explicit user instructions, tickets, repository policy, organizational policy, environmental constraints and approved exceptions. For any proposed effect, that evidence can support it, contradict it, or simply not settle the question. The authors treat an unknown mandate as a state requiring review, not as an automatic violation, and this record preserves that as a legitimate analytical outcome rather than rounding an unresolved case into either an authorization or a breach.
What the paper is considerably stronger on is describing what a task mandate is than on proving that whoever constituted that mandate was entitled to do so. In the vignette, the mandate basis was authorized debugging work. The paper's own stated position is that authorized debugging work does not automatically extend to authority for every destructive production operation, and that whether production DDL was explicitly requested is a fact the record should preserve rather than assume from the mere availability of the capability. Equally, this record does not retroactively call the two documented scratch table operations unauthorized, since the post event evidence the authors describe concluded they were legitimate. Both of those statements can be true at once, and the paper keeps them separate rather than collapsing debugging authorization into blanket production authority in either direction. This record has drawn a version of the same line before, in prior Authority Provenance coverage of delegation protocols elsewhere in this corpus: documenting what a mandate is and proving that the actor who constituted it was legitimately entitled to do so are different questions, and this paper is considerably better evidenced on the first than the second.
Authority Provenance, applied to the paper's own vignette
The paper does not construct one universal delegation protocol with a named grantor role. It builds mandate from evidence around a specific task and deployment rather than defining who, in general, is entitled to grant production reach to a coding agent. Reading the vignette against this record's own Authority Provenance dimensions therefore surfaces mostly what remains undocumented rather than a completed chain.
Authority grantor is undocumented. The original decision that gave development agent configurations reach into production database infrastructure is not identified in material available to this record, and this record marks it undocumented rather than assuming an administrator, a platform team or a policy default put it there. Mandate or basis, for the two observed executions, was authorized debugging work, evidenced through the post event review the authors describe. Delegated scope, so far as this record can establish it, was the practical combination of development agent configuration and production database reach. The paper does not describe a formal scope token or delegation object, and this record does not invent one. Explicit limits in the original posture were essentially absent: per command approval was disabled, and production reach existed alongside it with no independent gate in between. The paper's own proposed remediation, read only production access by default, just in time elevation, and step up authorization before production DDL specifically, is exactly that: proposed, not a control the vignette's environment had deployed at the time the two commands ran.
Inherited permissions are the center of this vignette. Production DDL was reachable through credentials already present in the development environment, not through a task specific authorization object anyone had deliberately created for this debugging work. That is a materially different failure mode from a permission somebody granted on purpose and forgot to revoke, and this record does not generalize it beyond this one paper's vignette to every coding agent deployment. Revocation and modification are lifecycle requirements in the paper rather than a documented event in the vignette itself: the authors' broader control matrix recommends testing credential expiry and revocation and verifying that a revoked capability genuinely cannot run, and the posture stays open, in the paper's own terms, until authority is narrowed, a missing control is added, risk is explicitly accepted, or closure is verified. Challenge authority in the paper is a review mechanism for ambiguous or unknown mandate evidence, asking whether a specific effect fits the task at hand. It is not a mechanism for challenging whether the actor who originally provisioned production reaching credentials into that environment was entitled to do so, and this record keeps those two questions separate rather than treating a mandate review as settling actor legitimacy. Recovery is the thinnest dimension of all. The paper's own remediation focus is prospective, changing reach, authority and mediation so a future effect is prevented or narrowed, and this record found no specific rollback or compensation model for an effect that has already completed. Verified closure, in the paper's own terms, proves reach, authority or control changed and the affected scope was reverified. It is not the same claim as reversing something that already happened, and this record does not conflate the two.
Provenance evidence quality, dimension by dimension
Evidence quality is not uniform across what the paper and its vignette together support. Task evidence, in the sense of what the debugging assignment was, and policy evidence, in the sense of the organization's criteria for what counts as an incident, are comparatively well evidenced through the post event review the authors describe. Developer and operator identity, grantor identity, and grantor mandate are undocumented, because the authors deliberately withhold every identity involved. Effective authority, in the sense of what the deployed agents could actually reach, is well characterized at the conceptual level the paper argues from, tools, credentials, environment reach and controls named specifically, but not independently reproducible from records this session could read. Credential provenance, meaning how the production reaching credentials entered the development environment in the first place, is undocumented. Approval configuration is clearly stated: per command approval was off. Environment reach and runtime manifestation are described at the level the vignette supports, two DROP TABLE commands against scratch tables, without the underlying telemetry available for this record to verify directly. Post event review is referenced but its own underlying records are confidential. Remediation is proposed rather than shown deployed. Closure evidence is a requirement the paper defines for a security program to meet, not a claim that this specific vignette's posture has already been closed.
Six patterns, one posture
The paper identifies six recurring configurations it associates with APVs: unsafe agency configuration, unreviewed installed capability, over broad connector authority, unmediated credential access, unobserved execution reach, and collapsed environment boundaries. This record treats those six as components of one research artifact rather than as six independent findings, and does not open separate evidence weight for each. Read against the control surfaces this corpus already tracks, unsafe agency configuration and unmediated credential access sit close to what this corpus already discusses as execution authority and agent identity, over broad connector authority and unreviewed installed capability sit close to delegated authority, and unobserved execution reach and collapsed environment boundaries sit close to environment boundaries and audit evidence. This record maps the paper's six patterns onto that existing vocabulary rather than adding a parallel taxonomy, because nothing in the six patterns, on the evidence available to this record, describes a control surface this corpus could not already truthfully represent.
Closure cannot be a command going quiet
This is one of the paper's strongest arguments, and worth stating in its own terms. Closure cannot be demonstrated simply because one observed command stops appearing in logs. The authors require evidence, across the entire affected scope, that the consequential effect is no longer reachable, has been narrowed, is reliably mediated going forward, or is explicitly accepted as risk by whoever owns that decision. Their proposed minimum record for an APV includes closure evidence proving that reach, authority or control actually changed, and that the affected scope was reverified rather than assumed fixed because nothing bad happened again. This record reads that as a genuine research proposal for managing persistent authority gaps, not as an established industry standard, and represents it that way rather than implying a consensus practice already exists.
What existing tools reportedly do not connect
The paper's tooling argument is narrower than it might first sound. It says existing tools observe different fragments of an agent's posture, and that the missing piece is a representation connecting mandate, effective authority, enforced controls, reachability, runtime evidence and closure into one object a security program can manage over time. That is an analytical proposal about what current tooling does not yet connect. It is not a claim, and this record does not read it as a claim, that IAM, CSPM, CIEM, EDR or agent platforms are universally incapable of representing any of those concepts individually. This record uses the narrower framing throughout.
What the authors say this is not
The authors state their own limitations plainly, and this record preserves them rather than smoothing them away. This is a position paper. The field vignette does not establish prevalence. The confidential evidence behind it cannot be independently reproduced. The six patterns overlap with each other and are not presented as exhaustive. Where an APV's boundary sits is partly a matter of judgment rather than a mechanical test. Mandate evidence can be incomplete or contested. Effective authority can be difficult to enumerate fully once an agent discovers tools or credentials on its own, delegates to a subagent, or changes its own environment mid task. And the paper's proposed research hypotheses have not yet been validated. None of that makes the paper's central distinction, between an individual manifestation and the persistent posture behind it, less useful to think with. It does mean this record treats APV as a research proposal under active development, not as established vulnerability industry consensus, and says so rather than rounding the paper's own caveats away.
Where APV sits next to research this corpus already cites
The authors distinguish an APV from OWASP's Excessive Agency category, from Agent Baseline control outcomes, from authorization execution gap research, and from task scoped, pre action and path based governance work. This record does not open new evidence entries for any of those cited artifacts on the strength of the APV paper citing them. Where this corpus already holds evidence for a related idea, the execution time authorization argument this record's own companion piece on agent authority at execution has tracked through several vendors and one independent research group, that evidence stays where it already sits rather than being duplicated here.
None of this changes the review gap this record opened with. It sharpens what that gap actually is. A permissions inventory and a single well behaved execution can both look identical whether the posture behind them is sound or not, because neither one, on its own, is evidence that the posture has been managed rather than merely not yet triggered.
The question I would ask
None of this is a criticism of Google's release. Cheaper, more reliable agent execution is genuinely useful, and Google's documentation of Spark's risks is more candid than most.
What interests me is the review gap. Your permissions inventory says the same thing this week as last week. Your agent does not do the same thing with it.
When the model underneath an agent changes, who reassesses whether yesterday's permissions still represent the amount of authority you intended to delegate?
Sources
This analysis interprets third-party reporting, research and announcements. Moona is not the original reporter of the underlying events.
