For most of the chatbot era, the central security question sounded familiar: What information should we trust an AI system with? That question is already obsolete. The more urgent question is what authority we are willing to give it.
An ordinary chatbot returns words. An agent can open a document, browse a website, call an API, query a database, send a message, modify a record, write code or execute a command. Connect enough tools and the model stops being a passive information appliance. It becomes an operator inside the system.
That changes the security boundary. If an AI agent can authenticate, retrieve private data or produce effects outside its own conversation window, it should be treated as an endpoint: separately identified, minimally authorized, continuously observed and immediately revocable. Giving an agent your own credentials because it is acting “for you” is the digital equivalent of hiring an assistant and issuing them your passport, master key, checkbook and signature stamp on the first morning.
The warning is no longer confined to laboratory demonstrations. Spain’s data-protection authority has disclosed a breach report in which an AI agent allegedly identified vulnerabilities, accessed a system, altered personal data and viewed billing records with minimal human intervention. The investigation remains open, so the incident should not yet be inflated into a final verdict about the first fully autonomous cyberattack. The narrower conclusion is already serious enough: autonomous software can now accelerate the familiar chain from discovery to access to action. The underlying weaknesses are old. The tempo is new.
The model is only one component in the chain.
“The model was not compromised” is an inadequate defense. An agent does not need to be corrupted at the model level to become dangerous. It can receive malicious instructions directly from a user or indirectly through the material it reads. A poisoned email, web page, document, issue ticket, source-code comment or image can contain instructions aimed not at the human reader but at the model processing it. OWASP identifies this as indirect prompt injection: the agent mistakes hostile data for an instruction and carries it into the rest of the workflow.
The real damage depends on the tools and permissions waiting downstream. An agent that can summarize an untrusted page may produce a bad summary. An agent that can summarize the same page while holding access to email, cloud storage, source control and payment systems may be maneuvered into disclosing data or taking action. Intelligence does not create the blast radius. Authority does.
Every agent needs its own identity.
Traditional endpoint security developed around machines and human users. Each received an identity. Access was restricted according to role. Actions were logged. Credentials could be revoked. Networks were segmented because no device was assumed to remain trustworthy forever. Agent security demands the same discipline, with one additional complication: agents ingest natural-language material from outside the security perimeter and then decide how to use tools.
The first rule is therefore simple: every agent needs its own identity. It should not silently inherit a human operator’s full session, share a general administrator account or use one permanent API key across unrelated jobs. A distinct identity makes permission boundaries possible and creates an audit trail that can answer two different questions: which person authorized the work, and which agent performed the action? Google’s agent-identity guidance describes this separation explicitly, while NIST has warned that current identity and authorization practices create substantial security challenges for agentic systems.
Identity without restraint is merely a name on a loaded gun. The second rule is least privilege, followed by what Microsoft calls “least agency.” Least privilege limits which resources an agent can reach. Least agency limits which effects it can produce, how independently it may produce them and how long that authority lasts. An editorial research agent may be allowed to read public sources and create a draft, but not publish it. A maintenance agent may inspect a repository and propose a patch, but not deploy to production. A financial agent may categorize transactions, but not initiate one. A home-automation agent may report that a door is unlocked, but not unlock it for an unidentified caller.
Separate proposing from executing.
The difference between proposing and executing is one of the most valuable safety barriers available. High-consequence actions should cross a human approval boundary: deletion, publication, credential changes, money movement, external messaging, production deployment, physical access and control of machinery. Approval must describe the actual effect in plain language. A vague “continue?” button is not informed control when the hidden operation is about to delete a directory or send private material to a third party.
The third rule is to distrust whatever the agent reads. Prompt injection cannot be solved by telling the model to ignore prompt injection. Retrieved material must be treated as data, not authority. Tools should validate inputs independently. Destinations should be allowlisted where practical. Sensitive values should not be exposed to the model unless they are required for the immediate task. Commands, URLs, filenames and recipient addresses derived from untrusted content deserve deterministic checks outside the model. The agent may reason about what ought to happen; hardened code should still decide what is permitted to happen.
Logs turn strange behavior into evidence.
The fourth rule is complete observability. Every consequential agent action should produce a durable record containing the agent identity, authorizing user, requested objective, tool invoked, resource touched, result and time. Logs should live somewhere the agent cannot casually rewrite. Without that record, a failure becomes folklore: the system did something strange, nobody knows exactly why, and the only witness is the same probabilistic mechanism under investigation.
Logs also make behavioral limits possible. A research agent that suddenly requests a credential store, reads thousands of unrelated files or contacts an unfamiliar domain has departed from its assigned role even if each individual call appears syntactically valid. Rate limits, resource boundaries and anomaly detection can stop a chain of small authorized actions from becoming one large unauthorized outcome.
The fifth rule is revocation. Agents need a kill switch that does more than close a browser window. It must terminate active sessions, revoke tokens, disable scheduled work, block tool access and preserve the evidence needed to reconstruct what happened. Short-lived credentials reduce the amount of forgotten authority left behind. For higher-risk systems, automatic revocation should follow inactivity, scope violations, abnormal transfer volume or loss of the human controller.
Practical architecture for ordinary people.
None of this requires rejecting agents. It requires refusing to confuse convenience with trust. A capable agent operating inside a narrow sandbox can be more useful than a supposedly “safe” agent holding broad credentials. Security should not depend on the model remaining obedient under every possible input. It should depend on the surrounding system making disobedience containable.
For an individual, small shop or home laboratory, the practical architecture can remain simple. Put experimental agents in a virtual machine or separate device. Create dedicated accounts instead of sharing personal ones. Mount working directories rather than entire drives. Keep sensitive archives offline or read-only. Require confirmation before external communication, deletion, purchasing or deployment. Back up important data independently. Record tool activity. Make credential revocation fast enough that it can be done while tired, sick or under pressure.
The same principles apply to a publication. A research agent may collect leads. A drafting agent may assemble prose. A verification process should inspect sources and claims. A human editor should approve the final object. Publishing credentials should remain outside the research environment, and the system that creates an article should not automatically possess the authority to make it public. Separation of duties is not bureaucratic ornament. It prevents one contaminated input from controlling the entire chain.
Agent developers often speak of autonomy as if more were always better. It is not. Autonomy is delegated power, and delegated power needs jurisdiction, duration, records and consequence. We learned that lesson for employees, service accounts, networked machines and industrial controllers. AI does not repeal it.
The agent is now an endpoint. Secure its identity. Limit its reach. Separate recommendation from execution. Record what it does. Keep the power to shut it down outside its hands.
Anything less is not automation. It is unattended authority.
Sources and further reading
- Spanish data watchdog publicises first AI agent-linked data breach report — Reuters
- AI Agent Security Cheat Sheet — OWASP
- LLM Prompt Injection Prevention Cheat Sheet — OWASP
- Top 10 for Agentic Applications for 2026 — OWASP GenAI Security Project
- Why Agentic AI Needs a Strong Identity and Authorization Foundation — NIST
- Agentic AI security guidance — Microsoft Learn
- How Google secures AI agents — Google Cloud

