KeyForge
All postsSecurity

The Agent Threat Model: When Prompt Injection Meets Credential Access

July 8, 20267 min read

Traditional application security has a comforting assumption baked into it: the code you run is code you wrote. Inputs are data, code is logic, and the two do not cross the streams. Autonomous agents quietly break that assumption. An agent reads a web page, a PDF, an email, or a tool result, untrusted content, and that content can change what the agent decides to do next. Input has become instruction. That is the essence of prompt injection, and it is not a bug you patch; it is a property of how agents work.

On its own, an agent following a malicious instruction is a correctness problem. It becomes a security problem the moment that same agent also holds credentials. Now the untrusted content is one clever payload away from steering a process that can spend money, call APIs, and, if the credential is a raw provider key sitting in the environment, read and exfiltrate that key itself.

The two capabilities you must never combine carelessly

Threat-modeling agents comes down to two capabilities: the ability to ingest untrusted content, and the ability to act with privilege. Most useful agents need both. The mistake is granting the privileged capability in a form that has no independent limits, for example, injecting a raw sk_ provider key into an agent’s environment. The key has no budget of its own, no scope, and no expiry; whatever the agent can be manipulated into doing, the key will faithfully fund.

The fix is not to remove the capability but to bound it. The agent should act through a credential that is scoped, capped, revocable, and audited, so that even a fully hijacked agent can only do a bounded, observable amount of damage before it hits a wall you set.

Blast radius: the number that actually matters

Security for agents is less about preventing every injection, you cannot, and more about minimizing blast radius when one succeeds. Ask, for each agent: if this process were fully controlled by an attacker for an hour, what is the maximum harm? With a raw provider key, the answer is “unbounded spend, plus theft of the key, plus whatever else the key unlocks.” With a scoped virtual key, the answer is “at most the key’s remaining quota and dollar cap, and nothing after I revoke it.”

That reframing is the whole game. A vk_ virtual key has no mathematical relationship to the underlying provider credential, so a compromised agent cannot leak what it never sees. It carries its own request quota and dollar cap, so a hijacked agent hits a hard ceiling. And it can be revoked in one click without touching any other key, so containment does not require a fleet-wide credential rotation at 2 a.m.

Detection is half the model

Prevention bounds the damage; detection tells you it happened and proves what occurred. This is where an audit trail stops being a compliance checkbox and becomes an incident-response tool. If every request an agent made is recorded and hash-chained, then after a suspected compromise you can reconstruct exactly which key called which model, when, and at what cost, and you can prove the record was not altered.

KeyForge HMAC-chains every entry to the previous one, so any insertion, deletion, or edit breaks the chain and is mathematically detectable. During a postmortem, that is the difference between “we think the agent did X” and “we can demonstrate the agent did exactly X, and here is the verifiable evidence.” The March 2026 LiteLLM supply-chain compromise, where trojanized packages harvested credentials from gateway machines, is a reminder that the environments around agents get compromised too, not just the agents themselves. Assume breach; instrument accordingly.

A practical checklist

Put the threat model into practice with a few concrete rules. First, agents never hold raw provider keys, they authenticate with scoped virtual keys and the real credential is injected server-side. Second, every agent key has a request quota, a dollar cap, and an expiry sized to its job, not to convenience. Third, one key per agent (or per task), so isolation and revocation are surgical. Fourth, every request is written to a tamper-evident audit trail you can verify and export.

None of this makes prompt injection go away, nothing does. What it does is convert a catastrophic, open-ended credential risk into a bounded, observable, recoverable one. That is what a mature agent threat model looks like, and it is the model KeyForge is built to enforce. You can issue your first scoped, capped, audited keys free, no card required.

Ready to forge your first virtual key?

3 virtual keys, 1,000 requests a month, and the full HMAC audit chain — free.