Pilot analysis AI system security

The prompt filter is not your agent's security boundary

A CISO's warning about filter-first defenses becomes more useful when tested against agent attack data: protect the action path, not just the text entering the model.

Safety Signal Editorial · 2026-09-27 · 3 min read

Pilot analysis · Archived from the previous site. This piece predates the automated three-per-week selection cadence. Its lead podcast was published on 2026-09-03.

If an AI agent can read a webpage and then move money, the security question is not whether a prompt filter recognizes a suspicious sentence. It is whether untrusted text can cross the boundary into an authorized action.

The signal

In a September 3 AI Security Podcast conversation, Standard Chartered Group CISO Cezary Piekarski argued that the industry's focus on distinguishing instructions from data inside prompts may be a brittle way to defend systems. He used the history of buffer overflows as an analogy: durable progress often came from changes in system architecture, not ever more elaborate input filters. That is an expert's interpretation, not a measured result from the episode.

The podcast matters because agent workflows make the trust boundary operational. A retrieved page, email, or repository file is meant to inform a task. The same text can carry attacker instructions. If the agent also has tools, a successful injection can become a search, file write, message, purchase, or data transfer. The failure is no longer confined to an answer on a screen.

What the evidence supports

A large public competition studied indirect prompt injection against agents exposed to external content. Its authors report that every evaluated model was vulnerable in their setup, with attack success rates ranging from 0.5% for Claude Opus 4.5 to 8.5% for Gemini 2.5 Pro. Those figures describe the study's tasks, attacks, and scoring rules; they are not an enterprise-wide probability of compromise. The researchers also released the competition environment and documented attacks that did not transfer from Qwen to closed models. Transferability and deployment context therefore matter.

OWASP's prompt-injection guidance treats malicious instructions in untrusted material as a core application risk. Its agentic guidance also points builders toward practical system controls. Neither source proves that prompt classifiers have no value. A filter can be one layer. The mistake is to make it the only barrier between third-party content and a privileged action.

My synthesis: measure the action boundary

For a tool-using agent, I would define risk as the probability that untrusted content causes an unauthorized action, multiplied by the impact of that action. A detection score alone misses the second term. The evaluation should replay realistic tasks with adversarial webpages, documents, and messages, then observe tool calls, permission checks, and final state—not merely whether the final answer looks suspicious.

This suggests a layered design. Give each agent only the tools and data needed for the current task. Validate a proposed action against the user's intent and a deterministic policy outside the model. Require approval for high-impact steps; keep an audit trail of the retrieved content, tool request, and decision. Test whether the agent can recover safely when hostile text is encountered. These are design recommendations inferred from the sources, not claims that a particular control has been proven sufficient.

The decision for a CXO

Ask the team to demonstrate one end-to-end failure test before expanding agent permissions: Can a malicious document persuade the agent to use a tool beyond the user's request? What blocks that action, what gets logged, and how quickly can access be revoked? If the answer is only 'we filter prompts,' the control story is incomplete.

The open question is how much risk moves when permissions, approvals, and task design change together. That should be measured in your own workflow. A benchmark result or a compelling podcast analogy can set the hypothesis; your action-level test must decide whether the system is ready.

Sources and limitations

Podcast assertions are practitioner judgments. Competition rates depend on the benchmark's setup and do not estimate risk for a specific enterprise deployment. The recommended control design is an inference, not a tested guarantee.