Luca Beurer-Kellner
Big tech
Agent security researcher, Snyk / Invariant Labs
PhD · ETH Zürich
Secure agent design patterns and skill supply-chain research.
His coauthored secure-design patterns connect model isolation, permissions and information flow. Later skill-file research examines extensions that agents treat as reusable instructions. That makes the extension installation path part of a personal agent’s security perimeter.
PhD candidate, UC Berkeley
First author of StruQ, SecAlign and Meta-SecAlign.
StruQ, SecAlign and Meta-SecAlign develop interfaces and training for separating legitimate instructions from injected data. The latest Meta-SecAlign version also addresses training shortcuts that harmed task utility. This is model hardening within a broader security stack.
Zhaorun Chen
AcademiaFrontier lab
PhD student, University of Chicago · Automated agent red-teaming lead, Meta Superintelligence Labs
First author of AgentPoison and DTap.
AgentPoison targets retrieval-backed memories, and DTap constructs executable red-team workflows. As first author of both, he connects a concrete persistence attack surface to richer tests of whether agent actions follow the user’s intended task.
Agent security researcher, Current affiliation unconfirmed
PhD · ETH Zürich — Florian Tramèr
First author of AgentDojo and CaMeL.
As first author of AgentDojo and CaMeL, he helped establish both an executable test environment and an architecture with enforced data-flow policies. His browser work investigates how untrusted page regions can be kept from influencing privileged agent decisions. His current affiliation is unconfirmed following AI Sequrity Company’s indexed wind-down announcement.
Security engineer, Anthropic
Coauthored Claude Code’s OS sandbox and credential-proxy design.
His coauthored Claude Code engineering work describes filesystem and network isolation plus cloud credential proxies. These mechanisms constrain executable actions independently of model intent. Their practical strength depends on the mounted data, allowed destinations and tools that remain outside the sandbox.
Research engineer, Snyk
PhD · ETH Zürich — SRI Lab
AgentDojo and secure agent tooling from Invariant Labs.
His contributions to AgentDojo and secure agent design patterns bridge adversarial testing and agent construction. Invariant Labs’ tooling lineage is especially relevant to mediation of tool calls and untrusted tool responses in deployed assistants.
Matt Fredrikson
StartupAcademia
Co-founder and CEO, Gray Swan · Associate professor, Carnegie Mellon University
Co-founded Gray Swan; coauthored public agent red-team competitions.
Gray Swan’s co-founder and CEO coauthored ART and the large-scale indirect-injection competition study. These contests expose failure modes across tool, coding and computer-use agents. They provide adversarial evidence, with results bounded by the contest surfaces and judges.
Researcher, OpenAI
PhD · Cornell University — Kilian Weinberger and Karthik Sridharan
Secure model training, instruction hierarchy and agent privacy evaluation.
His coauthored work includes instruction-boundary training, privacy evaluation and the WASP web-agent benchmark. This combination matters when an assistant must use private context to navigate websites without treating a hostile page as authorization to disclose it.
Wenbo Guo
AcademiaFrontier lab
Assistant professor; Zhu Chair, UC Santa Barbara · AI research scientist, Meta Superintelligence Labs
Tool privilege control and the agent attack/defense landscape.
Progent’s programmable privileges constrain tool calls through explicit rules; his coauthored system research examines how agent components interact. This is directly relevant to restricting account actions beyond the coarse permissions of a connected application.
AI control team, Anthropic
Built Claude Code auto-mode safeguards; coauthored monitor red teaming.
His work on Claude Code auto mode adds an independent action classifier to autonomous coding. The September red-team paper he coauthored shows how transcript formatting, unreviewed edits, compaction and delegation can defeat monitors. The paper excludes sandbox protections, so its attack rates are not full-product compromise rates.
Professor, University of Wisconsin–Madison
PhD · Carnegie Mellon University
Formal policy enforcement for agentic systems.
FORGE, the current version of the formal-policy work, uses an external reference monitor and explicit environment contracts. Its relevance is enforceable authorization over real actions, with guarantees dependent on complete mediation and accurate facts about the environment.
Assistant professor, University of Illinois Urbana-Champaign
PhD · Stanford University — Peter Bailis and Matei Zaharia
InjecAgent and adaptive attacks against agent injection defenses.
InjecAgent evaluates data theft and harmful actions through malicious tool outputs. His coauthored adaptive follow-up tests defenses tailored to the attacker’s knowledge. The resulting evidence helps distinguish apparent robustness from resilience under deliberate probing.
David Kohlbrenner
Academia
Assistant professor, University of Washington
Coauthored the study of same-origin-policy failures in agentic browsers.
His study with Franziska Roesner treats a browser agent as a privileged cross-origin observer. A malicious page can ask the agent to read another origin and relay information. The historical measurements highlight an architectural risk without proving every browser remains vulnerable today.
Zico Kolter
AcademiaStartup
Professor; Director, Machine Learning Department, Carnegie Mellon University · Co-founder and Chief Scientist, Gray Swan
Co-founded Gray Swan; coauthored large-scale indirect-injection evaluation.
Alongside his CMU role, Gray Swan’s co-founder and Chief Scientist coauthored the public agent-security competitions. The agent-specific work measures whether hostile external content causes unauthorized behavior; the company builds red-teaming and runtime-classification products around that problem.
Research associate professor, University of Chicago
PhD · Vanderbilt University
AgentPoison and the DecodingTrust-Agent red-teaming platform.
Her coauthored AgentPoison and DTap work examines compromised retrieval and controllable agent workflows. DTap’s executable task judges help test whether an agent both completes the user’s job and avoids unauthorized actions in the same environment.
Research scientist, Anthropic
Adaptive attack evaluation and instruction-hierarchy training data.
His coauthored adaptive-attack research challenges defenses with knowledge of their design. For personal-agent evaluation, the key implication is to test the whole data-to-action path under optimization, rather than relying on a static collection of injection phrases.
Principal research manager, Microsoft Security Response Center
PhD · University of Oxford — Andrew Martin and Ian Brown
Secure agent architecture patterns and the LLMail-Inject challenge.
His coauthored secure-design patterns and Microsoft system-defense work connect injection risks to software security boundaries. For personal agents, that means limiting action authority and data movement even when the model encounters hostile content.
Security researcher and author, Embrace The Red
Practical prompt injection chains and persistent memory exfiltration.
His published demonstrations trace injected content into durable memories and downstream actions. The useful deployment lesson is persistence: stopping the current response does not necessarily remove an instruction that has already entered the agent’s future context.
Franziska Roesner
Academia
Professor, University of Washington
Agentic-browser origin boundaries and user permissions from interface to enforcement.
Her browser-origin study with David Kohlbrenner examines how an agent can bridge boundaries the web normally enforces. Her permissions survey with Alexandra Michael connects user-facing choices to actual policy enforcement and revocation, a central issue for assistants acting on private accounts.
Software engineer and VP, Meta Superintelligence Labs
Authored Muse’s published isolation, authorization and credential-brokering architecture.
His Muse architecture article describes a host-enforced security layer around a persistent personal agent: isolation, credential substitution, action authorization and independently checked browser use. It is an implementation account, with planned confidential VMs clearly distinguished from launch controls.
AI security researcher, Current affiliation unconfirmed
Coauthored CaMeL, ceLLMate and MultiCaMeL isolation architectures.
He coauthored CaMeL, ceLLMate and the October 2026 MultiCaMeL preprint. Together they address data-flow enforcement, browser action mediation and trust preservation across delegated agents. Their guarantees depend on the specific mediated environment and stated threat model. He founded AI Sequrity Company, whose indexed site announces a wind-down; his current affiliation is unconfirmed.
Researcher; Alignment Training co-lead, OpenAI
PhD · UC Berkeley — Dan Klein and Dawn Song
Training models to prioritize privileged instructions.
Instruction-hierarchy training asks models to respect the source of a request rather than following whichever text appears last. His coauthored hierarchy work provides a model-level layer for agents, with system enforcement still needed for consequential actions.
Founder and Co-Director, Objective-See Foundation
Published a Muse local-client dictation-channel proof of concept.
His Muse proof of concept examines the local client that connects to a powerful cloud assistant. It requires an already-running local user process and a dictation trigger; it is not a remote entry or a VM escape. The case shows why the desktop control plane belongs in the agent’s threat model.
Assistant professor, University of Georgia
PhD · Pennsylvania State University — David J. Miller and George Kesidis
Retrieval poisoning and query-only memory injection.
His coauthored work on AgentPoison and DTap studies corrupted retrieval and controlled adversarial workflows. These contributions help evaluate personal agents whose accumulated knowledge can influence later tool use.
Chaowei Xiao
AcademiaBig tech
Assistant professor, Johns Hopkins University · Researcher, NVIDIA Research · Visiting researcher, Stanford University
PhD · University of Michigan
AgentPoison and adaptive evaluation of system defenses.
His agent-security work spans retrieval poisoning and system containment. AutoDojo generates attacks for the particular defense under test, including ambiguity in user requests. It brings adaptive pressure to evaluations that otherwise overfit a fixed attack set.
Assistant professor, Stanford University
PhD · Carnegie Mellon University
PrivacyLens and contextual privacy norms for agent actions.
PrivacyLens evaluates whether an agent discloses private information in context, rather than merely detecting sensitive strings. Her coauthored architectural work connects privacy expectations to the tools and resources a personal assistant can access.
Arman Zharmagambetov
Frontier lab
Research scientist, Meta Superintelligence Labs / FAIR
PhD · UC Merced — Miguel Á. Carreira-Perpiñán
Preference-based injection defenses and web-agent data minimization.
Meta-SecAlign strengthens model resistance to injected instructions, while WASP supplies web-agent attack scenarios. His coauthored work connects trained defenses to the browser environment where personal agents encounter hostile content.
Co-founder, Gray Swan
Co-founded Gray Swan; first author of the ART competition report.
Gray Swan’s co-founder is first author of the ART competition report and a coauthor of its indirect-injection follow-up. The agent work deserves attention beyond his earlier jailbreak research because it tests deployed-style tools and computer actions. His two linked reports distinguish direct harmful requests from attacks through untrusted external content.