Security for Personal AI Agents (2026)

In shortStart with the security stack around the agent: isolated execution, brokered credentials, scoped actions and recipients, independent review, and revocable access to data. This radar covers 41 researchers and builders, eight research areas, and the documented boundaries of Muse, Instinct, Hermes, Claude Code/Cowork, ChatGPT/Codex and Grok Bot. General jailbreaking, AI-control theory and agents used for cyber offense receive less weight here.

What security does a personal agent need?

Security for personal AI agents is the layer that lets an autonomous assistant use private files, messages, accounts and credentials while keeping its actions within the user’s authority. An assistant may read an attacker’s email and then act through the user’s authenticated browser or shell, so limits on what it can read, change, send and retain must hold even when its reasoning is wrong. The layer spans isolation, credential handling, enforceable permissions, browser boundaries, prompt injection resistance, persistent memory, delegation and the lifecycle of personal data. A cloud sandbox, a robust model and an approval prompt each cover different parts of this problem.

Personal agent security researchers by subfield

SubfieldLeading researchersWhy it mattersNotable work
Isolation, credentials & egressTarek Sheasha · David Dworken · Kai Greshake · Andrew PaverdKeep private files, secrets and outgoing traffic behind boundaries the reasoning model cannot rewrite.Beyond permission prompts: making Claude Code more secure and autonomous
Action authorization & user intentDawn Song · Tianneng Shi · Wenbo Guo · Somesh Jha · Nils Palumbo · John Hughes · Franziska RoesnerTranslate the requested job into limits on resources, operations, recipients, amounts and grant duration.Progent: Securing AI Agents with Privilege Control
Untrusted content & model robustnessSimon Willison · Sahar Abdelnabi · Sizhe Chen · David Wagner · Eric Wallace · Arman ZharmagambetovSeparate outside instructions from the user’s authority and harden models without losing legitimate task utility.Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
Browser & computer-use boundariesEdoardo Debenedetti · Nicholas Carlini · Ilia Shumailov · Franziska Roesner · David Kohlbrenner · Egor Zverev · Chuan GuoPreserve origin separation and constrain actions across webpages, screenshots and authenticated sessions.ceLLMate: Sandboxing Browser AI Agents
Persistent memory, skills & toolsLuca Beurer-Kellner · Marc Fischer · Johann Rehberger · Zhaorun Chen · Zhen Xiang · David Schmotz · Florian TramèrPrevent hostile content from becoming durable instructions, installed capabilities or scheduled work.Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
Delegation & agent identityIlia Shumailov · Andrew Paverd · Edoardo Debenedetti · Tianneng ShiKeep authority and data provenance intact through child agents, shared workspaces and connected accounts.Can CaMeLs Talk? Securing Multi-Agent Systems Against Indirect Prompt Injection Attacks
Privacy, retention & third-party dataDiyi Yang · Yijia Shao · Chuan GuoControl appropriate disclosure, provider processing, collected copies and deletion after access is revoked.PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
Adversarial testing & operational containmentMatt Fredrikson · Zico Kolter · Andy Zou · Qiusi Zhan · Daniel Kang · Bo Li · Chaowei Xiao · Milad Nasr · Patrick Wardle · John HughesTest unauthorized actions, monitor bypasses, cost amplification and the client that controls the agent.How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition

How the personal-agent security layer is being built

  • Muse — Meta describes a per-user cloud VM with host-side Sentinel authorization, surrogate credentials, taint-aware egress, a constrained browser broker and trusted approval UI. Independent classifiers and model hardening add another layer. Confidential VMs were planned, not a launch guarantee. Architecture.
  • Instinct — Its personal-assistant remit includes bookings, purchases and scheduling. The privacy policy distinguishes connector revocation from deletion of collected data; its training opt-out has exceptions, with a separate Google Workspace exclusion. The reviewed public material does not establish a comparable runtime security architecture. Product · Data policy.
  • Hermes — Security depends on the deployment backend. Command approvals, protected file writes, skill scanning and sanitized child environments are useful controls; they are not a universal OS boundary. Explicit skill/terminal exceptions can pass real secrets to processes. Persistent memory and cron jobs enlarge the lifecycle to secure. Security · Secrets.
  • Claude Code / Cowork — Claude Code combines OS filesystem/network sandboxing with action permissions and optional auto-mode classification. Cowork’s VM boundary still needs careful mounts, host-side MCP and credential mediation. Anthropic describes a historical exfiltration path through an approved API domain into an attacker’s account—a hostname allowlist alone was insufficient. Sandbox · Containment.
  • ChatGPT / Codex / dots — Codex separates command sandboxing from approval of boundary escalations; cloud, local, browser and connector paths have distinct controls. OpenAI’s environment guidance describes vault placeholders and host-side credential insertion for builders. Dots’ controls distinguish pausing a run from stopping delegated work and schedules. Sandbox · Environment guidance · Dots controls.
  • Grok Bot — The official x.ai docs describe Cursor-hosted per-user Firecracker computers. All of a user’s Bots share files, browser sessions and logins; separate Bots are not separate security boundaries. Auto Review is model-based, local execution has its own approval policy, and destination allowlists are Enterprise-only. Blocking a plugin leaves the website path open unless separately restricted. Security FAQ · Approvals.
PioneersField-definers · 7

Sahar Abdelnabi

Institute

Principal investigator, ELLIS Institute Tübingen · Research group leader, Max Planck Institute for Intelligent Systems · Research group leader, Tübingen AI Center

PhD · CISPA — Mario Fritz

Indirect prompt injection, adaptive email attacks and poisoned agent skills.

Her work follows attacker-controlled content through the surfaces personal agents actually consume: retrieved documents, email and skill files. LLMail-Inject and Skill-Inject connect realistic attacker access with measurable outcomes, rather than treating every attack as a malicious user prompt.

Nicholas Carlini

Frontier lab

Research scientist, Anthropic

PhD · UC Berkeley — David Wagner

CaMeL and rigorous adaptive evaluation of LLM defenses.

CaMeL places security policy outside the reasoning model; his adaptive-attack work asks whether a defense survives an attacker who understands it. Both matter for deployment: a successful fixed benchmark is weaker evidence than an enforced boundary tested by an adaptive adversary.

Kai Greshake

Big tech

AI security researcher, NVIDIA

Coauthored the foundational indirect prompt injection study.

The early indirect-injection study established that a malicious webpage or retrieved document can redirect an assistant without controlling the user’s request. His coauthored system-defense work moves the discussion toward architectures that limit what such redirection can accomplish.

Dawn Song

Academia

Professor, UC Berkeley

Agent memory poisoning, programmable privileges and agent security systematization.

Her coauthored agent work connects retrieval-memory poisoning with programmable tool privileges. AgentPoison studies targeted corrupted memories; Progent constrains actions through explicit policies. The practical question is which private resources a compromised agent can still reach.

Florian Tramèr

Academia

Assistant professor, ETH Zürich

PhD · Stanford University — Dan Boneh

AgentDojo, CaMeL and persistent-memory attack evaluation.

AgentDojo measures useful tool work alongside injection success. CaMeL separates untrusted data processing from policy enforcement, while Trojan Hippo Bench evaluates delayed memory attacks. These contributions address the transition from one chat turn to a persistent agent handling private data.

David Wagner

Academia

Professor, UC Berkeley

Structured-query defenses, secure preference training and memory attack analysis.

StruQ and SecAlign improve the distinction between instructions and data through interfaces and training. Trojan Hippo Bench extends the question to memories retained across sessions. These are complementary to runtime limits on files, credentials and outgoing actions.

Simon Willison

Startup

Developer and researcher, Independent open source · Part-time contributor, Prime Radiant

Named prompt injection and proposed the Dual LLM pattern.

Prompt injection becomes consequential when a reader can also use the user’s accounts. His terminology and Dual LLM proposal helped frame that boundary: quarantine outside text while a separate privileged component handles authorized work. This remains a useful starting point for personal-agent architecture.

Leading researchersShaping the field today · 28

Luca Beurer-Kellner

Big tech

Agent security researcher, Snyk / Invariant Labs

PhD · ETH Zürich

Secure agent design patterns and skill supply-chain research.

His coauthored secure-design patterns connect model isolation, permissions and information flow. Later skill-file research examines extensions that agents treat as reusable instructions. That makes the extension installation path part of a personal agent’s security perimeter.

Sizhe Chen

Academia

PhD candidate, UC Berkeley

First author of StruQ, SecAlign and Meta-SecAlign.

StruQ, SecAlign and Meta-SecAlign develop interfaces and training for separating legitimate instructions from injected data. The latest Meta-SecAlign version also addresses training shortcuts that harmed task utility. This is model hardening within a broader security stack.

Zhaorun Chen

AcademiaFrontier lab

PhD student, University of Chicago · Automated agent red-teaming lead, Meta Superintelligence Labs

First author of AgentPoison and DTap.

AgentPoison targets retrieval-backed memories, and DTap constructs executable red-team workflows. As first author of both, he connects a concrete persistence attack surface to richer tests of whether agent actions follow the user’s intended task.

Edoardo Debenedetti

Agent security researcher, Current affiliation unconfirmed

PhD · ETH Zürich — Florian Tramèr

First author of AgentDojo and CaMeL.

As first author of AgentDojo and CaMeL, he helped establish both an executable test environment and an architecture with enforced data-flow policies. His browser work investigates how untrusted page regions can be kept from influencing privileged agent decisions. His current affiliation is unconfirmed following AI Sequrity Company’s indexed wind-down announcement.

David Dworken

Big tech

Security engineer, Anthropic

Coauthored Claude Code’s OS sandbox and credential-proxy design.

His coauthored Claude Code engineering work describes filesystem and network isolation plus cloud credential proxies. These mechanisms constrain executable actions independently of model intent. Their practical strength depends on the mounted data, allowed destinations and tools that remain outside the sandbox.

Marc Fischer

Big tech

Research engineer, Snyk

PhD · ETH Zürich — SRI Lab

AgentDojo and secure agent tooling from Invariant Labs.

His contributions to AgentDojo and secure agent design patterns bridge adversarial testing and agent construction. Invariant Labs’ tooling lineage is especially relevant to mediation of tool calls and untrusted tool responses in deployed assistants.

Matt Fredrikson

StartupAcademia

Co-founder and CEO, Gray Swan · Associate professor, Carnegie Mellon University

Co-founded Gray Swan; coauthored public agent red-team competitions.

Gray Swan’s co-founder and CEO coauthored ART and the large-scale indirect-injection competition study. These contests expose failure modes across tool, coding and computer-use agents. They provide adversarial evidence, with results bounded by the contest surfaces and judges.

Chuan Guo

Frontier lab

Researcher, OpenAI

PhD · Cornell University — Kilian Weinberger and Karthik Sridharan

Secure model training, instruction hierarchy and agent privacy evaluation.

His coauthored work includes instruction-boundary training, privacy evaluation and the WASP web-agent benchmark. This combination matters when an assistant must use private context to navigate websites without treating a hostile page as authorization to disclose it.

Wenbo Guo

AcademiaFrontier lab

Assistant professor; Zhu Chair, UC Santa Barbara · AI research scientist, Meta Superintelligence Labs

Tool privilege control and the agent attack/defense landscape.

Progent’s programmable privileges constrain tool calls through explicit rules; his coauthored system research examines how agent components interact. This is directly relevant to restricting account actions beyond the coarse permissions of a connected application.

John Hughes

Big tech

AI control team, Anthropic

Built Claude Code auto-mode safeguards; coauthored monitor red teaming.

His work on Claude Code auto mode adds an independent action classifier to autonomous coding. The September red-team paper he coauthored shows how transcript formatting, unreviewed edits, compaction and delegation can defeat monitors. The paper excludes sandbox protections, so its attack rates are not full-product compromise rates.

Somesh Jha

Academia

Professor, University of Wisconsin–Madison

PhD · Carnegie Mellon University

Formal policy enforcement for agentic systems.

FORGE, the current version of the formal-policy work, uses an external reference monitor and explicit environment contracts. Its relevance is enforceable authorization over real actions, with guarantees dependent on complete mediation and accurate facts about the environment.

Daniel Kang

Academia

Assistant professor, University of Illinois Urbana-Champaign

PhD · Stanford University — Peter Bailis and Matei Zaharia

InjecAgent and adaptive attacks against agent injection defenses.

InjecAgent evaluates data theft and harmful actions through malicious tool outputs. His coauthored adaptive follow-up tests defenses tailored to the attacker’s knowledge. The resulting evidence helps distinguish apparent robustness from resilience under deliberate probing.

David Kohlbrenner

Academia

Assistant professor, University of Washington

Coauthored the study of same-origin-policy failures in agentic browsers.

His study with Franziska Roesner treats a browser agent as a privileged cross-origin observer. A malicious page can ask the agent to read another origin and relay information. The historical measurements highlight an architectural risk without proving every browser remains vulnerable today.

Zico Kolter

AcademiaStartup

Professor; Director, Machine Learning Department, Carnegie Mellon University · Co-founder and Chief Scientist, Gray Swan

Co-founded Gray Swan; coauthored large-scale indirect-injection evaluation.

Alongside his CMU role, Gray Swan’s co-founder and Chief Scientist coauthored the public agent-security competitions. The agent-specific work measures whether hostile external content causes unauthorized behavior; the company builds red-teaming and runtime-classification products around that problem.

Bo Li

Academia

Research associate professor, University of Chicago

PhD · Vanderbilt University

AgentPoison and the DecodingTrust-Agent red-teaming platform.

Her coauthored AgentPoison and DTap work examines compromised retrieval and controllable agent workflows. DTap’s executable task judges help test whether an agent both completes the user’s job and avoids unauthorized actions in the same environment.

Milad Nasr

Frontier lab

Research scientist, Anthropic

Adaptive attack evaluation and instruction-hierarchy training data.

His coauthored adaptive-attack research challenges defenses with knowledge of their design. For personal-agent evaluation, the key implication is to test the whole data-to-action path under optimization, rather than relying on a static collection of injection phrases.

Andrew Paverd

Big tech

Principal research manager, Microsoft Security Response Center

PhD · University of Oxford — Andrew Martin and Ian Brown

Secure agent architecture patterns and the LLMail-Inject challenge.

His coauthored secure-design patterns and Microsoft system-defense work connect injection risks to software security boundaries. For personal agents, that means limiting action authority and data movement even when the model encounters hostile content.

Johann Rehberger

Security researcher and author, Embrace The Red

Practical prompt injection chains and persistent memory exfiltration.

His published demonstrations trace injected content into durable memories and downstream actions. The useful deployment lesson is persistence: stopping the current response does not necessarily remove an instruction that has already entered the agent’s future context.

Franziska Roesner

Academia

Professor, University of Washington

Agentic-browser origin boundaries and user permissions from interface to enforcement.

Her browser-origin study with David Kohlbrenner examines how an agent can bridge boundaries the web normally enforces. Her permissions survey with Alexandra Michael connects user-facing choices to actual policy enforcement and revocation, a central issue for assistants acting on private accounts.

Tarek Sheasha

Big tech

Software engineer and VP, Meta Superintelligence Labs

Authored Muse’s published isolation, authorization and credential-brokering architecture.

His Muse architecture article describes a host-enforced security layer around a persistent personal agent: isolation, credential substitution, action authorization and independently checked browser use. It is an implementation account, with planned confidential VMs clearly distinguished from launch controls.

Ilia Shumailov

AI security researcher, Current affiliation unconfirmed

Coauthored CaMeL, ceLLMate and MultiCaMeL isolation architectures.

He coauthored CaMeL, ceLLMate and the October 2026 MultiCaMeL preprint. Together they address data-flow enforcement, browser action mediation and trust preservation across delegated agents. Their guarantees depend on the specific mediated environment and stated threat model. He founded AI Sequrity Company, whose indexed site announces a wind-down; his current affiliation is unconfirmed.

Eric Wallace

Frontier lab

Researcher; Alignment Training co-lead, OpenAI

PhD · UC Berkeley — Dan Klein and Dawn Song

Training models to prioritize privileged instructions.

Instruction-hierarchy training asks models to respect the source of a request rather than following whichever text appears last. His coauthored hierarchy work provides a model-level layer for agents, with system enforcement still needed for consequential actions.

Patrick Wardle

Institute

Founder and Co-Director, Objective-See Foundation

Published a Muse local-client dictation-channel proof of concept.

His Muse proof of concept examines the local client that connects to a powerful cloud assistant. It requires an already-running local user process and a dictation trigger; it is not a remote entry or a VM escape. The case shows why the desktop control plane belongs in the agent’s threat model.

Key work

Zhen Xiang

Academia

Assistant professor, University of Georgia

PhD · Pennsylvania State University — David J. Miller and George Kesidis

Retrieval poisoning and query-only memory injection.

His coauthored work on AgentPoison and DTap studies corrupted retrieval and controlled adversarial workflows. These contributions help evaluate personal agents whose accumulated knowledge can influence later tool use.

Chaowei Xiao

AcademiaBig tech

Assistant professor, Johns Hopkins University · Researcher, NVIDIA Research · Visiting researcher, Stanford University

PhD · University of Michigan

AgentPoison and adaptive evaluation of system defenses.

His agent-security work spans retrieval poisoning and system containment. AutoDojo generates attacks for the particular defense under test, including ambiguity in user requests. It brings adaptive pressure to evaluations that otherwise overfit a fixed attack set.

Diyi Yang

Academia

Assistant professor, Stanford University

PhD · Carnegie Mellon University

PrivacyLens and contextual privacy norms for agent actions.

PrivacyLens evaluates whether an agent discloses private information in context, rather than merely detecting sensitive strings. Her coauthored architectural work connects privacy expectations to the tools and resources a personal assistant can access.

Arman Zharmagambetov

Frontier lab

Research scientist, Meta Superintelligence Labs / FAIR

PhD · UC Merced — Miguel Á. Carreira-Perpiñán

Preference-based injection defenses and web-agent data minimization.

Meta-SecAlign strengthens model resistance to injected instructions, while WASP supplies web-agent attack scenarios. His coauthored work connects trained defenses to the browser environment where personal agents encounter hostile content.

Andy Zou

Startup

Co-founder, Gray Swan

Co-founded Gray Swan; first author of the ART competition report.

Gray Swan’s co-founder is first author of the ART competition report and a coauthor of its indirect-injection follow-up. The agent work deserves attention beyond his earlier jailbreak research because it tests deployed-style tools and computer actions. His two linked reports distinguish direct harmful requests from attacks through untrusted external content.

Rising starsEmerging first-authors · 6

Nils Palumbo

Academia

PhD student, University of Wisconsin–Madison

First author of FORGE’s formal policy-enforcement framework.

FORGE’s first author places a policy monitor outside the agent and makes environmental assumptions explicit. The approach is relevant to constraining real file and network actions while allowing useful work inside a verified policy boundary.

David Schmotz

Institute

PhD researcher, Max Planck Institute for Intelligent Systems · PhD researcher, ELLIS Institute Tübingen · PhD researcher, Tübingen AI Center

First author of Skill-Inject’s poisoned skill-file benchmark.

Trojan Hippo Bench evaluates attacks that persist through memory and activate later. Its latest version covers multiple memory backends and measures utility costs, making it relevant to assistants that keep working across sessions.

Tianneng Shi

Academia

PhD student, UC Berkeley

First author of Progent’s programmable privilege controls.

As first author of Progent, he develops programmable privilege control for agent tools. Narrowing privileges can be automated; expansion needs approval. The critical implementation question is whether every route to the protected resource passes through the policy layer.

Qiusi Zhan

Academia

PhD candidate, University of Illinois Urbana-Champaign

First author of InjecAgent and adaptive agent-defense attacks.

InjecAgent’s first author helped define measurable attacks through tool responses, including private-data theft. Her adaptive follow-up emphasizes attacks optimized against a known defense, a necessary counterpart to claims based on fixed benchmark performance.

Egor Zverev

InstituteBig tech

PhD student, ISTA · Research intern, Microsoft Research Cambridge

Instruction–data separation measurement, ASIDE and web-content masking.

His work on instruction–data separation and untrusted-content masking examines what a browser agent can safely read. The boundary between page content and privileged instructions becomes especially important when the browser already holds the user’s authenticated sessions.

Companies building in personal agent security

These companies primarily build security, identity, authorization or testing layers. Many serve enterprises; their controls can inform personal-agent stacks, but availability for individual users and supported integrations vary. Product capabilities below are vendor-described. Personal-agent builders appear in the implementation comparison above.

  • AembitX ↗ — Workload identity and policy-controlled credential exchange for agent/MCP access.
  • AstrixX ↗ — Agent identity and access lifecycle controls, including scoped and short-lived credentials.
  • Auth0 / OktaX ↗ — Token Vault for third-party OAuth credentials and user consent in agent applications.
  • DescopeX ↗ — Agentic Identity Hub for agent identity, authorization and credential management.
  • Gray SwanX ↗ — Shade automated red teaming, Cygnal runtime classification and Arena human red teaming.
  • Lasso SecurityX ↗ — Agent runtime monitoring and controls for actions, tools and sensitive-data flows; enterprise-oriented.
  • MindgardX ↗ — Adaptive red teaming and agent/tool/permission testing to find exploitable defensive gaps, with runtime protection options.
  • Noma SecurityX ↗ — Enterprise-adjacent agent discovery, access controls, testing and runtime detection.
  • Operant AI — Inline allow/block/redact decisions, endpoint protection and MCP access controls for agent workflows; enterprise-oriented.
  • OsoX ↗ — Authorization policy and visibility for agent and tool activity.
  • Pillar Security — Agent discovery, multi-turn red teaming, runtime guardrails and endpoint/MCP/skill controls; enterprise-oriented.
  • PromptArmorX ↗ — Supply-chain risk intelligence on agent vendors, connectors, permissions and private-data exposure; complementary to runtime enforcement.
  • Snyk / Invariant LabsX ↗ — Agent guardrails and research on tool poisoning, skills and secure agent architecture.
  • WitnessAI — MCP/tool access policies, runtime defense, sensitive-data redaction and audit trails for supported enterprise agent workflows.
  • ZenityX ↗ — Enterprise-adjacent stateful agent threat detection and runtime controls.

Frontier work that matters for personal agents

The next layer of research is about preserving authority across a persistent, connected system—not only refusing an injected sentence. The links below establish concrete failure modes or proposals; deployment conclusions remain bounded by each source’s threat model.

  • Bind actions to the intended account and recipient — An allowed API host can still receive private data in an attacker’s account. Policy needs to bind the operation, tenant, credential, recipient and payload—not just the hostname. Anthropic’s containment case.
  • Make permissions understandable and continuously revocable — The interface, inferred policy and enforcement path can disagree. Long-lived grants need inspectable scope and expiration; stopping a task, disconnecting an app and deleting copied data are distinct operations. Permissions survey · Instinct policy.
  • Preserve browser origins and data provenance — An agent reading multiple origins can become a bridge across browser security boundaries. Browser mediation and taint tracking must cover authenticated sessions, screenshots and actual network effects. Origin study · ceLLMate.
  • Secure memory, skills and policy changes — A one-turn injection can persist in memory or a reusable skill. Test the write path, later retrieval and scheduled activation, and keep security configuration outside the agent’s mutable workspace. Trojan Hippo Bench · AgentWorm.
  • Carry trust through delegation and compaction — A child agent must not treat untrusted parent data as fresh user authority. Likewise, a summary must not invent authorization. MultiCaMeL studies hierarchical composition; recent monitor red teaming exposes compaction and workflow gaps. MultiCaMeL preprint · Monitor study.
  • Include the desktop client in the trust boundary — Wardle’s Muse PoC requires existing local code execution and a dictation trigger. It tests the client control channel, not a cloud escape. A strong remote sandbox still depends on how the local app authenticates requests and handles tokens. Researcher PoC.
  • Stop the whole workflow and recover what can be recovered — Kill switches must cover child agents, routines and credential access. Dots documents separate controls; Grok Bot’s September changelog adds stop propagation to delegated work. Deleting a Bot does not clear its shared computer’s sessions or files. Dots · Grok Bot changes.
  • Bound cost and make actions auditable — A hostile page can increase work without causing an obvious task failure. Resource budgets and useful action logs belong alongside data-loss controls. Grok Bot separates administrative logs from opt-in action recording; sanitized logs can still contain sensitive context. CostBomb · Logging controls.

Frequently asked questions

What is the core security problem for a personal AI agent?

The agent combines private context, untrusted input and authority to act. The security layer must control reads, writes, recipients, credentials and retained state even when a webpage, email, tool result or mistaken inference redirects its reasoning.

Is prompt injection resistance enough?

No. Model training and classifiers reduce bad decisions, while operating-system boundaries, credential brokers and authorization controls constrain what those decisions can do. Each still needs the right scope: an allowed destination, mounted directory or connected account may contain too much authority.

Who are key researchers and builders in this field?

Debenedetti, Tramèr, Carlini and Shumailov contribute isolation and data-flow architectures. Sheasha, Dworken and Hughes describe implemented agent safeguards; Roesner and Kohlbrenner study permissions and browser boundaries. Gray Swan co-founders Matt Fredrikson, Zico Kolter and Andy Zou contribute large-scale agent red-teaming research.

Why are memory, deletion and delegation security topics?

Personal agents retain context and keep working beyond a single chat. A poisoned memory can affect later actions; a child agent can inherit excessive authority. Pausing work, revoking a connector and removing already-collected data require different controls.

Which companies provide this security layer?

The agent builders implement their own controls. NVIDIA and Snyk develop execution or tool-security layers; Gray Swan supplies adversarial testing and runtime classification. Identity and authorization vendors provide another part of the stack. These are different roles, not interchangeable proofs that a personal agent is secure.

How was this radar researched?

This is a research snapshot checked October 7, 2026. Papers, current product documentation, author biographies and code disclosures support the profiles and architecture summaries. Two last30days searches covered September 7–October 7 across available Reddit, Hacker News, YouTube and GitHub sources; news leads were checked against primary material. Tiers are editorial contribution groups, not a numerical ranking. This guide emphasizes agents handling personal data and deprioritizes general jailbreaks, broad AI-control theory and cyber-offense agents.

More on the radar