The Agent With Permission
Why this matters
For most of the history of artificial intelligence, the machine answered. We asked a question, it returned an answer, and a human decided what to do. AI agents change that boundary: they can use tools, access information, act across systems and continue work with less direct human intervention.
The system
USER → AI AGENT → MEMORY / DATA → TOOLS / CONNECTORS → CREDENTIALS → CONTROL LAYER → INTERNET / EXTERNAL SYSTEMS → HUMAN APPROVAL → REAL-WORLD CONSEQUENCE
SERE test
SERE examines the intended state, actual state, assumptions, dependencies, boundaries, deviation, propagation, controls, response, outcome, evidence and validation. The question is not simply whether an agent is capable. It is whether the organisation can demonstrate where that capability stops.
Case material
Meta's public Muse documentation describes a secure virtual machine, browser access, user-controlled permissions, approval for sensitive actions, audit trails and a separate control layer for internet access. Public documentation establishes the intended architecture; it does not by itself establish how every control behaves under every real-world condition.
Wider pattern
Recent AI security evaluations and the NATS 2026 air-traffic incident point to a related systems question: a local capability or defect can interact with surrounding systems in ways that matter far beyond the component itself.
SERE finding
Capability is not control. A boundary is meaningful only when the organisation can explain the boundary, test it, observe it, respond when it fails and preserve evidence of what happened.
What remains uncertain
Public sources cannot establish the complete operational effectiveness of private controls. This paper therefore distinguishes documented architecture from independently verified performance. It is not a cybersecurity test, AI safety certification or guarantee.