Service · AI security
AI Defence and Agent Containment
You are about to give an agent access to production. This engagement decides what happens when that agent is wrong, manipulated or compromised, and builds the environment that keeps the answer survivable.
Who it is for
Written for the CTO signing off agent access to production.
An agent is a process that reads untrusted content, holds credentials and takes actions. Every control you would apply to a privileged service account applies here, and almost none of them are applied by default.
The question this engagement answers is not whether the agent can be tricked. Assume it can. The question is what it reaches when that happens, how fast you can stop it, and whether you can prove afterwards what it did.
This is the defensive counterpart to the AI Offensive Security Review. Testing tells you where it breaks. This builds the environment where breaking it does not matter as much.
What changed in 2026
Two regulatory events moved AI from policy to proof.
Boards used to ask whether there was an AI policy. They now ask who has tested the AI and who is watching it.
Letter to industry on artificial intelligence
AI is not a separate regime. AI-enabled services must be managed under CPS 230 and CPS 234, and APRA expects continuous monitoring rather than a point-in-time audit.
Read the source ›Careful adoption of agentic AI services
Least privilege, a distinct identity per agent, continuous logging, red teaming through the lifecycle, and human approval for irreversible actions.
Read the source ›Half one · Defence
Hardening the AI workload and the platform under it.
Detection and response for systems whose failure mode is a confident, plausible, wrong action rather than a crash.
Guardrails and filtering
Input and output classifiers sized to your risk, with the residual rate stated honestly. Guardrails reduce the rate of a bad outcome. They are not a boundary, and we do not sell them as one.
Monitoring and detection for AI workloads
Prompt, retrieval, tool call and output treated as privileged activity and surfaced where your detection can see it. Detections for agent identities behaving outside their baseline.
Secrets and key handling
Credentials kept out of prompts, context and logs. Short-lived tokens brokered at call time rather than long-lived keys embedded in agent configuration.
AI incident response playbooks
What to do when an agent acts on an injected instruction: how to stop it, what to preserve, how to determine reach, and who decides. Written before you need it, and exercised.
Hardening the platform underneath
The identity, network, logging and endpoint controls the AI runs on. Most AI incidents are ordinary platform failures with an AI component attached.
Model and artefact provenance
Weights pinned by revision, safe serialisation formats, an allowlist for agent skills and MCP servers, and an inventory that includes what your build pipeline cannot see.
Half two · Containment
The environment an agent runs inside.
Containment is engineering work we design and stand up inside your environment, on your infrastructure, using components you already run. It is not a product and there is nothing to licence from us.
Isolated execution environments
Agents run inside containers or microVMs sized to the blast radius you are willing to accept, not on a developer laptop with ambient access to everything that laptop can reach.
Default-deny network egress
An agent reaches the endpoints it needs and nothing else. This single control turns most exfiltration and callback payloads into a blocked connection and an alert.
Scoped credentials per agent
A distinct identity per agent, least privilege, short lifetime. An agent that inherits an engineer's personal token has that engineer's blast radius.
Tamper-evident action logs
An append-only record of what the agent did, written somewhere the agent cannot reach. If the agent can edit its own audit trail, you do not have one.
Kill switches
A tested way to stop an agent and revoke its credentials in seconds, owned by someone who is on call. Untested kill switches do not work on the day.
Human approval for irreversible actions
Payments, deletions, production changes, external communications and access grants stay with a person. The joint agency guidance is explicit on this, and it is the control that survives a guardrail failure.
Control mapping
Mapped to the frameworks you already report against.
Agent containment is not a new control family. It is existing controls applied to a new kind of principal, which is what makes it defensible in an audit.
| Control | ASD Essential Eight | NIST CSF 2.0 |
|---|---|---|
| Isolated execution and egress control | Application control, restrict admin privileges | PR.PS, PR.IR |
| Scoped agent credentials | Restrict administrative privileges, MFA | PR.AA |
| Tamper-evident action logging | Regular backups, monitoring | DE.CM, PR.PS |
| Patching of AI platform components | Patch applications, patch operating systems | ID.RA, PR.PS |
| Human approval for irreversible actions | Restrict administrative privileges | GV.RR, PR.AA |
| AI incident response playbook | Regular backups | RS.MA, RC.RP |
What you get
Deliverables.
- Threat model for your agent estate, with the blast radius written down per agent
- Containment architecture, designed and stood up in your environment
- Agent identities, scoped credentials and egress policy in place
- Detections for agent activity, delivered into your existing tooling
- AI incident response playbook, exercised with your team once
- Control mapping to Essential Eight and NIST CSF 2.0 for your evidence pack
- Runbooks and handover to named internal owners
Investment
Scoped per engagement
Scope depends on how many agents you run, what they reach, and how much of the platform already exists. A scoping call establishes that, and a fixed-fee statement of work follows within two business days.
Book a scoping callGet started
Decide what an agent reaches before it reaches it.
A scoping call establishes your agent estate, what each one can touch, and the realistic shape of a containment engagement.