Skip to content
IC

Service · AI security

AI Offensive Security Review

We attack your AI system the way an adversary would, then tell you what broke and how to fix it. Manual testing of the integration, data, tools, prompts and identity layer around the model, delivered by a certified offensive practitioner.

Who it is for

You have AI in production and a board asking who tested it.

This is for security and technology leaders whose organisation has already shipped something: a customer-facing assistant, an internal copilot with access to real data, a retrieval system over corporate documents, or agents with permission to act in production systems.

The governance work tells you what your obligations are. It does not tell you whether your system can be broken. Those are different questions and they need different evidence. This engagement produces the technical evidence.

If you have not deployed yet, the useful engagement is Secure AI Enablement. If you need the governance position first, start with the AI Governance Posture Assessment.

The problem

Your AI system fails in ways your pen test does not look for.

A conventional application test looks for injection, broken access control and misconfiguration. An AI system adds a component that follows instructions, holds credentials, calls tools and reads content an attacker can write. None of the usual controls constrain it, and none of the usual tests find it.

Direct and indirect prompt injection

Instructions hidden in the content your system reads: documents, emails, tickets, web pages, log fields. We test whether attacker-controlled text reaches the model with enough authority to change what it does.

Tool and function-call abuse

Every tool you expose is an API the model can be talked into calling. We test argument injection, chained calls that combine into an action nobody authorised, and tools that trust model output as validated input.

Agent privilege escalation

Agents aggregate permissions across systems. We map what the agent can reach, then test whether a single manipulated instruction turns limited access into broad access.

RAG and data poisoning

Retrieval indexes and memory stores accept content from sources you do not control. We test whether poisoned content survives ingestion and shapes later answers for other users.

Model and supply chain tampering

Weights, adapters, agent skills and MCP servers all execute in your environment. We test provenance, pinning and what a malicious component would reach.

Exfiltration through model output

Rendered markdown, auto-fetched images, link previews and tool responses are all outbound channels. We test whether data inside the model context can be moved to somewhere you do not control.

Jailbreak resistance

Guardrails are probabilistic filters, not walls. We test the families that transfer across vendors rather than a list of published prompts, and report the residual rate honestly.

Identity abuse for non-human actors

Agents authenticate, hold credentials and act. We test whether an agent identity can be impersonated, whether its credentials are scoped, and whether its actions are attributable in your logs.

Methodology

How an engagement runs.

Five phases. The full methodology, including what we ask for at scoping and what your team needs to prepare, is on the methodology page.

01

Scoping and threat model

Targets, methods, time windows, escalation contacts and out-of-scope assets agreed in writing. Statement of work and a separate authorisation letter signed before anything is touched.

02

Attack surface mapping

We map the real system: prompts, tools, data sources, retrieval indexes, agent identities, permissions and the trust boundaries between them. Most findings start here.

03

Adversarial testing

Manual testing against each attack path in scope. Injection through the channels your system actually reads, not a generic prompt list run through an automated harness.

04

Exploitation and chaining

Individual weaknesses are combined to demonstrate real impact. A low-severity injection plus an over-permissioned tool is a high-severity finding, and we prove it rather than assert it.

05

Reporting and retest

Critical findings reported within 24 hours of discovery. Final report with reproductions, severity and remediation, a technical debrief, and a free retest within 60 days.

What a finding looks like

Findings prove impact, not theory.

Every finding carries a reproduction, evidence, a severity rating and a fix. The example below is an illustration of the format, not a real client finding.

High CVSS 8.1 OWASP LLM01 ATLAS AML.T0051

Indirect prompt injection in a shared document leads to cross-tenant data disclosure

Summary
A document uploaded by any authenticated user is indexed into the shared retrieval corpus. Instruction text placed in that document is retrieved and followed when a different user asks a related question.
Impact
An attacker with the lowest level of access causes the assistant to disclose retrieved content from another business unit, and to append it to an outbound link. Confidentiality boundary between tenants is not enforced at retrieval time.
Reproduction
Supplied as a step-by-step sequence with the exact payload, the account used, timestamps, and screen captures of the assistant response. Reproducible by your team without our involvement.
Remediation
Enforce tenant scoping at retrieval rather than at presentation. Treat retrieved content as untrusted data in the prompt. Remove automatic link rendering from assistant output.

Deliverables

What you take away.

  • Written report with executive summary and technical findings
  • Reproduction steps and evidence for every finding
  • Severity ratings using CVSS v3.1 with environmental scoring
  • Mapping to OWASP Top 10 for LLMs and MITRE ATLAS
  • CPS 234 reporting mapping for APRA-regulated clients
  • Technical debrief with the engineering team
  • Free retest of all findings within 60 days

Engagement parameters

Investment.

Single AI application or assistant
AUD 18,000 to 35,000
2 to 3 weeks. Scoped to the number of tools, data sources and user roles in the system.
Agent platform or multi-agent estate
AUD 35,000 to 75,000
3 to 6 weeks. Includes agent identity, tool permission and containment testing across the platform.
Retest
Included
All findings retested within 60 days of report delivery, at no additional cost.

Fixed fee, agreed before work starts. A scoping call establishes the surface, and the statement of work follows within two business days.

Standards and methodology

What we test against.

Findings map to published frameworks so the report is auditable and the method is defensible to a regulator, an auditor or your own engineers.

OWASP Top 10 for LLM Applications OWASP Top 10 for Agentic Applications MITRE ATLAS NIST AI RMF CISA and ASD ACSC agentic AI guidance APRA CPS 234 CVSS v3.1

Practical notes

A few things worth knowing.

What is in scope, and what is not

We test your integration, data, tools, prompts and identity layer. We do not attack the hosted model itself. That keeps the engagement inside the model vendor's terms of service and inside your contract with them, and it is where your exploitable risk actually sits. The rules of engagement and the authorisation letter set that boundary in writing.

We do not test what we built

If we designed or deployed the system, we do not review it. Where a client wants both, the review is run by an independent tester we bring in, or the build goes to a partner. Independence is what makes the report usable as assurance.

The AI extension of the penetration testing practice

This is the same discipline, applied to a new class of system. Where an engagement spans both, we run it as one piece of work. See Penetration Testing for conventional application, cloud, network and red team scopes.

Built for procurement

Independence, data handling, scoping and reporting are documented for enterprise buyers. See Working with enterprise.

Get started

Find out what your AI system does under attack.

A scoping call establishes the surface and the rules of engagement. Fixed-fee statement of work within two business days, then we test.