AI security · Prompt injection, data leakage, tool abuse
Your AI feature is a new attack surface.
An assistant with access to internal documents, a support bot that can issue refunds, an agent that can call your API. We test what happens when untrusted text reaches a system with real permissions.
Recon · Why it matters
The model is rarely the weakest part. What you connected it to is.
The risk in most deployments is not that the model says something odd. It is that a document, a ticket or a web page can carry instructions into a system that can act, and the model is happy to follow them. Attack surface here is defined by permissions, not by parameters.
Coverage · What we test
What is in scope.
01
Prompt injection
Direct and indirect, including instructions hidden in documents, web pages and tickets your model ingests.
02
Data leakage
System prompt extraction, retrieval and training data exposure, and cross-user leakage in shared context.
03
Tool and agent abuse
What the model can call, what it can be talked into calling, and whether authorisation is enforced anywhere other than the prompt.
04
Guardrail bypass
How far your filters hold under pressure, and what the system does when they fail.
Execution · How we work
How the engagement runs.
Step 01
Map what the model can reach
Every tool, data source and downstream service, with the permission behind each one.
Step 02
Attack the boundary
We treat every input the model consumes as attacker controlled, because in most deployments it is.
Step 03
Test the system, not the sentence
A jailbreak that produces rude text is a curiosity. One that moves money is a finding.
Step 04
Report what held
Including the controls that worked, so you know what to keep as the product changes.
Debrief · What you get
What lands on your desk.
Every engagement ends with something your engineers can act on and your auditors can accept.
- Reproducible payloads for every successful injection, with the exact context that triggered it
- Permission review of every tool the model can invoke and the authorisation behind it
- Architecture recommendations for keeping untrusted input away from privileged action
- Retest after your fixes, because guardrails regress with every prompt change
Questions · Straight answers
Common questions.
Can you test a model we did not build?
Yes. Most engagements involve a commercial model wrapped in your own data, tools and prompts, which is where the risk usually sits.
Is this just jailbreaking?
No. Producing disallowed text is the easy part. We focus on whether an injection can reach a system that does something consequential.
Need this scoped? Let's talk.
Tell us what you need tested and when. A senior tester reads every request and replies within an hour with scope, timing and price.
Reply within an hour · NDA on request · Scoped by a senior tester, not sales