Skip to content

AI security · Prompt injection, data leakage, tool abuse

Your AI feature is a new attack surface.

An assistant with access to internal documents, a support bot that can issue refunds, an agent that can call your API. We test what happens when untrusted text reaches a system with real permissions.

Recon · Why it matters

The model is rarely the weakest part. What you connected it to is.

The risk in most deployments is not that the model says something odd. It is that a document, a ticket or a web page can carry instructions into a system that can act, and the model is happy to follow them. Attack surface here is defined by permissions, not by parameters.

Coverage · What we test

What is in scope.

01

Prompt injection

Direct and indirect, including instructions hidden in documents, web pages and tickets your model ingests.

02

Data leakage

System prompt extraction, retrieval and training data exposure, and cross-user leakage in shared context.

03

Tool and agent abuse

What the model can call, what it can be talked into calling, and whether authorisation is enforced anywhere other than the prompt.

04

Guardrail bypass

How far your filters hold under pressure, and what the system does when they fail.

Execution · How we work

How the engagement runs.

Step 01

Map what the model can reach

Every tool, data source and downstream service, with the permission behind each one.

Step 02

Attack the boundary

We treat every input the model consumes as attacker controlled, because in most deployments it is.

Step 03

Test the system, not the sentence

A jailbreak that produces rude text is a curiosity. One that moves money is a finding.

Step 04

Report what held

Including the controls that worked, so you know what to keep as the product changes.

Debrief · What you get

What lands on your desk.

Every engagement ends with something your engineers can act on and your auditors can accept.

  • Reproducible payloads for every successful injection, with the exact context that triggered it
  • Permission review of every tool the model can invoke and the authorisation behind it
  • Architecture recommendations for keeping untrusted input away from privileged action
  • Retest after your fixes, because guardrails regress with every prompt change

Questions · Straight answers

Common questions.

Can you test a model we did not build?

Yes. Most engagements involve a commercial model wrapped in your own data, tools and prompts, which is where the risk usually sits.

Is this just jailbreaking?

No. Producing disallowed text is the easy part. We focus on whether an injection can reach a system that does something consequential.

Need this scoped? Let's talk.

Tell us what you need tested and when. A senior tester reads every request and replies within an hour with scope, timing and price.

Request a quote

Reply within an hour · NDA on request · Scoped by a senior tester, not sales