Independent AI security assessment · est. 2025800+ TESTS · 25 CATEGORIES · EU · ISO · NIST · OWASP

Ship AI
with proof,
not promises.

An independent adversarial assessment of your AI endpoint. 800+ tests across the OWASP LLM and Agentic Top 10, each run 5+ times to measure how often it actually fails. Every finding evidenced; one report your board and security team both trust.

Start an auditFixed price · no subscription · free scoping call first, no commitment.See a sample report
How it works · explainer
ScopeBlack-box · endpoint only
Coverage800+ tests · 25 categories
Rigour5+ runs per test · 4,000+ attempts
FrameworksEU AI Act · ISO · NIST · OWASP

Coverage is capability-aware: every applicable test runs against your endpoint; tests that don't apply to your architecture are marked N/A, never padded into the score.

01 / DeliverablesWhat you walk away with

A binder your board, your security team and your engineers can all read.

You don't get a verdict. You get a number. Every finding reports how often it actually fired, across repeated runs.

01 · Co-primary deliverable

The Technical Assessment Report

Every finding: the exact prompt that triggered it, how often it fired, mapped to all four frameworks.

PDF · HTML, JSON, MD on request
02 · Co-primary deliverable

The board-ready executive summary

The one-to-two-page credential your CISO forwards to the CEO. Written to be quoted.

One to two pages · quotable
03

An evidence package, quantified

For every finding: how often it fired, confidence interval, trial count, and the triggering prompt.

Frequency · CI · trial count
04

Three compliance dossiers, pre-mapped

EU AI Act Article 15, ISO 42001, NIST AI RMF, cross-tagged on every finding. Included.

Cross-framework
05

A bounded, fix-verification re-test

You patch. We re-run failed tests, same version, within 30 days. Included.

Closure proof · bounded

"We don't just tell you an attack succeeded. We tell you how often."

02 / CoverageWhat every audit looks for

25 categories. 800+ ways in. Every applicable one tested, repeatedly.

Coverage is capability-aware: every applicable test runs against your endpoint; tests that don't apply to your architecture are marked N/A, never padded into the score.

25 attack categories · 800+ tests · 5+ runs each

OWASP LLM Top 10
LLM01
Prompt Injection
Instructions hijacked through crafted input: direct, indirect, and multi-turn.
LLM02
Sensitive Information Disclosure
System prompts, credentials, and other sensitive content surfacing in responses.
LLM03
Supply Chain Vulnerabilities
Compromised dependencies, fine-tunes, plugins, and third-party model components.
LLM04
Data and Model Poisoning
Manipulated training, fine-tuning, or retrieval data skewing model behaviour.
LLM05
Improper Output Handling
Unsanitised model output passed downstream into code, queries, or rendered content.
LLM06
Excessive Agency
Tools and permissions that let the model take actions beyond its intended scope.
LLM07
System Prompt Leakage
Extraction of the instructions and guardrails meant to stay hidden.
LLM08
Vector and Embedding Weaknesses
Retrieval and embedding pipelines manipulated to poison or exfiltrate context.
LLM09
Misinformation
Confident, plausible, and wrong: outputs presented as fact without basis.
LLM10
Unbounded Consumption
Resource exhaustion and uncontrolled inference cost from adversarial load.
OWASP Agentic Top 10 (ASI01–ASI10)

Tested via adaptive, multi-iteration attack chains, standard on every audit. Inapplicable tests marked N/A, never padded.

ASI01
Agent Goal Hijack
Injected context that redirects an autonomous agent's objective mid-task.
ASI02
Tool Misuse & Exploitation
Coercing an agent's connected tools into unauthorised or destructive calls.
ASI03
Identity & Privilege Abuse
Impersonation or privilege escalation across agent identities and permissions.
ASI04
Agentic Supply Chain Vulnerabilities
Compromised tools, plugins, or sub-agents introduced into the chain.
ASI05
Unexpected Code Execution (RCE)
Prompts that coerce an agent into executing unintended code or commands.
ASI06
Memory & Context Poisoning
Persistent memory or session state corrupted to influence later behaviour.
ASI07
Insecure Inter-Agent Communication
Trust and validation gaps between agents that communicate with each other.
ASI08
Cascading Failures
A single compromised step propagating errors through a multi-agent chain.
ASI09
Human-Agent Trust Exploitation
Social-engineering an agent, or the human relying on it, into an unsafe action.
ASI10
Rogue Agents
An agent that deviates from its intended scope and acts against instructions.
TestMy.AI proprietary bands

Five TestMy.AI bands beyond the OWASP Top 10, three of them also mapped to responsible-AI frameworks.

SCENARIO
Realistic Abuse Scenarios
Multi-step adversarial chains modelled on how real attackers actually combine techniques.
PRODUCTION
Disclosed-Incident Patterns
Attack patterns drawn from publicly disclosed real-world AI security incidents.
SAFETY
Content Safety
Probes that push the model to generate toxic, hateful, or otherwise harmful content.
BYPASS
Linguistic Reframing Jailbreaks
Past-tense, hypothetical, narrative, and authority-persona reframings that slip past guardrails.
DATAEXT
Training Data Extraction
Probes for memorized training data via repetition divergence, verbatim completion, and membership inference.

Breadth answers what did you test. Depth answers how confident are you.

The OWASP LLM test catalog is open-source. Inspect it on GitHub.

03 / FrameworksIncluded, not the headline

Also mapped to four frameworks, included, not upsold.

Every finding is pre-mapped to all four frameworks, so the same evidence answers the regulator, procurement, and the security review. If and when you need it.

European Regulation

EU AI Act, Article 15

Accuracy, robustness and cybersecurity for high-risk AI. Evidence mapped to the article, not a compliance certification.

15.115.215.315.4
International Standard

ISO/IEC 42001

The recognised standard for AI management systems and governance maturity.

A.6A.7A.8A.9A.10
US Framework

NIST AI RMF 1.0

The risk-management framework referenced across US federal and enterprise AI procurement.

GovernMapMeasureManage
Industry Standard

OWASP LLM & Agentic Top 10

The de-facto security checklists for any application built on, or acting as, a large language model agent.

LLM01–10ASI01–10
04 / Deliverable previewThe artefact you open

The report your board and your engineers both open.

Cover sheet, executive summary, findings ledger, evidence appendix. Scannable in sixty seconds; full detail one page deeper.

SAMPLE EXTRACT · FINDINGS LEDGER

Every finding, with its frequency.

Severity, category, and how often it fired, so a non-technical reader can scan the risk surface in under sixty seconds. Evidence and trial-level detail sit one page deeper.

PDFHTMLJSONMarkdown
Technical Assessment Report
Endpoint · api.acme.com/v1/chat
REF · TMA-2026-0714
ART. 15 · ISO 42001 · ASI
CRTF-001Agent goal hijack via injected task contextASI0118/20
CRTF-002System prompt disclosure under role-play pressureLLM077/20 (35%)
HIF-003Unauthorized tool invocation reported as completed (model claim)ASI025/20
HIF-004Indirect injection via retrieved documentLLM019/20
MEDF-005Improper output encoding in rendered responseLLM0511/20
LOF-006Verbose refusal leaks internal policy structureLLM073/20
743 of 818 applicable tests run · 75 marked N/A · 6 findings
See a sample report

Every report is signed by Burcin Sarac, independent lead auditor.

05 / EngagementsOnce, or ongoing

One audit to start. Continuous coverage once you ship again.

There is exactly one audit product, and it always runs the full catalog: full breadth, full adaptive depth, full multi-trial depth. No lite tier, no scope-reduced version. Continuous Assurance is the destination once you're testing every release.

◗ Founding cohort · 5 slots · through Dec 31, 202601 / AI Security Audit

AI Security Audit

The entry point. A standalone, one-time engagement: no lock-in, no subscription. The full catalog, every time.
12,000 USD/ endpoint · standard rate
6,500 USD/ endpoint
Founding rate for the first five audits, in exchange for a reference.
Report in 10 business days · one endpoint
  • All 25 categories: OWASP LLM Top 10, OWASP Agentic Top 10, and five TestMy.AI bands
  • Full adaptive, multi-iteration attack chains on every agentic category, standard, not an add-on
  • Every applicable test run multiple times; findings reported as observed frequency, confidence interval, and trial count
  • Board-ready executive summary and three compliance dossiers included
  • One bounded re-test within 30 days, same version, failed tests only, included
Start an audit
The destination after audit one◗ Founding cohort · year one02 / Continuous Assurance

Continuous Assurance

A point-in-time audit describes a system that has already changed: a new prompt, a new model version, a new tool. It decays in weeks. We keep you tested every time you ship.
30,000 USD/ endpoint / year · standard rate
18,000 USD/ endpoint / year
Founding cohort rate, locked for year one, converting to standard at renewal.
Four quarterly full audits · regression re-testing on material change
  • The same full audit, every quarter, no reduced scope
  • Regression re-testing whenever the system materially changes: new model version, new prompt, new tool
  • Dossiers and board summary kept current between quarters
  • Multi-endpoint pricing available
Talk to us about ongoing coverage
Included in every audit

Three compliance dossiers (EU AI Act Article 15 · ISO/IEC 42001 · NIST AI RMF) and the board-ready executive summary. No separate governance tier.

Add-ons
  • Additional endpoint 7,000 USD 5,000 USD
  • Extra re-test 2,500 USD 2,000 USD
  • Rush turnaround +30%
Scoping

What counts as one endpoint? See /audit.

06 / FAQDirect answers

The plain version.

Plain answers, written for the people who'll actually read the report: security leads, compliance officers, executives, and the engineers on the receiving end of the remediation.

The deliverable
Q · 01
What does an audit deliver?
A Technical Assessment Report, a board-ready executive summary, and a full evidence package: for every finding, observed frequency, confidence interval, trial count, and representative evidence. Plus three compliance dossiers (EU AI Act Article 15, ISO/IEC 42001, NIST AI RMF), prioritized remediation, and one bounded 30-day re-test. Delivered as a PDF. HTML, JSON, and Markdown are available on request.
Q · 13
What can we tell our board and customers after the audit?
You get a board-ready executive summary written to be quoted, and a pre-approved sentence you're licensed to publish: "[Company]'s AI assistant is independently security-tested by TestMy.AI against 800+ adversarial scenarios spanning the OWASP LLM and Agentic Top 10, with quarterly re-testing." We report what was tested and found. We don't certify systems as "safe", and neither should you.
Q · 15
Who is the report written for?
Security teams validating AI before production; compliance officers preparing evidence for ISO 42001 or NIST filings; legal teams preparing EU AI Act dossiers; procurement teams responding to enterprise security reviews; and executives who need to report on AI risk upward. The report has a section for each audience.
How the testing works
Q · 03
Will testing damage production?
Tests run against the endpoint you designate, whether staging, shadow, or production, at a throttle you set. We coordinate test windows, respect your rate limits, and pause on the first 5xx pattern. No destructive payloads, no data exfiltration beyond what your model itself surfaces.
Q · 04
What access do you need?
Just an endpoint URL and authentication. No source code, no model weights, no infrastructure access. The audit is performed black-box, exactly the way an external attacker would see your AI.
Q · 08
If we re-run a finding ourselves, will we see the same result?
Not necessarily, and that's expected. AI systems are non-deterministic; a finding that occurs 35% of the time may not reproduce on one manual attempt. That's precisely why every attack is run multiple times and reported as a frequency rather than a yes/no. Every finding ships with its trial count and observed rate so your team knows exactly what to expect on re-runs.
Q · 09
What does it mean when a test shows no findings?
Not "safe", and not "passed". It means the attack was not observed to succeed across N attempts, and testing at that depth reliably detects behaviours occurring more often than roughly 1 in 7. Rarer intermittent behaviour cannot be excluded, and we say so rather than implying a clean bill of health.
Q · 10
Do you test AI agents, not just chatbots?
Yes. The OWASP Agentic Top 10 (ASI01–ASI10) is covered via adaptive, multi-iteration attack chains on every audit. Full agentic coverage requires tool-calling, RAG, or memory on the target; tests that don't apply to your architecture are marked N/A, not padded into the score.
Q · 11
Do you verify what happened in our backend?
No. Testing is black-box: we observe what your endpoint returns, not what your systems did. Where the model claims to have performed an action, we report it as a model claim with supporting or contradicting signals, never as a confirmed database event.
Commercial & compliance
Q · 02
Is this a certification?
No. TestMy.AI is an independent technical assessor, not a certification body. The deliverable is a Technical Assessment Report with evidence designed to support your compliance filing alongside qualified legal counsel. The EU AI Act conformity-assessment framework for Article 15 is still being established; no one yet holds formal certification authority.
Q · 05
What about data retention?
Test artefacts are retained encrypted for 90 days post-delivery, then destroyed on written request, or extended for your audit-trail retention period. We sign mutual NDA before scoping and never use client data to train or tune models.
Q · 06
Which frameworks does it map to?
Included by default, on every audit: EU AI Act Article 15, ISO/IEC 42001, NIST AI RMF, and OWASP. Every finding is pre-mapped to all four the moment it's written up, so the same evidence answers the regulator, the procurement team, and the security review, if and when you need it.
Q · 07
How long does it take?
A single audit: report delivered within 10 business days of scoping, for one endpoint. Rush turnaround compresses the window at +30%.
Q · 12
What happens after we patch?
One re-test is included within 30 days, same version, failed tests only. You get a verified fix rate and the re-issued evidence. Extra re-tests are 2,000 USD during the founding window, 2,500 USD standard.
Q · 14
Do you offer ongoing testing?
Yes. Continuous Assurance is the same full audit, quarterly, plus regression re-testing on material change (new model version, prompt change, new tool), with dossiers and the board summary kept current. A point-in-time audit decays in weeks; quarterly re-testing is what keeps the statement "our AI is independently tested" true.
Begin · Endpoint to report, in time for the next board review

An audit before the next
board review.

Hand us an endpoint and an auth header. We hand you a report your security team, your executives, and your board can all open.

Start an auditSee a sample reportFixed price · no subscription · free scoping call first, no commitment.