Bounty platforms were designed for web applications. AI systems fail differently — through language, context, memory and delegated authority — and a finding is only useful if it can be reproduced, scored, turned into a runtime control, and shown to be fixed. That is the whole method.
Four services, in the order we recommend adopting them. Each produces findings in the same structure, so a result from a managed engagement, a private program or an automated run can be compared, retested and attested the same way.
A scoped engagement against one system — an LLM application, a RAG pipeline, a tool-using agent, a multimodal model or a classical ML decision system. Our team tests through the interfaces a real user or attacker has, reproduces every finding, and hands you a report with severity, reproduction steps and a recommended runtime control for each.
A two-week, fixed-price test of one agent environment that claims to confine what its agents can reach or change. We test the assumptions behind the confinement — reads that cannot write, egress policy that sees true names, a resolver the agent cannot alter, credentials it cannot borrow — through the interfaces the agent itself has, and rank every path by what it could change.
An invitation-only program in which vetted external researchers test defined AI assets under our code of conduct and your rules. We run triage, reproduce submissions, score them on our severity scale, and pay researchers for validated findings. You see one queue, one format, one standard.
Automated attack suites run against every model or agent release in scope. Results are versioned, so a regression shows up as a difference between two runs rather than an incident in production, and a confirmed attack path becomes a Sentinel Runtime policy that blocks it before it executes.
On the roadmap, not yet offered: agent bounties — an open marketplace where researchers are rewarded for validated AI security findings. We will open it when there are enough programs and researchers to make it a real market; until then, private programs are the path.
Seven steps, and the order matters: nothing is reported until it is reproduced, nothing is closed until it is retested, and the retest is attested in the same format as a runtime decision.
The taxonomy follows the layers of the AI attack surface, from the model to the systems it can reach. Every category has documented patterns that are attempted and recorded, so coverage is a fact about the engagement rather than a claim.
Direct instructions that override system policy or extract it.
Instructions carried in documents, web pages, tickets or emails the system reads as data.
Planting content in a corpus so that it is retrieved and trusted later.
Persisting an instruction across sessions so the attack lands long after it was placed.
Tool-description drift, malicious servers, over-broad manifests, and tool-name collisions.
Chaining permitted tools into an unpermitted outcome, or calling a tool outside its scope.
Actions the agent can take that no one decided to grant it — the gap between permission and intent.
Secrets, PII or internal data leaving through outputs, tool arguments or logs.
Rate, scope and data-boundary violations driven through the agent's own integrations.
Recovering weights, behavior or training data through query access.
Unverified checkpoints, tampered fine-tunes, shadow models, and unpinned dependencies.
Systems that behave differently under test than in production, and evals that can be gamed.
Systems in scope: LLM applications, RAG pipelines, agentic and multi-agent systems, MCP servers and clients, multimodal models, and classical ML decision systems. Categories are mapped to the OWASP Top 10 for LLM Applications and to the attack-surface chain on our platform overview.
An illustrative finding in the format every engagement uses. The target and details are synthetic; the structure is exact.
treasury.transfer tool), staging environmenttreasury.transfer call with the injected destination. Reproduced 5/5 across two model versions.treasury-v3: destination must be on the approved counterparty list; instructions originating in retrieved content cannot authorize a transfer; matches injection signature IPI-0417. Enforced pre-execution.Private programs are staffed by researchers we have vetted and who have agreed to our code of conduct: scope is a contract, prove don't exploit, report promptly and completely, and the customer owns confidentiality. Researchers are paid for validated findings and credited when the customer permits.
If you do AI security research and want to be considered for private programs, tell us what you have found before — a public writeup, a CVE, a disclosed finding — and what kinds of systems you want to test.
Apply to the researcher program →A managed assessment is the first half of every SentinelPoT pilot: we find the highest-risk actions, then instrument them with Sentinel Runtime so the same finding cannot recur.