B2B sales-intelligence SaaS · multi-tenant · anonymised
ConfirmedThey had been pentested. The AI had not.
RAG pipelines, a multi-agent LLM orchestrator, sequential-ID data APIs. Conventional pentests had already been commissioned. Nothing had ever treated the AI itself as the attack surface. Six criticals — every one reachable from a free trial seat.
The journey
How a tested platform stayed wide open.
-
01
A modern AI stack, sold on a promise of confidentiality
Each tenant uploads proprietary knowledge — RFP responses, competitive research, deal intelligence — and the AI layer works on top of it. The subscription guarantees a customer's intelligence stays theirs.
-
02
Conventional pentests had already been run
The company had commissioned them before. The application had been tested and the report was in hand.
-
03
Three assumptions looked safe
- The UI is the boundary. Human-review gates, which agents a user can reach, which account a report belongs to.
- The system prompt is a secret. Guardrails, routing logic, tool access and tenant-scoping rules all live inside model instructions.
- Tenant isolation is handled. Ownership checks assumed uniform across hundreds of routes.
-
04
None of them had been tested as an AI system
The prompts, the agents, the retrieval pipeline, the tenant boundaries — none had ever been the attack surface. The API beneath accepts whatever a scripted client sends, and the AI's tool calls open paths the authorization design never anticipated.
Not exotic bugs. The predictable failure modes of the current build pattern.
-
05
We tested from an ordinary user's seat
Two tenants, one a free trial. Authenticated, unprivileged, under written authorization — no insider access, no social engineering.
We mapped the live API surface — over 300 requests — then drove every boundary from both directions with tagged canary documents. Every probe replays byte for byte, and every headline finding was reproduced against a known-good control.
-
06
What was sitting there
The account-research endpoint looks up a report by account_id taken straight from the URL, and never checks that the account belongs to the caller's tenant.
Evidence · DATA-01 · BOLA 200 OK GET /v1/tenants/201/users/7206/accounts/8200/research_report
# the S3 path's tenant-shard segment does NOT match our tenant
# fetched with NO Authorization header, NO cookie, NO token:
$ curl -I "$pdf_url" → 200 · 3,491,517 bytes
PDF /Title: " fortune-500 fin-services co Account Research Report"Account IDs are sequential. The 95-account ceiling was self-imposed, not a technical limit.The same pattern held everywhere beneath the interface:
- The download link is a pre-signed URL good for 7 days — a bearer credential that outlives the revoked trial.
- The system prompt was refused on request, then handed over when we asked for it as JSON: routing logic, tool schemas, five agent prompts.
- One agent was reachable from the router but advertised nowhere: natural language to SQL, run against production, read-only enforced by prompt text alone.
- Attacker instructions written to the preference store survived session termination.
- An arbitrary document uploaded to the RAG knowledge base was retrievable within seconds.
- One request set status: APPROVED and require_human_review: false together. Accepted.
- Nothing rate-limits any of it, which turns every finding above into an unattended loop.
-
07
What the client walked away with
- Six criticals with a byte-for-byte reproducer each — proof, not conjecture.
- Root cause and a fix for every finding, down to the missing authorization middleware and the absent server-side SQL guard.
- A map of the AI attack surface they did not know they had.
- Every finding mapped to OWASP LLM, MITRE ATLAS and quantified penalty exposure.
- A clear statement of what was built correctly, so effort goes where it counts.
What held
The file-sharing layer was engineered correctly.
We ran the same forgery playbook against the RFP file-sharing layer: a canary planted
in one tenant, then read attempts from the other using the victim's file reference, a
mutated tenant prefix, a foreign session ID and an altered storage path.
Every variant was blocked. The server normalised each request to the caller's own
bucket (400 "No valid files found") and rejected cross-tenant URLs at the
middleware (401 "unauthorized tenant").
One asymmetry was worth watching — a route that silently accepts foreign session IDs but returned no data. Filed as a hardening item, not a breach.
A report that cries wolf on everything is worthless. The credibility of the six criticals rests on our willingness to say this part was built right.
Why the board cares
Regulatory exposure, quantified.
Cross-tenant report access invalidates the platform's core commercial promise, and personal data flows through every layer.
GDPR
€20M or 4% of global turnover
Art. 5, 25, 32, 33 — cross-account disclosure, no access-control-by-design, breach-notification triggers.
EU AI Act
€15M or 3% of turnover
Art. 50 — AI-generated certifications published with no provenance. Recipients cannot tell them from authentic ones.
India DPDPA
up to ₹250 crore
Sec. 8(5) and 8(7) — adversarial payloads held indefinitely alongside personal data.
US state privacy
per-record penalties + AG action
Unthrottled enumeration of customer records mirrors every major ID-cycling breach.
If you ship agents that call tools, you are standing on the same three assumptions.
We test them from an ordinary user's seat. Twenty minutes is enough to scope it.
Anonymised under NDA and published under written authorization. Client identity, hostnames, storage buckets, data subjects and tenant names redacted · No client data redistributed