Regulated consumer lending · AI voice assistant · anonymised

Confirmed

The guardrail held. The name field didn't.

Asked directly, the agent refused to reveal its instructions — correctly, every time. So we stopped asking in the conversation and put the request in the caller's name instead. It read its own system prompt aloud.

Sector
Consumer lending · regulated
Stack
Realtime voice · token proxy · LLM
Surface
Token minting, voice agent, agent conduct
Access
Unauthenticated
0Credentials needed to reach any of it
8.1CVSS on the top finding
4Confirmed findings at Medium-High or above
16/16Privilege-injection attempts correctly refused

The journey

Everything guarded the conversation. Nothing guarded the metadata.

  1. 01

    The conversational guardrail was doing its job

    We asked the agent to reveal its instructions. It refused. We asked for another customer's data. It refused. We asked for a written guarantee of returns. It refused, and it described the product's illiquidity accurately when pressed.

    On the channel everyone tests, this agent behaved.

  2. 02

    Getting to it cost nothing

    The endpoint that mints session tokens required no login, no key and no credential of any kind — and accepted calls from any web origin. A page a victim merely visits could mint one silently in their browser.

  3. 03

    The caller's name was never treated as input

    Every control sat on the conversation. The name field attached to the session was treated as a label, not as untrusted data — so instructions placed there arrived inside the agent's context having passed nothing.

    It disclosed its verbatim system prompt, its complete internal tool inventory, and the identity of the foundation model underneath — the three things it was explicitly instructed to keep confidential. It could also be driven off its intended script on demand.

    Same request. Same agent. A different door.

  4. 04

    Memory answered to a name, not to a person

    The agent remembers callers between sessions, keyed to the name it is handed. Nothing verified that name. Asserting somebody else's returned their remembered financial context — including an investment figure from an earlier conversation.

  5. 05

    Chained, it needs no click at all

    Any site the victim visits mints a token, injects through the name, and reaches a regulated lender's production agent — prompt, tools and model out, behaviour under external control, and a named customer's remembered context resumed. All from an ordinary browser, with no account anywhere.

    Sessions are server-isolated, so no attacker could listen to somebody else's live call. The blast radius is capped per session. Every step to reach it needed zero credentials.

  6. 06

    Two kinds of finding, and only one has an attacker

    Separately from anything we broke, the agent introduced itself with an advisory title and promoted historical return figures with thin risk disclosure — behaving exactly as designed. No attacker, no exploit. The design itself was the exposure.

    We reported the two classes separately, because they are fixed by different people. One is an engineering defect. The other is a decision about what the agent is told to say.

  7. 07

    Fixed, then retested

    The token endpoint now authenticates and no longer answers to arbitrary origins, session metadata is treated as untrusted input, and remembered context is bound to a verified identity rather than an asserted name. Each fix was re-attacked before it was written down as closed.

The two classes, side by side

Security

An attacker and an exploit

  • Prompt injection through session metadata — prompt, tools and model disclosed, behaviour overridden. High · CVSS 8.1
  • Token minting with no authentication.
  • Token minting accepted from any web origin.
  • Remembered context keyed to an unverified name.

Conduct & compliance

No attacker — the design is the risk

  • An advisory persona promoting historical returns with thin risk disclosure.
  • Remembered financial data re-served against an unverified identifier, with no consent binding.
  • The agent instructed not to volunteer that it is an AI.

The two connect: the conduct findings are standing risks by design, and the injection makes them attacker-controllable on demand. One flaw, told from both ends.

What held

Published so it does not get regressed. The token infrastructure underneath was genuinely well built, and we said so.

  • A 16-field privilege-injection matrix was refused outright. Room, identity and grants stayed server-authoritative; the server never trusted a client-supplied privilege.
  • Token forgery failed. The signing secret was strong, and both unsigned and wrong-secret tokens were rejected.
  • Input validation was robust across the endpoints we exercised.
  • The conversational channel refused everything asked of it directly — exfiltration, third-party data, and a written return guarantee.

Where it bites

India DPDP Act, 2023
Financial history re-served against an unauthenticated, self-asserted identifier with no consent or identity binding — Sec. 8(5) reasonable security safeguards, with Sec. 8(6) notification in scope.
Sectoral lending rules
Marketing and disclosure obligations that govern how a regulated lender may present its product, and how it may not present itself.
Consumer Protection Act, 2019
Return-marketing to consumers without commensurate risk disclosure.
Threat frameworks
OWASP LLM Top 10 — LLM01 prompt injection, LLM07 system-prompt leakage, LLM06 sensitive-information disclosure. MITRE ATLAS — AML.T0053, AML.T0007, AML.T0035.

The lesson that generalises

  • A voice agent's attack surface is not only what the caller says. It is every field attached to the call.

    What this engagement changed

  • The guardrail was real and it worked. It was simply mounted on one channel, and the agent had two.

    Root cause, not a model failure

Findings fixed and re-attacked before this was published. Reference available on request.

Shipping an AI agent that talks to customers?

Twenty minutes is enough to scope it.

Anonymised and published under written authorization, after the findings were remediated and retested · No client data redistributed · Technique and reproduction steps withheld