Research

What we learn when the attacker is patient.

Assessments tell us what broke in one system. Research is what we do when the same thing breaks in several — the failure mode gets written down, measured, and published so someone else can check it.

Papers

  1. Paper · LLM authorization

    Submitted

    Adversarial inputs don't fail quietly. They fail confidently.

    An intent classifier guarding an LLM agent doesn't just mislabel a manipulated prompt — it mislabels it with higher confidence than it gives clean traffic, and the system auto-allows on that confidence. Binary accuracy metrics cannot see this. We measured it, then built a calibration that closes it.

    BSides Bangalore 2026 3,000 prompts 37 intent classes 94% accuracy P95 100ms

    Read the summary →

Field notes

Shorter write-ups — what breaks in production AI systems, what the fix actually costs, and which controls hold up when someone is trying. Drafted as we find them, published when they are reproducible.

In preparation

We also publish two free resources — a DPDP readiness check and a live AI threat briefing. Both are here →

The paper above was submitted under the Trampolyne AI affiliation and is published here unchanged, byline included. Korelex is the assessment practice the same team runs at Vaultmark Private Limited. No client system, client name or client data appears in this research — the evaluation deployment is our own.