A. De Souza
Case file
No. 03 / 11

red-agent

2026  ·  Adversarial LLM harness · TypeScript
Re
red-agent
Date
2026
Discipline
Adversarial LLM harness · TypeScript
Stack
TypeScript · Claude · Prisma · Next.js · YAML

Stress-testing framework that injects YAML-defined attacks into tool-using agents and uses a Claude LLM-judge to classify compromise vs. safe refusal vs. detection. 3 victim models, 5 attack categories, 4 safety metrics, persisted via Prisma with a Next.js dashboard.

Back to the index