Heads up This site is currently under heavy development.
← all datasets
◆ Evals & benchmarks · Prompt injection & jailbreaks

AgentDojo

MIT

A benchmark of agent tasks with embedded prompt-injection attacks, measuring both task success and attack success.

Summary

Formats
python
License
MIT
Added
2026-09-10

Get it

From the project’s own instructions where it documents any; otherwise a plain clone.

  1. shell
    $ pip install agentdojo

    Add the transformers extra (agentdojo[transformers]) for the prompt-injection detector.

Contents

36,860 files 37.8 MB repository

  • .json 36,680
  • .py 122
  • .md 24
  • .yaml 17
  • (no ext) 6
  • .sh 3
  • .ipynb 2
  • .bib 1

Use it with

Commands are curated, not yet run by us.

Inspect AI — Run an Inspect task file; the benchmark has to be wrapped as an Inspect task first.

task.py
$ inspect eval task.py --model openai/gpt-4o-mini
my-toolchain — 0 tools
paste an install list to detect your tools

A brew list, a Brewfile, requirements.txt, a Dockerfile — or just the product names, free-form. Nothing leaves your browser.

    browse all tools →

    Send us feedback