Test datasets
Public agent instructions, skills, MCP servers, prompt-injection corpora and eval benchmarks you can use to test an AI product or build agents — curated, with safety notes. Click any column heading to filter by it — type a name, pick a type, format or license — click the arrows to sort, and click a row for details.
A benchmark of agent tasks with embedded prompt-injection attacks, measuring both task success and attack success.
Open full page →Anthropic's official Agent Skills (SKILL.md packages) for documents, design and development — the reference implementation of the format.
Open full page →100+ ready-to-install Claude Code subagent definitions across engineering, data and operations roles.
Open full page →Hundreds of real .cursorrules files by framework and language — a corpus for comparing agent-instruction styles.
Open full page →A standardised red-teaming evaluation of harmful-behaviour refusals across attack methods and models.
Open full page →164 hand-written Python problems with unit tests. Old and heavily contaminated — a smoke test, not a benchmark.
Open full page →15,000+ real jailbreak prompts collected from Reddit, Discord and prompt sites (CCS 2024).
Open full page →An open benchmark and artifact repository for jailbreak attacks and defences, with a public leaderboard.
Open full page →The Model Context Protocol's reference server implementations — the baseline for testing MCP clients and gateways.
Open full page →A framework plus registry of community evals for LLMs and LLM systems.
Open full page →A widely adopted skills library for coding agents covering brainstorming, planning, TDD and debugging workflows.
Open full page →Real GitHub issues with tests, used to measure whether a coding agent can resolve them end to end.
Open full page →A large collection of Claude Code subagents, plugins and slash commands organised by domain.
Open full page →