Heads up This site is currently under heavy development.
← all datasets
◆ Evals & benchmarks

Terminal-Bench 2.0

Apache-2.0

Hard, realistic command-line tasks, each in its own container with tests, measuring whether an agent can finish real terminal work end to end.

Summary

Terminal-Bench is a popular benchmark for measuring the capabilities of agents and language models to perform valuable work in containerized environments. Tasks include assembling proteins for synthesis, debugging async code, and resolving security vulnerabilities. It is used by virtually all frontier labs.
from the project README
Formats
tomldockerfile
License
Apache-2.0
Added
2026-09-10

Get it

From the project’s own instructions where it documents any; otherwise a plain clone.

  1. shell
    $ uv tool install harbor

    Terminal-Bench 2.0 runs on the Harbor harness.

  2. shell
    $ uv run harbor run --dataset [email protected] --agent oracle --n-concurrent 4

    The oracle agent runs the reference solutions, so it needs no API key. Every task runs in its own Docker container.

Contents

1,038 files 45.7 MB repository

  • .sh 184
  • .md 182
  • (no ext) 181
  • .py 135
  • .scm 126
  • .toml 90
  • .c 17
  • .jpg 13
my-toolchain — 0 tools
paste an install list to detect your tools

A brew list, a Brewfile, requirements.txt, a Dockerfile — or just the product names, free-form. Nothing leaves your browser.

    browse all tools →

    Send us feedback