FreeBench

Evidence before rankings

FreeBench's planned agent workflow evaluation: frozen campaigns, trusted tools, protected artifact checks and complete denominators. Rankings remain pending.

FreeBench v2 agent benchmark rankings are pending execution qualification, calibration and publication review. No approved public model cohort exists.

Source-reported facts

Catalog facts come from stored OpenRouter records. Every model page identifies its sync date and source. Listed tool support is a declaration; current endpoint availability and task performance are not established.

Planned agent evaluation

The first scope is agent workflows. Actual artifacts must pass protected mandatory behavioral checks. Model verdicts and fabricated test output cannot prove success.

Complete campaigns

Identities, weights, repeats and limits must be frozen before execution. Every scheduled episode is retained, including failures, aborted attempts and retries. Related repeats do not count as independent samples.

Publication requirements

The exact runtime requires qualification, followed by fresh calibration and scoped evidence review. Unsupported comparisons remain unresolved. Missing provenance, incomplete campaigns and runner defects block affected claims.