Agent evaluation fails when a team tests only one impressive demo task. A good dataset should separate task domains, tool permissions, hidden state, expected evidence, pass/fail oracles, safety probes, latency and cost logging, and regression cadence.
This checklist connects RepoDaily-covered tools — Agent-Reach, Firecrawl, Playwright, Jina Reader, codebase memory, Cognee, skills, SkillSpector, AI Website Cloner Template, and DeerFlow — to public evaluation patterns from SWE-bench, WebArena, VisualWebArena, OSWorld, GAIA, τ-bench, MCP-Bench, and MCP-specific eval notebooks.
The goal is not to crown one universal benchmark. The goal is to pick the smallest dataset that matches the behavior under review: code patches, web tasks, GUI control, tool use, memory recall, permission boundaries, collaborative user interaction, or production regression.