Create a repeatable baseline for the tool-selection scenarios supplied in this MCP-focused benchmark. The dataset is intended for developers evaluating how their own agent workflow chooses among the tool options represented by the fixtures.
What this download contains
Synthetic cases are not an independently validated model leaderboard or a universal measure of safe tool use. The benchmark’s documented task definitions and expected outcomes determine what a result can legitimately demonstrate.
Working with this edition
Review the supported scenarios, map them to your test harness and retain the case version with each run. Keep evaluation isolated from real external side effects and compare changes against the same documented expectations.
Product questions
Does the package run an MCP server for me?
No hosted service is promised. The product is a benchmark dataset; use the included documentation to identify any supplied local tooling.
Does a passing score prove an agent chooses tools safely?
No. It describes performance on the included synthetic cases, not every possible real-world situation.






Reviews
There are no reviews yet.