TAPBench
TAP-Bench: Think-Aloud Protocol + Benchmark. An instrument for measuring knowledge in configuration space.
TAP-BENCH · THINK-ALOUD PROTOCOL
Can we verify knowledge and capability without ground truth, fact checks, or hallucination tests?
Checking that a model stated a fact is a different job from verifying a body of knowledge, or verifying what it means to know something. TAPBench exists because we need agents that can work in knowledge configuration space and locate the regions where understanding actually sits. Some models will be better, geometrically, at finding those regions. The purpose of this benchmark is to better understand what harnesses and what models are better at building knowledge regions.
A Think-Aloud Protocol (TAP) asks someone to speak every thought while solving a problem or working on an exercise or question: guesses, dead ends, the click. Uncertain Systems records that trail as proof-of-work and embeds it in a high-dimensional mathematical space: an attempt to map someone's knowledge as a state. The person becomes a pin on a map of ways of knowing the topic.
We can use agents to simulate the human think-aloud process, thus generating a cloud of points in that space as benchmark runs. Those points mark a neighborhood: the agent's map of a knowledge-configuration region.

The ScoreBoard tab is the leaderboard. Issue a TAPBench key or download skills.md from a row. How to run is the overall process: key, think-alouds, snapshots, region.