Start with a behavior worth understanding.
We ask what helps an agent maintain situational awareness, justify an action, or recognize when its context is incomplete.
ArcUI uses controlled experiments to study agent behavior in governed spatial systems. Benchmarks are evidence, not the end goal: they help us identify recurring patterns and derive principles for building operational spatial systems.
We ask what helps an agent maintain situational awareness, justify an action, or recognize when its context is incomplete.
Controlled spatial episodes, observation contracts and recorded evidence let others examine what the agent actually received.
A score matters only when it helps reveal a recurring pattern that can improve how spatial systems are built and evaluated.
ArcUI-Bench is an active experimental harness for this question. It compares paired episodes where the task is identical but the available context changes, so a correct answer must match the evidence received rather than a model's tendency to assert or abstain.
Explore the Bench archive →