A measured verdict needs a sample. Early products do not have one, and waiting months for min_n is not a product. Expert reads are how Sightspool produces evidence at n=0 — and human escalation is where the judgement calls it should not make alone actually go.
An expert read is evidence, not a verdict
When an assumption is UX-shaped, a routed lens gives a qualitative read: what it judges to be true, how confident it is, and why. This is stored entirely separately from the verdict machinery. An expert read can never become a verdict, and no amount of agent confidence can rule one. Judgement and measurement are different things, and the product keeps them apart structurally rather than by convention.
Hard calls escalate to a human
When the agent's read is uncertain or low-confidence, it does not pick an answer. The read is flagged as a hard call and enters a human queue, where a senior practitioner records their own read. The human read is stored the same way, graded the same way, and carries the practitioner's name.
The five things that escalate
- A consequential call — where the downside of being wrong is large.
- An ambiguous read — where the agent cannot confidently settle the question.
- A research-integrity question — participant safety, sensitive topics, recruitment.
- An unresolved tension — the lenses disagree and no evidence settles it.
- A proposed action — anything the agent wants to put in front of your customers.
What happens to an escalation
Each one opens a single auditable event recording what triggered it and why. It is assigned to a named practitioner, who is emailed, works it from a real queue and records the minutes it took. Nothing terminates silently — an escalation that nobody handled is visible as exactly that.
You choose the resolution path
- In-house — your team handles it; the escalation closes with that recorded.
- Measure it — turn it into an assumption to be settled by evidence instead.
- Commission the practitioner — have the senior reviewer do the work.
- Dismiss with a reason — a recorded decision, not a silently ignored item.
Practitioners advise; you decide
A practitioner assigned to your workspace can read your evidence, design a study and advise on a call. They cannot rule a verdict, approve an action against your product, or make a decision on your behalf. That boundary is enforced in the application's permission layer, not in a contract — the same guard fires even when a practitioner is actively working inside your workspace.
Judgement is graded by outcome
When an assumption that someone judged is later settled by measurement, that read is scored against the measured result: matched, contradicted, or still unsettled. This applies to the agent's reads and to human ones, and it is how expertise is evaluated here — not by seniority claimed, but by accuracy recorded. Grading happens once, automatically, and never touches the verdict itself.