A UX recommendation often sounds plausible before it is useful. Add a few familiar heuristics, mention user friction and finish with a tidy prioritised list. A capable language model can produce that shape immediately.
The problem is not that every sentence will be false. The problem is that the reader cannot tell which sentence came from customer evidence, which came from the implemented product, which is a professional prior and which is simply a confident completion.
Trust needs a product structure
Prompts asking an agent to cite sources are helpful. They are not enough. A trustworthy system should make a sourceless consequential claim difficult to publish and impossible to convert directly into action.
No proof, no post is not a writing preference. It is a constraint on what the product allows to count as a finding.
In Sightspool, a proof-gated finding refers to the evidence read that supports it. The tool call, source and relevant identifiers remain available for review. An intervention requires a proven finding and separate human approval. Closing a response loop requires the agent to have actually read the responses it claims to synthesise.
Proof still needs limitations
Citing a source does not make the conclusion inevitable. A session replay can demonstrate that a person encountered a flow; it may not establish why they left. Five interviews can reveal a useful pattern; they do not establish a population rate. A revenue field can expose stakes; it does not make a UX change causal.
Every consequential read should carry the strongest evidence against it, the confidence the evidence earns and the downside the team still needs to carry.
Measurement needs its own truth states
The same principle applies to quantitative evidence. An event absent from the taxonomy cannot be treated as a measured zero. A result below the declared minimum sample cannot quietly become a verdict. A product change affecting the flow should re-open the belief instead of inheriting the old answer indefinitely.
- Unknown event: the signal is not grounded.
- Gathering: the signal exists, but the sample is not ready.
- Holds or refuted: the declared measure crossed its threshold.
- Verified: a human-owned judgement, never an agent’s self-awarded status.
A team should be able to disagree with the agent
Trust is not the absence of disagreement. It is the ability to inspect the basis of a recommendation and record what the customer did with it. Accept, amend, challenge and dismiss are all useful dispositions when the reason remains visible.
Those dispositions are also evaluation data. A system can learn that a particular class of read is frequently amended, that one specialist escalates too little or that a recommendation is persuasive but rarely implemented. Outcome review adds the final correction: even accepted work can be wrong.
The standard is not certainty
Product teams rarely receive complete evidence. A useful UX agent should move the work forward without pretending the uncertainty disappeared. The target is a defensible next action with its source, limitation and downside intact.
