Find the belief that could change the decision
Start with the choice in front of you: build a feature, change onboarding, or explore a customer segment. List what must be true for that choice to make sense. “Customers want this” is too broad. Separate the problem they experience, the behavior you expect, your ability to deliver, and the effort needed to support it.
Pick one belief that would materially change the choice if it were wrong. Prefer a gap you can investigate with the people and resources you can legitimately reach. Do not invent a precise probability simply to make the list look scientific.
Make the question observable
A useful question names a situation and an action. For example: “Can a founder find the reason for a previous pricing decision in their existing notes?” You can observe that task, ask what information is missing, and record where the attempt stops. You cannot infer willingness to pay from someone saying the idea sounds useful.
- A conversation can explore the problem
- Ask about a recent real occasion and the workaround used. Preserve what the person said separately from your explanation. A story can suggest another question; it does not establish how common the problem is.
- A task can reveal a specific difficulty
- Let the participant attempt a clearly defined activity with permitted material. Note prompts or help you supplied, because they change what the observation can support.
- A prototype can test a proposed interaction
- Label what is simulated. Completing a prototype task does not establish repeated use, production reliability, willingness to pay, or demand across a market.
Write a small test brief
Agree on the interpretation before collecting results. A threshold here is a rule for your next action, not automatic proof that a business is viable. Keep inconclusive results available as an honest outcome.
Decision this test informs: Belief we are uncertain about: People and situation included: People and situations excluded: Observable task or behavior: Recruitment method and likely selection bias: Permitted inputs and participant agreement: Observation to record; help we will provide: Result that supports the next bounded step: Result that changes or stops the plan: What would make the result inconclusive: Owner, time limit, and review date: After the test: What actually happened, with dated evidence: What remains unexplained: Response and smallest next step:
Store identifying details only where you have permission to keep them. A public summary can describe the method and limits without reproducing someone’s private documents or implying they endorsed your product.
A worked example: finding a past decision
Fictional example: a founder is considering a decision-history feature. They invite five consenting founders to find the reasoning for one past decision in notes those participants choose to share. The task ends after five minutes. The brief says that three people unable to find the reasoning would justify a second round of interviews, not a feature launch.
In the imagined result, three find it, one cannot, and one decides the notes are too private to use. Record those outcomes separately. The privacy-related withdrawal is neither task failure nor success. This test has not reached its chosen next-step rule, and the recruited group cannot represent all founders.
A reasonable response is to clarify whether the important problem is retrieval, missing reasoning, or the willingness to keep a separate record. Do not change the threshold after seeing the results and then describe the original hypothesis as proven. A revised question starts a new, explicitly revised brief.
Keep an observation separate from a decision
Write three short paragraphs: what happened, what it may mean, and what you will do next. Include contrary observations and missing data. “We did not observe it” is different from “it never happens,” just as a positive reaction differs from actual adoption.
Add the brief and result to your decision journal. Use a weekly decision review to check whether the next question was answered. If the test concerns an AI workflow, use the evaluation scorecard for output quality; customer interest and technical quality answer different questions.