Idea & Market Validation

Idea Evaluation: Set Evidence Gates Before You Build

Before you build a scheduling tool for independent clinics, decide what you need to learn: for example, interview 15 clinic managers and advance only if at least 8 describe missed appointments as a costly, recurring problem they already try to solve. That threshold turns customer discovery from a string of agreeable conversations into a decision with consequences.

The point is not to make early-stage uncertainty disappear. It is to decide, in advance, which evidence would earn another investment of time—and which result would send you back to the drawing board.

An evidence gate is a decision, not a questionnaire

Founders often enter interviews with a list of questions and leave with a collection of quotes. Both can be useful. Neither, by itself, tells you whether to build. An evidence gate supplies the missing piece: a pre-agreed rule that connects observations to a next step.

A practical gate names four things: the assumption under examination, the people or context that can test it, the observable signal that would count, and the threshold required to proceed. It also states what happens if the result falls short. Without that last part, a “test” can become a ritual whose outcome never changes the plan.

For a clinic scheduling concept, the risky assumption might be: independent clinics lose meaningful revenue because patients miss appointments, and current reminders or booking tools do not adequately reduce the problem. That is more useful than “clinics need better software.” It is specific enough to investigate, and it can be wrong.

Suppose the team sets a first gate of 15 interviews with clinic managers, with at least 8 describing missed appointments as frequent, costly, and unresolved by their current process. The number is not a universal benchmark or a statistical proof of market size. It is a working decision rule for a small, directional discovery exercise. Its value lies in being explicit before the team hears the answers.

Start with the assumption that could break the business

A startup idea usually contains several assumptions: that a particular buyer has a problem, that the problem is urgent, that an available solution is inadequate, that the buyer can and will pay, and that you can reach enough buyers at an acceptable cost. Testing all of them at once makes results difficult to interpret.

Write each assumption as a statement that could be contradicted. “Clinics want a modern experience” is vague and hard to disprove. “Managers at independent clinics spend at least two hours a week manually filling cancellations caused by no-shows” is much more testable. It identifies a customer, a behavior, and a scale of pain.

Then rank assumptions by two factors: how damaging it would be if the assumption were false, and how little evidence you currently have. The combination helps separate the foundational unknowns from details that can wait. If a clinic manager’s willingness to pay is uncertain, polishing the interface does not reduce the central risk.

For the scheduling example, a useful sequence might be:

  • Problem: Do missed appointments happen often enough to matter financially or operationally?
  • Existing behavior: Do clinics already spend staff time or money trying to prevent them?
  • Buyer: Is the clinic manager able to choose or influence a purchase?
  • Solution fit: Would a different scheduling workflow address the cause of missed visits?
  • Commercial case: Is the improvement valuable enough to justify a paid product?

Test the first assumptions before asking customers to react to a detailed product concept. A polished prototype can make an idea easier to imagine, but it can also pull the conversation toward interface preferences before you know whether the underlying problem deserves a product.

Design customer discovery around evidence

Customer discovery is not a pitch disguised as a conversation. The founder’s job is to understand what people did, what it cost them, and what they do now—not to persuade them that a proposed solution is attractive.

Ask about recent, specific events. “Tell me about the last time a patient missed an appointment” is more revealing than “Do patients ever miss appointments?” Follow with questions about frequency, consequences, the people involved, and what happened next. If the manager says the clinic uses reminders, ask when they were introduced, who administers them, and whether the clinic can compare missed-appointment rates before and after.

Keep questions neutral. “Would an automated tool that fills open slots be useful?” invites a pleasant hypothetical. “What did your team do the last time a patient cancelled on the day of the appointment?” asks for behavior. People can usually recall a workflow more reliably than predict whether they will buy a product they have never used.

Separate the interview notes into observations and interpretations. “The manager checks a spreadsheet each morning for cancellations” is an observation. “The clinic needs an AI scheduling platform” is an interpretation. Keeping those apart makes it easier to revisit the evidence if several explanations fit the same behavior.

Ask for artifacts when appropriate and with permission: an anonymized workflow, a blank reminder template, a description of the reports the clinic reviews, or an example of how staff record cancellations. Evidence of a real process is more informative than a respondent’s general enthusiasm. Do not request sensitive patient information; protect privacy and follow applicable data rules.

Polite interest is not the same as demand

Early interviews are emotionally flattering. A potential customer says the idea sounds smart, volunteers a feature, or promises to “keep in touch.” These reactions may signal curiosity, but they are weak evidence of urgency, purchase intent, or product-market fit.

Stronger signals require more effort or consequence from the customer. A manager sharing the current cancellation workflow, introducing the founder to the person who controls purchasing, agreeing to examine anonymized data, or committing staff time to a pilot has done more than offer a compliment. A paid pilot is stronger still, although the amount, scope, and terms matter.

Think of evidence as a ladder rather than a binary. At the bottom is a favorable opinion. Above it sits a description of a recent problem; then documented frequency or cost; then an existing workaround with time or money attached; and, higher still, a concrete commitment such as a scheduled pilot or payment. The right rung depends on the decision. Interviews may be enough to justify a prototype, while a hiring plan or substantial engineering commitment should require stronger evidence.

Do not treat every signal as equivalent. A clinic’s use of a generic calendar may show a workaround, but not necessarily a desire to replace it. A manager’s willingness to test a free tool may establish usability interest, but not a viable business model. If the product depends on recurring subscription revenue, eventually you need evidence about who pays, what budget it comes from, and what measurable outcome could justify the fee.

A practical follow-up to a positive interview is not “Would you use this?” Try asking, “What would have to be true for you to test it with real appointments?” Then look for a concrete next step. A response such as “Send something over” is different from a named contact, a date on the calendar, and a defined pilot condition.

Compare the idea with the alternatives customers already use

Your competitor is not only another startup with a matching feature list. It may be an established scheduling system, a phone call, a text reminder, a paper list, a staff member’s manual follow-up, or the decision to tolerate empty appointment slots. The status quo has an advantage: customers already understand it, even when it is inconvenient.

For each interview, record the current alternative, who operates it, its direct cost, the staff time it consumes, and the conditions under which it fails. If a clinic says missed appointments are painful but has made no attempt to change anything, explore why. The issue may be low priority, limited authority, a lack of budget, integration concerns, or a belief that every alternative would create more work.

This comparison prevents a common misreading: assuming that a problem automatically creates a market. A problem can be frequent and still not support a business if customers will not pay, if the buyer cannot approve a purchase, or if existing systems solve the issue well enough. Conversely, a modest problem can be commercially interesting when an existing workaround is expensive and a clearly defined buyer has budget authority.

Look for a specific advantage, not a grand claim of disruption. In the clinic example, a scheduling tool might be valuable if it helps staff fill cancellations faster without adding data entry or replacing a system they rely on. Whether that is plausible is a question for evidence, not a headline.

Set thresholds before the first interview

A threshold keeps enthusiasm, fatigue, and a particularly charismatic respondent from rewriting the standard halfway through discovery. It should be precise enough to guide a decision, but not dressed up as scientific certainty when the sample is small.

For the first discovery round, a team might write:

  • Sample: 15 managers at independent clinics that schedule recurring appointments.
  • Pass signal: At least 8 describe missed appointments as recurring and materially costly, and can explain their current response.
  • Stronger signal: At least 5 provide a concrete example of a workaround involving staff time, direct spending, or measurable lost capacity.
  • Fail or revise signal: Fewer than 5 describe the issue as a current priority, or most say existing tools handle it adequately.
  • Timebox: Complete the interview round and synthesis within three weeks.
  • Decision: Advance to a narrow prototype, revise the target customer or problem, or stop work on the concept.

These counts are illustrative, not an industry standard. A small, deliberately selected sample does not estimate the views of every clinic. It can still expose a weak assumption, reveal repeated workflows, and help a founder decide whether a more expensive test is warranted.

Define the terms that carry weight. What counts as “costly”? Is staff time enough, or must there be lost revenue? What qualifies as “recurring”—weekly, monthly, or simply more than once? Ambiguous criteria make it easy to declare success after the fact. A short scoring rubric, applied consistently, is often more useful than an elaborate spreadsheet.

Record dissent as carefully as support. If seven people describe a serious issue and eight say it is minor, the result is not a clean pass merely because a threshold was nearly met. Examine whether the two groups differ by clinic size, specialty, patient mix, or current software. The contradiction may point to a narrower customer segment worth testing.

Read the outcome without moving the goalposts

When the interviews are complete, compare the evidence with the gate exactly as written. Do not upgrade a weak result because the interviewees were friendly, or downgrade a useful finding because it challenges the original product idea. The purpose of a gate is not to protect the idea. It is to protect the quality of the next decision.

A pass does not mean “build the whole product.” It means the evidence supports the next, limited investment. If managers repeatedly report a costly problem and can show an existing workaround, the next test might examine whether the proposed workflow is feasible. That could be a clickable prototype, a manually operated service, or a tightly scoped paid pilot—chosen to answer the next riskiest question.

A fail is not automatically a verdict on the founder or even the broad market. It might show that independent clinics are the wrong segment, that no-show reduction is not a priority, or that the current systems already address the problem. Preserve what was learned. Change one important assumption at a time so the next test can explain why the result changed.

An inconclusive result deserves its own category. Perhaps the interviews reached front-desk staff rather than buyers, or respondents could not estimate costs without records. Do not quietly relabel “we did not find out” as “customers are interested.” Decide whether a better sample or a different method is worth the added time. If it is, set a new, bounded gate. If not, move on.

Use the smallest test that can settle the next question

Once a problem gate passes, the team does not need to leap directly into a production MVP. Choose a test proportional to the uncertainty. If the open question is whether managers can understand a proposed workflow, a low-fidelity prototype may be enough. If the concern is whether staff will use it during a busy clinic day, observation or a carefully controlled pilot may be more informative.

If you test a pilot, specify in advance what success means: the clinic type, the duration, the users, the workflow, and the outcome to track. For instance, you might ask whether staff can manage cancellation follow-up with fewer manual steps over four weeks. Avoid promising a reduction in no-shows before you have a baseline and a credible way to attribute change to the product.

Distinguish product evidence from commercial evidence. A user completing a task in a prototype shows that the interaction may be understandable. It does not show that a clinic will buy the product, that the product can integrate with its systems, or that acquisition costs will support a sustainable business. Those questions belong to later gates.

Similarly, early customer discovery cannot establish CAC or LTV. Customer acquisition cost depends on real acquisition spend and attributable customers; lifetime value depends on revenue, margin, retention, and the time horizon used. Interviews can reveal likely channels and purchasing constraints, but founders should not turn hypothetical answers into metric forecasts.

Common ways evidence gates go wrong

  • Testing the feature instead of the problem. A reaction to a screen cannot establish that the underlying need is urgent. Learn about recent behavior first.
  • Recruiting only friendly respondents. Existing contacts may be generous or unusually invested. Seek participants who fit the customer profile but have no reason to reassure you.
  • Counting repeated opinions as independent proof. Several managers in the same clinic group may share one policy or purchasing process. Note where participants come from.
  • Changing the threshold after hearing the results. If the team changes what counts as a pass, document why and run a new test rather than rewriting history.
  • Confusing a waitlist with willingness to pay. An email address is a low-cost expression of interest. It is not a purchase commitment.
  • Collecting more data without a decision attached. Another dozen interviews are not automatically useful. Name the uncertainty they will resolve and the choice that follows.
  • Ignoring the buyer. A daily user may feel the problem acutely while someone else controls procurement. Map both roles before inferring demand.

A compact evidence-gate worksheet

Before scheduling your first conversation, write a one-page brief. Keep it visible while recruiting and interviewing; it is a guardrail against turning customer discovery into a product demonstration.

  • Decision at stake: What will you do if the evidence is strong, weak, or inconclusive?
  • Riskiest assumption: What must be true for this idea to deserve another investment?
  • Target respondent: Who experiences the problem, who uses a solution, and who can approve spending?
  • Observable evidence: What recent behavior, record, workaround, cost, or commitment would support the assumption?
  • Threshold: How many qualifying observations are enough for this next decision?
  • Counterevidence: What result would make you pause, narrow the segment, or stop?
  • Method and timebox: How will you gather evidence, and by what date will you make the call?
  • Next test: If the gate passes, what is the smallest experiment that answers the next important question?

For the clinic example, this brief might authorize a prototype only after the team hears repeated accounts of costly missed appointments and sees a real workaround. It would not authorize a full scheduling platform. That distinction is the quiet discipline behind an evidence gate: every result earns only the next sensible step.

Make the decision before the idea gets expensive

Evidence gates are most valuable before sunk costs begin to accumulate. Once a team has spent months building, recruited early users, and developed a polished identity, contradictory findings become harder to accept. A written threshold established before discovery gives founders something sturdier than confidence to consult.

For a scheduling tool aimed at independent clinics, 15 interviews and an 8-of-15 problem threshold are not a market verdict. They are a deliberate first test. Stronger evidence might justify a constrained prototype; a weak result might point toward a different customer or a different problem. Either outcome is useful if it changes what the team does next.

Build when the evidence says the next experiment belongs in the product—not because the idea has become familiar. A good gate does not promise certainty. It makes uncertainty legible, keeps compliments in their proper place, and lets a founder spend the next dollar, week, or engineering sprint on a question that matters.

Theme