Methodology

How assessments are calculated

StartupX AI turns founder context, attached evidence and experiment history into structured findings. It does not prove market demand or guarantee business success.

Evidence quality

Higher-quality inputs carry more weight: customer interviews, completed experiments and dated source links qualify for scoring. Founder context explains what to test, but it is not independent proof.

Evidence quantity

A score is less confident when only a few items support a claim. Small samples can point to a useful next test, but they should not be treated as proof of demand.

Contradictions

Supporting and contradicting evidence are reviewed separately. Contradictions reduce confidence and help identify what needs another interview, source or experiment.

Freshness

Recent evidence is generally more useful than stale inputs, especially for pricing, competitor positioning and customer acquisition assumptions.

Verified, unverified and missing evidence

Founder context
Customer interview
Experiment result
Public URL
Imported source
AI suggestion - unverified

AI-assisted suggestions are treated as generated assessments until a founder links real evidence. If evidence is weak, absent or one-sided, the product shows insufficient evidence and recommends the next validation action instead of presenting a precise conclusion.

What affects confidence

number of evidence items linked to the assumption

source quality and attribution

supporting versus contradicting evidence

completed experiment outcomes

missing interviews, public sources or measured results

Implemented calculation model

  • Evidence Engine assessments evaluate ten dimensions, but a dimension receives no numerical score until its independent-evidence threshold is met.
  • Founder descriptions, target customers, business models, known competitors and assumptions are stored as context. They have zero independent-evidence weight.
  • Qualified evidence includes verified public evidence, customer research and completed experiment results. Planned experiments and generated assessments do not raise scores or confidence.
  • Scored dimensions show component-level calculations using source quality, freshness, relevance, evidence direction and supporting-versus-contradicting signals.
  • An overall Evidence Score appears only when enough dimensions are assessable. Otherwise the product shows Evidence Coverage and the missing requirements.

The result screen includes calculation details and threshold status directly, so founders can see why a dimension is scored or why it remains unassessed.

Start assessment