How assessments are calculated
StartupX AI turns founder context, attached evidence and experiment history into structured findings. It does not prove market demand or guarantee business success.
Evidence quality
Higher-quality inputs carry more weight: customer interviews, completed experiments and dated source links qualify for scoring. Founder context explains what to test, but it is not independent proof.
Evidence quantity
A score is less confident when only a few items support a claim. Small samples can point to a useful next test, but they should not be treated as proof of demand.
Contradictions
Supporting and contradicting evidence are reviewed separately. Contradictions reduce confidence and help identify what needs another interview, source or experiment.
Freshness
Recent evidence is generally more useful than stale inputs, especially for pricing, competitor positioning and customer acquisition assumptions.
Verified, unverified and missing evidence
AI-assisted suggestions are treated as generated assessments until a founder links real evidence. If evidence is weak, absent or one-sided, the product shows insufficient evidence and recommends the next validation action instead of presenting a precise conclusion.
What affects confidence
number of evidence items linked to the assumption
source quality and attribution
supporting versus contradicting evidence
completed experiment outcomes
missing interviews, public sources or measured results
Implemented calculation model
- Evidence Engine assessments evaluate ten dimensions, but a dimension receives no numerical score until its independent-evidence threshold is met.
- Founder descriptions, target customers, business models, known competitors and assumptions are stored as context. They have zero independent-evidence weight.
- Qualified evidence includes verified public evidence, customer research and completed experiment results. Planned experiments and generated assessments do not raise scores or confidence.
- Scored dimensions show component-level calculations using source quality, freshness, relevance, evidence direction and supporting-versus-contradicting signals.
- An overall Evidence Score appears only when enough dimensions are assessable. Otherwise the product shows Evidence Coverage and the missing requirements.
The result screen includes calculation details and threshold status directly, so founders can see why a dimension is scored or why it remains unassessed.