MindProof · Puzzle guides
Are online IQ tests accurate?
What accuracy means for a cognitive test, which parts of it free online tests can and cannot provide, and how to read an online result without over-reading it.
Usually not in the way the word implies — but the interesting question is which part of accuracy is missing, because online tests fail at some things and are perfectly serviceable at others.
“Accurate” bundles together several separate properties. Pulling them apart makes it obvious what a free browser test can honestly offer.
What accuracy actually requires
Standardisation
Everyone takes the test under the same conditions: same instructions, same time limits, same environment, no interruptions. A supervised administration controls this. A browser tab cannot. Nothing stops a test-taker pausing, retrying, looking something up, or being interrupted, and none of that is visible in the result.
A norming sample
A score is a comparison, so it needs something to compare against: a sample of people, of known composition, who took the same items under the same conditions. Building one is the expensive part of test construction, and it is the part most free online tests skip. Without it, a percentile is a presentation choice rather than a finding.
Reliability
The same person should get a similar result on equivalent versions. Every test has a standard error of measurement, and a report that gives a single integer with no confidence band is claiming more precision than any test delivers.
Validity
The test should measure what it claims to. This is the hardest property to establish and the one most rarely demonstrated by a free test, because it requires evidence beyond the test itself.
Where online tests specifically go wrong
- Self-selected norms. When an online test builds its comparison group from its own visitors, the group is whoever chose to take an internet IQ test — not the general population. Any percentile derived from it describes that audience.
- Practice and familiarity. Matrix and sequence puzzles are learnable formats. Someone who has seen a dozen scores higher than someone who has not, independently of reasoning ability.
- Too few items. Short tests are noisy. A handful of questions cannot separate ability from luck, and the fewer the items the wider the true uncertainty around the result.
- Uncontrolled conditions. Tiredness, phone screens, interruptions and untimed retries all move the number.
- Incentives. A test that produces a flattering number gets shared more. That is a commercial pressure on the scoring, not a psychometric one, and it points in a predictable direction.
What they are genuinely useful for
All of which is an argument against over-reading the number, not against taking the test. An online reasoning test can reasonably tell you:
- Which puzzle formats you find easy and which you do not — useful and quite stable information.
- Whether you enjoy this kind of thinking, which is the honest reason most people click.
- Practice with item types that appear on aptitude tests used in hiring, where familiarity genuinely helps.
- A worked explanation for each answer, which teaches something a score never does.
That is a real set of benefits. It is simply not the same as measurement.
How to read an online result
- Look for the norming claim. If a test reports a percentile, it should say who it compared you against. If it will not say, treat the percentile as decoration.
- Look for a confidence band. A single integer is a presentation choice. A range is a measurement.
- Discount a repeat attempt. A second run on the same items measures familiarity as much as reasoning.
- Check what happens to a skipped question. Whether it counts as wrong changes the score and is rarely stated.
- Be sceptical of a high number. A test that tells almost everyone they are gifted has told you about its scoring, not about you.
Where this site stands
The five-question preview on this site reports how many you answered correctly, out of five. It does not convert that into a score, a percentile or a grade, because five items cannot support that claim. The longer assessment in the app does produce an estimate on the familiar scale, and the methodology page states plainly how: questions carry an authored difficulty, the weighted share of correct answers is compared against an estimated distribution, and — the part most tests leave out — that mapping currently rests on published test-score distributions rather than a calibration sample of our own. That is why every score is shown with a confidence band and described as an estimate.
By the standards set out above, that is a test that is honest about lacking the norming sample, not one that has it. We would rather say so than imply a precision we cannot support.
If what you want is the practice rather than the verdict, the puzzle guides cover solving methods, the free preview explains every answer, and what a good IQ score means covers how the scale is constructed in the first place.