Why reliability and validity matter
When an assessment decides who gets support, being right is not optional
Get it wrong and you mislabel a learner, or leave one who needs support without it. Cognitive assessment reliability and validity are what stand between a defensible decision and a guess.
Here is why the science matters, and Cognassist's documented proof, honest limits included.
Reliability and validity: why a cognitive assessment must show both
Two properties decide whether a score can carry a decision. Most look-alike tools can show neither, which is exactly why you should ask for both.
Reliability is consistency
Would the same learner get the same result if they sat the assessment again? A score that swings from one day to the next cannot carry a support decision.
Validity is truth
Does it measure cognition, not digital literacy or motor speed, and does it predict what matters? Without validity, the number is just a number.
The evidence behind the score
The Cognassist Validity Series
An assessment is its evidence, not its interface. Anyone can rebuild a screen of timed tasks; no one can rebuild the norm sample, the validation and the years of item refinement behind the score. Here is ours, paper by paper, each stated with its limits.
Paper 1
How the assessment was built, and why its scores can carry a decision
Paper 1 sets out the psychometric foundations: built to established theory, tested by factor analysis, normed on a large, population-matched sample, and reliable where a decision rests on it. It is our science team's internal development paper (the Cognassist Assessment Technical Manual v1.8), honestly framed, not peer-reviewed.
- Normed on 26,301 UK apprentices, distinct from the 400,000+ learners assessed on the platform. That is more than ten times the standardisation sample of the Wechsler Adult Intelligence Scale, fifth edition (WAIS-5, n = 2,020) - a matter of scale, not a like-for-like comparison.
- Grounded in the Cattell-Horn-Carroll (CHC) model and the Baddeley-Hitch working-memory model. Factor analysis extracts a three-factor structure, and inter-task correlations of .01 to .59 confirm each domain measures a distinct ability.
- The two primary decision indexes, Language & Numbers and Executive Function / Working Memory, meet the British Psychological Society (BPS) high-stakes reliability standard of r >= 0.8. Across the indexes, internal consistency runs .78 to .93 and corrected 30-day test-retest .68 to .87 - coefficients we report in full, not a single summary number.
- Nine Performance Validity Tests guard every result, one per subtask.
Paper 2
It reflects the cognitive signatures associated with real learning differences, and stays specific
Paper 2 tests known-groups validity: does the assessment measure the cognitive differences the clinical literature documents for specific learning differences, and does it stay close to typical where none is expected? A real cognitive assessment does both; a tool that flags everyone does neither. It is doctoral research at Newcastle University, independently supervised and in preparation, not peer-reviewed; diagnoses were self-reported and the sample is not nationally representative.
- The largest, most predictable effects land where theory predicts: the cognitive differences associated with dyscalculia show on Numeracy (Cohen's d = 1.37), and those associated with dyslexia on Reading Decoding (d = 0.73).
- It does not flag ADHD or autistic learners as having a learning-difference profile. The discriminant result matters as much as the convergent one.
- It measures the cognitive differences associated with specific learning differences to inform professional judgement. It does not detect or diagnose them.
Paper 3
The cognitive profile predicts withdrawal, better than a diagnosis label
Paper 3 tests criterion validity against a real outcome: apprenticeship withdrawal. Early assessment scores predict it, and the cognitive profile predicts it better than a diagnosis label does - which is why measuring cognition directly reaches the many learners with a real barrier but no diagnosis. It is doctoral research at Newcastle University, not peer-reviewed and directional: modest but consistent.
- Learners flagged in two cognitive areas withdrew at 86%, against 77% for those with no flag (p = .012; n = 588 across three providers). This is a subgroup finding, not a national completion rate.
- Weighing 20 variables, the cognitive scores carried more predictive signal than the diagnosis flags.
- The link is probably an underestimate: where a flag triggers support, that support shrinks the very gap being measured.
Why this matters in education
The learner an assessment misses is the one it costs the most
Criterion validity sounds like a technical property. In a college or a training provider it is the difference between finding a learner and missing one. A false negative is a learner who has a genuine barrier and is never identified by the assessment, so nothing follows: no adjustment, no plan, no support. They are not recorded as a problem. They are recorded as fine.
- The learner still struggles, but the cause is never named. They meet the same difficulty in every module, and the most available explanation is that they are not cut out for this.
- Withdrawal is the end of that road, and it is why criterion validity is the strand that matters most in education: an assessment that cannot predict the outcome cannot help you prevent it.
- Your staff pay for it too. Tutors can see a learner is struggling and will spend real time trying to help, but without an identified need that effort is untargeted - support given in good faith, aimed at the wrong thing.
It became a battle. The specific gap the EPAO pointed to was processing speed. There were no numbers, no testing. They couldn't see how we could validate that a learner needed extra time.
Why Cognassist is the scientifically credible choice
Weigh what a board, an auditor or a tribunal would weigh - the things a new entrant or an AI cannot reproduce.
A documented evidence base, not a screen
Three strands of evidence - reliability, known-groups validity and criterion validity - built on a large, population-matched normative sample and years of item refinement a new entrant cannot shortcut. Much of the market can show none of it.
Independent scientific oversight
An independent Science Advisory Board - Professor John R. Crawford, Dr John Welch, Dr Clive Skilbeck and Professor Chris Petkov - actively oversaw the assessment's development.
Reviewed against recognised standards
Reviewed by PATOSS (Professional Association of Teachers of Students with Specific Learning Difficulties) against the scientific-validity requirements of JCQ (Joint Council for Qualifications), and built to BPS high-stakes standards, aligned to the international standards for educational and psychological testing (the AERA, APA and NCME standards) and to the International Test Commission (ITC).
Honest limits
Where the assessment stops, stated plainly
Credibility comes from what we concede. Three limits, stated confidently.
- It is not a diagnosis. It measures the cognitive differences associated with specific learning differences to inform professional judgement; it does not detect or diagnose dyslexia, and does not by itself establish disability under the Equality Act 2010.
- Formal exam access arrangements need a separate route: a JCQ Form 8 assessment completed and certified by a Level 7 specialist assessor. Cognassist is a recognised tool toward Form 8 - use it for inclusive profiling and everyday reasonable adjustments, and a Level 7 assessment where formal exam access is required.
- Two of the three validity strands are doctoral and directional, in preparation, and not peer-reviewed. We say so.
FAQ
The questions a prospect actually asks
Why do reliability and validity matter in a cognitive assessment?
Because the output changes a learner's life. Reliability means the result would hold if the learner sat the assessment again; validity means it measures cognition and predicts what matters. Without both, a support decision is a guess you cannot defend.What is the Cognassist Validity Series?
The Validation White Paper Series (Cognassist Science, 2026): Paper 1 (development and psychometric foundations), Paper 2 (known-groups validity), Paper 3 (cognition and apprenticeship completion), alongside the Cognassist Assessment Technical Manual v1.8.Is the Validity Series peer-reviewed?
No, and we do not claim it is. Paper 1 is our internal development paper; Papers 2 and 3 are doctoral research at Newcastle University, independently supervised and in preparation. Separately, the assessment has been reviewed by PATOSS against JCQ requirements - a professional-association review, not peer review.What is a false negative in a cognitive assessment, and why does it matter?
A false negative is a learner who has a genuine barrier but is not identified by the assessment, so no adjustment, plan or support follows. It is the costliest kind of error in education because nothing flags it: the learner is recorded as fine, keeps meeting the same difficulty, and the most available explanation becomes that they are not suited to the course. Staff effort is wasted too, because support given without an identified need is aimed at the wrong thing. This is why criterion validity - whether early scores actually predict an outcome such as withdrawal - is the strand that matters most in an education setting.Is Cognassist scientifically reliable and valid?
Yes, with the limits stated. The two primary decision indexes meet the BPS high-stakes reliability standard (r >= 0.8), and validity rests on three strands: construct, known-groups and criterion. Two of those strands are doctoral and directional.Is it a diagnosis?
No. It measures the cognitive differences associated with specific learning differences to inform professional judgement. It is not a diagnosis and does not by itself establish disability under the Equality Act 2010; formal exam access needs a separate Level 7 (Form 8) assessment.
Go deeper
Start here, then inspect the evidence yourself, or apply a neutral test to every vendor.
See the assessment behind the evidence
You have seen why the science matters and the evidence behind it. Now see the assessment at work.