Raven Progressive Matrices Test: A Complete Guide

The Raven Progressive Matrices is an untimed, nonverbal test of abstract reasoning and fluid intelligence consisting of 60 visual pattern items, first published in 1938 and still widely used in clinics and schools. It asks the test taker to identify the missing part of a visual pattern rather than answer questions about vocabulary, facts, or arithmetic.

A familiar referral can look like this: a student solves construction puzzles quickly, understands new games after one demonstration, and spots relationships other children miss, yet performs poorly on a language-heavy cognitive screen. The same mismatch appears with an adult patient whose limited English, aphasia, or communication disability makes verbal reasoning tasks a poor proxy for underlying problem-solving ability.

The Raven Progressive Matrices test can reduce language demands, but it isn't a complete intelligence assessment. Its value depends on selecting the right version, administering it consistently, using an appropriate norm group, and interpreting the result alongside the person's history and other evidence.

Why the Raven Progressive Matrices Test Matters

A school psychologist may meet a student who understands visual patterns quickly but struggles to explain an answer in English. A clinician may see the same issue in an adult whose speech or communication disability makes verbal tasks difficult to interpret. In both cases, Raven can provide evidence about problem solving with fewer language demands.

John C. Raven's test was first published in 1938 as a nonverbal measure of abstract reasoning and fluid intelligence. The first full standardisation involved 1,407 children in Ipswich, England, and the standard form contains 60 items arranged in five groups of 12, according to the American Psychological Association's description of Raven's Progressive Matrices. The APA describes it as a prototypical measure of general intelligence.

Practical rule: A nonverbal format reduces language load, but it does not remove the need for professional judgment.

Raven also fits practical school and clinical workflows. Depending on the form, administration windows range from about 15 to 60 minutes. California districts have used Raven measures in gifted screening and in research involving linguistically and culturally diverse student groups. Those applications show why a compact visual task can be useful when a district needs an initial screening measure, while also showing that screening is not the same as a complete evaluation.

Parents often meet the test during gifted-identification discussions. Clinicians may consider it when language, reading, or expressive communication could affect a broader cognitive assessment. A result still represents one type of performance under specific testing conditions. It cannot, by itself, explain achievement, learning history, motivation, attention, or communication needs.

The interpretation traps begin after the answer sheet is scored. A culture-reduced task is not automatically culture-free. Norms can become less suitable as populations and testing practices change, and newer adaptive formats may not be interchangeable with older paper administrations. This guide therefore moves from one matrix puzzle to version choice, administration, percentile interpretation, cultural limits, and format changes. For broader context on how cognitive tests fit within professional evaluation, see this guide to cognitive assessment.

What the Raven Test Actually Measures

Start with the item itself. Imagine a 3 by 3 grid containing eight completed squares and one empty square in the lower-right corner. Each row contains shapes that rotate, add elements, or change shading. Your task is to infer the rule and select the answer option that completes the pattern.

A careful test taker might work through it like this:

  1. Inspect each row. Perhaps the first shape contains one line, the second contains two, and the third contains three.

  2. Check the columns. The same relationship may appear vertically, but with rotation instead of addition.

  3. Separate the rules. One feature may change across rows while another changes down columns.

  4. Test every option. The correct answer must satisfy the whole grid, not just one attractive part of it.

  5. Choose the missing piece. The response reflects the rule that best explains the relationships already present.

Think of the pattern as a lock and the answer options as keys. The test taker doesn't need to name the shapes or explain the rule in words. They need to discover which key fits the lock.

An infographic detailing the cognitive skills and abilities measured by the Raven Progressive Matrices intelligence test.

This is why Raven is associated with fluid intelligence, the ability to identify relationships and solve unfamiliar problems. It relies less on crystallised knowledge such as vocabulary, general information, or memorised school content. A person can know many facts and still find novel visual rules difficult. Another person can have limited vocabulary in the test language and reason effectively through the matrices.

The distinction matters in practice. A high score supports an interpretation of strong performance on nonverbal abstract reasoning under the test conditions. It doesn't establish reading achievement, verbal comprehension, creativity, social judgment, practical wisdom, or subject knowledge. A low score doesn't prove low general ability either. Visual impairment, fatigue, unfamiliarity with multiple-choice testing, motor difficulties, anxiety, or an unsuitable version can all affect performance.

The test is untimed in the Canadian clinical description. That changes the meaning of a slow but accurate performance. If a student needs more time yet reaches correct solutions, the clinician shouldn't automatically treat that pattern as evidence of weak reasoning. Processing speed may require separate measurement, depending on the referral question. A broader discussion of visual reasoning appears in this spatial abilities test guide.

Comparing the Three Versions SPM CPM and APM

A school team may begin with a child who cannot yet manage the standard matrix format, while a clinician may assess an adult whose reasoning is expected to be unusually strong. Both cases involve visual pattern solving, but the appropriate form differs. Selection depends on developmental level, expected ability, sensory or cognitive limitations, and the referral question.

Start with the item itself. The examinee studies a visual pattern with one piece missing, compares the response choices, and identifies the option that completes the relationship. CPM uses simpler material and colour to make this reasoning task more accessible to younger children and people who may find the standard form too demanding. It can reduce language demands, although it does not remove the need for adequate vision, attention, and understanding of the response process. Canadian clinical materials describe this version (Pearson Clinical Canada).

Standard Progressive Matrices, or SPM, is the broad-use form. The standard form contains 60 items in five groups of 12. The groups generally build from more direct relationships toward more complex combinations, so an examinee may solve early items while reaching a ceiling later. That pattern is informative for choosing norms and interpreting the level of discrimination available. Canadian product information describes this version (Pearson Clinical Canada).

Advanced Progressive Matrices, or APM, is intended for higher-ability adolescents and adults. It is appropriate when SPM may be too easy to distinguish among stronger performers. A high score on SPM can therefore reflect genuine ability while still providing limited separation at the upper end. APM is better suited to referrals where that upper range matters.

Version

Target population

Administration time

Best fit

CPM

Younger children or people needing a less demanding format

About 15 to 30 minutes

Early screening and accessible nonverbal reasoning assessment

SPM

General child, adolescent, and adult populations

About 20 to 45 minutes

Broad assessment of abstract reasoning

APM

High-ability adolescents and adults

About 40 to 60 minutes

Differentiating stronger levels of nonverbal reasoning

The table supports planning, not final selection. A second-grade class receiving a universal screen may need a different starting point from an adult applicant expected to perform well above average. The same choice affects interpretation of culture-reduced reasoning, since familiarity with testing, visual conventions, and the selected norm group can still influence results.

Selection principle: If the form is too difficult, the task can create a floor effect. If it is too easy, scores may cluster near the top and lose useful discrimination.

School teams may pair Raven with a pattern-recognition measure, but the referral should define what each tool contributes. This pattern recognition test resource helps distinguish pattern reasoning from the wider abilities required in a full evaluation.

How Administration and Scoring Work in Practice

A real administration begins before the first matrix. The examiner selects the correct version and edition, prepares the response materials, checks visual access, and explains how to record an answer without revealing the pattern rule. For a group screen, instructions and timing must be delivered consistently. An individual assessment also allows the examiner to observe concentration, hesitation, misunderstanding, and response style.

Administration differs by edition and setting. Some forms support group testing with shared instructions and separate response recording. Others are given individually, allowing closer monitoring of access needs and task engagement. A school screening and a clinical evaluation therefore produce different kinds of observational information, even when both use Raven materials. The report should identify who administered the test, how responses were recorded, and whether any assistance or interruption occurred.

From raw response to norm-referenced result

A raw score is the number of correct responses. It becomes interpretable only after the examiner applies the scoring rules and relevant norm tables. The resulting score may be reported as a percentile rank or another norm-referenced value, depending on the edition and available reporting system.

The conversion resembles placing a runner in the correct race before comparing finishing positions. The examiner must match the raw score to the exact form, age range, norm group, and administration conditions. The same number correct can carry a different meaning under another norming context. A report that omits those details makes the result difficult to evaluate.

The format is also changing. The Raven's 2 Canadian flyer describes clinical editions and the shift from fixed paper workflows toward newer adaptive formats, with paper versions slated for discontinuation in 2025. Scores from an older paper form should not automatically be treated as interchangeable with results from a newer adaptive or clinical edition. The item route, scoring process, and norm reference may differ.

A short pre-administration checklist

  • Materials: Confirm the version, edition, response method, and current scoring resources.

  • Setting: Reduce visual distractions, position materials clearly, and address vision-related access needs.

  • Instructions: Demonstrate answer selection without coaching pattern-solving strategies.

  • Observation: Record fatigue, guessing, confusion, hesitation, and requests for clarification.

  • Reporting: Name the form, norm source, score type, testing conditions, and interpretation limits.

An infographic explaining percentile ranks for the San Diego Unified Gifted Program using a bar chart.

Interpreting Scores and Percentile Ranks

A child selects one tile to complete a matrix. The number correct is the raw score. The percentile rank then places that score beside results from the relevant norm group. A 75th percentile means the child performed as well as or better than 75 percent of that comparison group. It does not mean the child answered 75 percent of the items correctly.

San Diego Unified School District shows how a school system can turn a percentile into a decision. Its gifted programme used Raven for screening, with scores at or above the 98th percentile qualifying for the Cluster programme and scores at the 99.9th percentile qualifying for the Seminar programme (San Diego Unified gifted-program account). Those cutoffs belong to that district. They are not universal definitions of giftedness.

A student who struggles with vocabulary-heavy testing may still reach a nonverbal screening threshold when pattern reasoning is strong. That finding supports further review. It does not replace evidence from achievement, classroom functioning, language development, attention, or other relevant assessments.

What a district workflow can look like

Oxnard School District provides a practical sequence. The district administers Raven's Coloured Progressive Matrices in January to all second graders, then gives Raven's Standard Progressive Matrices to GATE-referred students in grades 3 through 8 (Oxnard School District assessment information). The first stage casts a wide net; the second examines older students who have been referred.

Dry Creek Joint Elementary School District describes Raven as a non-reading measure of abstract reasoning, generally administered in a group setting, taking about 45 to 60 minutes, with 60 multiple-choice items. Its materials explain that percentile scores are age-based and that the test targets fluid intelligence rather than achievement, memorised facts, or vocabulary (Dry Creek Joint Elementary School District GATE materials).

These workflows clarify the difference between screening and diagnosis. A district may use a percentile as one gate in an identification process. A clinician may use the result to investigate whether language-heavy findings deserve closer examination. Neither use makes the Raven score a complete description of ability.

Norms can shift the meaning

The same raw score can carry different meaning under another norming context. In a Canadian Edmonton sample using the 1947 Raven Progressive Matrices, raw-score means were 5.8, 5.9, and 7.1 across three child groups differing in grade and socioeconomic composition, while scaled-score means were 15.5, 24.8, and 25.5 (University of British Columbia archive). The group composition helps explain why apparently similar results should not be compared without context.

Before interpreting a score, ask which norms were used, whether they fit the person's age, language, education, and setting, and whether the selected form was appropriate. Check the testing conditions and consider evidence that supports or contradicts the conclusion. A percentile is a comparison point, not a diagnosis.

An infographic explaining the difference between raw test scores and percentile ranks with a grading guide.

Is the Raven Test Culture-Fair?

A student who has recently entered an English-language classroom may understand a visual pattern yet struggle with the examiner's instructions or answer format. Raven reduces language and culturally specific knowledge, but performance still reflects the conditions in which the test is taken.

Schooling quality, familiarity with multiple-choice tasks, comfort with abstract visual puzzles, sustained attention, and opportunities to practise can all matter. Selecting one answer from several similar options is itself a learned convention. A learner may have strong reasoning potential without much experience with that convention.

Canadian assessment guidance describes comparable nonverbal assessment as culture-reduced rather than culture-free, while emphasising appropriate Canadian norms and bilingual or linguistically diverse administration contexts (Canadian WNV brochure). That wording is more useful than assigning the test a simple fairness label.

A young student looking thoughtfully at a Raven's Progressive Matrices test sheet during a cognitive evaluation.

When the format helps

Raven can be a sensible first-line option when language might otherwise dominate the result. A Canadian school-system comparison selected it for students with limited English, showing how a visual format can fit a real district screening workflow. For Francophone, Anglophone, and allophone learners, pattern reasoning may be easier to demonstrate through a matrix than through a vocabulary-heavy task.

The format can also help when expressive language does not reflect classroom problem-solving. A California research record describes Raven-based comparisons across White, Black, and Chicano students in grades K through 8, illustrating its use for examining nonverbal reasoning in varied school populations (ERIC record).

When corroboration is essential

Raven alone cannot support a high-stakes conclusion when:

  • Language exposure is uneven: The student recently entered the instructional language or has experienced interrupted schooling.

  • Test conventions are unfamiliar: The person needs repeated clarification about choosing and recording answers.

  • Visual access is uncertain: Uncorrected vision, visual-perceptual difficulty, or motor demands may affect responses.

  • The result conflicts with history: Classroom work, adaptive functioning, or other assessments show a substantially different pattern.

  • The decision is broad: Placement, diagnosis, or intervention planning also requires verbal, memory, attention, achievement, or executive-function information.

A high score may say little about practical judgment, communication, or day-to-day independence. A low score may reflect fatigue, distress, unfamiliarity with the format, or a poorly matched version rather than limited reasoning alone.

For decisions involving language, consult the language of assessment guide and record why the selected format was appropriate. Fairness depends on the person, instructions, norms, testing conditions, and decision being made.

Choosing Between Raven and Other Cognitive Assessments

Raven is efficient when the referral question is narrow: How well does this person solve novel visual pattern problems with limited language demand? It becomes insufficient when the question concerns a full cognitive profile.

A language-loaded IQ assessment can provide information about verbal comprehension, working memory, processing speed, and other domains that Raven doesn't sample. It may be the better choice when language itself is central to the referral, provided the examiner accounts for bilingualism, language proficiency, and educational history.

Digital cognitive platforms offer another option for broader screening and monitoring. Orange Neurosciences describes cognitive skills assessments covering attention, memory, executive function, perception, processing speed, and eye-hand coordination, with multilingual delivery and reporting designed for repeat evaluations. Its examples of cognitive assessment help illustrate how a multi-domain profile differs from a single matrix score.

A comparison chart showing features of Raven, WAIS, MMPI, and MoCA cognitive assessment tests.

For families, a parent-friendly explanation of the broader evaluation process is available in this Children Psych parent-friendly guide. It can help parents understand why one nonverbal score shouldn't carry the entire interpretation.

Use this checklist before selecting a measure:

  • Match the referral question: Choose Raven for nonverbal fluid reasoning, not as a complete cognitive profile.

  • Verify the norms: Confirm that the norm group and edition fit the person's age and background.

  • Check administration conditions: Record language, setting, vision, familiarity, fatigue, and assistance.

  • Corroborate high-stakes decisions: Add achievement, language, adaptive, or broader cognitive evidence when placement or diagnosis depends on the result.

  • Plan repeat testing carefully: Don't assume paper, clinical, and adaptive scores are directly interchangeable.

Orange Neurosciences offers multilingual digital cognitive assessments that profile several domains in an under-30-minute workflow, alongside reporting and reassessment tools for clinical and educational settings. If Raven is answering only one part of your referral question, visit Orange Neurosciences to explore a broader assessment and monitoring option.

Orange Neurosciences' Cognitive Skills Assessments (CSA) are intended as an aid for assessing the cognitive well-being of an individual. In a clinical setting, the CSA results (when interpreted by a qualified healthcare provider) may be used as an aid in determining whether further cognitive evaluation is needed. Orange Neurosciences' brain training programs are designed to promote and encourage overall cognitive health. Orange Neurosciences does not offer any medical diagnosis or treatment of any medical disease or condition. Orange Neurosciences products may also be used for research purposes for any range of cognition-related assessments. If used for research purposes, all use of the product must comply with the appropriate human subjects' procedures as they exist within the researcher's institution and will be the researcher's responsibility. All such human subject protections shall be under the provisions of all applicable sections of the Code of Federal Regulations.

© 2026 by Orange Neurosciences Corporation