Separate Face Blindness From Forgetting Famous Names
Score the 18 stronger PI20 items, examine why two questions are weak, and learn when face-recognition trouble—not name recall—needs assessment.

Two of the 20 questions in the standard PI20 prosopagnosia screening test add little useful information. In a study of 599 Japanese adults, items 3 and 13 showed comparatively weak, overlapping response patterns; the other 18 formed the stronger version of the questionnaire (Behavior Research Methods study). That does not invalidate the PI20, but it does mean every item should not be treated as equally informative—and a result influenced by those two questions deserves caution.
Prosopagnosia, often called face blindness, concerns recognizing facial identity. Looking at Taylor Swift, knowing who she is and blanking on her name is a name-recall failure. Looking at a familiar person and experiencing the face as unfamiliar until they speak is more relevant to face-recognition screening.
Why The Full 20-Item PI20 Became The Standard
The received view has a sound basis. The PI20 is a brief self-report questionnaire designed to capture lifelong face-recognition problems in ordinary life. Its 20 statements cover recurring failures, compensatory strategies and perceived ability. Oppositely worded items are reverse-scored so every response contributes in the same direction.
That standard total is easy to interpret. PI20 scores run from 20 to 100, and a score above 65 has been used to flag possible developmental prosopagnosia. It is a screening threshold, not a diagnosis (published PI20 adaptation).
The questionnaire also measures something a laboratory task can miss: years of recognizing people at work, in shops, after haircuts and outside their usual context. Treating all 20 items consistently made sense while the questionnaire was being used and validated as a complete scale.
The consensus is therefore right about the PI20’s broad value. It is a useful symptom screen, and the 2026 findings do not justify discarding it. What the new item analysis challenges is the narrower assumption that every question contributes comparable screening information.
Items 3 And 13 Are The Weak Links
The 2026 study, published in Behavior Research Methods on August 12, examined PI20 responses using classical test theory and item response theory. Classical test theory considers how the scale performs overall. Item response theory examines how well individual questions and response options distinguish between modeled levels of difficulty.
Most PI20 items showed orderly changes as reported face-recognition difficulty increased. Items 3 and 13 did not. Their response patterns were comparatively weak and overlapping, according to reports on the study from News-Medical and Medical Xpress.
One flagged item concerns recognizing people who have distinctive facial features. That is a weak discriminator because unusual faces may be memorable to people across the ability range. The other concerns recognizing your own face in photographs. Most people with developmental prosopagnosia do not report consistent difficulty recognizing themselves, so agreement with that statement does not track the broader condition particularly well (Behavior Research Methods study).
The remaining 18-item version showed excellent reliability and validity under both statistical approaches. The supplied reports do not provide the numerical reliability estimates or new score bands, however. Nor do they establish a validated diagnostic cutoff for an 18-item total.
The defensible verdict is narrower than saying the original PI20 is wrong. Items 3 and 13 can add noise and may push some totals higher or lower without adding much discrimination. The study did not report how many people would change screening category after removing them, so it does not establish a population-wide overcount rate.
Score The 18 Stronger Items Without Inventing A New Cutoff
The original questionnaire’s exact wording, scoring directions and validated translations should come from a credible PI20 source. The tool below therefore does not reconstruct the statements. It lets you transfer the already direction-scored values for the 18 retained item numbers from a PI20 you have completed.
Each value should run from 1 for the lowest reported difficulty to 5 for the highest after any required reverse-scoring. The raw 18-item range is 18 to 90. For comparison only, the calculator converts that result to the original 20–100 scale by multiplying the 18-item total by 20 and dividing by 18.
That proportional result is not a newly validated cutoff. It is shown beside the established original flag of above 65 so you can see whether the two weak items materially change which side of that historical line your answers occupy.
Transfer your direction-scored PI20 responses; the result shows whether items 3 and 13 change the screening side.
Enter already direction-scored values: 1 means lower reported face-recognition difficulty and 5 means higher difficulty. Neutral is 3.
These are all original PI20 item numbers except the two flagged by the study.
Item 3 concerns distinctive facial features. Item 13 concerns recognizing your own face in photographs. Add their direction-scored values to compare the original 20-item total.
The proportional estimate is 54 × 20 ÷ 18 = 60.0. This comparison does not create or validate an 18-item clinical cutoff.
An all-neutral response produces an original total of 60 and an 18-item proportional estimate of 60, both below the historical flag. That illustrates why occasional name blanks or uncertainty about a few faces should not automatically become “face blindness.” A persistent pattern of failing to recognize familiar people matters more than a label generated from one quiz.
A PI20 Result Is A Symptom Report, Not A Diagnosis
Self-report can capture experiences that no short performance task reproduces. It can also be influenced by self-awareness, interpretation, memorable mistakes, translation and response style. That is why a high PI20 result means that your reported experiences resemble the difficulties the questionnaire was designed to flag; it does not confirm developmental prosopagnosia.
Language versions matter as well. A Mexican Spanish adaptation evaluated in 333 adults reported internal consistency of .84 and test-retest reliability of .81 (published adaptation). Those results support that adaptation in the population studied. They do not validate every informal translation or establish identical performance in every Spanish-speaking population.
The 2026 analysis also has boundaries. Its 599 participants were Japanese adults, so the results do not by themselves establish performance across every language, culture, age group or clinical population. News-Medical and Medical Xpress summarize the same underlying paper rather than two independent replications. The released data and analysis code allow scrutiny and future work, but open materials do not themselves prove that a revision will classify people more accurately.
There is no universally agreed single test or cutoff that establishes prosopagnosia. A peer-reviewed review therefore recommends structured assessment rather than reliance on one task (review of prosopagnosia testing).
Face Recognition And Name Recall Split Before The Score
The most useful distinction happens before any questionnaire. Ask whether the face itself supplied identity.
If a face feels familiar and you know the person is an actor from a particular film but cannot retrieve the name, facial identity has been recognized. If the person remains visually unfamiliar until their voice, clothes, gait or location gives them away, the failure is more directly relevant to prosopagnosia.
Recurring examples include walking past a close friend, recognizing a colleague only at their usual desk, losing track of screen characters who share similar hair or clothing, or needing someone to speak before identity becomes clear. Occasional mistakes in poor lighting, at a distance or after a dramatic appearance change occur in people with otherwise typical recognition. The stronger signal is a persistent, disproportionate pattern involving familiar people and meaningful social or practical consequences.
This also defines what a celebrity-name game can and cannot test. In the NameIt timed famous-name game, players begin with 30 seconds and gain two seconds for each accepted name. Because no faces are presented, the game measures retrieval of famous names rather than recognition of facial identity. It has no clinical validation, diagnostic threshold, sensitivity or specificity for prosopagnosia.
Performance Tests Answer A Different Question
A questionnaire reports lived experience. A controlled task measures performance under specified conditions. Neither result replaces the other.
| Tool | What It Measures | Result | Main Limitation |
|---|---|---|---|
| PI20 | Lifelong everyday difficulty | 20–100; above 65 has been used as a flag | Self-report, with two weak items identified |
| CFMT | Learning and recognizing unfamiliar faces | Implementation-specific performance score | Still images and test strategies may not reflect daily life |
| UNSW Face Test | Learning, recognition and comparison across photos | Comparison with other participants | No clinical cutoff is supplied on the university page |
The Cambridge Face Memory Test is one of the best-known research measures of unfamiliar-face memory. A version described by Face Blind UK takes about 20 minutes and recommends a PC rather than a phone or tablet. Its still, black-and-white faces may not represent moving and expressive faces, and some participants may solve items by memorizing an eyebrow or another isolated feature.
The UNSW Face Test samples learning, recognition and comparison across changing photographs or viewing conditions. It requires a desktop or laptop and compares performance with other participants. The supplied university page does not identify a validated clinical cutoff or describe its reference sample in enough detail to treat that comparison as diagnostic.
Results can disagree because these tools measure different constructs. In one self-referred cohort, 56% of participants who believed they had prosopagnosia did not meet the study’s CFMT criterion. The excluded group nevertheless reported substantial symptoms and scored below average on some other face measures (peer-reviewed study).
That figure is not a universal 56% false-negative rate. The participants were self-referred, not independently confirmed cases that the CFMT subsequently missed. The finding supports the narrower point that CFMT-only classification can exclude people with meaningful symptoms.
A high PI20 result with typical objective performance may reflect difficulties in dynamic, familiar-face or real-world situations that the task did not sample. Poor objective performance with little everyday difficulty may warrant checking attention, vision, device conditions, demographic familiarity and comparison norms. A normal result on either type of test should not erase persistent, consequential experiences.
Lifelong And Newly Acquired Problems Need Different Responses
Developmental prosopagnosia is generally lifelong, although compensatory habits may keep it unnoticed until adulthood. Acquired prosopagnosia begins after previously typical face recognition and can follow neurological injury or disease.
Sudden, new, progressive or worsening difficulty recognizing faces requires prompt medical evaluation rather than repeated online testing. Assessment may need to investigate an acquired neurological cause, particularly when the change begins rapidly or after a head injury or other neurological event (Cleveland Clinic guidance).
For a lifelong, stable concern, record concrete incidents: whom you failed to recognize, whether the setting was expected, which cue finally revealed the identity, how often it happens and what effect it has on work or relationships. A clinician or research assessment can then compare that history with more than one face-recognition measure and consider vision, attention, general memory, broader cognition and object recognition.
The 18 stronger PI20 items offer a cleaner view of reported face-recognition difficulty than assuming all 20 questions are equally informative. They still cannot decide on their own whether someone has developmental prosopagnosia. The decisive distinction remains whether identity fails at the face—or whether the face is recognized and only the name refuses to arrive.