
Facial Comparison: What Can the Images Support?
A grainy CCTV face and a clear reference image can look easy to compare. Research shows why unfamiliar-face comparison is difficult, why experience alone does not establish skill, and why the field cannot yet provide a measured method error rate. This reading explains the evidence and the limits an examiner should be ready to discuss in court.
Recognition and comparison are different tasks
You can recognise a close relative across a crowded street in poor light and from an awkward angle. Facial comparison asks a different question: do two images show the same person when the examiner is unfamiliar with that person?
Vicki Bruce, Zoe Henderson, Craig Newman and Mike Burton tested the difference in 2001. They filmed University of Glasgow staff with a building-entry CCTV camera, then placed each video beside a high-quality photograph. Viewers could pause and replay the clip, both images remained in front of them, and there was no time limit.
Glasgow participants who knew the staff scored about 92% correct, with discriminability (A prime) of 0.93. Participants from the University of Paisley, who knew none of the faces, scored 70% overall, with an A prime of 0.75. On the hardest trials, where the video target was paired with a similar-looking stranger, the unfamiliar group scored 56%. Performance was only six percentage points above chance despite the favourable viewing conditions.
Familiar viewers can recognise identity rather than relying only on a comparison of visible features. Bruce and colleagues addressed the courtroom implication directly. Comparing a defendant with CCTV of an unfamiliar person is a practice that "should be avoided, as resemblances between images of otherwise unfamiliar faces can be misleading and such judgments are highly prone to error."
Jenkins and colleagues demonstrated the same divide in 2011 with two Dutch television celebrities. They collected twenty internet photographs of each celebrity and shuffled the forty images together. British students, who did not know the faces, were asked to sort the photographs by identity. Although the images showed only two people, the students produced a median of seven and a half identities. None of the twenty students sorted them into the correct two groups.
Dutch viewers knew both celebrities. Given the same forty photographs, they sorted every image into the correct two identities. Familiarity, rather than image quality or available time, accounted for the difference.
Photographs of one person can vary more than photographs of two different people. Lighting, expression, camera position and age can alter appearance substantially. Jenkins's team concluded that face photographs are "unsuitable as proof of identity," because variation within one person's photographs was large compared with variation between people.
An examiner comparing CCTV with a reference image is usually performing the unfamiliar-face task. The examiner does not know either person, and the images often differ in quality, angle and lighting. The research shows why that task needs more caution than everyday recognition.
“Photos of the same face were often deemed too dissimilar to go together, leading participants falsely to fractionate a single identity into several identities.”

The hard direction
Counsel turns from the CCTV still to the examiner, calm and deliberate.
“You had never met the person in the CCTV or the defendant before this case, so this was an unfamiliar-face comparison. I’m instructed that untrained viewers can perform close to chance on difficult comparisons of that kind. What evidence shows that your method and training improve on that?”
A job title does not prove accuracy
David White and colleagues tested 30 Sydney Passport Office staff in 2014. The officers averaged eight and a half years in the role, and all but three had completed the office's identity-verification training module. Their work required them to decide whether the person in front of them was the person shown in a photograph.
On a live person-to-photo test, the officers made errors on 10% of decisions. They accepted 14% of fraudulent photographs as authentic, even though the impostors had been selected from a small, diverse student pool rather than chosen for a close resemblance.
White's team also compared the officers with first-year university students on a photo-to-photo task. There was no significant main effect of group. On the Glasgow Face Matching Test, officers scored 79.2%, which was statistically indistinguishable from the published norm of 81.3%. Length of service did not predict accuracy.
White, Towler and Kemp later reviewed 29 professional-versus-novice comparisons from 12 papers involving more than 1,600 practitioners. Twelve comparisons found no significant difference. Bank tellers and notaries in Papesh's study made errors on 25% of trials. Police officers in Towler's training study were less accurate than students. Passport officers reviewing face-recognition candidate lists made errors on about half of the trials and selected the wrong face 40% of the time. Across the review, the group described as "facial reviewers" exceeded novice performance by about 1.5 percentage points.
Other specialist groups performed much better. Phillips and colleagues tested 57 trained forensic facial examiners from five continents in 2018. Participants received difficult image pairs and up to three months to complete the work. Examiners reached a median AUC of 0.93, compared with 0.68 for students. In White, Towler and Kemp's review, forensic examiners exceeded novices by 13 percentage points and police super-recognisers exceeded them by 14 percentage points. These results support measurable expertise in tested examiners and super-recognisers.
The examiner result came with important limits. Performance was better with long study times than with short ones. The high-confidence false-alarm rate was about 0.9%, and some examiners scored below the median student. The best result in the Phillips study came from combining an examiner with the leading 2017 algorithm; that pairing outperformed two examiners working together.
The evidence therefore supports ability that has been measured in particular people under particular conditions. A job title, years of service or routine exposure does not establish that ability. An examiner who relies on specialist status should be able to explain how their performance has been tested on a relevant task.
“Trained passport officers also perform poorly when matching unfamiliar faces. High error rates were consistent across three tests... length of time employed as a passport officer did not predict accuracy.”

Better than a stranger off the street?
Counsel sets down the witness CV and asks why the qualification should carry weight.
“You hold a recognised qualification and you compare faces every day at work. On what evidence should this court treat your conclusion as more reliable than that of an untrained member of the public, given that passport officers with eight years' experience perform no better than first-year students?”
A face-recognition result is an investigative lead
Randal Quran Reid was pulled over outside Atlanta in November 2022 and arrested on theft warrants from Louisiana, a state he had never visited. He spent six days in the DeKalb County jail. A Jefferson Parish detective had submitted a store-surveillance face to Clearview AI, and the search returned Reid. The warrant affidavit described the identification as coming from a "credible source" without explaining that the source was an algorithm. Reid was released after his lawyer identified a mole on Reid's face that the thief did not have.
Detroit police arrested Porcha Woodruff for carjacking in 2023 while she was eight months pregnant. A face-recognition search had returned 73 candidates from a database of millions of mugshots. Woodruff's old arrest photograph appeared in the results. Police placed her in a six-person lineup, and the complaining witness selected her. As the ACLU's brief explains, the algorithm first selected a person who resembled the suspect and the lineup then directed the witness back to that candidate. The two steps did not provide independent corroboration.
Two features of the process matter to an examiner.
First, algorithmic error is not distributed evenly across demographic groups. Buolamwini and Gebru's 2018 Gender Shades audit tested three commercial gender classifiers. Error rates reached 34.7% for darker-skinned women, compared with a maximum of 0.8% for lighter-skinned men. Gender classification is not identification, but the result showed why a single accuracy figure cannot describe performance across groups.
NIST's later demographic testing found the same problem in face recognition. The figures cited in the Woodruff brief report that some algorithms misidentified Asian and African American faces up to 100 times more often than white men. In some tests, older Black women were more than 3,000 times more likely to produce a false positive than younger Eastern European men. Howard and colleagues (2019) also found that comparisons within the same race, gender and age group raised the false-match rate by more than 400 times compared with cross-group comparisons.
Second, a candidate list contains people selected because they resemble the probe. Parasuraman and Manzey describe automation bias as a shift from vigilant analysis towards deference to computer advice. The reviewer does not see the other near-matches discarded by the search, and human operators selecting from candidate lists make errors on roughly half of trials. The ranked list can therefore influence the human review rather than provide an independent result.
Clearview CEO Hoan Ton-That said in 2023 that an arrest should never rest on a face-recognition search alone. The result is an investigative lead that requires independent inquiry. It is not an identification by the examiner who later reviews it.
“Even if Clearview AI came up with the initial result, that is the beginning of the investigation by law enforcement to determine, based on other factors, whether the correct person has been identified.”

Two identifications, or one?
Counsel frames the search and the review as if they were two separate witnesses agreeing.
“The algorithm returned my client's photograph, and you later concluded that the images showed the same person. Why should the court treat those as two independent identifications?”
What you know before you compare
Itiel Dror and David Charlton tested the effect of context on fingerprint decisions in 2006. They returned previous casework comparisons to the same experts several months later. This time the experts were told that the prints came from the Madrid bombing case in which the FBI had made an erroneous identification. Several experts changed their previous decisions even though the prints had not changed. The different case information produced a different assessment of the same material.
Rebecca Heyer and Carolyn Semmler applied the forensic confirmation-bias framework to facial image comparison in 2013. They noted that ground truth is usually unknown, there is no statistical basis for quantifying a facial-comparison conclusion, and examiners may receive contextual information before or during the examination.
Heyer and Semmler also reported results from 149 experienced facial-comparison specialists in Australian government agencies. The group was well calibrated on average, with confidence tracking accuracy. Individual performance varied widely. Some specialists were 20% more confident than accurate, while others were 20% less confident. The group average did not describe the calibration of any particular examiner.
Dror, Wertheim, Fraser-Mackenzie and Walajtys studied ranked candidate lists in the casework of 23 court-qualified fingerprint examiners in 2012. The examiners averaged more than 19 years of experience and did not know they were participating in a study. The dataset covered 3,680 candidate lists and 55,200 comparisons.
The researchers moved the mated print to different positions in the algorithm's ranked list. False identifications clustered in positions one and two, including cases where the mated print appeared lower in the same list. Examiners also spent less time on lower-ranked candidates and made more missed identifications there. The position assigned by the system influenced the examiner's decision.
A face-recognition candidate gallery creates the same concern because rank one arrives with an implied recommendation. Heyer and Semmler note that facial-review performance drops as candidate lists grow. They call for in-casework testing of the kind used by Dror's team so the effect can be measured in facial comparison rather than assumed from another discipline.
Heyer and Semmler recommend procedural controls. In a linear examination, the examiner evaluates and records the questioned face before placing it beside the reference image. Some agencies separate the evaluation and comparison stages between specialists. Linear sequential unmasking applies the same principle to case information by delaying exposure until the examiner has committed the initial assessment.
Awareness of bias does not provide the same protection. Pronin's work on the bias blind spot shows that learning about bias can increase confidence that it has been overcome without reducing exposure to it.
In court, the relevant questions concern sequence. Did the examiner learn the suspect's identity, arrest status, algorithmic rank or investigator's theory before forming the conclusion? The report should disclose that sequence and any procedure used to control the influence of context.
“These findings are not a function of the print itself; the same print is considered differently when presented at a lower position on the ranked list.”

What did you know first?
Counsel sets aside the comparison itself and asks only about the order of events.
“Before you concluded that these images showed the same person, what had you already been told: the suspect's identity, the algorithm's rank-one candidate, the fact of an arrest, or the investigating officer's theory of the case?”
The missing method error rate
A cross-examiner can ask for the measured error rate of the method used in the case. For forensic facial comparison as practised, no method error rate has been measured. Examiner-performance studies cannot supply that figure unless they tested the same procedure under conditions relevant to the case.
Photoanthropometry, also called facial mapping, marks anatomical landmarks, measures the distances between them and converts those distances into proportionality indices. Reuben Moreton and Johanna Morley, then at the Metropolitan Police Service, tested the method in 2011. They selected 25 people from the Home Office multipose database, each photographed at high resolution from 20 camera angles.
Every proportionality index changed significantly when the camera moved 10 degrees vertically. Moreton and Morley concluded that variation caused by camera angle within one person's images could be as large as variation between different people. On that evidence, photoanthropometry as practised was unsuitable even for elimination.
Image quality added another problem. In Moreton and Morley's CCTV footage, two people shared an identical nose-width index at 2.5 metres. At 5 metres, four people shared the same value. As resolution fell, measurements from different faces became indistinguishable. The authors described the comparison as subjective and noted that it provided no estimates of error.
The morphological method compares features such as the eyes, ears, nose and chin, guided by FISWG and OSAC documents. FISWG's 2010 training guideline tells trainees to be "aware of available and relevant statistics regarding facial shapes and relative frequency of occurrence." For most features, those frequency statistics do not exist. Without them, an examiner cannot quantify how many other people may share the observed features.
OSAC 2022-S-0008 requires a facial-comparison report to disclose the "absence of citable empirical measures of performance" in section 4.3.6.1. The document is a reporting standard, not a validation study. It requires the examiner to record that no empirical performance figure is available.
The National Research Council in 2009, PCAST in 2016 and the UK Forensic Science Regulator have all called for validity and error rates established through empirical testing. That work has been done for DNA. For forensic facial comparison, the method-level evidence remains incomplete. An examiner asked for a method error rate should say that none has been measured rather than borrowing a figure from a different task or study.
“Absence of citable empirical measures of performance.”

Where is your number?
Counsel asks the validity question that has nothing to do with whether the examiner got this case right.
“You concluded that these images show the same person. What is the validated error rate of the method you used, and what published study established it?”
What does "strong support" measure?
Robert Neave had worked in facial comparison for about twenty years when he gave evidence in R v Atkins. He was a medical artist at Manchester with forty years of experience. Neave spent sixteen hours comparing indistinct CCTV from a west London armed robbery with Dean Atkins. He also excluded about twenty known burglars and Atkins's brother.
Neave used a scale of his own design. Level 0 "lends no support" and level 5 "lends powerful support." He placed the case between the top of level 3 and level 4, moving from "lends support" towards "lends strong support." He accepted that no database supported the levels; they reflected his experience. The Court of Appeal admitted the evidence in 2009 on the condition that the jury was told the labels were subjective.
R v Tang had described such a scale as "no more than a series of convenient labels, arranged in an ascending hierarchy, that state a conclusion." In R v Gray, Mitting J said that without a national database or agreed formula, the degree of support "must be only the subjective opinion" of the witness, adding that "this court doubts whether such opinions should ever be expressed." Atkins allowed the scale but required its subjective basis to be explained. A phrase such as "strongly supports" therefore describes the examiner's judgement rather than a measured probability.
Kristy Martire and colleagues tested how jurors interpret this language in 2013. They gave 494 mock jurors the same larceny case with forensic evidence pointing towards the defendant. The researchers changed only the way the expert expressed evidential strength: either as a number or as a corresponding verbal label such as "weak or limited support" or "strong support."
Numerical strength affected the direction of jurors' decisions. The verbal labels did not behave as intended. When the expert said the evidence lent "weak or limited support" to the prosecution, 61.7% of jurors moved towards innocence. Nearly a quarter changed from "more likely guilty" to "more likely not guilty" after receiving evidence that favoured the prosecution. Martire's team called this the weak evidence effect.
Jurors also struggled to distinguish large numerical differences. Evidence expressed as 495,000 times more likely if the defendant was the source produced about the same average change in belief as a figure of 1.4. Responses to "450 times more likely" and "495,000 times more likely" were not statistically different, despite the figures being more than a thousandfold apart.
Edmond, Biber, Kemp and Porter identified the underlying problem in 2009. Facial mapping had not been validated, and no study had measured examiner hit rates and false-positive rates. In Tang, the court had treated reliability as "an extraneous idea." A verbal scale could therefore order the examiner's conclusions without measuring the strength of the evidence.
If a validated error rate exists for the method as practised, the examiner should state it. For facial comparison, no such method error rate has been measured. The examiner should describe "strong support" as a judgement based on experience rather than a probability. The weak evidence effect also makes it necessary to explain what the phrase is intended to convey and what it cannot quantify.
“A majority of those in the low/verbal condition (61.72%) responded in a manner incongruent with the evidence provided by the expert (taking inculpatory evidence to be exculpatory).”

Each phrase claims more certainty than a method with no measured error rate can support, or may be interpreted differently from what the examiner intended. Replace it with language that states the limit. Grounded in the facial-comparison research.
- 01Identify the task you performed. Recognising someone you know is easier and more reliable than deciding whether two images show the same unfamiliar person, even with both images side by side and unlimited time.
- 02A job title is not evidence of skill. Passport officers averaging eight years in the role performed no better than first-year students. Tested forensic examiners and super-recognisers outperformed novices under the study conditions, so any claim to added expertise should rest on measured performance rather than position or years of service.
- 03Treat a face-recognition hit as a lead to investigate, never as an identification. The software's error rates differ by race, sex and age, its candidate list is by design a gallery of look-alikes, and a human reviewing its top pick is not an independent second opinion.
- 04State the order in which you received case information. If you saw the suspect's name, the algorithm's first pick or the investigator's theory before comparing the images, that context may have shaped the conclusion. A linear procedure records the questioned image before revealing the reference or case theory.
- 05No method error rate has been measured for forensic facial comparison. Facial measurements vary with camera angle, feature-frequency statistics are unavailable, and the field's reporting standards require disclosure that empirical performance figures are absent.
- 06"Strong support" is a phrase rather than a measurement. Courts treat it as personal judgement, and research shows that jurors may interpret "weak support" for the prosecution as a point for the defence. Explain the basis and limits of the phrase.
What does "strong" rest on?
You are on the stand. Counsel has saved the conclusion language for last.
“You told this jury the similarities offer 'strong support' that the man on the camera is the defendant. What is the validated error rate behind that phrase, and if you have none, can you rule out that some of these jurors will read 'support' as helping my client rather than the Crown?”
Still have questions about the research?
Ask anything about the forensic facial-comparison research. The tutor answers from the document itself — and keeps one eye on how it might come up under cross-examination.
- Bruce, V., Henderson, Z., Newman, C., & Burton, A. M. (2001). Matching identities of familiar and unfamiliar faces caught on CCTV images. Journal of Experimental Psychology: Applied, 7(3), 207–218.
- Jenkins, R., White, D., Van Montfort, X., & Burton, A. M. (2011). Variability in photos of the same face. Cognition, 121(3), 313–323.
- White, D., Kemp, R. I., Jenkins, R., Matheson, M., & Burton, A. M. (2014). Passport officers’ errors in face matching. PLoS ONE, 9(8), e103510.
- White, D., Towler, A., & Kemp, R. I. (2021). Understanding professional expertise in unfamiliar face matching. In M. Bindemann (Ed.), Forensic Face Matching: Research and Practice. Oxford University Press.
- Phillips, P. J., Yates, A. N., Hu, Y., Hahn, C. A., Noyes, E., Jackson, K., ... & O’Toole, A. J. (2018). Face recognition accuracy of forensic examiners, superrecognizers, and face recognition algorithms. Proceedings of the National Academy of Sciences, 115(24), 6171–6176.
- Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 1–15.
- Grother, P., Ngan, M., & Hanaoka, K. (2019). Face Recognition Vendor Test (FRVT) Part 3: Demographic Effects. NIST Interagency Report 8280. National Institute of Standards and Technology.
- Howard, J. J., Sirotin, Y. B., & Vemury, A. R. (2019). The effect of broad and specific demographic homogeneity on the imposter distributions and false match rates in face recognition algorithm performance. 2019 IEEE 10th International Conference on Biometrics Theory, Applications and Systems (BTAS).
- Dror, I. E., & Charlton, D. (2006). Why experts make errors. Journal of Forensic Identification, 56(4), 600–616.
- Dror, I. E., Wertheim, K., Fraser-Mackenzie, P., & Walajtys, J. (2012). The impact of human-technology cooperation and distributed cognition in forensic science: Biasing effects of AFIS contextual information on human experts. Journal of Forensic Sciences, 57(2), 343–352.
- Heyer, R., & Semmler, C. (2013). Forensic confirmation bias: The case of facial image comparison. Journal of Applied Research in Memory and Cognition, 2(1), 68–70.
- Moreton, R., & Morley, J. (2011). Investigation into the use of photoanthropometry in facial image comparison. Forensic Science International, 212(1–3), 231–237.
- Martire, K. A., Kemp, R. I., Watkins, I., Sayle, M. A., & Newell, B. R. (2013). The expression and interpretation of uncertain forensic science evidence: Verbal equivalence, evidence strength, and the weak evidence effect. Law and Human Behavior, 37(3), 197–207.
- Edmond, G., Biber, K., Kemp, R., & Porter, G. (2009). Law’s looking glass: Expert identification evidence derived from photographic and video images. Current Issues in Criminal Justice, 20(3), 337–377.
- R v Atkins & Atkins [2009] EWCA Crim 1876 (Court of Appeal, England and Wales).
- National Research Council. (2009). Strengthening Forensic Science in the United States: A Path Forward. Washington, DC: The National Academies Press.
Voice on Trial: Forensic Speaker Comparison
Counsel is briefed on this literature. Take it into the witness box and practise facial comparison.