
Gait, Body and Clothing Comparison
A CCTV comparison may combine gait, build, clothing and estimated height. Each component has different sources of variation, and none has a validated casework error rate for source identification.
Four comparisons, no casework error rate
A low-frame-rate CCTV clip may lead to comparisons of gait, build, clothing and estimated height. Presenting four observed similarities together may sound like independent confirmation, but the discipline still lacks a validated casework error rate for a source conclusion.
Macoveciuc, Rando and Borrion said in their 2019 review that, almost two decades after the first UK forensic gait evidence, the methods "remain insufficiently robust for use in court." They identified no ground-truth validation, standardised feature list, population-frequency data or finalised code of practice. They advised courts to treat gait evidence "with caution, as they should any other form of evidence originating from disciplines without fully established codes of practice, error rates, and demonstrable applications in forensic scenarios."
Gait evidence first appeared in R v Saunders in 2000. Macoveciuc and colleagues observe that "forensic gait analysis appears to have become a forensic science field because of this first case in which it was applied," rather than following validation and peer review. In R v Otway in 2011, the evidence remained admissible "despite being qualitative in nature, with no empirical support for the applied method and for the conclusions drawn."
The Royal Society's 2017 Primer for Courts classifies the support available from current gait methods as "weak" on a verbal likelihood scale, in part because there is no large UK population database. Van Mastrigt and colleagues (2018) reach a similar conclusion about validity. The issue is the scientific support for the method, not the examiner's personal competence.
In a Danish bank-robbery case, police considered that a disguised robber "had a unique gait." Larsen and colleagues analysed short footage in which the hands were in pockets and the feet were partly obscured. They reported a corresponding limp using variability data from eleven people, while stating that the analysis "cannot establish the identity of the individual." The defendant was convicted. The example shows why limitations should appear in the conclusion itself and why each component of a combined comparison needs separate support.
“Forensic gait analysis appears to have become a forensic science field because of this first case in which it was applied.”

Four matches, where is the number?
Counsel lists the four claims back to the witness, slowly, then asks for the one thing none of them carries.
"You told this jury the walk, the build, the clothing, and the height all 'match' the accused. For each of those four matches, what is your method's error rate, and where is it published?"
A feature that will not hold still
A feature used for identification must be measurable, reasonably stable within a person and sufficiently variable between people. Gait, build, height and clothing all present difficulties, beginning with the recording itself.
Birch et al. (Science & Justice, 2014) filmed one man wearing an ankle-foot orthosis at 25 frames per second alongside Qualisys motion capture, then reduced the footage to nine frame rates, down to one frame every four seconds. Twelve experienced podiatrists scored visible gait features. Accuracy increased strongly with frame rate (r = 0.868, p = 0.002), which explained about three quarters of score variation. The authors note that apparent movement in CCTV is reconstructed from still frames through persistence of vision and the phi phenomenon: "all perception of movement from CCTV footage is therefore illusionary." Submitted custody footage has been as low as 2 fps. At that rate, an arm swing or knee flexion may occur entirely between captured frames.
Scoleri, Lucas and Henneberg (Forensic Science International, 2014) asked three assessors to estimate stature from images of the same men shirtless and wearing a black shirt, striped shirt or padded leather jacket. Relative to the shirtless estimate, the padded jacket produced a 37.8 mm bias and the striped shirt 36.4 mm. In airport footage, technical error exceeded 60 mm for weaker assessors. Stature can also reduce by up to 28.1 mm during a day.
Thakkar, Pavlakos and Farid (CVPR Workshops, 2022) tested a 3D body model on simulated figures with known dimensions. Even with perfect scale, a person at 175 cm and 90 kg produced 95% ranges of 170–177 cm and 69–103 kg. With realistic scale estimation, performance was no better than predicting the sex-group average. Spinal compression may change height by up to 2 cm daily, weight may vary by 2.25 kg, and posture or walking may change apparent height by up to 6 cm. Changing from neutral pose reduced body-shape classification from 95% to about 46%.
The report should identify which observations describe recorded movement and which describe a body position captured in a frame. Clothing, posture, time of day and frame rate may change the apparent feature before any comparison between people begins.
“All perception of movement from CCTV footage is therefore illusionary, the brain making a series of assumptions as to the way in which an object or any part, thereof, gets from one location to another.”

Movement, or a frozen posture?
Counsel has already established the footage runs at 2 frames per second, and turns that fact on the gait opinion itself.
"You called this a gait feature, but you have agreed the footage runs at 2 frames per second. At that rate, the movement you described to the jury happens between the frames the camera recorded, so isn't your opinion really a description of a frozen posture passed off as motion?"
Two examiners, one walk, two answers
The Sheffield Features of Gait Tool contains 113 gait features and variants across 14 sections, developed from 51 forensic gait reports. In a 2019 study, fourteen experienced gait analysts used it to assess computer-generated avatars with fixed gait, clothing, high frame rate, good resolution and good lighting.
Within-examiner repeatability ranged from 68.35% to 94.65%, with a mean of 79.54%. Between-examiner reproducibility averaged 73.45%. The authors described these as "good" levels for an initial structured tool, comparable with clinical tools such as the Salford Gait Tool and Edinburgh Visual Gait Score. For forensic use, however, the figures show that a material proportion of feature ratings changed across examiners or occasions even under favourable conditions.
Birch's 2020 study compared eleven trained gait analysts with nineteen members of the public on the same CCTV identification task. Analysts were 71.64% correct and lay participants 64.42%; the difference was not statistically significant (p = 0.29). Lay participants were significantly more confident than experts on both correct and incorrect answers (p < 0.05). Confidence therefore did not provide a reliable indication of accuracy.
Nakhaeizadeh, Dror and Morgan (2014) gave 41 forensic anthropologists the same skeleton with different contextual information. With no context, 31% classified the remains as male. Told that DNA indicated male, 72% did so. Told that DNA indicated female, none did. The authors concluded that context can override "the actual physical evidence present."
Gait assessment remains subjective. Even a structured tool showed limited between-examiner reproducibility, and contextual information can affect related visual assessments. Independent procedures and careful qualification are therefore necessary.
“In the participant group where the context was that the remains were of a female, 0% of the participants concluded that the remains were male.”

Told the answer before you looked
Counsel ties together what the officer said beforehand and the absence of any reliability figure, and asks for a number.
"Mr Examiner, before you ever looked at the questioned footage, the officer told you the suspect had an 'asymmetrical limp,' didn't he? And you found one. Can you tell this jury, with a number, how reliably a second qualified examiner with no knowledge of the case would reach the same conclusion you did?"
A defensible height estimate is a range, not a point
Alberink and Bolck of the Netherlands Forensic Institute demonstrated a scene-specific method in a cinema-robbery case. Six people of known height were returned to the scene and positioned near the offender's location and pose. A 3D room model was built from photographs and fixed reference points. Four operators measured each person three times in random order.
The offender's mean measured image height was 166.4 cm. Measurements of test participants were systematically 6.3 cm below true height, with standard deviation 1.7 cm, giving a corrected estimate of 172.7 cm. Because there were six test participants, a Student's t distribution produced a 95% interval of 168–177.5 cm. The suspect, at 176 cm, fell within the interval. That means the height did not exclude him; it did not identify him.
Alberink and Bolck also calculated a likelihood ratio using population height and proximity to the suspect's height. The LR was about 2, which they described as "very weak evidence against the suspect." For common heights, the maximum obtainable LR was near 6. The method accordingly produced wide intervals and modest evidential support.
Edelman, Alberink and Hoogeboom later tested related methods using four offenders and a camera moved between the event and reconstruction. When the same camera correspondence was reused, projective geometry and 3D modelling performed accurately. When operators had to redraw vanishing points or camera correspondences, projective geometry failed: fourteen of sixteen test sets were rejected and some repeated intervals did not overlap. The 3D model was more stable, with one significant result among sixteen, consistent with chance at the 5% level. They note that lens distortion and changed camera orientation occur "very regularly," requiring validation experiments for the scene and setup.
A statement such as "about 180 cm, the same as the accused" omits reconstruction, test participants, uncertainty interval, population comparison, posture, footwear and possible camera movement. A defensible estimate reports a range, the correction and uncertainty used, and any LR that the validated method supports.
“This constitutes very weak evidence against the suspect.”

Where is your confidence interval?
Counsel lays the rigorous photogrammetric method beside the witness's casual point estimate and asks for everything the loose version skipped.
"You told the jury the figure was 'about 180 centimetres, the same height as the accused.' Did you return to the scene, record people of known height where the offender stood, and correct for posture and footwear? What is the uncertainty interval for your estimate?"
What the “one in 1.27 billion” figure measures
The frequently cited one-in-1.27-billion figure for barefoot impressions comes from an RCMP study by Robert Kennedy and colleagues. It used 5,755 pairs of inked impressions and analysed 19 of 119 possible measurements.
Massey and Kennedy explain that impressions were measured and searched against the database. With exact measurements, every foot appeared new after three to five measurements, including repeat impressions from the same donor. The researchers therefore widened the tolerances. At 5 mm, searches still returned only same-donor pairs. At 15 mm, some different donors were returned, although an examiner could distinguish their impressions visually. The reported uniqueness therefore concerns failure to find another donor in that database under selected tolerances. It is not a casework error rate for attributing a crime-scene impression to a person.
The database used complete, distinct walking impressions made on paper by cooperative donors; smeared and unclear impressions were excluded. Kennedy acknowledged that crime-scene impressions may not meet those conditions, which limits use of the numerical figure in casework.
Hu, Arnold, Causby and Jones reviewed 1,340 records in 2018 and retained eleven studies. They rated overall quality from "Poor" to "Fair," with nine rated Poor. Two studies describing the Kennedy method were excluded because they did not report measurement reliability. The underlying data were unavailable for independent calculation. The reviewers concluded that "additional testing is required to determine the reliability of foot impression measurements informing this technique, as to date this has not been reported."
Nirenberg's 2016 case study describes the first Daubert hearing centred on forensic podiatry, State of Wisconsin v Travis Petersen. The evidence involved a sock-clad bloody footprint and eleven linear measurements within a recognised 5 mm tolerance. It was admitted. The tolerance concerned individual measurements, not a validated error rate for a common-source conclusion. The defence reported finding no published article refuting the method.
The important distinctions are the database against which uniqueness was assessed, the repeatability of the measurements and the absence of a casework source-attribution error rate.
“additional testing is required to determine the reliability of foot impression measurements informing this technique, as to date this has not been reported.”

A database that stopped, not an error rate
Counsel walks the witness back through how the "unique" number was actually produced, then asks for the real error rate.
"Your one-in-1.27-billion figure: that came from searching a curated database of cooperative donors and not finding a second match within a chosen tolerance, didn't it? So what is the measured error rate of your conclusion that this crime scene print and the defendant share a source?"
How verbal conclusions may be misunderstood
Daniel Aitken was convicted of first-degree murder in 2009 after a 41-day jury trial and 40 days of pre-trial argument. Podiatrist Haydn Kelly compared six seconds of night CCTV showing a shooter in loose track pants and slip-on sandals with covert footage of Aitken. He described the likeness as "very strong," the second-highest point on his scale below "extremely strong likeness." Cunliffe and Edmond examined the admission of that evidence in their 2014 Canadian Bar Review paper, "Gaitkeeping in Canada."
Kelly said he had not performed blind testing, had not published the method, and had not had a report verified by anyone "other than myself." When asked about error rate, he relied on "20-odd years of examining people's gait." He could not identify a database supporting a claim that features occurred in 1% of the population. Satanove J excluded the frequency claim but admitted the remainder. The BC Court of Appeal treated the evidence as "specialized knowledge gained through experience and specialized training," making peer review, error rate and validation of "limited relevance." Chin and Dallen later criticised this as "a dangerously facile approach towards scientific evidence: admitting it without scrutiny by dressing it up as specialized knowledge." Nirenberg, Vernon and Birch (2018) discuss a similar approach in R v Otway, where experience was accepted as a basis for judgement.
Cunliffe and Edmond note that jurors do not interpret expressions such as "very strong," "lends support" and "consistent with" in a uniform way. Drawing on Martire and colleagues, they discuss the weak-evidence effect, in which verbal evidence may even shift belief in an unintended direction. The passage from R v DD quoted by the court warns that jurors may "abdicate their role as fact-finders and simply attorn to the opinion of the expert."
There is no validated casework error rate for gait, body or clothing source comparison. Nirenberg and colleagues say that, without population databases, inferences are "currently limited to the experience" of the analyst. The conclusion should therefore state the observed similarities, recording conditions and method limitations, and should not be framed as identification.
“a dangerously facile approach towards scientific evidence: admitting it without scrutiny by dressing it up as specialized knowledge”

Each phrase below is one a gait, body or clothing examiner might offer, and each claims more than an unvalidated visual comparison can carry. Swap it for language that states the observed similarity and its limits, so the jury cannot run it to identification.
- 01"Match" on gait, build, clothing or height is a feature comparison with no validated casework error rate. Roughly two decades in, the discipline still cannot give you the number a court will ask for.
- 02The feature will not hold still. Frame rate, garment, posture and time of day move it before you compare two people, and below a few frames a second the "movement" you describe is the brain filling gaps the camera never recorded.
- 03The read does not reproduce. Even a 113-feature structured tool on pristine avatars left about a quarter of the score unstable, trained experts barely beat members of the public, and the most confident examiner is often the wrong one.
- 04Context bends the answer. Tell an anthropologist the sex of a skeleton and the call swings from 0% to 72%. Know what you were told about the case, and when you were told it.
- 05Height can be done rigorously, as a scene reconstruction with test persons, a wide confidence interval, and a small likelihood ratio. "About 180, the same as the accused" is a point estimate passed off as a match.
- 06A "unique" footprint usually means "I searched a curated database and did not find a second match within a chosen tolerance," not a measured error rate, and the measurements underneath were never shown to be reliable.
- 07Your words get misread. "Consistent with" can settle in a juror's head as near-certainty. State the limitation on the face of the conclusion, and put the brake on before the jury runs it to identification.
What does "consistent with" mean?
You are on the stand. Counsel saves the hardest question about your own words for last.
"You told this jury the gait is 'consistent with' the accused. Tell us the error rate for that conclusion in casework of this kind, and if you cannot, explain how the jury is supposed to know whether 'consistent with' means near-certain or close to worthless."
Still have questions about the research?
Ask anything about Forensic gait analysis and the limits of identifying a person from an image. The tutor answers from the document itself — and keeps one eye on how it might come up under cross-examination.
- Macoveciuc, I., Rando, C. J., & Borrion, H. (2019). Forensic gait analysis and recognition: standards of evidence in forensic gait analysis. Journal of Forensic Sciences, 64(5), 1294-1303.
- Birch, I., Vernon, W., Burrow, G., & Walker, J. (2014). The effect of frame rate on the ability of experienced gait analysts to identify characteristics of gait from closed circuit television footage. Science and Justice, 54(2), 159-163.
- Scoleri, T., Lucas, T., & Henneberg, M. (2014). Effects of garments on photoanthropometry of body parts: Application to stature estimation. Forensic Science International, 237, 148.e1–148.e12.
- Thakkar, K., Pavlakos, G., & Farid, H. (2022). The reliability of forensic body-shape identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 829-837.
- Birch, I., Vernon, W., Walker, J., Young, M., & Saxelby, J. (2019). The development of the Sheffield Features of Gait Tool for use in forensic gait analysis. Journal of Forensic and Legal Medicine, 67, 49-54.
- Birch, I., Raymond, L., Christou, A., Fernando, M. A., Harrison, N., & Paul, F. (2020). The identification of individuals by observational gait analysis using closed circuit television footage. Science and Justice, 60(3), 285-291.
- Nakhaeizadeh, S., Dror, I. E., & Morgan, R. M. (2014). Cognitive bias in forensic anthropology: visual assessment of skeletal remains is susceptible to confirmation bias. Science and Justice, 54(3), 208-214.
- Alberink, I., & Bolck, A. (2008). Obtaining confidence intervals and likelihood ratios for body height estimations in images. Forensic Science International, 177(2-3), 228-237.
- Edelman, G., Alberink, I., & Hoogeboom, B. (2010). Comparison of the performance of two methods for height estimation. Journal of Forensic Sciences, 55(2), 358-365.
- Hu, A., Arnold, J., Causby, R., & Jones, S. (2018). The reliability of measurements taken from podiatric data used in identification: a systematic review. Forensic Science International, 287, 71-81.
- Nirenberg, M. S. (2016). Forensic methods and the courts: a daubert ruling on forensic gait analysis. Podiatry Management.
- Cunliffe, E., & Edmond, G. (2014). Gaitkeeping in Canada: mis-steps in assessing the reliability of expert testimony. Canadian Bar Review, 92(2), 327-368.
- Nirenberg, M., Vernon, W., & Birch, I. (2018). A review of the historical use and criticisms of gait analysis evidence. Science and Justice, 58(4), 292-298.
- Royal Society. (2017). Forensic gait analysis: a primer for courts. London: The Royal Society.
Image Authentication: What Can the File Establish?
Counsel is briefed on this literature. Take it into the witness box and practise gait & body comparison.