On airBulletin №23 · The Dark Side of AI Companionship: A Cautionary Study|Up next · Rejected in 30 Seconds: When No Human Reads Your Application
From the journal

Too Average to Be Real: How Generated Faces Skew Police Lineups

Published
25 September 2026

Consider the most consequential three seconds in a criminal case. A woman looks up. There is a man in her kitchen doorway. The light is bad, the adrenaline is worse, and her visual system is doing something it evolved to do badly: encoding an unfamiliar face under threat. Three seconds later he is on her, and the encoding stops.

Six weeks after that, in a small room at a police station, she is shown six photographs on a screen. One of them is a man the police have arrested. The other five are what the science calls fillers: known-innocent people placed in the array so that the array functions as a test rather than a suggestion. She looks. She points. She says she is certain.

Everything downstream of that moment, the charge, the plea negotiation, the jury's verdict, the years, depends on what those other five faces were. And there is now a serious, funded body of work asking what happens when the answer is: nothing. Not photographs of anyone. Not people with addresses and dental records and mothers. Pixels sampled from a latent space in the time it takes to refill a coffee, resembling human beings the way a convincing forgery resembles a signature.

On 2 September 2026, East Texas A&M University announced that a team led by Curt Carlson, professor of experimental psychology, had won a 100,000 dollar Research Excellence Fund grant from The Texas A&M University System to find out. The project pairs conventional eyewitness identification experiments with Tobii eye-tracking glasses, in collaboration with Dawn Weatherford at Texas A&M University-San Antonio, Alyssa Jones at Tarleton State University and Maria Carlson at East Texas A&M. The intention is to record not just which face a witness picks but how the eye arrives there: dwell times, saccades between images, whether attention settles on internal features such as the eyes and nose or drifts to external ones such as hair. It underwrites a planned 1.3 million dollar National Science Foundation proposal for a three-year expansion.

Carlson's worry is the sort that only occurs to someone who has spent a career thinking about how lineups fail. “Having a real suspect's photo in there with AI fillers could be problematic,” he said. “Someone's got to do the research to see, wait a second, are you putting the suspect in a problematic situation there because he's in there with five AI-generated fillers?”

Put it plainly. If a machine-made face carries some residue of its manufacture, some statistical signature that a visual system registers without the person noticing, then the one authentic photograph in a six-pack becomes the odd one out. The witness does not consciously think “that one is real”. She simply finds her eye returning to it. And a lineup in which the suspect draws the eye for reasons unrelated to memory is not a test. It is a nudge with an evidentiary receipt.

What a lineup is for, and why fillers are the whole apparatus

The lineup is one of the few genuinely experimental instruments the criminal justice system operates, and almost nobody in the system treats it as one.

Its logic is diagnostic. Simply presenting a suspect and asking “is this him?” is a procedure with no error control at all: a witness motivated to help, primed by the fact that police have arrested someone, will often say yes. So the suspect is embedded among known-innocent fillers. If the witness picks the suspect, that choice carries information, because she could have picked five other people and did not. If she picks a filler, the police learn their witness is guessing, and the record shows it. Fillers are the control condition. They are, in a real sense, the entire scientific content of the procedure.

Which is why the discipline has spent five decades building measurement tools for filler quality. Malpass introduced effective size in 1981, using mock witnesses, people who never saw the crime but are given the verbal description, to estimate how many lineup members are plausible candidates. Tredoux later refined the statistics into the measure now known as Tredoux's E. Functional size counts how many mock witnesses pick the suspect from description alone; if the answer is materially more than chance, the lineup is biased and the identification is worth less than it looks. These are the reason a defence expert can run a mock-witness study on the actual array used and demonstrate that a six-pack had a functional size of two.

There is also a long-running methodological argument about how fillers should be chosen. Match-to-suspect selection picks fillers who look like the person in custody, which sounds fair and can backfire: make everyone a near clone and you suppress the witness's ability to discriminate the guilty from the innocent. Match-to-description selects fillers who fit the witness's original account, which preserves discriminability but risks the suspect standing out on any feature the description did not mention. The 2014 National Research Council report “Identifying the Culprit”, still the most authoritative synthesis of this science, worked through these trade-offs alongside recommendations that have since become reform orthodoxy: double-blind administration, unbiased pre-lineup instructions, and an immediate confidence statement taken in the witness's own words before anyone says a thing about whether she got it right.

Three years later the science was reframed by John Wixted and Gary Wells, who argued that eyewitness confidence does predict accuracy, but only under what they called pristine conditions. There are five. Only one suspect in the lineup. Unbiased instructions. Double-blind administration. Confidence recorded at the moment of the identification. And, first on the list, a fair lineup in which the suspect does not stand out.

That is the sentence to hold onto through everything that follows. The scientific case for trusting a confident eyewitness rests on a conditional, and the condition at the top of the list is filler quality. Break the fillers and the confidence-accuracy relationship does not merely weaken. It stops being interpretable.

The old failure mode, wearing new clothes

None of this is hypothetical. The Innocence Project reports that 62 per cent of the exonerations in its casework involved eyewitness misidentification, making it the single largest contributing factor in its data. The National Registry of Exonerations found mistaken witness identification implicated in 26 per cent of exonerations recorded in 2024 alone, across a much broader population of cases than DNA exonerations.

The case reformers keep returning to is Ronald Cotton's, because it demonstrates every mechanism at once. In July 1984, Jennifer Thompson was raped in her home in Burlington, North Carolina. She made a deliberate effort to memorise her attacker's face. She identified Cotton in a photo array and then in a live lineup and told the court she was completely certain. Cotton served more than ten years. In 1995 DNA testing excluded him and matched Bobby Poole, a man produced in court during Cotton's 1987 retrial and whom Thompson, looking directly at her actual attacker, told the court was not the man. The charges were vacated on 30 June 1995. Poole pleaded guilty on 11 July.

The point is not that Thompson lied or was careless. She was neither. The point is that a memory can be overwritten, that certainty can be manufactured by procedure, and that a witness's confidence at trial often measures everything that has happened since the identification rather than what happened during the crime.

Which brings us to why a suspect standing out matters so much, and why the reasons are usually stupid. Lineups have been criticised because the suspect's photograph had a different background colour, because his was the only image with a date stamp, because five fillers were shot on a modern booking camera and the suspect's came from a decade-old arrest at lower resolution, because his was the only face not smiling, because the compression artefacts differed. None of these bears on whether the man committed the crime. All of them can steer a hesitant witness towards the one image that feels different, and once she has chosen it, the feedback loop does the rest.

Synthetic fillers, if Carlson's concern proves out, would be a new species of exactly this old problem, with one alarming property the old versions lacked. Background colour and JPEG artefacts are visible to a defence lawyer with a printout. A subtle statistical divergence between a generated face and a photographed one, detectable by the visual system but not nameable by the person doing the detecting, would be a biasing feature that nobody in the room could see, including the officer who built the array in good faith and the witness whose eye it was steering.

Why departments want the machine faces anyway

It would be easy to write this as a story about police reaching for a shiny tool. It is not that story, and pretending otherwise misses what makes it hard.

Filler selection is genuinely difficult, genuinely laborious, and genuinely bad in a lot of jurisdictions. An investigator working from a description of a man with a distinctive facial scar, or an unusual combination of age, build and ethnicity, has to trawl a local mugshot database that may hold a few thousand images, most of them of the wrong demographic and years out of date. A department in a county whose arrest population is overwhelmingly of one ethnicity will struggle to construct a fair array for a suspect from a minority group, which is precisely the case in which fairness matters most. And investigators are not paid to be experimental psychologists.

Then there is the objection almost nobody outside the field raises, which happens to be the most interesting one. Every real filler photograph is a picture of a real person who was arrested for something at some point, and whose booking image is now being shown to a rape victim in connection with a crime he had nothing to do with. He did not consent. He was not asked. He will never be told. The research group at Heinrich Heine University Düsseldorf that published on AI-generated fillers in Scientific Reports in 2024 named this directly, arguing that generating fillers from text descriptions “avoids the violation of identity rights of natural persons who are not suspects” and removes the constraint of being tied to a limited database.

That is a reform argument, not a surveillance argument, and it deserves to be taken seriously by anyone inclined to reject synthetic fillers on reflex.

The Düsseldorf findings were also, on their face, encouraging. Their text-to-image lineups were, in the authors' words, perfectly fair, and produced less biased suspect selection than the database-derived filler photographs used in their earlier experiments. Bell and colleagues had, though, named Carlson's worry themselves before they went looking for it: the risk exists, they wrote, that using AI-generated filler images provokes more biased selection of the suspect if eyewitnesses are able to distinguish those images from the photograph of the suspect's face. Their data said it did not. But the concern was theirs first, which makes it harder to file away as idiosyncratic. A separate study by Rachel Greenspan and Amanda Bergold, published in Memory in 2025, went further: participants largely failed to detect any difference between the real photograph of the suspect and the AI-generated fillers, and across identification outcomes and decision criteria the researchers found no evidence of differences between lineups built with real fillers and those built with synthetic ones.

The most technically sophisticated attempt so far comes from the da/sec biometrics group at Hochschule Darmstadt, published in Frontiers in Artificial Intelligence on 7 July 2026 under the title “Towards synthetic fillers for fair photo lineups”. Rather than generating faces from a text prompt, the team inverted a suspect's mugshot into StyleGAN2's extended latent space, then injected weighted noise into the intermediate layers that carry identity-related traits, leaving the layers encoding demographic characteristics comparatively undisturbed. The result is a family of faces that are demonstrably not the suspect but plausibly of the same age, sex and ethnicity. Across 452 participants, 44 biometric experts and 408 recruited through Prolific, suspect identification came out at 68.6 per cent for experts and 70.1 per cent for non-experts, neither so high that the suspect was obvious nor so low that the task was impossible. The authors concluded that synthetic fillers contribute to fair and balanced lineups.

So the empirical picture that reaches the reform community is broadly positive, and Carlson's project is in part a check on that optimism. It is worth being precise about why the check is needed.

The evidence points in two directions at once

Carlson's hypothesis is that a real face among synthetic ones will stand out. There is a substantial body of work suggesting the opposite may be closer to the truth, which is not reassuring, because the opposite is arguably worse.

In 2022, Sophie Nightingale and Hany Farid published a study in the Proceedings of the National Academy of Sciences that has aged into a landmark. Face synthesis engines, they found, had passed through the uncanny valley: participants could not reliably distinguish AI-synthesised faces from photographs of real people, and rated the synthetic faces as slightly more trustworthy than the real ones.

The following year, Elizabeth Miller and colleagues sharpened the finding into something they named AI hyperrealism. In their first experiment, with 124 participants, white AI-generated faces were judged to be human more often than actual photographs of human beings. And the kicker: the people most likely to be fooled were the least likely to believe they could be fooled.

In 2026 a team spanning UNSW Sydney and the Australian National University added the mechanism. Writing in the British Journal of Psychology under the title “Too good to be true”, they showed that synthetic faces are hyper-average, occupying a more central position in face-space than real faces do, which is exactly what you would expect from a generator trained to produce plausible samples from a distribution. Super-recognisers, people with unusual face-processing ability, outperformed typical participants by around 15 per cent at spotting AI faces, and their advantage tracked their sensitivity to that hyper-averageness. Everyone else, super-recognisers included, was equally overconfident about their own ability to tell.

Run this through the lineup and the problem inverts elegantly. If synthetic faces are more average than real ones, then in a six-pack containing one photographed suspect and five generated fillers, the suspect is the face with the idiosyncrasy: the asymmetry, the odd nose, the thing that makes a face memorable rather than merely plausible. A generative process that produces faces clustered around the population mean does not just fail to hide the suspect. It systematically makes the suspect the most distinctive item in the array, which is the definition of a biased lineup, arrived at by a route no existing fairness measure was designed to detect.

There is a second, harder-edged version of the same worry. Hu, Li and Lyu demonstrated in 2021 that GAN-synthesised faces frequently show inconsistent corneal specular highlights between the two eyes, because the generator has no model of a physical light source and no reason to enforce one. Faces produced by StyleGAN2 are also almost perfectly aligned on facial landmarks, an inherited property of the alignment applied to the FFHQ training set, which is why superimposing a set of them produces eyes in the same place every time. Nobody claims a witness consciously audits corneal highlights. But eye-tracking is precisely the instrument for asking whether the visual system is doing something with information the person cannot report, which is what Carlson's Tobii glasses are for.

So the honest state of the evidence is three published studies suggesting synthetic fillers behave acceptably, at least one of which tested only a white male target, against a parallel literature on face perception suggesting a mechanism by which they might fail in a direction that lineup fairness statistics would not catch. That question has barely been asked, and it is being asked considerably later than the tools have been available.

The demographics of the distractor

If there is a place where this goes badly wrong first, the literature already tells you where to look.

The own-race bias in face recognition is one of the most replicated findings in the field. Meissner and Brigham's meta-analysis, covering 39 articles, 91 independent samples and close to 5,000 participants, established that people are markedly worse at recognising faces of another race than their own, with the effect showing up in both hit rates and false alarms. Cross-race identifications are, before you add any technology at all, the highest-risk category of eyewitness evidence in the system, and they are disproportionately the identifications that produce wrongful convictions of Black defendants.

Now add generative models, which are known to be worst at exactly the demographics that matter most here.

The Darmstadt team's own results contain the warning. Their method preserved ethnicity in 89.4 per cent of cases overall, a substantial improvement on prior approaches, with a mean absolute age error of 7.2 years and gender preserved in around 85 per cent of cases. But that aggregate conceals a chasm. Broken out by group, the method preserved ethnicity in 98.1 per cent of cases for white individuals, 78.9 per cent for Black individuals and 53.0 per cent for Asian individuals, a spread the authors attribute to biases inherited from StyleGAN2 itself. They further flag the other-race effect as a reason for caution in cross-ethnic lineups. To their considerable credit, they say all of this in the paper. It is not a hostile reading.

Sit with the arithmetic, because the headline figure is doing a great deal of work it has not earned. Nine cases out of ten is the performance a white suspect gets. For a Black suspect, roughly one filler in five drifts off his apparent ethnicity. For an Asian suspect it is closer to a coin toss than to nine in ten, which in a six-pack means that on average more than two of his five fillers are not, to a witness's eye, of his ethnicity at all. One drifted filler is not a rounding error. It is a face a mock-witness study would immediately flag as implausible given the description, which reduces the functional size of the array, which increases the probability that the suspect is picked for reasons unrelated to memory. Two or three of them and the array has stopped being a test of anything. The failure concentrates on exactly the suspects who are already most likely to be wrongly identified.

The broader generative-model literature makes the same point from the other end. AlDahoul, Rahwan and Zaki, publishing on AI-generated faces and their social effects, documented significant racial homogenisation in text-to-image systems, including the depiction of nearly all Middle Eastern men as bearded, brown-skinned and in traditional attire, alongside the reinforcement of occupational gender stereotypes. Separate work presented at the AAAI/ACM Conference on AI, Ethics and Society found that Stable Diffusion XL produces roughly 30 per cent less variability in skin tones than earlier versions of the model, and between about 19 and 56 per cent less variability than human face datasets, with the greatest homogenisation applied to racial and ethnic identities.

Less variability in skin tone is not a neutral aesthetic property when the artefact you are building is a lineup. It means the generated fillers for a Black suspect are drawn from a narrower slice of appearance space than the real population that suspect belongs to. It means the suspect's actual complexion has a higher chance of sitting outside the range spanned by his fillers. And it means the AI hyperrealism finding, which Miller and colleagues explicitly tied to white faces and to the disproportionate representation of white faces in training data, may not transfer at all: synthetic non-white fillers may look synthetic in ways synthetic white fillers do not, reintroducing Carlson's original worry, but only for non-white suspects.

Every failure mode in this story, on current evidence, points the same way. That is the part of the file that should stop a policy-maker cold.

A lineup that cannot be re-examined

Suppose the science comes back clean. Suppose synthetic fillers turn out to be fair, robustly, across demographics. There is a whole second body of problems waiting, and it belongs to lawyers.

American courts still assess suggestive identifications largely through Manson v. Brathwaite, decided by the Supreme Court in 1977, which asks whether an identification was nonetheless reliable by reference to five factors: the witness's opportunity to view the criminal, her degree of attention, the accuracy of her prior description, her level of certainty, and the time between crime and identification. Scientists and legal scholars have criticised that test for decades, on the grounds that several of its factors are themselves corrupted by suggestive procedures, most obviously certainty, which post-identification feedback reliably inflates. In 2011 the New Jersey Supreme Court in State v. Henderson became the first American court to reject Manson outright, replacing it with a structure built around system variables, the things police control, and estimator variables, the things they do not.

Notice that a synthetic filler is a system variable of a novel kind, and that neither framework has any idea what to do with it.

Start with disclosure. Would the defence be told that five of the six faces were generated? In most American jurisdictions there is no rule requiring it, because no rule contemplated it. An identification is documented by preserving the array, and preserving six images does not tell you which of them ever corresponded to a person.

Then reproducibility. Traditional lineup fairness litigation depends on the array being examinable after the fact: a defence expert runs a mock-witness study on the actual images, calculates functional size and Tredoux's E, and testifies. That still works with synthetic images. But the deeper question, whether the generation process itself was biased, cannot be answered from the images alone. It requires the model, its version, the random seed, the latent vector derived from the suspect's booking photograph, the noise weights, the layer selection, and, in the Darmstadt method, the record of any images an operator discarded by hand before assembling the final six. That last one is the killer. A procedure in which a human generates a batch of candidate faces and picks which to include reintroduces every subjective bias the automation was supposed to remove, and leaves no trace unless someone requires one.

Are those parameters discoverable? Are they preserved? Is there a chain of custody for a face that was never in anyone's custody? If a department uses a commercial tool, is the generator's configuration a trade secret, as the source code of proprietary forensic software has repeatedly been held to be? The forensic-software litigation of the past decade offers a discouraging preview of how that argument usually ends.

And then the question with no clean answer at all. A traditional lineup can be interrogated by finding the fillers. They are people. They have booking records. You can establish who they were, what they looked like, whether they matched the description. A synthetic filler cannot be produced, deposed, photographed from another angle or compared against a contemporaneous record, because there is no fact of the matter about it beyond the file. The image is not evidence of anything in the world. It is an output. If the file is lost, or the tool is updated, or the seed is not recorded, the array is gone in a way a set of mugshots never is.

What England and Wales quietly built instead

There is an instructive contrast sitting on the other side of the Atlantic, and it is instructive precisely because it solved the same problem twenty years ago without generating anything.

Under Code D of the Police and Criminal Evidence Act 1984, identification procedures in England and Wales are governed in considerable procedural detail, and the default method is video identification. A parade normally consists of nine moving images: the suspect plus eight others of broadly similar appearance. The images are drawn from national libraries maintained under the VIPER system, developed by West Yorkshire Police and now run as a managed national service used by forces across England and Wales as well as the Police Service of Northern Ireland and Police Scotland, or from the parallel PROMAT system.

The critical design decision is that the library is a curated database of real volunteers, filmed to a standard specification, who consented to be there. It solves the small-jurisdiction problem, because a force in Cornwall draws on a national pool rather than its own custody suite. It solves the standardisation problem, because every clip is captured under the same conditions, eliminating the background-colour and resolution artefacts that plague American six-packs. It solves the dignity problem the Düsseldorf group identified, because nobody's arrest photograph is conscripted into a stranger's prosecution. And every filler remains a documented human being whose image can be retrieved and examined years later.

That is four of the five arguments for synthetic fillers, answered by institutional infrastructure rather than by a generative model. The remaining argument, that a generator can produce a face matching an arbitrarily unusual description no volunteer library contains, is real but narrow. It is a case for a targeted exception with heavy safeguards, not for a general-purpose tool in every investigator's browser.

The lesson is not that Britain got it right and America got it wrong. It is that the problem synthetic fillers claim to solve is a resourcing and standardisation problem, and such problems have unglamorous, expensive, boring solutions that do not require anybody to reason about the epistemology of manufactured evidence.

What a defensible standard would have to require

If synthetic fillers are coming, and the practical pressures suggest they are, the minimum conditions for their use are not mysterious. They follow directly from what has already been established.

Mandatory fairness testing, per array rather than per method: a documented mock-witness study on the specific array used, reporting functional size and effective size, stratified by the suspect's demographic group. The Darmstadt findings on differential ethnicity preservation make aggregate performance figures actively misleading.

Mandatory disclosure. The defence must be told, in every case, which images were generated, by what system and version, and by whom. Anything less makes it impossible to raise the argument at all, which in practice means the argument is never raised.

Mandatory retention and reproducibility. Model identifier and version, seed, latent vector, noise parameters, the complete set of candidate images generated, and a log of any human selection among them. A synthetic lineup that cannot be regenerated exactly is not evidence anyone can test. If an operator curates the output, that curation is part of the procedure and must be recorded and blinded in the same way lineup administration is blinded.

And a presumption against use in cross-race identifications until the demographic performance question is settled, which on current evidence it is not remotely. This is the recommendation most likely to be resisted and the one most clearly supported by what is already known.

None of these exist anywhere. There is no American federal standard for synthetic fillers, no state statute, no model policy from any of the bodies that have issued model eyewitness policies. The Department of Justice memorandum issued by Deputy Attorney General Sally Yates in January 2017, which established department-wide procedures for photo arrays including blind or blinded administration and contemporaneous confidence statements, predates the technology entirely. The American statutory landscape remains a patchwork in which many states have adopted no reforms at all, and those that have wrote them in the language of an era when a filler was a photograph of someone.

Nor does the legal scholarship reach these questions. The most recent legal treatment of synthetic fillers, Elena Mione's April 2026 article in the Michigan Technology Law Review, argues that AI-generated images could improve lineup fairness and spare real people the conscription of their booking photographs, works carefully through Manson v. Brathwaite, and counsels caution and further study across racial demographics before any large-scale implementation. It proposes no specific safeguards. And it does not reach disclosure to the defence, discoverability, or the retention of generation parameters at all.

The regulatory gap is not a gap. It is the entire field.

What the witness is actually being asked

Strip away the machinery and there is a question at the bottom of this that the psychology does not currently have an answer to, and it is the question that makes the story genuinely strange rather than merely worrying.

A recognition test measures the difference between a target and a set of distractors, and its diagnostic value depends on both being drawn from the same population, so that the only thing separating them, from the witness's point of view, is her memory. Nobody ever wrote that assumption down, because for a hundred and fifty years it was unfalsifiable in practice: every face in a lineup was a face, produced by a human being existing and being photographed.

Synthetic fillers break the assumption in a way that has no precedent. The target is a photograph, an optical record of light reflected from a person who was standing in front of a camera. The distractors are samples from a learned distribution, produced by a process with different statistics, different failure modes and different demographic coverage. They are not worse faces. In several measurable respects they are better ones: more average, more symmetrical, more trustworthy-looking, more likely to be judged human than actual humans. But they are not the same kind of object, and the test assumed they were.

What the witness is being asked to do is therefore not quite what anybody thinks. She believes she is being asked which of these six people she saw. She is in fact performing a discrimination between one photographic record and five generative samples, and the psychology tells us, from at least three directions, that this discrimination is not perceptually neutral: hyper-averageness is detectable, super-recognisers detect it, everyone is overconfident about detecting it, and the detectability itself varies by race.

Carlson's instinct, that the real face might stand out, is likely to prove correct in a more complicated way than he framed it. The real face will stand out not because it looks fake but because it looks specific, and specificity is precisely what a face has to have to be remembered and precisely what a generative model tends to average away.

There is a version of this that ends well. A national library of consenting volunteers, standardised capture, generation reserved for the genuinely impossible description, fairness tested per array, disclosed as a matter of course, reproducible from a retained seed. That version is boring and expensive, which is why it is unlikely to be the one that happens. The version that happens will be a browser tab, a booking photograph, five faces in nine seconds, and no line in the file recording that anything unusual took place.

And somewhere at the end of that, a woman who saw a man for three seconds will look at six faces and point at one, and say she is certain, and be believed. What she will not be told, and what nobody will think to write down, is that she was the only person in the room, in the array, or in the entire procedure whose certainty was about a human being.

Previous

Deepfakes Did Not Steal Your Face

The Dark Side of AI Companionship: A Cautionary Study
Bulletin №23