DECISION RESEARCH DOSSIER · VOL. I PERSON → EVIDENCE → EVALUATOR → DECISION
Who Decides · Who Gets Believed · What It Costs

THE HUMAN VERDICT

An evidence dossier on discretion, bias, and consequence across twenty-seven categories of gatekept decisions — visas, jobs, courtrooms, road tests, hospitals, and beyond

Framework 27 Decision Types
Evidence Base 40+ Primary Studies
Core Pipeline Person → Judgment → Fate
5% vs 88% Same law, diff. asylum judge
+50% Callbacks, white-sounding name
14.2 yrs Avg. wrongful imprisonment
+13.4% Sentence gap, Black vs white men
1.9× False "high risk" algorithm flags
$500K Lifetime earnings lost, 1 conviction
Contents
COMPILED FROM PUBLIC RESEARCH · PEER-REVIEWED & GOVERNMENT SOURCES FOR RESEARCH & EDUCATIONAL USE
44.9% → 0% Asylum grant rate range, by judge
65% → ~0% Parole grant rate, within one session
32% vs 56% UK driving pass rate, groups compared
2.0× Raw mortgage denial gap, Black vs white
50% Med. students holding false pain beliefs
Wrongful conviction risk, Black defendants
Part 00 · Framework

The Pipeline: How a Person Becomes a Decision

Every one of the twenty-seven categories in this dossier — from a driver's test to a death-penalty appeal — reduces to the same four-step chain: a person presents themselves, evidence is gathered, a human evaluator interprets that evidence, and a decision is rendered into one of a small set of binary outcomes. This report treats that chain as the unit of analysis.

00.1

The Chain and Its Failure Points

The chain looks simple — Person → Evidence → Human Evaluator → Decision — but each arrow is a place where something other than merit can enter. The evidence stage can omit relevant facts or admit irrelevant ones (a photograph, an accent, a name). The evaluator stage is where discretion, mood, fatigue, and unconscious pattern-matching operate. The decision stage compresses a complex human situation into a binary: admit or reject, grant or deny, pass or fail, release or detain.

// the four-step chain, annotated
PERSON        —> who they are, what they present
EVIDENCE      —> what the evaluator is allowed / able to see
HUMAN EVALUATOR —> discretion, training, fatigue, bias enter HERE
DECISION       —> binary output: admit/reject, grant/deny, pass/fail...

Two people can enter this chain with identical evidence and exit with opposite decisions purely because of which evaluator they were randomly assigned. That single fact — documented over and over in the sections that follow, in courts, hospitals, hiring offices, and orchestra pits — is the empirical core of this entire dossier.

Framing note A raw outcome gap between two groups (e.g., "Group A is denied twice as often as Group B") is not, by itself, proof of bias. Groups can legitimately differ in the underlying evidence they present. Distinguishing a real gap from a legitimate gap is the central methodological problem in every section below, and this report flags it every time the distinction matters.
00.2

The Twenty-Seven Categories, Grouped

The source framework for this research names twenty-seven domains where a human being holds discretionary power over another person's fate. They cluster into seven families by what is actually at stake:

ClusterCategoriesTypical decision verbs
Belonging & statusImmigration & border, citizenship, military admission, membership/communityAdmit / Reject, Grant / Withhold
LivelihoodEmployment, professional licensing, business approval, financial creditHire / Reject, License / Refuse
LibertyCriminal justice, corrections, legal/judicial rulingsRelease / Detain, Appeal upheld / rejected
BodyHealthcare eligibility & treatment, insuranceTreat / Don't treat, Approve / Deny
Shelter & welfareHousing, government benefits, tax/public financeApprove / Deny, Fund / Don't fund
Mind & meritEducation, research participation, arts/entertainment, sportsSelect / Don't select, Pass / Fail
Everyday competenceDMV road test, permits & regulatory approval, identity/civil statusPass / Fail, Clear / Don't clear

What unites them is not subject matter but structure: a closed room (literal or figurative), one or a few evaluators, incomplete visibility into the reasoning, and a decision that is nominally appealable but practically difficult to reverse. The DMV road test belongs on the same page as an asylum hearing because both share that structure exactly — "no one knows what happened in that car except the two [people in it]."

Part I · The Appearance Question

Does Your Face Decide?

Before touching any single institution, the research asks the most uncomfortable question first: does what an evaluator sees — not what a person has done — move the needle? The answer, replicated across seventy years of psychology and economics, is an unambiguous yes, with real financial and legal magnitude attached.

I.1

The Halo Effect

In 1920, psychologist Edward Thorndike noticed that military officers who rated a soldier highly on one trait — physical fitness, say — tended to rate that same soldier highly on unrelated traits like leadership and intelligence. He called it the halo effect: a single positive impression bleeding into every other judgment about a person, with no evidentiary basis for the bleed. In 1972, Karen Dion, Ellen Berscheid, and Elaine Walster gave the effect its modern, best-known form: participants shown only a photograph assumed attractive people had better personalities, happier lives, and more successful marriages than unattractive people — despite having zero information about any of those things.

The effect has since been demonstrated in hiring interviews, courtroom sentencing, university grading, credit approval, and political voting. A 2018 meta-analysis spanning more than a hundred studies found the effect holds consistently across cultures.

9%Wage penalty, below-average looks
5%Wage premium, above-average looks
10%Higher lawyer earnings, +1 attractiveness rank, 5 yrs out
25%Overweight workers reporting overlooked promotions

Economists Daniel Hamermesh and Jeff Biddle formalized the wage numbers above using interviewer-rated attractiveness across three large national surveys, holding education and experience constant — meaning two people with identical résumés earned different amounts purely as a function of how the interviewer had rated their looks years earlier. Follow-up work found the effect for lawyers specifically grows with time in the profession, and that body weight and height explain wage variation on top of — not merely as a proxy for — attractiveness.

Where the evidence is thinner The classic 1991 Eagly meta-analysis found attractiveness has only a weak link to perceived intelligence specifically, even though it strongly predicts perceived competence, trustworthiness, and moral character. The halo effect is real but not uniform across every trait — a nuance worth preserving rather than flattening into "attractive people always win everything."
I.2

The Horn Effect — What Happens to People Judged Unattractive

The mirror image of the halo effect is the horn effect: a single negative trait — often weight, disability, visible scarring, or non-conforming grooming — contaminates every other judgment downward. In simulated interviews, good grooming alone outweighed job qualifications in producing favorable hiring decisions, even though the same interviewers insisted appearance played almost no role in their choices. That gap between what evaluators believe about their own objectivity and what actually drives their decisions recurs throughout this dossier — in courts, hospitals, and licensing boards alike.

75%

of hiring managers make a poor decision traceable to the halo/horn effect

A bad hire driven substantially by first-impression bias costs roughly 30% of that role's annual salary — meaning appearance bias is not just an ethical problem but a measurable institutional cost.

Part II · Case Study: Belonging

Visas & Asylum: The Adjudicator Lottery

No category in this framework demonstrates "same evidence, different verdict" more starkly, or with higher stakes, than asylum and immigration adjudication. The applicant cannot choose their judge. The law does not change from room to room. The outcome does.

II.1

Refugee Roulette

A landmark study analyzed 133,000 decisions by 884 U.S. asylum officers, 140,000 decisions by 225 immigration judges, 126,000 Board of Immigration Appeals rulings, and over 4,000 federal appellate decisions. Its findings became known as "refugee roulette."

Same courthouse, worst case
Miami, Colombian Applicants
Judge A grant rate5%
Judge B grant rate88%
Same building?Yes
Same law applied?Yes
Regional office variance
Chinese National Applicants
Lowest officer grant rate0%
Highest officer grant rate68%
60% of officers deviated>50% from mean
Current data, 2026
All U.S. Immigration Judges
Judges >500 decisions, grant >30%17
Judges >500 decisions, grant <3%378
Highest single-judge rate44.9%

The Kahneman/Sibony/Sunstein "Noise" framework gives this pattern a name: level noise — some evaluators are simply, permanently more generous or harsher than others, independent of the merits in front of them. What made the original study notable is that it also found correlations between grant rates and the judge's own gender and pre-appointment career background, and confirmed that having legal representation strongly predicted a favorable outcome — meaning the applicant's fate depended not only on the assigned judge, but on whether they could afford a lawyer at all.

Why this domain sits at the top of the severity list Unlike a road test, an asylum denial can mean deportation to danger. Unlike a hiring rejection, there is often no "reapply next quarter." Unlike most legal proceedings, U.S. immigration courts use a single judge with no jury and no panel of peers to catch an outlier ruling — the credibility call rests on one person's read of one person's testimony, exactly like the sealed-car dynamic of a driving test, except the stakes are a matter of safety and survival rather than a delayed commute.
II.2

What Recovery Looks Like After a Denial

A negative asylum decision is technically appealable through the Board of Immigration Appeals and, beyond that, the federal circuit courts — but appeals take years, cost money most applicants do not have, and are themselves subject to circuit-by-circuit variance in how readily they overturn the immigration court below. For an applicant without representation, the practical difference between "technically appealable" and "effectively final" collapses to almost nothing.

Part III · Case Study: Livelihood

Work & Hiring: The Résumé Test

Employment discrimination is the most heavily audited domain in this entire dossier — because it is the easiest to test experimentally. Send identical fictitious applications, vary only one signal, count the callbacks. The results are simultaneously the field's strongest evidence and its messiest replication history.

III.1

The Name on the Résumé

Economists Marianne Bertrand and Sendhil Mullainathan sent nearly 5,000 fictitious, otherwise-identical résumés to real job postings in Boston and Chicago, randomly assigning each one a stereotypically Black-sounding or white-sounding name.

Callback rate by name signal — Bertrand & Mullainathan (2004)
White-sounding name
9.65 / 100 sent
Black-sounding name
6.45 / 100 sent

White-sounding names received roughly 50% more callbacks — a gap the authors calculated as equivalent to what eight additional years of work experience would buy an applicant. Discrimination was uniform across occupation, industry, and employer size, and federal contractors and self-declared "Equal Opportunity Employers" discriminated just as much as everyone else. Better résumés helped white-named applicants far more than Black-named applicants — a 30% callback boost from a stronger résumé for white names, versus a much smaller boost for Black names.

Replication is messier than the headline number A later, larger study of college-graduate hiring (Nunley et al.) found a much smaller gap of 2.6 percentage points — less than half the size of the original. A well-known 2016 American Economic Review paper by Deming, Yuchtman, Abulafi, Goldin & Katz, using a similar design, did not replicate the original effect size at all. The direction of the bias is one of the most robust findings in labor economics; the exact magnitude is genuinely contested and appears to shrink in tighter labor markets and among more credentialed applicant pools.
III.2

The Blind Audition — A Fix That May Not Have Worked

This is the single most-cited "we found the bias and fixed it" story in all of social science, and it deserves to be told with its full, more complicated ending.

The famous finding When U.S. symphony orchestras began auditioning musicians behind a screen that concealed their identity, the probability a woman advanced past preliminary rounds rose by roughly 50%. The switch to blind auditions was credited with 30–55% of the rise in the proportion of women hired into major U.S. orchestras between 1970 and 1996. It has been cited by Malcolm Gladwell, invoked in a Justice Ginsburg dissent, and spun into commercial "blind hiring" software.
The complication Researchers who went back to the raw, unadjusted tabulations found that women actually did worse behind the screen in the raw data — the famous 50% figure only appears after a statistical adjustment applied to a small subsample. A large 2017 randomized controlled trial by the Australian government, explicitly designed to test the same mechanism at scale with 2,000+ managers, found the opposite result: de-identifying candidates reduced the odds women were shortlisted.

The lesson for a research project like this one is not "blind auditions don't work" — it's that even the field's most celebrated success story turns out to be genuinely unsettled on close inspection, and repeating it as proven fact without the caveat is itself a form of the overconfidence this whole dossier is trying to document.

Part IV · Case Study: Liberty

The Courtroom: Noise, Sentencing, and the Judge You Get

If asylum courts show variance between evaluators, the criminal justice system adds a second axis: variance within the same evaluator, across the same day, depending on factors that have nothing to do with the case.

IV.1

The Hungry Judges

In 2011, researchers analyzed 1,112 parole rulings made by eight experienced Israeli judges — averaging over twenty years on the bench — across ten months. Each day was split into three sessions by two food breaks.

Favorable parole ruling rate, across one session (Danziger, Levav & Avnaim-Pesso, 2011)
Right after a break
~65%
Mid-session
~30%
Right before next break
~0%

The same prisoner, with the same file, appearing at a different point in the judge's session, faced a dramatically different probability of release. The finding became one of the most cited results in behavioral decision science — and one of the most contested. Critics, including Cornell's Jeffrey Rachlinski, argued the pattern could be an artifact of case scheduling (harder cases pushed later, or attorneys arranging dockets non-randomly) rather than true mental fatigue. The original authors maintain the case order was arbitrary. Both readings remain live in the literature — a useful reminder that a dramatic, easily-quotable finding about bias is not automatically a settled one.

IV.2

"Noise": When Two Judges See the Same File Differently

Nobel laureate Daniel Kahneman, with Olivier Sibony and Cass Sunstein, coined the term system noise for exactly this problem: unwanted variability in judgments that should, in principle, be identical. Their book documents it everywhere discretion exists — two doctors giving different diagnoses to identical patients, two underwriters quoting prices 55% apart on the same case file, a quarter of interview panels disagreeing internally about who the best candidate was, despite watching the same candidates in the same room.

Level noiseSome judges are simply, consistently harsher or more lenient than their peers, across every case type.
Pattern noiseAn evaluator reacts idiosyncratically to specific case features — their own personal "triggers" — that other evaluators don't share.
Occasion noiseThe very same evaluator would rule differently depending on time of day, hunger, mood, or whether their sports team lost last night.
IV.3

Federal Sentencing: The Government's Own Audit of Itself

The U.S. Sentencing Commission's 2023 report — not an advocacy document, but the Commission's own internal statistical review — examined over 300,000 federal sentences from 2017–2021.

Group (vs. white men)Sentence length gap (incl. probation decision)Less likely to receive probation
Black men+13.4%−23.4%
Hispanic men+11.2%−26.6%
Hispanic women (vs. white women)+27.8%

The Commission's own analysis located where the disparity lives: overwhelmingly in the binary decision of whether to incarcerate at all, more than in the length of the sentence once incarceration is decided. That's an important structural finding — it means the highest-leverage point of bias in federal sentencing is a single yes/no gate, not a continuous number a judge slides up or down.

Part V · Case Study: Everyday Competence

Behind the Wheel: The DMV Examiner

The road test is the smallest-stakes decision in this dossier and, in some ways, the hardest to study — because, as you noted, no one else is in the car. There is no transcript, no recording, no second observer. It is the purest version of the sealed-room problem this whole framework is built around.

V.1

What the Available Data Shows

A UK Freedom of Information request into the national driving-test agency (DVSA) found stark differences in pass rates by demographic group:

UK driving test pass rate by group, DVSA data 2008–2017
White men
56%
Black women
32%

Only about 21% of UK driving examiners are women, and the agency maintains that all candidates are scored against identical, safety-based criteria. The source reporting this data is explicit that a raw pass-rate gap alone cannot prove examiner bias — the same caution that applies to every raw-gap number in this report.

A genuinely mixed evidence base Studies of examiner bias in other subjective, one-on-one licensing exams cut both ways. A review of 1,790 UK medical examiners found only three showed statistically detectable ethnic bias relative to peers scoring the same candidates. A separate UK study of GP licensing exams found ethnic background did not predict pass rates once entry-test scores and other legitimate factors were controlled for. Meanwhile, a study of international medical graduates found examiner-to-examiner stringency alone explained 16% of the variance in domain-level scores — meaning which examiner you got mattered a great deal, independent of any demographic pattern, purely because some examiners are systematically "hawks" and others "doves."
V.2

Why the DMV Case Is the Cleanest Illustration of the Whole Problem

The road test compresses every mechanism in this dossier into one twenty-minute window: a single evaluator, full discretion over pass/fail, no witness, a plausible safety justification for every judgment call, and a candidate with almost no leverage to contest a specific ruling ("you hesitated at the roundabout" cannot be independently verified after the fact). It is structurally identical to an asylum credibility hearing or a parole interview — just with lower stakes and, consequently, far less research funding directed at it. That gap between stakes and scrutiny is itself a finding: the domains where bias is hardest to detect are not the ones where it would do the least damage — they're simply the ones nobody has funded a large audit study for.

Part VI · Case Study: Shelter & Credit

Credit & Shelter: Lending and Housing Decisions

Mortgage lending is the best domain in this whole dossier for showing exactly how much a number changes depending on whether you control for legitimate factors — because researchers have run the identical dataset through both a naive and a rigorous lens and published both results.

VI.1

Raw Gap vs. Controlled Gap

Observed / raw denial rate
No controls applied
White applicants13.6%
Black applicants27.1%
Ratio2.0×
"Real denial rate" (Urban Institute)
Adjusted for credit profile
Ratio, Black vs white1.2×
Ratio, Hispanic vs white1.1×
Full individual-level controls (Fed Reserve)
Income, debt, credit history held constant
Remaining rejection ratio1.6×
Statistically significant?Yes

A 2024 Federal Reserve Bank of Philadelphia paper pushed the question further: what if lenders used only a race-blind automated underwriting recommendation? Even in that scenario, the data suggest roughly a 9 percentage-point gap would remain — implying part of the disparity is embedded in the underlying financial profile itself, a residue of generations of unequal access to education, employment, and housing wealth, rather than something a lending officer decides in the room.

Why both numbers matter Citing only the 2.0× raw gap overstates in-the-moment lender bias. Citing only the smallest controlled number understates the compounding effect of historical inequality baked into "objective" credit data. A rigorous research project reports both, explicitly labeled, every time.
VI.2

Where the Gap Is Largest

The disparity is not uniform across loan types. Home-improvement loans had the highest denial rate of any category studied in 2020 (38.8% overall) — and more than half of Black (63.0%) and Hispanic (56.6%) applicants seeking to repair or renovate their homes were denied, despite living in older housing stock more likely to need exactly that work. For manufactured homes, over three-quarters of Black borrowers' applications were denied even when the home was secured by land.

Part VII · Case Study: The Body

The Exam Room: Medicine's Judgment Gap

Healthcare is the domain where bias is least likely to look like hostility — and that makes it, in some ways, the hardest to fix, because it can operate through sincerely held false beliefs rather than conscious prejudice.

VII.1

False Beliefs, Real Undertreatment

A 2016 study published in PNAS tested whether racial bias in pain treatment traces back to specific false beliefs about biology — for instance, that Black people's skin is thicker or that their blood coagulates faster than white people's. Both statements are false.

70%Laypeople endorsing ≥1 false belief
50%Med. students/residents endorsing ≥1
Pain rated lower for Black patients, by believers
Treatment accuracy for Black patients, by believers

Participants who endorsed the false beliefs rated an identical vignette patient's pain as lower, and recommended less accurate treatment, when the patient was described as Black rather than white. A separate analysis of nearly one million pediatric appendicitis cases found Black children were less likely than white children to receive any pain medication for moderate pain, and less likely to receive appropriate opioid treatment for severe pain.

The unsettling part Researchers studying this pattern note it does not appear to trace back to overt hostility — it operates through beliefs sincerely held by people who would likely describe themselves as caring, unbiased clinicians. That makes it structurally different from, and arguably harder to correct than, a case of deliberate discrimination.
Part VIII · The Machine Alternative

When the Judge Is a Machine

A natural response to everything documented so far is: replace the biased human with an algorithm. The evidence says that doesn't dissolve the problem — it relocates it into a definitional argument about what "fair" even means.

VIII.1

COMPAS: A Case Study in Two Correct Answers

In 2016, ProPublica's "Machine Bias" investigation examined COMPAS, a proprietary algorithm used across U.S. courts to predict a defendant's likelihood of reoffending, feeding into bail, sentencing, and parole decisions.

ProPublica's finding Among defendants who were not re-arrested within two years, Black defendants were misclassified as "high risk" by the algorithm nearly twice as often (45%) as white defendants who were also not re-arrested (24%).
Northpointe's (the vendor's) rebuttal Among defendants who received the same risk score, Black and white defendants were, in fact, equally likely to actually reoffend — meaning the tool was well "calibrated" even while its error rates were unequal across groups.

Both claims are statistically true simultaneously, and that is not a data error — it is a proven mathematical result. Except in special cases, it is mathematically impossible for a predictive tool to have both equal false-positive rates across groups and equal calibration across groups whenever the two groups have different underlying base rates of the outcome being predicted. Choosing which fairness definition to prioritize is not a technical fix; it's the same value-laden choice a judge or an examiner makes, just made once, upstream, by whoever configures the software — and then applied at scale to everyone who comes after.

Why this matters for the whole dossier "Get rid of human discretion" is a common proposed fix for every category in this framework. COMPAS shows that removing the human doesn't remove the value judgment — it just makes the value judgment invisible, embedded in code that few of the people affected by it, and not all of the people who deploy it, can actually inspect.
Part IX · Consequence

After the Verdict: The Cost of "No"

This section answers the question underneath your original one most directly: is the damage irreparable, or does it just feel that way? For one category — wrongful criminal conviction — the answer is measured in years and dollars with unusual precision, because it is the one type of "wrong decision" that gets independently, retrospectively confirmed at scale.

IX.1

Wrongful Conviction: A Timeline of What Is Actually Lost

Year 0
Conviction on flawed evidence
Leading causes across the National Registry of Exonerations: official misconduct (71–85% of cases depending on crime type), mistaken witness identification (26%), perjury or false accusation (up to 72% of cases in some years), false or misleading forensic evidence, and false confessions.
Liberty lost
~13–16 years later
Average time to exoneration
Persons exonerated in 2025 lost an average of 14.2 years each; in 2024, 13.5 years each. Black murder exonerees wait roughly three years longer than white murder exonerees before release — about 16 years versus 13, regardless of whether they were sentenced to death.
Prime working years gone
Extreme cases
Decades, not years
In the 2025 cohort, 37% of exonerees lost more than 20 years; 13 lost more than 30 years; 3 lost more than 40 years. Since 1989, wrongful convictions have cost exonerees a combined total exceeding 21,000 years of life.
Effectively irreversible
Post-release
Compensation, where it exists
States have paid over $5.5 billion in compensation and civil damages to exonerees since 1989 — but compensation statutes exist in only 38 states plus D.C., and typically require the exoneree to affirmatively prove innocence, not merely have their conviction vacated.
Partial, uneven repair
Who bears this risk disproportionately Black Americans are 13.6% of the U.S. population but 53% of all registry exonerations, and are roughly seven times more likely than white Americans to be wrongfully convicted of a serious crime. In wrongful drug-crime convictions specifically — where Black and white Americans use illegal drugs at similar rates — innocent Black people are about nineteen times more likely to be wrongfully convicted than innocent white people, largely traced to racial profiling in stops and, in some documented cases, officers fabricating evidence.
IX.2

Collateral Consequences: When One "No" Becomes Many

Even a conviction that was not wrongful — and even, in many U.S. jurisdictions, a mere arrest that never led to conviction — can trigger consequences that compound across every other domain in this dossier simultaneously.

Lifetime financial impact of one incarceration — Brennan Center analysis
AVG. EARNINGS REDUCTION, POST-RELEASE−52%
UNEMPLOYMENT RATE, FORMERLY INCARCERATED27%
SHARE ACTIVELY WORKING / SEEKING WORK93%
LIFETIME EARNINGS LOST, ON AVERAGE−$500,000

That 27% unemployment rate is not a motivation problem — 93% of formerly incarcerated working-age people are actively working or looking for work. It is a gatekeeping problem: 9 in 10 employers, 4 in 5 landlords, and 3 in 5 colleges and universities now run background checks that can screen an applicant out before a human ever evaluates their actual qualifications. Between 70 and 100 million American adults carry some form of criminal record — including arrests that never resulted in a conviction — meaning a negative decision in the justice system routinely becomes a standing negative decision in housing, education, and employment for the rest of that person's life, applied automatically, by institutions that were never part of the original case.

Part X · Methodology

Detecting Bias: Tools, Fixes, and Their Limits

"The outcome gap looks unfair" is a hypothesis, not a finding. This section covers how researchers actually try to move from a raw gap to a defensible claim of bias — and how often even careful attempts get contested.

X.1

The Benchmark Test, the Outcome Test, and the Threshold Test

The simplest approach — the benchmark test — compares decision rates across groups after controlling for every legitimate factor researchers can measure. Its weakness is baked in: it is essentially never possible to observe and control for every legitimate factor, so an apparent gap can always be explained away as an omitted variable.

The outcome test tries to sidestep this by looking not at how often a group is screened (searched, denied, flagged) but at the hit rate — how often that scrutiny actually turns up what it was looking for. If officers search minority drivers more often but find contraband at a lower rate than when they search white drivers, that's suggestive of a lower evidentiary bar being applied to minority drivers specifically.

The newer threshold test, developed to fix a known flaw in the outcome test (infra-marginality — the risk that the "marginal" cases searched aren't comparable across groups), was applied to 4.5 million North Carolina traffic stops and found real evidence of a lower search threshold applied to Black and Hispanic drivers — while also demonstrating, in the same paper, that the simpler outcome test can produce a false finding of bias against the majority group when applied carelessly to a single department. The lesson: the statistical tool you choose can change your conclusion, even applied to the identical underlying data.

X.2

Interventions That Have Actually Been Tested

InterventionDomain testedResultConfidence
Blind auditions (physical screen)Orchestra hiring+50% advancement for women (original study)Contested — RCT found opposite
De-identified applicationsAustralian public-service hiringReduced, not increased, women's shortlistingLarge RCT, high confidence
Algorithmic underwritingMortgage lendingShrinks but does not eliminate racial gap (~9pp residual)Moderate confidence
Risk-assessment algorithmsPretrial/sentencing (COMPAS)Relocates bias into definitional choice, doesn't remove itHigh confidence, proven mathematically
Controlled resume auditsHiringConfirms bias direction; magnitude shrinking across replicationsHigh confidence on direction only
Honest takeaway Almost nothing in this literature is a clean, unambiguous "problem solved" story. The most responsible research posture is to treat every widely-repeated success story — blind auditions above all — as provisional until you've checked whether it replicated.
Part XI · Where the Research Goes Next

Research Agenda: Open Questions

A synthesis of where the evidence base is strong, where it is thin, and where a new research project — rather than a summary of existing work — could add genuine value.

XI.1

What's Well-Established

  • Appearance measurably shifts judgments across hiring, wages, legal outcomes, and medical pain assessment — this is among the most replicated findings in social psychology.
  • Inter-evaluator variance for the identical case, in judicial, asylum, and parole systems, is large enough to functionally decide outcomes by lottery of assignment.
  • Wrongful conviction produces quantifiable, severe, partially irreparable harm — and falls disproportionately on Black defendants by a wide, consistent margin.
  • Raw demographic gaps in lending, sentencing, and healthcare persist, in reduced but real form, even after rigorous statistical controls are applied.
XI.2

What's Genuinely Contested

  • The exact magnitude of name-based hiring discrimination (shrinking across replications, not disappearing).
  • Whether blind/de-identified hiring processes reduce bias at all (directly contradicted by a large government RCT).
  • Whether an algorithm can be made "less biased" than a human in any definition-independent sense (proven mathematically impossible to satisfy every fairness criterion at once).
  • Whether the "hungry judges" effect reflects true decision fatigue or an artifact of case-ordering — both readings remain live in the literature.
XI.3

The Gap Worth Filling

Where a new study could matter most Almost no existing research measures appearance bias and race/gender bias simultaneously, in the same dataset, to see how they interact — most audit studies isolate a single variable by design. Domains like the DMV road test, most professional licensing boards, and most of the twenty-seven categories in this framework have thin or nonexistent large-sample data, unlike hiring, asylum courts, and federal sentencing, which are comparatively saturated with research. A serious next step is either (a) an intersectional re-analysis of existing datasets that already contain both appearance and demographic variables, or (b) a new audit study — using paired photographs and identical case files — in one of the under-studied domains, most plausibly DMV/licensing exams, where the sealed-room structure makes bias hardest to detect and therefore most valuable to finally measure.
XI.4

Sources Consulted

  • Dion, Berscheid & Walster (1972); Eagly et al. (1991) — foundational halo-effect studies
  • Hamermesh & Biddle, "Beauty and the Labor Market," American Economic Review (1994) & NBER Working Paper 4518
  • Ramji-Nogales, Schoenholtz & Schrag, "Refugee Roulette," Stanford Law Review (2007); OpenImmigration.us judge-variation data (2026)
  • Danziger, Levav & Avnaim-Pesso, "Extraneous Factors in Judicial Decisions," PNAS (2011); critique by J. Rachlinski / Weinshall
  • Kahneman, Sibony & Sunstein, Noise: A Flaw in Human Judgment (2021)
  • U.S. Sentencing Commission, "2023 Demographic Differences in Federal Sentencing"
  • National Registry of Exonerations, Annual Reports 2023–2025; "Race and Wrongful Convictions" (2022)
  • Bertrand & Mullainathan, "Are Emily and Greg More Employable Than Lakisha and Jamal?," AER (2004); replication debate (Nunley et al.; Deming et al.; DataColada #51)
  • Goldin & Rouse, "Orchestrating Impartiality," AER (2000); critique (AEI, 2022) and Australian Public Service RCT (2017)
  • Angwin et al., "Machine Bias," ProPublica (2016); Northpointe/Equivant response; Kleinberg, Chouldechova on fairness impossibility
  • Urban Institute, "Real Denial Rate" mortgage analyses; Federal Reserve Bank of Chicago & Philadelphia mortgage-discrimination studies; HMDA data
  • Hoffman, Trawalter, Axt et al., "Racial Bias in Pain Assessment," PNAS (2016)
  • Chohlas-Wood, Goel et al., "threshold test" for discrimination in policing, North Carolina traffic-stop data
  • Brennan Center for Justice; Prison Policy Initiative; U.S. Commission on Civil Rights — collateral consequences data

This dossier synthesizes publicly available peer-reviewed research, government reports, and investigative journalism current as of August 2026. Figures are attributed to their original studies; where findings have been contested or failed to replicate, that is noted alongside the original claim. This document is intended as a research aid, not a legal or statistical authority in itself — primary sources should be consulted directly for any citation in further work.

❖ ❖ ❖