For most of hiring's history, bias was something organizations could claim they were working on. AI has made that harder. When an algorithm screens 4 million job applications and produces documented racial disparities (adverse impact against Black applicants in 1 in 10 roles, and against Asian applicants in 1 in 20), "we're working on it" stops being a strategy and starts being a liability.
That is the central finding of a landmark study published in May 2026 by researchers at Stanford's Human-Centered AI Institute (HAI) and collaborators. It is the largest examination of AI hiring algorithms ever conducted. The study analyzed 4 million applications across 156 employers, all using the same games-based behavioral assessment platform. It also found that 42 algorithmic scoring models were shared across different employers. What that means, and why it matters, is the starting point for any serious conversation about AI hiring tools.
Those are serious findings that the hiring industry cannot deflect with marketing language. But the most important question for TA leaders isn't "can AI be biased?" It clearly can. The question is: what specific design choices increase bias, which design choices reduce it, and how do you tell which kind of a tool you are evaluating?
A caveat matters up front: no tool removes bias entirely. It can still surface in how a hiring process is structured and how people and systems act on a score, so the realistic goal is to reduce and manage bias, not to declare it solved.
The Stanford study focused on a specific class of tool: games-based behavioral assessment, where candidates complete a series of cognitive and personality tasks and their scores are compared against a profile built from a company's existing employees in the role. TestGorilla does not offer game-based assessments of this kind. They have been widely criticized for weak face validity: they don't clearly measure the skills that actually predict job performance. TestGorilla also does not use 'positive training examples,' where a handful of an organization's top performers are used to calibrate a model.
That model has a structural problem worth spelling out. If your top-performer benchmark is built from a workforce that already skews toward a dominant demographic group, as is often the case at large employers, then the algorithm learns to replicate that workforce. The AI is accurate at what it was trained to do. The problem is what it was trained to do.
Taking a step back, there is a problem that comes even earlier: the same test is used to evaluate everyone, whatever role they are applying for. This is where a job-relevant approach helps most. TestGorilla offers tests built for specific roles and skills, so candidates are measured against what a given job actually requires rather than against a single generic yardstick. Multi-measure assessment reinforces this: combining several tests surfaces the different strengths each candidate brings, giving more people a way to show what they are good at. A process built this way is far less likely to produce the kind of uniform, repeated rejection the Stanford study describes.
It is worth being precise about the two forms of discrimination, because the study captures only one of them: it documents adverse (disparate) impact, not disparate treatment.
Disparate treatment is intentional discrimination that treats individuals differently because of who they are. Disparate impact (also called adverse impact) is unintentional: a seemingly neutral selection procedure that systematically favors some groups over others.
What the Stanford study documents is adverse impact. No one programmed the system to screen out Black applicants. The outcome emerges through mechanisms that are already well understood: chiefly a reliance on cognitive tasks known to produce group differences, and, potentially, the historical data these models are trained on. That distinction matters legally and practically, but it does not make the outcome less discriminatory under US federal law, and it does not make it less real for the candidates affected.
A second mechanism the study identified: model sharing across employers.
In a TestGorilla-style process, each employer selects and configures tests based on the specific role they are hiring for. The scoring model is not a fixed generic template. It is anchored to the job-relevant criteria the hiring team defines.
The Stanford study found the opposite design in the platform it examined: 42 algorithmic models shared across the platform's client base. A candidate applying at two companies using the same model received the same algorithmic evaluation at both, whether they knew it or not. The researchers found that 4% of applicants who applied to 10 roles were recommended for rejection across all of them, and that figure is more striking than it first appears. If each of those applications were a genuinely independent evaluation, being rejected by all 10 would be vanishingly unlikely. It happened because the same 42 models sat behind decisions at employer after employer. What they experienced as 10 independent evaluations was, in effect, one evaluation, repeated. Beyond fairness, this is a practical problem for the candidate: if they had known that a single platform's model would determine their fate across multiple applications, they might have made different choices about where and how to invest their time.
The study examined one vendor, one modality (games-based assessment), and one time window. With a very precise scope, the researchers explicitly caution against generalizing to all algorithmic screening. The mechanism of model sharing is real and worth interrogating directly for any platform you evaluate, but the conclusion that "all AI hiring is broken" overstates what the evidence shows.
There's a second limit worth naming. The study measures outcomes, not their cause. Adverse impact in a hiring funnel often comes down to where an employer sets the cut-off, not the assessment score underneath it. A rigid pass mark, applied high in the funnel with no human review, can manufacture disparity from a score that would clear the bar if it were used differently. The Stanford data can't separate the two effects. How a tool gets used is only part of what you're evaluating.
A third mechanism is specific to AI video interview tools, and the known problem is worth naming directly. Some platforms score candidates on facial expression analysis, vocal tone, or speaking pace. These signals sit far from the skills a role actually requires, and the scientific validity of scoring them is weak: the American Psychological Association has raised serious concerns (APA, 2023). The disparity risk is also high, because expression- and tone-based scoring tends to track culture, language, and individual style rather than ability. TestGorilla deliberately avoids this modality: where we use AI video scoring, it reads only the transcript of what a candidate says, assessed against job-relevant criteria, and ignores facial expression, tone, and pace.
Every hire is a decision made under conditions of radical information asymmetry. The candidate presents the best version of themselves. The employer describes the culture they aspire to, not quite the one they have. Neither is lying. Both are responding rationally to genuine uncertainty.
This asymmetry has direct implications for every hiring tool, whether or not it uses AI. When application documents such as resumes and cover letters are the primary screen, you're betting on signals whose informational content has been degrading for years. Credential signals were never as reliable as hiring processes treated them, and they have weakened further, in part because the signaling arms race of elite credentialing has decoupled many credentials from actual ability. And AI writing tools are now eroding the application document signal further: a polished, tailored resume no longer separates a strong candidate from a weak one.
Assessment doesn't eliminate this information asymmetry, but it changes what the decision rests on. A resume, a cover letter, or an interview answer is a claim about ability; a well-designed assessment is a sample of it. Moving the basis of judgment from a document anyone can manufacture to a task the candidate has to actually perform produces a signal that is harder to game, easier to compare fairly across candidates, and tied directly to what the role requires. That is a more reliable basis for a decision, not a perfect one, and it's where a skills-based assessment earns its place over document screening.
A separate but related risk is upstream of any assessment tool: bias introduced by AI resume-screening tools. Research from the University of Chicago's Becker Friedman Institute has documented significant name-based discrimination in resume screening: distinctively Black names reduced the probability of employer contact by 2.1 percentage points relative to distinctively white names across a sample of Fortune 500 companies.
Resume screening tools that parse applicant names, or proxy signals that correlate with names, or more broadly with protected characteristics, inherit this risk. A skills-first process that shifts the scored output to demonstrated task performance reduces exposure to this class of problem.
Biweekly updates. No spam. Unsubscribe any time.
The bias mechanisms above share a common root: the AI is trained on or applied to inputs that carry historical discrimination. Well-designed skills assessment avoids this by starting from a different premise. Instead of asking "does this candidate resemble our current top performers?", it asks: "can this candidate do the job?"
The practical differences worth understanding:
In most major jurisdictions, including the US (UGESP), UK, South Africa, and the EU, employment discrimination law evaluates whether a selection tool is sufficiently related to the job it is used for.
A test that measures coding ability, numerical reasoning, or situational judgment for a specific role measures something directly relevant to job performance. A test that scores how closely a candidate's micro-expressions or reaction-time patterns match the behavioral profile of existing employees measures resemblance, not ability, and is far harder to defend in a legal challenge.
TestGorilla's library of 350+ tests is built on systematic job analysis, drawing on the US Department of Labor's O*NET and the EU's ESCO frameworks, so test content maps to the knowledge, skills, and abilities each role requires.
In games-based behavioral assessment, candidate scores are typically calibrated against a benchmark built from existing employees' performance profiles, which means the scoring algorithm learns to replicate the demographic composition of that group.
TestGorilla's skills tests do not use employer-specific top-performer benchmarks. For multiple-choice and knowledge tests, a candidate's final score does not depend on who your current employees are. The final percentile score depends on the norm group (TestGorilla's global talent pool) and how a candidate performed relative to it, not on who your current employees are.
Our recommended metric is the percentile score, which compares a candidate's performance against TestGorilla's global talent pool of all candidates who have taken that test, a far more representative and diverse benchmark than any single organization's incumbent workforce.
For open-response questions, AI video interviews, and resume scoring, TestGorilla uses rational keying: scoring against explicit, expert-defined criteria rather than patterns mined from historical decisions.
TestGorilla provides a library of ready-made questions and scoring criteria vetted by IO psychologists and subject-matter experts, and customers can use these as-is, adjust them, or define their own with their own experts. The AI scores how well a candidate's response maps to expert-defined criteria; it does not learn which candidates to favor by mining statistical patterns in historical hiring decisions.
This is a meaningful structural difference from empirical keying (the "black box" approach) and is what enables the scoring to be auditable, explainable, and defensible. Customers can review, adjust, and override any AI-generated score.
When a hiring decision rests on a single score from a single algorithmic model, any bias in that model is unchecked. Combining multiple independent measures (a skills test, a cognitive reasoning test, a structured behavioral question, a work sample) means no single score determines an outcome.
Each test type captures a different facet of the candidate and the role. This also reduces the structural risk of any one measure: if a single test has a modest adverse impact footprint, the other tests in the battery may provide compensating information, and a candidate who underperforms on one dimension may demonstrate clear strength on others.
A candidate who receives a low score on a games-based "trust and care" behavioral test has no meaningful way to understand or appeal that outcome.
TestGorilla's internal research shows that candidates rate our tests as accurately reflecting the skills they were told the test was measuring. Candidates can see their overall scores on tests they complete (where the hiring team enables this). That transparency isn't just a regulatory concern; it's a basic fairness requirement and materially affects candidate experience and your employer brand.
When every employer in a sector uses the same algorithmic model to filter candidates, the sector stops producing independent judgments. It produces one judgment, repeated. An employer that builds its own skills-based process, anchored to what the specific role actually requires, produces a genuinely independent evaluation. That is better for hiring quality and structurally fairer to candidates.
Here is an honest account of what TestGorilla does and where our monitoring still has limitations. Skills-based assessment is not free of bias risk. It can enter at test design, question selection, norm-setting, or scoring calibration, so the task is to actively manage and disclose that risk at each stage, which is what the rest of this section describes.
Every skills test in TestGorilla's library is developed by subject matter experts in the relevant domain, using a seven-phase development lifecycle that includes test design, independent SME review, copy editing for inclusive language, quality assurance, and ongoing monitoring.
We do not use employer-specific top-performer benchmarks. Our scoring is anchored to demonstrated performance on the task, benchmarked against our global talent pool.
Fairness is designed in from the start, not added at the end. Our SMEs complete training on avoiding cultural idioms, regional slang, and references that disadvantage non-native speakers.
Every item undergoes an inclusive language audit using a proprietary style guide, with a diverse multicultural name bank for any items that involve named scenarios. The result: in a large-scale audit of 6,886 items across 119 skills-based tests, only 0.61% of items (42 in total) were flagged for a large gender differential item functioning (DIF) effect size. The direction of that small effect was not systematic: among the flagged items, 54% favored men and 46% favored women. The handful of items showing large DIF did not lean consistently toward either gender, so there is no systematic gender bias baked into the item bank.
TestGorilla runs adverse impact analysis on test score distributions across demographic groups using the four-fifths (80%) rule, consistent with the Uniform Guidelines on Employee Selection Procedures (UGESP). We additionally monitor score patterns using Cohen's d, the standardized mean difference metric, to separate two kinds of gaps.
Some group differences are expected: in certain skill domains, research consistently finds small differences that reflect broader factors such as unequal access to education, and we monitor these rather than mask them. Others are unexpected: differences that diverge from established benchmarks or lack support in outside research. These are flagged as problematic and trigger an immediate Science team review for construct-irrelevant variance or systematic bias.
When aggregate monitoring surfaces a group difference large enough to raise an adverse-impact concern, we investigate at the item level using DIF analysis to see whether specific questions are driving it. Items flagged for DIF are removed from the active item bank and reviewed by our psychometricians so the same issue can be avoided in future test development.
We offer local adverse impact studies for our eligible customers. We don't publish these openly, because demographic distributions and hiring contexts vary significantly across organizations, and context matters for interpretation. We encourage customers to run their own adverse impact studies on their specific hiring pipelines, and we actively support that.
Every critique of AI scoring rests on a quiet assumption: that the thing it replaces is fair. It isn’t.
The realistic alternative to AI scoring is a human interview panel, and human panels carry well-documented bias. Reviewers drift with fatigue, mood, and affinity for candidates who remind them of themselves.
When human reviewers score the same thing, they tend to agree only 30–50% of the time: under 50% agreement among 244 recruiters inferring personality (Cole et al., 2009), about 41% agreement across three raters scoring 61 resumes (Mu et al., 2025), and around 33% agreement in a meta-analysis of journal peer reviews spanning 48 studies and more than 19,000 rating items (Bornmann et al., 2010). Two qualified reviewers watching the same answer frequently disagree.
So perfection is the wrong bar to hold our AI to. The fair test is whether it achieves the same level or surpasses the human review it aids. On the evidence, it does. Across 21,000+ ratings, our optimized AI configuration triggered 16% fewer adverse impact flags than human expert raters under identical conditions. It treats the 500th candidate of the week exactly as it treated the first, holding test-retest reliability at ICC(3,1) > 0.99. The fatigue, affinity bias, and inconsistency that creep into manual review are the very failures well-engineered AI scoring strips out.
For open-response questions, AI video interviews, and resume scoring, TestGorilla's AI acts as a synthetic expert rater, not an autonomous decision-maker. Here is the empirical evidence:
Validity: Benchmarked against 150 calibrated, pre-screened subject matter experts across 1,500 candidate responses, our AI scoring achieved Pearson correlations of r = 0.84–0.89 with expert judgment, a level of alignment that confirms the AI is identifying the same latent constructs as trained human experts.
Anti-gaming: Standard LLMs inflate scores for longer answers regardless of quality, a "verbosity trap" that our architecture specifically counters. In adversarial testing, our model maintained rank stability of r = 0.93 between original and artificially lengthened responses. Keyword stuffing produced zero score inflation.
For resume scoring, a masking study conducted on 115,924 profiles across 14 configurations identified an optimal configuration that removes high-bias, low-validity signals (prestige employers, elite university names, graduation dates, specific locations) while preserving behavioral evidence of skill application. This optimal masking configuration achieved a maximum effect size of 0.137 across age, gender, and ethnicity, and reduced bias in targeted scenarios by as much as 28% compared to unmasked review.
Crucially, the customer always has the final say. Every AI-generated score comes with a natural-language rationale. Customers can review that rationale, override the score, and add their own notes. The platform is designed to make automated-only decisions structurally difficult. Our recommendation is to use assessment scores as one significant input among several, with human review at the decision point.
TestGorilla's Science team is continually working to analyze and publish relevant validity and fairness data, and to broaden the evidence base across our test library.
We also partner closely with customers and prospects to meet their specific evaluation needs, including local validity and adverse impact studies, and we're committed to strengthening and sharing this work over time.
Candidates who complete TestGorilla assessments know what skills are being measured. The tests they complete relate to recognizable, job-relevant tasks. When the hiring team enables results sharing, candidates can view their overall test scores.
How scoring rationale reaches candidates has changed substantially. Candidates who create a TestGorilla profile after completing an assessment can access AI-generated feedback for each test result: a plain-language breakdown of what their score reflects, what they handled well, where the gap is, and one specific area to work on next. The feedback is private. No employer sees it.
The redesigned end-of-assessment experience also makes next steps explicit; a stepper guides candidates through profile creation to unlock their scores, replacing what had previously been a dead end.
Creel, K., Bomasanni, R., & Banna, S. (2026). Algorithmic monocultures in hiring. algorithmichiring.github.io.
Kline, P.M., Rose, E.K., & Walters, C.R. (2022). Systemic Discrimination Among Large US Employers. The Quarterly Journal of Economics, 137(4). (Note: covers name-based discrimination in resume screening across Fortune 500 companies.)
Bornmann, L., Mutz, R., & Daniel, H.-D. (2010). A reliability-generalization study of journal peer reviews. PLoS ONE, 5(12).
Cole, M.S., Feild, H.S., Giles, W.F., & Harris, S.G. (2009). Recruiters’ inferences of applicant personality. Human Performance, 22(3).
Mu, E., et al. (2025). Inter-rater agreement in resume screening. Working paper.
Sackett, P.R., Zhang, C., Berry, C.M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection. Journal of Applied Psychology, 107(11), 2040.
Sackett, P.R., Zhang, C., Berry, C.M., & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 1–18.
American Psychological Association (2023). APA guidance on AI in employment settings.
Uniform Guidelines on Employee Selection Procedures (1978). Federal Register, US EEOC.
TestGorilla Technical Manual for Skills-Based Hiring Assessments (2026). Internal documentation.
TestGorilla Technical Manual for AI-Driven Assessments (2026). Internal documentation.
Why not try TestGorilla for free, and see what happens when you put skills first.