Every assessment score on TestGorilla rests on two pillars. Measurement science that's predicted job performance for decades, and engineering that applies it consistently to every candidate, at any volume. See how your candidates' scores are actually calculated. No black boxes.

TestGorilla assessments measure job-relevant skills with validated tests and structured AI interviews. Our tests are grounded in subject matter expertise and calibrated before launch, then continuously monitored for reliability, validity, and fairness.
AI scores open-ended responses against structured rubrics, with the reasoning shown for every score. And your team stays in control. AI features are optional, and every AI score can be reviewed, adjusted, or overridden.
No test ships on someone's say-so. Our tests are built to help predict job performance. Each goes through a rigorous 5-stage process.
Every test starts with a precise definition of the skill it measures. We ground that definition in established frameworks like the US Department of Labor's O*NET skills database or the European Commission's ESCO framework. Then we sharpen it against what we hear from customers and see in the market, because some of the skills companies hire for hardest today didn't exist a year ago.
Test content is developed against that definition from a deep pool of subject matter expertise. Knowing a subject and knowing how to measure it are different skills. Our process requires both.
Before anything moves forward, the content is checked for accuracy, technical correctness, and alignment by experts independent of whoever built it. Reviewers have to agree on every point, and everything is copyedited until it's error-free.
Before a test goes live, we set its difficulty and scoring standards using structured expert-judgment methods like the modified Angoff method, real test-taker data where it's available, or both. That's how a percentile means something on day one, then gets sharper as real-world results come in.
After launch, our Science team runs ongoing psychometric analyses on live test data. Reliability, content validity, construct validity, and criterion-related validity all get monitored, alongside demographic fairness checks like group differences, differential item functioning, and adverse impact. The results are published in each test's fact sheet, so you can read the numbers rather than take our word for them. Questions that stop performing get fixed or retired. And when the data shows a skill is shifting, the findings feed straight back into stage 1.
Scores on TestGorilla are consistent, comparable, and fair. Same principles whether a candidate answers a multiple-choice question or records a video interview.
Each test draws from a large question bank. An item-drawing algorithm builds every test instance to be unique and balanced for difficulty. Scoring stays consistent and fair across candidates.
Within a single assessment, every candidate answers the same questions, making their scores comparable. Percentile scores are reported against calibrated global norms. A percentile tells you how a candidate performed relative to other candidates, adjusted for how hard their questions were — so a strong score means the same thing whichever version of the test someone sat.
Our AI scores candidate responses against structured rubrics. For interviews and questions from our library, those rubrics are built and verified by our Science team. For custom interviews you create, AI co-creates the scoring criteria with you and gives you full edit access. You stay in control.
Each criterion is scored on a defined scale, and the reasoning behind every score is shown alongside it. You see what was evaluated, how the response measured against each criterion, and why it got the score it did. Agree with it, adjust it, or override it. The AI drafts the evaluation. Your team owns the decision.
Scoring looks only at the content of what a candidate says. Tone, accent, appearance, and background never factor in.
Here's how our AI scoring works:
Our scoring is calibrated against evaluations from trained expert scorers. The standard the AI applies is the standard an expert panel would apply. Consistently from candidate number 4 to 4,000.
TestGorilla doesn’t train AI models on your candidates’ data. Personally identifiable information (PII) from customers, candidates, and job seekers is never shared with generative AI models. External model providers we work with are contractually prohibited from training on any data we share. Any assessment data we use to evaluate and calibrate scoring quality is anonymized first.
We test our scoring for bias and adverse impact across demographic groups before and after deployment. If scores differ between groups for the same quality of answer, our Science team investigates it thoroughly. When AI reads resumes, personal details such as names, photos, pronouns, and locations are automatically removed before scoring, so bias-prone inputs never reach the model.
And to state the obvious plainly: AI never makes a hiring decision on this platform. It scores, summarizes, and suggests. People decide.
Structured, validated tests give every candidate the same opportunity. The same conditions. The same standard. Our Science team monitors live test data for group differences, right down to single questions that behave differently for different groups. Anything that shows bias gets revised or removed.
give every candidate the same opportunity. The same conditions. The same standard.
monitors live test data for group differences, right down to single questions that behave differently for different groups. Anything that shows bias gets revised or removed.
are used to build and monitor, including the EEOC's Uniform Guidelines on Employee Selection Procedures (UGESP), the SIOP Principles for the Validation and Use of Personnel Selection Procedures, and the Standards for Educational and Psychological Testing.
Remote testing raises a fair question: now that AI can sit alongside any candidate, how do you know the score reflects the candidate? Most candidates are honest, and integrity protection exists to keep it that way. It works in two layers:
Question sets are cycled by algorithm. No two test instances are the same, even when different companies use the same test. Questions expire once they’ve been used too often, before they can leak. Candidates can’t sign up as customers to study the tests, since account creation requires a business email. And within your assessment, every candidate faces the same questions, so comparisons stay clean.
Formats like AI interviews and job simulations go further. A live, adaptive exchange is much harder to fake than a static quiz.
Every assessment runs inside the TestGorilla Trust Layer. This gives your assessment our built-in bundle of integrity safeguards. It doesn't record video or watch candidates. Signals roll up into three clear tiers, so a flag is a data point for your review, never an automatic disqualification. A human always makes the final call. None of this is about catching candidates out. It's about protecting the honest majority, because your strongest candidates lose the most when someone else fakes a score.
Candidate and customer data are handled under GDPR, and our security practices are SOC 2 Type II audited. We're actively preparing for the EU AI Act, and we track evolving AI regulation as part of our science and legal teams' standing work — not as a scramble when laws land.
The details live in our Trust Center, where procurement teams can review our security practices directly. We store images, video, and audio for up to 6 months, and other PII such as names and emails for up to 24 months. After that, the PII is destroyed. Anonymous data like platform interactions and responses is retained for benchmarking, psychometric analyses, and product improvements.
Candidates always know when AI is used in their assessment. Transparency isn't a setting. It's the default.
Want to dive deeper? Our Science team writes about how assessment actually works, including the parts most vendors avoid.
Build your first assessment and look at what comes back. Every score, every rubric, every piece of reasoning, right there to inspect.
Only if you turn it on. There's no mandatory use of AI on TestGorilla. When enabled, AI scores open-ended responses — like AI interview answers and custom questions — against structured rubrics. Multiple-choice skills tests are scored automatically against calibrated norms without AI. Every AI score comes with visible reasoning, and your team can review, adjust, or override any score.