Built on science. Run on evidence.

Every assessment score on TestGorilla rests on two pillars. Measurement science that's predicted job performance for decades, and engineering that applies it consistently to every candidate, at any volume. See how your candidates' scores are actually calculated. No black boxes.

Grounded in expertiseReliable, valid, fairHumans stay in control

How TestGorilla works

TestGorilla assessments measure job-relevant skills with validated tests and structured AI interviews. Our tests are grounded in subject matter expertise and calibrated before launch, then continuously monitored for reliability, validity, and fairness.

AI scores open-ended responses against structured rubrics, with the reasoning shown for every score. And your team stays in control. AI features are optional, and every AI score can be reviewed, adjusted, or overridden.

How a test gets made

No test ships on someone's say-so. Our tests are built to help predict job performance. Each goes through a rigorous 5-stage process.

01

Define what to measure

Every test starts with a precise definition of the skill it measures. We ground that definition in established frameworks like the US Department of Labor's O*NET skills database or the European Commission's ESCO framework. Then we sharpen it against what we hear from customers and see in the market, because some of the skills companies hire for hardest today didn't exist a year ago.

02

Build against the definition

Test content is developed against that definition from a deep pool of subject matter expertise. Knowing a subject and knowing how to measure it are different skills. Our process requires both.

03

Verify it independently

Before anything moves forward, the content is checked for accuracy, technical correctness, and alignment by experts independent of whoever built it. Reviewers have to agree on every point, and everything is copyedited until it's error-free.

04

Calibrate the difficulty

Before a test goes live, we set its difficulty and scoring standards using structured expert-judgment methods like the modified Angoff method, real test-taker data where it's available, or both. That's how a percentile means something on day one, then gets sharper as real-world results come in.

05

Let the data keep checking

After launch, our Science team runs ongoing psychometric analyses on live test data. Reliability, content validity, construct validity, and criterion-related validity all get monitored, alongside demographic fairness checks like group differences, differential item functioning, and adverse impact. The results are published in each test's fact sheet, so you can read the numbers rather than take our word for them. Questions that stop performing get fixed or retired. And when the data shows a skill is shifting, the findings feed straight back into stage 1.

What happens between an answer and a score

Scores on TestGorilla are consistent, comparable, and fair. Same principles whether a candidate answers a multiple-choice question or records a video interview.

Skills tests

Each test draws from a large question bank. An item-drawing algorithm builds every test instance to be unique and balanced for difficulty. Scoring stays consistent and fair across candidates.

Within a single assessment, every candidate answers the same questions, making their scores comparable. Percentile scores are reported against calibrated global norms. A percentile tells you how a candidate performed relative to other candidates, adjusted for how hard their questions were — so a strong score means the same thing whichever version of the test someone sat.

Open-ended responses and AI interviews

Our AI scores candidate responses against structured rubrics. For interviews and questions from our library, those rubrics are built and verified by our Science team. For custom interviews you create, AI co-creates the scoring criteria with you and gives you full edit access. You stay in control.

Each criterion is scored on a defined scale, and the reasoning behind every score is shown alongside it. You see what was evaluated, how the response measured against each criterion, and why it got the score it did. Agree with it, adjust it, or override it. The AI drafts the evaluation. Your team owns the decision.

Scoring looks only at the content of what a candidate says. Tone, accent, appearance, and background never factor in.

AI scoring

What our model is trained on

Here's how our AI scoring works:

It’s benchmarked against human experts.

Our scoring is calibrated against evaluations from trained expert scorers. The standard the AI applies is the standard an expert panel would apply. Consistently from candidate number 4 to 4,000.

Your candidates’ data doesn’t train AI models.

TestGorilla doesn’t train AI models on your candidates’ data. Personally identifiable information (PII) from customers, candidates, and job seekers is never shared with generative AI models. External model providers we work with are contractually prohibited from training on any data we share. Any assessment data we use to evaluate and calibrate scoring quality is anonymized first.

It’s tested for bias, on purpose.

We test our scoring for bias and adverse impact across demographic groups before and after deployment. If scores differ between groups for the same quality of answer, our Science team investigates it thoroughly. When AI reads resumes, personal details such as names, photos, pronouns, and locations are automatically removed before scoring, so bias-prone inputs never reach the model.

And to state the obvious plainly: AI never makes a hiring decision on this platform. It scores, summarizes, and suggests. People decide.

How we keep tests fair

Structured, validated tests give every candidate the same opportunity. The same conditions. The same standard. Our Science team monitors live test data for group differences, right down to single questions that behave differently for different groups. Anything that shows bias gets revised or removed.

Structured, validated tests

give every candidate the same opportunity. The same conditions. The same standard.

Our Science team

monitors live test data for group differences, right down to single questions that behave differently for different groups. Anything that shows bias gets revised or removed.

Recognized standards

are used to build and monitor, including the EEOC's Uniform Guidelines on Employee Selection Procedures (UGESP), the SIOP Principles for the Validation and Use of Personnel Selection Procedures, and the Standards for Educational and Psychological Testing.

Concrete evidence, not gut feel

How we protect test integrity

Remote testing raises a fair question: now that AI can sit alongside any candidate, how do you know the score reflects the candidate? Most candidates are honest, and integrity protection exists to keep it that way. It works in two layers: 

Test integrity

Designed to resist cheating from the start

Question sets are cycled by algorithm. No two test instances are the same, even when different companies use the same test. Questions expire once they’ve been used too often, before they can leak. Candidates can’t sign up as customers to study the tests, since account creation requires a business email. And within your assessment, every candidate faces the same questions, so comparisons stay clean.

Formats like AI interviews and job simulations go further. A live, adaptive exchange is much harder to fake than a static quiz.

The Trust Layer second

Every assessment runs inside the TestGorilla Trust Layer. This gives your assessment our built-in bundle of integrity safeguards. It doesn't record video or watch candidates. Signals roll up into three clear tiers, so a flag is a data point for your review, never an automatic disqualification. A human always makes the final call. None of this is about catching candidates out. It's about protecting the honest majority, because your strongest candidates lose the most when someone else fakes a score.

  • An honesty agreement candidates sign upfront
  • AI browser agents blocked before test content loads
  • Behavior monitoring for full-screen exits, tab switches, copy-paste, and tampering
  • Optional camera snapshots and ID verification
Fingerprint validation security
Security and compliance

We protect the data

Candidate and customer data are handled under GDPR, and our security practices are SOC 2 Type II audited. We're actively preparing for the EU AI Act, and we track evolving AI regulation as part of our science and legal teams' standing work — not as a scramble when laws land.

The details live in our Trust Center, where procurement teams can review our security practices directly. We store images, video, and audio for up to 6 months, and other PII such as names and emails for up to 24 months. After that, the PII is destroyed. Anonymous data like platform interactions and responses is retained for benchmarking, psychometric analyses, and product improvements.

Candidates always know when AI is used in their assessment. Transparency isn't a setting. It's the default.

Data decode
From our Science team

Make sense of the science

Want to dive deeper? Our Science team writes about how assessment actually works, including the parts most vendors avoid.

  • Why it's impossible to prevent cheating
  • How to interpret test fact sheets, part 1: reliability
  • How to interpret test fact sheets, part 2: validity
  • An introduction to percentile scores
  • Understanding adverse impact: examples, best practices, and FAQ
  • How to set cut-off scores correctly

See the evidence for yourself

Build your first assessment and look at what comes back. Every score, every rubric, every piece of reasoning, right there to inspect.

Frequently asked questions

Technology behind TestGorilla

Only if you turn it on. There's no mandatory use of AI on TestGorilla. When enabled, AI scores open-ended responses — like AI interview answers and custom questions — against structured rubrics. Multiple-choice skills tests are scored automatically against calibrated norms without AI. Every AI score comes with visible reasoning, and your team can review, adjust, or override any score.