TestGorilla LogoTestGorilla Logo
Pricing
homeLibraryBlog
June 26, 2025

No one wants to tell you this, but it’s impossible to prevent cheating

TG-Profile-Generic
Bogdan Radičević

Here’s the uncomfortable truth nobody in HR tech wants to say out loud: no assessment has ever been 100% cheat-proof. Not before AI. Not now.

Proctored exams have had answer keys floating around since the 1980s. Candidates have been exaggerating on resumes for decades. The problem isn’t new. What’s new is the scale and the speed. Every candidate now has a language model in their pocket that can draft a coherent answer to almost any text-based question in under three seconds.

So what do you do about it? Build higher walls? You’ll lose. LLMs are very good at scaling walls. So are people.

The smarter move is to stop building walls and start building assessments where AI assistance returns diminishing results.

Cheating was never the real problem

When home testing went mainstream in the early 2000s, the industry coined a new acronym: UIT, Unproctored Internet Testing. The concern was the same one you’re probably hearing in your team right now: “But candidates can cheat if they test at home!”

What research actually showed was that proctored and unproctored assessments produce similar levels of predictive validity. In other words, both predict job success. The "problem" of cheating, at scale, turned out to be smaller than the fear of it.

The numbers back this up. Internal TestGorilla shows only one in six candidates has attempted to cheat. Designing your entire hiring process around the behavior of a minority means creating friction, cost, and drop-off for the majority who are acting in good faith.

Today, the acronym has changed. Then it was UIT. Now it’s AI. The underlying logic hasn’t.

What detection actually buys you

Detection matters. It creates friction and leaves a trail. TestGorilla’s behavioral monitoring layer - webcam snapshots every 30 seconds, tab-switching detection, paste tracking, full-screen enforcement, and a three-tier flagging system - does exactly this. The structural layer adds another line of defense: randomized question banks with exposure limits, server-side timers that keep running even when a window closes, and honesty agreements surfaced at the point of submission.

None of this is magic. None of it is a guarantee. But it raises the cost of cheating enough to deter the opportunistic majority while giving you the data to ask better questions about outliers. It doesn’t and cannot tell you who cheated; it can point to behaviours which you can contextualise based on the role you’re hiring for and the corresponding skills assessment.

The problem is when detection becomes the whole strategy. Playing whack-a-mole with LLMs is a game you will not win. The tools are too fast, too accessible, and too good at text. A different frame is required.

The best insights on HR and recruitment, delivered to your inbox.

Biweekly updates. No spam. Unsubscribe any time.

The real breakthrough: modality switching

Think of the difference between an open-book exam and a live debate. You can Google facts all day. But you can’t outsource thinking on your feet when someone is watching you respond in real time.

That’s the logic behind Immersive Job Simulations and Conversational AI Video Interviews. These formats require live audio responses within 90-second windows. There is no pause button. There is no draft-and-paste.

LLMs are getting faster but even real-time voice assistance can’t replicate the cognitive overhead of thinking on your feet in a structured live format. The latency, the delivery constraints, and the absence of a text buffer make high-quality live responses far harder to outsource than written answers. 

Do candidates cheat? Of course. Second screens, reading answers, asking for questions to be repeated thrice to buy time. All of it. 

Our suggestion? Build an assessment that naturally filters for genuine capability rather than relying solely on detection.

This creates a structural barrier that detection tools can’t replicate because it doesn’t try to catch cheating. It makes cheating irrelevant.

Related posts

How to assess AI fluency for non-technical roles

Skills-based hiring is the best defense against AI

Defining “AI fluency” and why hiring rubrics fail

What the research shows

The validity argument here isn’t incidental. Multi-method batteries - combining cognitive tests with video simulations and structured interviews - increase predictive validity and AI-resistance simultaneously. Cheating one component doesn’t help with the others. The assessment battery is designed so that the components don’t correlate in ways a candidate can exploit.

TestGorilla’s Conversational AI Video Interviews are validated on the same psychometric framework as its custom questions: a scoring engine benchmarked against calibrated subject matter experts, built to evaluate behavioral evidence rather than surface-level signals.

Real-time, audio-required formats don’t just close the LLM loophole. They measure the skills that actually predict job performance: communication under pressure, adaptive problem-solving, and thinking in the moment.

The same logic applies to cognitive ability tests and situational judgment tests. These formats require candidates to reason through novel problems in constrained time. The signal they generate doesn’t improve much when a language model is assisting, because the bottleneck is reasoning speed, not content generation.

Trust, but verify

The practical framework for all of this is one that predates LLMs: trust, but verify. Trust your candidate scores, and build a multi-stage process that lets you confirm them.

In practice, that means:

  • Evaluate candidates holistically. Test scores alongside resume signals and any other early-stage data. Multiple data points give you a more accurate picture than any single measure.

  • Use behavioral flags as prompts, not verdicts. A high flag count is a reason to ask a question, not a reason to reject. Ask the candidate about their test-taking experience. A candidate who aced the coding tests but has a thin resume? Ask them why that might be.

  • Verify at the next stage. Short-list candidates get a second look: skills-based questions, a live conversation, or a task-based exercise. This is where suspicious scores surface naturally.

  • Design the battery for coverage, not redundancy. Different test types measuring different constructs reduce the overall surface area for gaming. No single method should carry the full weight of a hiring decision.

Trust but verify in practice graphic

None of this has made the hiring manager's job easier. You have a longlist where you used to have a shortlist or a single-digit candidate pool if the role is new enough that the keyword universe barely exists yet. You're measuring against seven to ten yardsticks simultaneously: cognitive scores, personality profiles, situational judgment, role-specific skills, and increasingly, AI fluency markers that weren't in most job specs two years ago. 

And on top of the measurement itself, you're doing correlation work. Did the personality outcome track with the behavioral interview score? Does the cognitive result hold up against how they reasoned through the simulation? A gap isn't automatically a red flag, but it's a question. The data is richer than it's ever been. The interpretive load has scaled with it. Building confidence in a hiring decision now requires more steps, more cross-referencing, and more judgment calls than it did when the shortlist arrived pre-filtered and the biggest variable was whether someone interviewed well on the day. What you need isn't more data. It's a clear pathway through the data you already have.

What to do about it

No platform can guarantee zero cheating. Anyone who says otherwise is selling you something.

What you can control is the design. Configure your assessment battery so that AI assistance hits diminishing returns at every layer: behavioral monitoring and structural safeguards at the base, format switching at the top. Layer in a Conversational AI Video Interview or Immersive Job Simulation that requires live, spoken responses within strict time windows. Now you're not chasing detection signals — you're building a system where genuine capability is the path of least resistance.

Use the correlation signals. A personality profile that doesn't track with a behavioral interview score isn't a verdict — it's a prompt. Ask the question. The data isn't there to make the decision for you. It's there to tell you where to look.

And when the longlist feels unmanageable, resist the instinct to collapse back to a single yardstick. The complexity is the point. A candidate who scores consistently well across cognitive, behavioral, and live-response formats, without obvious gaps, is exactly what a multi-method battery is designed to surface. The noise is how you find the signal.

The goal was never to build a perfect process. It was to build a better one than a resume and an unstructured interview. By that measure, skills-based assessment — done well, with the right format mix — still wins.

The walls were never the answer. The format was.

Thanks for reading! For related TestGorilla blog content, check out: 

Sources

  1. Beaty, J. C., Nye, C. D., Borneman, M. J., Kantrowitz, T. M., Drasgow, F., & Grauer, E. (2011). Proctored Versus Unproctored Internet Tests: Are unproctored noncognitive tests as predictive of job performance?, International Journal of Selection and Assessment, 19(1), 1-10. International Journal of Selection and Assessment, 19(1)

You've scrolled this far

Why not try TestGorilla for free, and see what happens when you put skills first.

Free resources

Skills-based hiring handbook cover image
Ebook
The skills-based hiring handbook
Ebook
How to elevate employee onboarding
Top talent assessment platforms comparison guide - carousel image
Ebook
Top talent assessment platforms: A detailed guide
The blueprint for boosting your recruitment ROI cover image
Ebook
The blueprint for boosting your recruitment ROI
Skills-based hiring checklist cover image
Checklist
The skills-based hiring checklist
Onboarding email templates cover image
Checklist
Essential onboarding email templates
HR cheat sheet cover image
Checklist
The HR cheat sheet
Employee onboarding checklist cover
Checklist
Employee onboarding checklist
Key hiring metrics cheat sheet cover image
Checklist
Key hiring metrics cheat sheet
Ending AI Arms Race in Hiring Webinar
Checklist
Ending the AI arms race in hiring
It's not you, it's your hiring process
The dream job equation
The State of Skills-Based Hiring 2024