A proxy test measures the proxy
Why a hiring exercise that is not the job tells you very little.
If a test doesn’t reflect the job you’ll be doing, it isn’t a test of whether you can do the job.
That sounds obvious and it’s routinely ignored, mostly because a test that’s genuinely like the work is expensive to build and awkward to mark. What you get instead is something that assesses a nearby skill and stands in for the real one. Everybody involved knows it’s a stand-in and everybody quietly assumes the correlation is strong enough not to matter.
The trouble is that a proxy test measures the proxy. If your exercise only touches the skill you care about at an angle, then the result gets shaped by all the other things it touches head on and you’ve no way to separate them afterwards. You end up with a number that feels like evidence and is mostly made of something else.
Whiteboard exercises are the obvious case. Unless remembering an algorithm and explaining it standing up is something the role actually involves, the exercise is not testing what you think. It sort of tests how well someone explains an algorithm. What it really tests is memory and presentation and how a person handles being watched while thinking. Those are real skills and some jobs need them. They’re not the same skill as writing good software and someone brilliant at a whiteboard isn’t necessarily anyone you want near your codebase.
Hiring somebody because you were impressed by how well they explained something in front of a board feels wrong and it feels wrong for a reason. It feels especially wrong to people who aren’t good at the tangential stuff, which is a large number of perfectly capable developers who now have to get good at a performance in order to be allowed to do a job that doesn’t contain one.
Notice too that whether something counts as cheating depends entirely on the job rather than on the test. If you’re hiring for a role where you’re expected to know an application cold, then looking in the help is cheating. If reaching for the documentation is normal in the actual work, then it isn’t cheating, it’s the job and penalising it means selecting for people who’ve memorised things you don’t need memorised.
Which is why the clever alternatives often aren’t better. Take an exercise that asks someone to find their way around an unfamiliar application and work something out from the help. It sounds like it’s testing resourcefulness and it is, a bit. It’s also testing how comfortable they are in software they’ve never seen and how they cope with an odd interface and whether they happen to have used that specific tool before. Pick something like Blender and you’ve added a large advantage for anyone who’s touched 3D software and a large penalty for everyone who hasn’t, neither of which you meant to measure.
None of this means testing is hopeless. It means the further your test sits from the work, the more of the result is made of things you didn’t intend to assess and can’t see. The closer it sits, the more expensive it gets and that trade is the real decision. Most companies make it badly and then describe the outcome as a hiring bar.
The honest test is boring and it’s the one almost nobody runs. Give someone a small piece of the work you actually have, with the resources they’d actually have, and enough time to do it properly, then look at what comes back and talk to them about it. That costs real money and real hours from people who are busy, which is precisely why the industry keeps reaching for something cheaper and then describing the cheap thing as rigorous.
A bar you can clear in forty minutes at a whiteboard isn’t a high bar. It’s a cheap one, and those aren’t the same thing however it feels from the other side of the table.