Conversations about hiring bias in engineering tend to split into two camps that don't actually talk to each other. One camp treats bias as mostly solved once you've trained interviewers and written a rubric. The other treats it as intractable — a function of broader industry dynamics that any individual team's process can't really move. Neither camp is right, and the productive middle gets underrepresented in most write-ups.
Bias in hiring isn't a single thing you can eliminate. It's a bundle of small inconsistencies in how candidates get read, and most of the levers that actually move the needle are structural — they change what a reviewer is looking at, rather than what they believe about bias in the abstract.
What bias actually looks like in practice
The clearest way to think about hiring bias isn't through the language of intent. Most engineers interviewing for their company have good intent. The issue is more mundane: reviewers read different candidates against different mental templates.
Candidate A walks in, speaks confidently, says they're from a big-name company, and gets benefit-of-the-doubt on answers that were actually pretty thin. Candidate B walks in, hesitates, says they're from a company nobody's heard of, and gets scrutinised for weaknesses in answers that were actually stronger. The interviewer leaves both sessions with a clean explanation of what they saw. They are probably not conscious of the reference shift.
This pattern shows up in enough interviews that anonymous scoring studies consistently find reviewer bias of this kind. You can spot it in your own reviews, if you look back at the criteria you applied to three different candidates you read this month. The criteria almost certainly shifted.
Levers that actually work
A few structural moves reliably reduce that kind of shift. None of them eliminate bias — the research is clear that nothing does — but they meaningfully narrow the variance between how candidates get read.
Shared artefacts to anchor disagreement. When a debrief is a memory negotiation, the senior voice tends to win regardless of the strength of the candidate. When it's a conversation anchored in a written submission or a transcript everyone can quote from, disagreements get grounded in specifics. That alone narrows the variance in outcomes between candidates reviewed at 10am and candidates reviewed at 5pm.
Written evaluation criteria before reading submissions. This is boring advice that most teams skip. Writing down what a good submission looks like before you open any specific candidate's work means you're reading against a pre-decided standard rather than calibrating against the last submission you happened to see. That reduces the drift that bias research shows is one of the largest sources of uneven outcomes.
Blind async review where feasible. Reviewing a candidate's written submission with the name and CV hidden is unusual in engineering hiring but works well. Not everywhere — you usually un-blind before the follow-up interview, because the conversation needs the full context. For the initial read of a take-home or async artefact, blinding is cheap and removes a large class of first-impression effects.
Pattern review after the fact. At least once a quarter, look back at which candidates got passed and which got rejected, and look for patterns correlated with things that shouldn't correlate. Accent. University. The company they came from. The length of their CV. If a pattern shows up, the process isn't doing what you think it's doing. This is the slowest lever but the one most likely to surface the kinds of bias a rubric can't catch.
Shared vocabulary across interviewers. "Strong hire" and "weak pass" mean different things to different reviewers unless you've calibrated them. Two interviewers using the same words to mean different things is one of the most common causes of uneven candidate reads. Calibrating the vocabulary is a cheaper fix than most teams expect.
Levers that don't
A few things are often proposed as bias reduction but don't reliably move the needle.
Unconscious-bias training alone. The evidence on training as a standalone intervention is weak — most programmes produce awareness without behaviour change. Training in combination with structural changes has some effect; training without them rarely does. It's useful, but not instead of the structural moves.
Rubric inflation. Adding more rows to your scoring rubric tends to increase the appearance of objectivity without changing the outcomes. Reviewers fill in the rubric according to their overall read of the candidate, and the overall read is where the bias lives. Rubric detail can actually obscure the bias by making the process look more quantitative than it is.
Scoring demographic targets. Setting hiring targets at the pipeline level is a legitimate organisational decision but doesn't reduce bias in the evaluation step itself. It changes who gets into the pipeline, which matters; it doesn't change how candidates get read once they're in.
What this looks like in a small team
You don't need a dedicated DEI team to run this. For a team hiring 5–15 engineers a year, the useful version of the above is:
Write evaluation criteria before reviewing any candidate's work for a given role. Review the first pass of async submissions without the candidate's name visible. Anchor debrief discussions in the submission itself rather than in memory. Once a quarter, pull the list of candidates you interviewed and check for patterns you didn't intend. Keep a shared glossary of what the vocabulary on your debrief means.
That's a lot, but the individual steps are each an hour's work. The cost is much lower than most teams fear going in.
At CriticCode we build around the artefact-anchored half of this by default — the submission includes answers, transcript, and paste events, which is what most debriefs end up pointing at. The rest is process. Process is where most bias-reduction work actually lives, regardless of tool.
The honest close
Bias reduction is ongoing work, not a project with an end date. Teams that treat it as something you can finish tend to stop improving after the first intervention, which is usually the weakest one. Teams that treat it as a standing commitment — review quarterly, edit as the team changes, keep the criteria fresh — tend to produce more consistent outcomes over time, and their hires tend to look different, in a useful way, after a couple of years.