AI for Hiring Decisions: Where Claude and ChatGPT Actually Help
A grounded look at what generic AI use gets right and wrong when you're deciding on a hire, four prompts that genuinely sharpen the decision, and what changes when the AI applies a documented evidence-based hiring framework instead of generic interview advice.
Ask AI to help with a hiring decision and the easy failure mode is asking it the wrong question: "here are two resumes, who should I hire." That produces a confident-sounding answer built on whatever the model associates with a strong candidate, polish, credentials, the right buzzwords, which is exactly the kind of surface-level signal that structured hiring research has spent decades trying to get interviewers away from.
The better use isn't asking AI to decide. It's asking it to structure your own evidence well enough that the decision becomes clear on its own, which is a genuinely different, more useful kind of help than a ranked shortlist ever provides.
What does asking AI "who should I hire" actually produce?
Give an AI two resumes or two sets of interview notes and ask which candidate is stronger, and you'll get a fluent, confident answer, usually favoring whichever candidate's materials read as more polished or used more of the language the role's description contained. It sounds authoritative. It's also reasoning from surface features that correlate weakly, at best, with actual job performance, the exact pattern structured-interview research has spent fifty years documenting as a problem with unstructured hiring judgment generally, AI-assisted or not.
The failure isn't specific to AI. It's the same failure unstructured human gut-feel hiring has always had, now delivered with more fluency and less visible hesitation, which is arguably worse: a hiring manager who pauses and second-guesses a snap judgment at least notices the uncertainty, while a fluent AI answer papers over the same uncertainty with confident, well-formed prose.
What can AI actually do well in a hiring decision, and what shouldn't you trust it for?
Unaided, it's genuinely strong at generating structured interview questions tied to a specific competency, drafting a scoring rubric so two interviewers can be checked against the same standard, and catching inconsistency in how you personally described two candidates for equivalent behavior.
It's weak at anything that requires having actually observed the candidate: tone, real-time problem-solving under pressure, how they handled an unexpected follow-up question, the small hesitations a good interviewer reads in the room. It also has no independent way to verify what you tell it happened in an interview. If your notes already favor one candidate for reasons that don't hold up, the AI has no way to know that unless you specifically ask it to check.
What actually works when you prompt AI for a hiring decision?
These four work in Claude, ChatGPT, or Gemini, no special setup required.
1. The rubric builder. "I'm hiring for [role]. Build me a structured interview rubric for [specific competency], with three questions and a description of what a strong, adequate, and weak answer actually sounds like for each." Why it works: a rubric written before the interview, not scored impressionistically afterward, is the single most-replicated fix for inconsistent hiring judgment in the structured-interview research.
2. The evidence audit. "Here are my interview notes on two candidates for the same role: [notes]. For each of my stated conclusions about them, tell me whether it's backed by a specific example in the notes, or whether it's an impression without a cited example." Why it works: this catches the gap between "she seemed really sharp" (impression, no example) and "she caught an edge case in the take-home the other two candidates missed" (evidence), which is exactly the distinction unstructured hiring judgment tends to blur.
3. The disconfirming-evidence check. "For the candidate I'm currently leaning toward, list every piece of evidence in these notes that argues against hiring them, even minor ones. Don't soften it." Why it works: once a preference forms, it's easy to unconsciously discount contrary evidence; explicitly asking for the case against your current favorite forces you to weigh it rather than skip past it.
4. The reference-question generator. "Generate five reference-check questions specific to [role and concern, e.g. 'whether they can operate with minimal oversight'], each designed to get a concrete example rather than a generic endorsement." Why it works: most reference checks default to "would you recommend them," which produces almost uniformly positive, low-signal answers; questions built around a specific concern and demanding an example produce far more usable information.
What's still missing once you've run all four?
Run all four and your process for this one hire will be noticeably more rigorous than most unaided hiring decisions, closer to what a well-run structured-interview process looks like than most solo hiring managers ever get to on their own. What you still won't have is consistency across every hire you make, next quarter, with a different urgency, a different role, a different set of interviewers who weren't part of building this particular rubric. The gap isn't the quality of any single prompt; it's a repeatable system that survives being run by someone other than you, under time pressure, without reinventing the rubric from scratch.
How does a documented framework actually change the interview process?
Laszlo Bock ran people operations at Google and documented the shift away from gut-feel hiring in Work Rules!: structured interviews, with pre-set criteria and a scoring rubric fixed before anyone sits down with a candidate, scored independently by each interviewer, compared only after every scorecard is submitted, so no interviewer's early impression contaminates another's.
Before, generic prompting: asked to help decide between two finalists, a generic AI response produces a pros-and-cons list built from your own framing, "candidate A seems more experienced, candidate B seems like a better culture fit," language that mirrors whatever impression you led with rather than testing it.
After, Bock's framework applied: the framework refuses to compare candidates holistically at all. It first asks for the fixed list of competencies the role actually requires, decided before either candidate was discussed, then asks for each interviewer's independently scored evidence against that same list, and only then allows a comparison, criterion by criterion, flagging any place where one interviewer's score and another's diverge sharply as a signal worth investigating rather than averaging away. "Culture fit," specifically, gets replaced with a named, evidence-checkable competency, because Bock's documented critique of the phrase is that it's frequently a proxy for hiring people who resemble the existing team rather than a real signal.
A generic prompt has no structural reason to enforce that sequence, criteria before candidates, independent scoring before comparison, evidence before impression, on its own initiative. The framework does, because the sequence itself, not any single well-worded question inside it, is what actually protects the hire from gut-feel bias.
What does it take to get this framework running in Claude, ChatGPT, or Gemini?
Claude's version of this is a Skill package: go to Settings → Skills → Add skill and upload the .zip. From there Claude decides on its own when to apply it, matching your question against the skill's description, though typing /laszlo-bock-framework at the start of a message forces it regardless. Anyone working in Claude Code gets the same thing by unzipping into ~/.claude/skills (available in every project) or into a project's local .claude/skills folder (that project only).
There's no Skills equivalent in ChatGPT, so the route there is a Custom GPT: Create a GPT → Configure → Instructions, and paste the plain .md content in. Skip the regular Custom Instructions fields for this, they're capped at 1,500 characters, well short of what a framework needs; a Custom GPT's Instructions field gives you roughly 8,000.
Gemini has no equivalent upload flow either. Paste the .md content into the system prompt if your interface has one, or open your first message with it instead: "Here is a thinking framework to apply throughout this conversation. Read it carefully, then answer using this framework's approach," with the framework text following.
Where should this take you next?
For the one hire you're deciding on right now, Hiring Decision is a $199 tool built to run this exact judgment call before an offer goes out. For the broader discipline, Bock's method sits in the People & Culture category with Amy Edmondson, Kim Scott, and others, all downloadable as .md files for Claude, ChatGPT, or Gemini.
Frequently asked questions
Can AI actually reduce bias in a hiring decision?
It can reduce specific, nameable biases if you ask it to check for them explicitly, comparing your notes on two candidates for language patterns that track demographics rather than substance, for instance. It does nothing automatically. If your own interview notes already contain a biased framing, praising one candidate's "culture fit" and another's "technical skill" for equivalent behavior, the AI reasons from what you gave it and will not catch a bias you didn't ask it to look for.
Is it appropriate to have AI screen resumes or applications?
Treat AI resume screening as a first-pass filter you personally audit, not an unsupervised gatekeeper. Automated screening at scale has a well-documented history of encoding bias present in historical hiring data, and using it to make a final rejection decision without human review raises both fairness and, in some jurisdictions, legal exposure. Use it to help you read faster and flag genuinely relevant signal, and keep a human reviewing every rejection it recommends.
What's the single biggest mistake people make using AI in hiring?
Asking it to rank candidates holistically instead of structurally. A prompt like "which of these two candidates is better" produces a confident-sounding answer built on whatever the AI's training data associates with a strong hire, which can quietly reward polish and confidence in the writing over the actual evidence in the interview notes. Structured, criterion-by-criterion comparison catches this; a single holistic ranking usually doesn't.
Should the hiring manager be the only person using AI in this process?
No, and that's actually a documented weak point of ad hoc AI-assisted hiring: one person's prompts and one person's judgment, run through a tool, still produce one person's decision. Laszlo Bock's structured-interview approach specifically pairs the AI-assisted structure with input from multiple, independent interviewers scored on the same rubric before anyone compares notes, precisely so no single reasoning chain, human or AI-assisted, determines the outcome alone.
Can AI help me write better interview questions?
Yes, and this is one of its more reliable uses. Ask it to generate behavioral questions tied to a specific competency, rather than generic ones, and to include a rubric for what a strong, adequate, and weak answer actually sounds like. The rubric matters more than the questions; without it, two interviewers hearing the same answer will still score it differently, AI-generated questions or not.
Is it honest to use AI to help write a rejection message?
Yes, and it's often kinder than the alternative. A rushed, unaided rejection message is more likely to be vague or unintentionally cold; asking AI to draft a clear, respectful, specific rejection, referencing what genuinely didn't fit rather than a generic template, usually produces something better for the candidate to receive, provided you still read it before sending and correct anything that misrepresents the actual reason.
How is a documented framework different from just asking AI good interview questions?
A framework fixes the entire sequence, not just the questions: structured, pre-defined criteria set before any interview happens, the same rubric applied by every interviewer independently, and evidence-based scoring compared only after everyone has scored separately. Ad hoc good questions improve one interview. The framework prevents the more common failure, structured questions asked, but scored by gut feel and unstructured discussion afterward, which quietly re-introduces the bias the structure was meant to remove.
What's the honest limit of using AI for a hiring decision?
It has never seen the candidate in the room, has no access to reference-check conversations unless you transcribe them, and cannot verify anything you tell it about the interview. Its output is only as reliable as your notes and your honesty in reporting what actually happened. Use it to structure the comparison and catch inconsistency in your own reasoning; the underlying judgment about a specific person still has to be yours.
Written by Gareth Hoyle. Last updated 26 September 2026. Part of the authority.md guides library.
More guides.
Bringing AI Into a Creative Project Without Ruining It
What generic AI brainstorming actually gives a creative project and what it quietly flattens, four prompts that help today without taking over, and what changes when the AI applies Rick Rubin's documented listening practice instead of generic idea generation.
Using AI to Prep for a Negotiation: What Actually Works
Generic AI negotiation advice is genuinely useful and genuinely limited. Here's what Claude and ChatGPT get right before a negotiation, four prompts that work today, and what changes once the AI has a documented framework instead of a generic one.
AI Won't Write Your Best Work, But It Can Improve It
What generic AI editing actually fixes and flattens in a piece of writing, four prompts that sharpen a real draft today, and what changes when the AI applies Joan Didion's documented precision instead of generic style advice.