
Make talent quality your leading analytic with skills-based hiring solution.

You can trust AI to help make hiring decisions when the result it produces is reviewable: you can see the evidence behind it, trace which inputs drove it, challenge any part of it, and produce a record of all three six months later. Trust is a property of what the model hands to the person who signs off, and of what that person is required to do with it.
Because confidence and accuracy are produced by different things. A model returns a score reflecting how well an input matched the patterns it was built on, and an input that matches a pattern badly can still return a high number if the tool has no way to say “I do not know.” On screen, both look like a decision.
This is why a reviewer cannot use the score to check the score. A recruiter seeing “87 percent match” has learned nothing they can verify unless the 87 arrives with the lines of evidence that produced it. Findem’s platform is built around that distinction: the data labeling side of the platform produces 2M+ labeled Success Signals with 75 to 100 attached per profile, so an attribute can be shown rather than asserted, and Findem does not make employment decisions, with agent output presented as a recommendation a human reviews. On the measurement side, how accurate AI candidate scoring really is covers what accuracy means for a candidate score and how it degrades.
A better model does not solve it. A stronger model shifts the error rate and leaves the structure untouched: you still cannot tell from the output alone which answers are the wrong ones. That is a process problem, and process problems get fixed with process.
A reviewable result shows its working. A verdict only result shows a conclusion and asks you to accept it, which means the reviewer’s approval carries no information about whether the reasoning was sound.
| Review Criteria | Verdict only result | Reviewable result |
|---|---|---|
| What you see | A score, a rank, or “recommended” | The evidence first, then the recommendation, with the attributes that drove it named |
| What you can trace | The timestamp and the user who opened it | The exact inputs, the requirement set in force that day, the tool and its version |
| What you can challenge | The conclusion, with nothing to point at | Any single attribute, the weighting, the requirement itself, or the inputs |
| What you can defend six months later | That a human clicked approve | What the tool concluded, on what basis, who reviewed it, what they wrote, and what notice the candidate got |
The fourth row is the one that matters in a dispute. “A human reviewed it” is not a defense if the record cannot say what the human was looking at. Getting to the right column is an interface and policy question: the reviewer needs the evidence in one place, which is what a consolidated candidate 360 view is for.
Often no, and the strongest version of that argument is worth stating plainly. Reworked published “Humans in the Loop Isn’t Stopping AI Hiring Bias,” arguing that human oversight in AI hiring frequently operates as theater: the reviewer defers to the machine, approves the list, and the sign off adds a name to the record without adding judgment. That is a fair description of how a lot of review actually works, and a page that skips past it is not worth reading.
Human factors researchers call the effect automation bias: people accept an automated recommendation more readily than the same one from a colleague, and check it less when it agrees with what they expected. Add throughput pressure and the failure is predictable. A recruiter with 300 applicants, a two day SLA and a ranked list has an obvious path of least resistance.
Four things drive it, and none of them are laziness. The interface shows the conclusion before the evidence, so the reviewer is anchored before they think. The reviewer has no time budget for the review, because the review was sold as the thing that saves time. Nothing is recorded when they agree, so agreement is free and disagreement costs an explanation. And nobody measures whether reviews change outcomes, so a reviewer who rubber stamps and a reviewer who digs look the same in the reporting.
Structural changes, not exhortation. Four of them do most of the work.
Where the tool itself can help is by refusing to answer. A system that abstains on cases outside its competence and routes them to a person is removing exactly the cases where rubber stamping does the most damage. What happens when those cases go unhandled is covered in what happens when an AI recruiting agent gets it wrong.
A named person, and the record should say which one. No regulator, court or candidate accepts “the system decided” as an account of a rejection, and no vendor takes on the employer’s liability for a selection decision.
In practice three roles share it. The recruiter or hiring manager who accepts or overrides the recommendation owns the decision. The TA leader who configured the requirements, the thresholds and the review policy owns the process. Legal or compliance owns the records and the notices. Findem’s platform structure reflects that division with a Trust Layer and an Agent Orchestrations Layer sitting under the named agents, including the Calibration Agent, Screening Agent and Scheduling Agent, and the Screening Agent and Scheduling Agent pages are published under Glider AI branding. The agents recommend. People decide. Write that into the policy with the name of the role that signs, not a team.
Yes, in several jurisdictions, and the requirements vary by jurisdiction and keep changing. The list below is accurate as of September 2026 and is not legal advice, so check anything that affects your process with your own counsel before you rely on it.
New York City. Local Law 144 has been in effect since July 5, 2023 and is enforced by the Department of Consumer and Worker Protection. If you use an automated employment decision tool to substantially assist a hiring or promotion decision for a role in New York City, you need an independent bias audit conducted within the prior year, you must publish a summary of the audit results, and you must notify candidates and employees who live in the city at least 10 business days before the tool is used. The DCWP publishes the requirements at the NYC Automated Employment Decision Tools page.
Federal, United States. Title VII of the Civil Rights Act applies to a selection procedure whether a person or an algorithm applies it, which is the EEOC’s stated position on algorithmic selection, published at the EEOC’s AI and algorithmic fairness resources. The agency’s published technical assistance on AI has been revised and withdrawn across administrations. The statute has not changed, so the conservative reading is that algorithmic screening is a selection procedure and is treated as one.
The Uniform Guidelines on Employee Selection Procedures, at 29 CFR Part 1607, are the older framework that still governs how adverse impact gets measured. Section 1607.4(D) sets out the four fifths rule: a selection rate for any race, sex or ethnic group that is less than four fifths, or 80 percent, of the rate for the highest scoring group will generally be regarded by federal enforcement agencies as evidence of adverse impact. It is a rule of thumb rather than a legal test, courts and agencies also use significance testing, and clearing it does not immunize a tool. Run the ratio on your own hiring data by stage, not just at the offer, because a screening step can produce impact that the final numbers hide.
European Union. The EU AI Act, Regulation (EU) 2024/1689, entered into force in August 2024 and phases in over several years. Annex III lists employment, worker management and access to self employment among the high risk uses, which brings obligations covering risk management, data governance, logging, technical documentation, human oversight and registration. The application dates for Annex III systems have been subject to amendment, so confirm the current schedule at the EU AI Act text and timeline rather than relying on a date in an article.
States. Colorado enacted SB 24-205, which places duties on developers and deployers of high risk AI systems used in consequential decisions including employment, with a reasonable care standard around algorithmic discrimination. Its effective date has moved once already, so confirm the current one. Illinois has had the Artificial Intelligence Video Interview Act, 820 ILCS 42, in force since 2020, requiring notice, an explanation of how the AI works and what characteristics it evaluates, consent, limits on who receives the video, and deletion within 30 days of a request. Illinois also amended the Illinois Human Rights Act through HB 3773, effective January 1, 2026, making it a civil rights violation to use AI that has a discriminatory effect in employment decisions and barring the use of zip code as a proxy for a protected class.
For voluntary structure rather than obligation, the NIST AI Risk Management Framework, released as AI RMF 1.0 in January 2023, organizes the work into four functions: Govern, Map, Measure and Manage. It is not enforceable and it is a reasonable spine for an internal policy, because an auditor recognises the vocabulary.
Keep enough that someone outside your team can reconstruct the decision without asking you what happened. Eight records cover most challenges.
Keep these as long as your jurisdiction’s recordkeeping requirement, and longer where a claim period runs past it. The item teams most often lack is number 5. Approval without a reason is the weakest record in the set, which is why the recorded reason earns its cost.
At the screening stage, because that is where volume meets consequence. A screen that runs across 5,000 applicants and cuts 4,000 of them affects far more people than the final interview panel, and it is the stage most likely to be automated and least likely to be documented. That risk concentrates in high volume hiring, where the review policy has to hold up across thousands of decisions rather than dozens.
Three places to look first. Any stage where a tool ranks or filters candidates before a person sees them. Any use of video or voice analysis, since Illinois already regulates video interview AI specifically and other states have looked at it. And any assessment that contributes to a cut score, where validation evidence matters and psychometric assessments for hiring decisions explains how a properly constructed instrument documents its own validity. On the bias question specifically, how AI recruitment can reduce bias in hiring covers what structured, evidence backed screening changes and what it does not, and enterprise hiring covers the governance side at scale.
This page is about standing behind a decision you have already made with AI help: reviewability, regulation, documentation and sign off. It does not cover choosing a vendor.
If you are still evaluating tools, the questions are different ones: what to ask about training data, how to test repeatability in a trial, what a weak vendor answer sounds like, and how to score the options. That lives on the three checks to run before you trust an AI recruiting tool, with the wider procurement view in the checklist for buying AI hiring tech. It also does not explain what the tools themselves hand back, which is set out in what AI agents in recruiting actually hand back, or whether to build your own instead of buying, which is weighed in build versus buy for AI recruiting tools. This page also does not assess any specific tool’s compliance posture. That is a question for the vendor and your counsel, with the audit evidence in front of both.
You can trust an AI assisted hiring decision when it is reviewable: the evidence behind it is visible, the inputs are traceable, a named person reviewed it, and the whole record can be retrieved later. Trust does not come from the model’s confidence score. It comes from what the process requires the reviewer to see and record.
It should not, and in most compliance frameworks the employer remains the decision maker regardless of what software was involved. Findem does not make employment decisions: agent output is a recommendation a human reviews. Write into your policy which role signs off and what they must record.
Four properties. You can see the evidence that produced the result, you can trace which inputs and requirements were in force, you can challenge any single component rather than just the conclusion, and you can retrieve all of it months later. A result that has three of the four is not defensible, because the missing one is where the challenge lands.
An independent bias audit of the automated employment decision tool conducted within the previous year, publication of a summary of the audit results, and notice to candidates and employees who live in New York City at least 10 business days before the tool is used. It applies to tools that substantially assist hiring or promotion decisions for roles in the city, and the Department of Consumer and Worker Protection enforces it.
A rule of thumb in the Uniform Guidelines on Employee Selection Procedures at 29 CFR Part 1607, section 1607.4(D). If a group’s selection rate is less than four fifths, or 80 percent, of the rate for the highest scoring group, federal enforcement agencies will generally treat that as evidence of adverse impact. It is a screening indicator, not a legal standard, and passing it is not a defense on its own.
Yes. Annex III of Regulation (EU) 2024/1689 lists employment, worker management and access to self employment among high risk uses, which brings obligations on risk management, data governance, logging, documentation and human oversight. The phase in schedule has been amended, so confirm the dates that apply to your systems with counsel rather than relying on a published timeline.
Not by itself. Reworked’s argument that human oversight often functions as theater is fair: a reviewer who sees the conclusion first, has no time budget and records nothing when they agree is adding a signature rather than judgment. What makes review real is structural: evidence before conclusion, a recorded reason on borderline cases, sampling and rechecking, and a monitored override rate.
The requirement set in force that day, the tool and version, the exact input, the output with its attached evidence, the reviewer’s name and timestamp, the reviewer’s recorded reason where one was required, the candidate notice and its date, and the most recent bias audit covering that stage. Keep the retention and deletion record too. The record teams most often lack is the reviewer’s reason.
There is no published benchmark, and any figure presented as an industry standard is invented. Use your own baseline: measure the rate per stage and per reviewer, then investigate movement. A rate at or near zero means the review is not functioning, and a very high rate means the tool or the requirements need recalibration.
The employer, in nearly every framework, because the employer made the selection decision and the anti discrimination statutes attach to that decision. A vendor contract may allocate some indemnity, and it does not move the statutory obligation. That is why the audit evidence and the documentation matter more than any assurance in a vendor’s marketing.

No. You do not need a data team, data scientists, or engineers to run an AI recruiting agent, as long as the agent you are being offered is the kind that matches the technical capacity you actually have. What decides the answer is not the size of your team, it is which of three setup […]

An AI agent that hands you a confident wrong answer is more dangerous than one that hands you nothing, because confidence is what gets acted on. The framework below is three questions you can ask of any agentic tool before you trust its output: what data did it reason over, whose method did it follow, […]

Yes, AI recruiting agents get things wrong, and the useful question is not whether it will happen but whether your process is built to catch it, explain it and let a person fix it before it reaches a real candidate or a real client. An agent can misread a work history, infer a skill nobody […]