You are looking at four take-home submissions and they are all good. Same clean structure, same defensible tradeoffs, same competent prose. The advice you keep getting is a tighter rubric and a disclosure line at the top of the brief. Neither fixes what actually broke.
What broke is the spread between submissions. No declaration of AI use repairs it. In a randomised trial published in Science, 444 college-educated professionals wrote mid-level business documents in two rounds, half of them given ChatGPT before the second. On those writing tasks, not on any hiring assessment, the correlation between a person's first-round and second-round grade was 0.49 in the control arm and 0.25 in the treatment arm, and the authors describe initial differences as half-erased (Noy and Zhang, 2023).
An instrument that compresses the distance between your strongest and weakest candidate has stopped doing its one job. Who declared what does not change that.
What a take-home was supposed to measure
The take-home earned its place on one comparison. An unstructured conversation tells you how well somebody talks about work, and an artifact tells you what they can build. That is still the pitch in the guidance genre.
Indeed's hiring-resources page recommends reserving take-homes for final-stage candidates, keeping them under an hour, and giving feedback whether or not you hire. Toggl walks through five task examples by role. iMocha's guide covers task design and evaluation, resting on one statistic (skills assessments are three times more effective than gut instinct) that it attributes to nobody. A practitioner guide to unmonitored challenges recommends them as an entry-level filter and tells candidates to submit about 70% of the way through the deadline window.
Those four pages were read at source on 4 September 2026. Not one mentions AI, ChatGPT, or generative assistance anywhere. Toggl's was updated in February 2026 and iMocha's in August 2026, both well over a year after ChatGPT shipped. The genre that tells employers how to run a take-home has not registered the thing that changed it.
The instrument was never a reading of output quality on its own terms. It was a proxy for one sentence: this person, working alone, unassisted, produced this. Every word of that sentence except "this person" is now optional.
What the evidence actually shows about AI and performance spread
Noy and Zhang recruited 444 professionals (marketers, grant writers, consultants, data analysts, HR staff, managers) and gave each two 20 to 30 minute writing tasks drawn from their own occupation. The tasks were fielded in March 2023. Half were offered ChatGPT before the second task. Two results carry the argument.
The first is the collapse in how well round one predicted round two. The correlation between a participant's two grades was 0.49 in the control group and 0.25 in the treatment group: a difference in slope of 0.243, a 95% confidence interval of 0.08 to 0.41, p=0.004. Once the tool was in the room, the first sample told you about half as much about the second.
The second result explains the first. Lower-scoring participants gained roughly 0.4 standard deviations in quality, while higher-scoring participants mostly gained speed. The tool lifted the floor and left the ceiling where it was. That is the shape that destroys a screening instrument: you keep the good submissions and lose the ability to tell them apart.
The extension to hiring is stated as one. A take-home is unsupervised, effectively untimed, and graded on the artifact rather than on the person. On that reasoning, the compression Noy and Zhang measured on professional writing tasks should be expected to appear at least as strongly on take-home submissions, and no study has measured it there.
No assessment vendor, platform report, or academic database publishes a before-and-after score distribution for hiring take-homes. The nearest published thing comes from a party with an interest in the answer. CodeSignal, which sells proctored assessments, reports score increases on its unproctored assessments more than four times larger than on proctored ones, and fraud flags on entry-level assessments rising from 15% to 40% year over year. That is detected cheating on one platform, audited by the vendor selling the alternative, and it measures flags rather than spread.
Why a disclosure policy does not fix it
Disclosure is the right default, and we have argued for it. The three-tier interview AI policy puts submitted work in the middle tier, allowed with disclosure, because you are assessing judgment and output and need to know what produced it. That holds for a portfolio piece and a live exercise. It fails for this one instrument, on arithmetic rather than ethics.
Disclosure tells you which submissions were assisted. Grant it perfect compliance: everybody declares honestly, nobody hides anything, the policy works as written. You now hold a stack of assisted submissions that are all good and you still cannot order them, because the variance you were reading is what the tool removed. Knowing the provenance of a compressed distribution does not decompress it.
Two escape hatches get reached for here and both leak. Banning AI in the brief turns a measurement problem into an enforcement problem you cannot win off camera. Karat's account of one unnamed employer, relayed by a company that sells interviewing, is that around 80% of candidates used a model on a top-of-funnel code test that forbade it, and that this was the point the company gave up enforcing. That figure is sourced in the three-tier interview AI policy piece linked above, not here.
Asking candidates to walk through their submission live is much better. It works because it is a live structured interview with homework attached. The signal has moved to the interview. The take-home is now the prompt for it.
The cost you are already paying: candidate hours, drop-off, and reputation
The instrument stopped discriminating. It did not stop costing.
interviewing.io surveyed about 700 candidates from its own user base. It sells interview practice and live technical screening. This is an interested party's survey of self-selected respondents, not a market measurement. Respondents reported that tasks advertised as one hour took about five hours.
Over 80% said a take-home should run four hours or less. And 58% believe they deserve to be paid for one against 4% who ever have been, a gap of 14.5 to one. The most-cited cause of a bad experience was neither the length nor the money. It was getting no feedback.
Individual accounts show the shape of the reaction rather than how common it is. On Blind, a candidate facing a third take-home in one loop submitted a deliberately minimal answer and disengaged. A marketing candidate was asked for a 30-day launch strategy for a real client, priced by the employer at two to three hours and estimated by the candidate at 15 to 20, then ghosted after a partial submission. One writeup describes 10 interview rounds plus four days producing a 40-page deliverable, then a phone rejection with no reason given and the candidate's own summary that an agency would have taken a month and $50k for the same work.
Three accounts are three accounts: they show that this complaint exists and what it sounds like, which is all a handful can do.
No published survey isolates candidate drop-off at the take-home stage. The closest funnel data folds assessments into "interview stage", so any take-home abandonment rate quoted at you was manufactured from something else. What you can price without a study is exposure: every candidate who spends five unpaid hours on your brief and hears nothing back can write about it. The people you are least able to source twice are the ones with the standing to be believed.
What still discriminates, with current validity coefficients
Selection validity was rewritten in 2022, and hiring content has not caught up. Sackett, Zhang, Berry and Lievens showed that the standard meta-analytic estimates, the ones traced back to Schmidt and Hunter in 1998, were inflated by an overcorrection for range restriction. Their revised table moves almost every predictor down, several of them sharply.
| Predictor | Schmidt and Hunter 1998 | Sackett et al. 2022 revision |
|---|---|---|
| Structured interviews | 0.51 | 0.42 |
| Work sample tests (proctored) | 0.54 | 0.33 |
| General mental ability | 0.51 | 0.31 |
| Assessment centers | 0.37 | 0.29 |
| Unstructured interviews | 0.38 | 0.19 |
| Empirically keyed biodata | 0.35 | 0.38 |
Two things fall out. Structured interviews are now the strongest widely used single predictor at 0.42, ahead of work samples at 0.33 and general mental ability at 0.31, and they beat unstructured interviews at 0.19 by a factor of 2.2. Work samples took one of the deepest cuts, losing 0.21 of validity where structured interviews lost 0.09. Biodata, at 0.35 to 0.38, is the only row that went up.
One error here is live in hiring content right now. A 2023 RecruitingDaily piece defending take-home assignments supports them by citing a 1998 APA study on assessment validity. That study is the Schmidt and Hunter meta-analysis. Its work-sample coefficient is the 0.54 that Sackett, Zhang, Berry and Lievens revised down to 0.33 in 2022, and it measured proctored work samples: supervised, standardised, timed, watched.
An assignment done alone at a kitchen table with any tool the candidate owns is a different instrument on a different population. Citing proctored work-sample validity to defend an unsupervised take-home swaps both the referent and the population. Defenders of this instrument keep making that error in print.
Three replacements, priced by the effort they cost your team
Every replacement below moves work off the candidate's evenings and onto your team's calendar. That is the trade. It is why take-homes lasted: they are the only assessment whose cost lands on somebody not in the room. The effort column holds design parameters you choose, not measurements anyone published.
| Replacement | What it reads | Validity on the 2022 revision | What it costs your team |
|---|---|---|---|
| Structured interview: fixed questions, scored rubric, two interviewers | Judgment under questioning, on the same scale for every candidate | 0.42, the strongest widely used predictor | An hour of two people per candidate, plus a day writing the questions and anchors once |
| Live work sample: one hour, paired, your codebase or your deck, tools allowed and visible | How somebody produces: what they prompt, accept and reject | 0.33, measured on supervised work samples | An hour of a senior person per candidate, and the exercise built once |
| Cognitive ability test at the top of the funnel | General problem solving, before anyone spends interview time | 0.31 | Near zero per candidate; real work in vendor choice and adverse-impact review |
The order matters. The structured interview is the one that pays. Its only real cost is that somebody writes the questions and the scoring anchors before the first call, which is exactly the work a team skips when it hands out an assignment instead. Our set of interview questions that predict performance is where to start.
The live work sample is the honest version of the take-home. Same task, one hour, you in the room, tools on the table and in view. You watch what gets prompted, what gets accepted uncritically and what gets caught, which reads better on how somebody works in 2026 than any finished artifact can. It also repairs the population swap above: the 0.33 was measured under supervision, so running it under supervision is the only way to claim it.
The cognitive test is the cheapest per candidate and the most legally exposed, and it belongs at the top of the funnel or nowhere. If you run one, pick a vendor that publishes adverse-impact data, and read what AI resume screening legally cannot do before you automate a rejection off the back of it.
When a take-home is still the right instrument
Three cases survive.
Paid work, scoped and priced at a market rate. At that point you are running a trial rather than an assessment. The compression problem stops mattering, because you are buying an outcome instead of inferring a person from an artifact. It is also the version 58% of interviewing.io's respondents said they wanted and 4% have seen.
A review of work that already existed. Asking for something the candidate built before your process began costs them nothing and costs you the interview time to interrogate it. Provenance is still open. It is settled in conversation: who else was in the room, what broke, what you would do differently now.
The version Indeed describes: one hour, final stage, feedback guaranteed. If the loop is already short, the candidate is already a finalist, and the artifact is a tie-breaker rather than a filter, a small assignment is defensible. Keep it under the hour, pay for it if it runs over, and send the feedback either way.
What is not on that list is the four-hour unpaid assignment mailed to five candidates at the screening stage. That one now tests who owns the better tool and who had the freer weekend, and you can guess how both correlate with the hire you want.