Every large professional-services firm now says it hires for skills rather than credentials. Almost none of them explain what that sentence costs, or what it replaced. The honest version is narrower and more interesting than the press release: the Big Four did not abandon selectivity, they moved it. The filter used to sit at the top of the funnel, in the form of a degree classification and a university name. It now sits in the middle of the funnel, in the form of timed assessments, structured behavioural evidence and, increasingly, a work sample that looks like the first week of the job. For a candidate, that shift changes what is worth preparing, what is worth signalling, and which parts of a CV have quietly stopped mattering.
This dossier explains the mechanism at Deloitte and its peers, tests the claim against the peer-reviewed selection literature rather than against firm marketing, and states plainly where the model is weaker than its advocates admit.
What the old screen actually did
Until the late 2010s, graduate entry into audit, tax and consulting in the UK and much of Europe ran through two hard gates: a minimum degree classification (typically a UK 2:1 or its continental equivalent) and, informally, a shortlist of target universities where the firms ran campus events. Neither gate was designed to predict performance. Both were designed to make an unmanageable application volume manageable. A firm receiving six figures of applications for a few thousand seats needs a cheap sort, and academic classification is the cheapest one available.
The problem is not that academic results are uninformative: they carry real signal about conscientiousness and reasoning. The problem is that they carry that signal together with a large amount of socioeconomic noise: school quality, family stability, whether the candidate worked twenty hours a week during term. When a firm screens on classification, it cannot separate the two. It selects for capability and for circumstance simultaneously, then discovers years later that its partner cohort looks like its intake cohort.
What replaced it, gate by gate
| Gate | Old design | Current design | What it is actually measuring |
|---|---|---|---|
| 1. Eligibility | Degree classification floor, target-school shortlist | Right to work, no academic floor at several firms; contextual data on school and postcode | Whether the firm can legally and practically employ you |
| 2. Screen | CV and cover-letter read | Online cognitive and situational-judgement assessment, often adaptive | Reasoning speed and decision preferences under time pressure |
| 3. Evidence | Competency interview against a rubric held loosely | Structured interview with fixed questions, anchored scoring, multiple raters | Past behaviour as a proxy for future behaviour, with rater bias suppressed |
| 4. Sample | Group exercise, sometimes a written case | Job simulation: client brief, real data, a deliverable, a defence of it | Whether you can do a scaled-down version of the work |
Read the table as a transfer of burden. The old process asked the education system to pre-sort candidates and then spent its own effort on interviewing. The new one absorbs the sorting cost itself. That is expensive, which is why it only became viable when assessment delivery moved online and scoring became partly automated.
Does it predict better? What the evidence says
Here the record is stronger than the marketing, but not in the way most articles claim. For twenty-five years the reference point was Schmidt and Hunter's 1998 meta-analysis in Psychological Bulletin, which put general mental ability at the top of the validity table. In 2022, Sackett, Zhang, Berry and Lievens re-ran the corrections in the Journal of Applied Psychology and showed that the older estimates had been systematically inflated by overcorrection for range restriction. The revised ordering matters for anyone designing a hiring process:
| Method | Schmidt & Hunter (1998), corrected r | Sackett et al. (2022), revised r | Reading |
|---|---|---|---|
| Structured interview | 0.51 | ~0.42 | Now the strongest single predictor |
| Job knowledge test | 0.48 | ~0.40 | Strong where the role has a defined body of knowledge |
| Work sample / simulation | 0.54 | ~0.33 | Lower than long assumed, still solidly predictive |
| General mental ability | 0.51 | ~0.31 | Displaced from the top of the table |
| Unstructured interview | 0.38 | ~0.19 | Weak; the format most firms defaulted to |
Three consequences follow. First, the single highest-leverage change any firm can make is not adding a game: it is imposing structure on the interview it already runs. Second, work samples remain worth their cost, but the case for them is now as much about candidate experience and fairness perception as about raw prediction. Third, and least comfortable for the sector, a cognitive test used alone is a weaker instrument than the industry believed when it bought one.
The corollary for a candidate is precise: the interview is not a formality behind the assessment. It is the part of the process with the most predictive weight, and structured interviews are the most preparable stage in the entire funnel, because their questions are fixed by design.
Where the 2022 revision moved each instrument
| Method | Rank in 1998 | Rank in 2022 | What the move means for a graduate gate |
|---|---|---|---|
| Work sample / simulation | 1st (0.54) | 3rd (~0.33) | Still worth running, no longer the strongest gate |
| Structured interview | joint 2nd (0.51) | 1st (~0.42) | The stage that decides the offer is also the best evidence |
| General mental ability | joint 2nd (0.51) | 4th (~0.31) | An aptitude battery alone cannot carry a screen |
| Job knowledge test | 4th (0.48) | 2nd (~0.40) | Content-based screens beat content-free ones |
| Unstructured interview | 5th (0.38) | 5th (~0.19) | The informal partner chat remains the weakest stage |
Sources 1 Journal of Applied Psychology · 2 Psychological Bulletin (American Psychological Association)
Why a firm like Deloitte moves first
Skills-based hiring is usually framed as a diversity initiative. In professional services it is at least as much an operational one, for four reasons specific to the model.
Audit quality is regulated and inspected. Audit files are reviewed by national regulators, and findings attach to the firm's licence to operate at scale. Anything that improves the average competence of a first-year on an engagement team has a direct compliance value, not only a cultural one.
Attrition is the dominant cost. Graduate intake in professional services has historically lost a large share of a cohort inside three years. Every departure destroys the training investment made in a person who had just become billable. A selection method that predicts fit to the actual work reduces the most expensive kind of mis-hire: the competent person who discovers in month eight that the job is not what they were shown.
The work itself changed. Data-analytics-led audit, tax technology and technology consulting need people who can manipulate structured data, not only people who can read a standard. That skill is poorly correlated with the degree classifications the old screen used, and it is well correlated with what a simulation can observe directly.
Apprenticeship routes created a control group. Firms that opened school-leaver and apprenticeship pathways ended up with two intakes doing similar early work, one selected on academic classification and one not. That comparison is uncomfortable to ignore.
~0.42revised validity of the structured interview
The highest coefficient in the 2022 table, and the stage the firms already own
Sackett, Zhang, Berry and Lievens (2022), a cross-occupational population estimate. It is evidence about a method, not about any firm’s implementation of it, and it does not measure fairness.
Where the model is genuinely weaker
A dossier that only lists advantages is advertising. Four real objections:
1. Preparation markets rebuild the advantage the reform removed. Within eighteen months of any assessment becoming standard, a paid preparation industry forms around it. Access to practice is unequal in the same direction as access to good schools. The gate moves; the gradient survives.
2. Automated scoring is hard to contest. A candidate rejected by a rubric-scored interview can be told which behavioural evidence was missing. A candidate rejected by a model score often cannot be told anything specific. Under the EU AI Act, employment-related assessment systems fall in the high-risk category with transparency and human-oversight duties attached, which is a legal problem for opaque scoring and a fairness problem regardless.
3. Simulations measure the job as currently designed. That is their strength for year-one performance and their weakness for ten-year potential. A process optimised for immediate deliverable quality can systematically under-select people whose value is judgement, client trust or dissent.
4. Published outcome figures are firm-produced. Retention and satisfaction improvements attributed to skills-based hiring are almost always self-reported, uncontrolled, and coincident with other changes: pay revisions, hybrid work, intake size. They are directional evidence, not causal evidence, and they should be read as such.
What to do with this if you are applying
- Stop optimising the CV for prestige signals and start optimising it for verifiable artefacts. A named deliverable (a model you built, an analysis you shipped, a process you fixed, with the size of the thing stated) survives a skills screen. A society presidency does not.
- Prepare the structured interview like an exam, because it is one. Fixed competencies mean a finite question set. Write six to eight situations from your own experience, each with the decision you took, the constraint you were under, the number attached to the outcome, and what you would change. Rehearse them out loud until the timing is stable.
- Treat the assessment stage as a speed problem, not a knowledge problem. Most candidates lose points to unfamiliarity with the interface and the clock, not to the underlying reasoning. One or two full-length timed run-throughs remove most of that penalty.
- In a simulation, state your assumptions in writing. Rubrics for work samples reward traceable reasoning, and a defensible wrong number scores better than an unexplained right one.
- Ask which stage carries the most weight. Recruiters will usually answer. If nobody will tell you how you are being scored, that is itself information about the employer.
The candidate-side conclusion
Skills-based hiring is not a softer process. It is a process that has moved the difficulty from things you inherited to things you can build, and from a one-line signal to a demonstrated one. That is a genuine improvement in fairness, and it comes with a cost: you can no longer coast on an institution's name, and you can no longer be surprised on the day. The candidates who now win these processes are the ones who have already done a scaled version of the work and can walk someone through how they did it.
How the same shift reads in France
The skills-based argument was written in an Anglo-Saxon labour market where the degree classification was the gate. France has a different architecture, so the reform lands differently, and candidates who import the British version of the advice get it wrong.
| Feature | UK / US pattern | French pattern | Consequence for a candidate |
|---|---|---|---|
| Primary sorting signal | Degree classification and university tier | The grandes écoles / université distinction, fixed early by competitive entry | The signal is set years before hiring, so a mid-career reset needs demonstrated output, not another diploma |
| Certified alternative route | Apprenticeship, still comparatively marginal | Alternance and apprentissage, industrial in scale and socially accepted | The strongest available entry path into a large firm, and structurally under-used by students who consider it second-best |
| Skills recognition | Employer-defined competency frameworks | National certification registry and recognised professional titles | A certified skill is legible to HR systems, which matters when the screen is automated |
| Assessment governance | Contract and case law | Labour code plus data-protection and AI-Act obligations on automated decisions | You have a documented right to information about automated assessment; asking for it is legitimate |
The practical translation is that in France, "skills-based" does not primarily mean dropping the diploma requirement. It means that a second, parallel ladder (alternance, certified titles, demonstrated deliverables) has become genuinely load-bearing alongside the initial-selection ladder. For anyone who did not win the competitive-entry lottery at twenty, that parallel ladder is the whole opportunity, and it rewards artefacts and certifications over restated ambition.
The one-line verdict
Skills-based hiring in professional services is a real improvement measured against the screen it replaced, and an unproven claim measured against its own marketing. The defensible position is narrow: structure your interviews, sample the work, publish what you can verify, and stop citing outcome percentages nobody audited. For a candidate, the operating rule is simpler still. Build one artefact that a stranger can inspect, learn to narrate the decisions behind it in four minutes, and apply through the route where that evidence is read rather than the route where a classification is read. That is the whole arbitrage, and it is available to anyone willing to do the work before the application rather than after the rejection.
Method and limits
Validity coefficients in this dossier come from two peer-reviewed meta-analyses, Schmidt and Hunter (1998), Psychological Bulletin 124(2), and Sackett, Zhang, Berry and Lievens (2022), Journal of Applied Psychology 107(11), 2040–2068, and are population estimates across occupations, not firm-specific figures. Process descriptions reflect publicly documented Big Four graduate pathways and firm-published human-capital research; exact stage weightings are not disclosed by any firm and are not claimed here. Retention, satisfaction and diversity outcomes attributed to skills-based hiring are firm-reported and uncontrolled; earlier versions of this article cited specific percentage improvements which we have removed because no independent verification exists. Regulatory statements refer to the EU AI Act's classification of employment-related assessment as high-risk. Where this dossier expresses a judgement (on preparation markets, on simulations under-selecting long-horizon potential) it is labelled as judgement, not finding.
