Every large strategy firm now screens candidates with a digital assessment before a human sees them. McKinsey runs Solve, an ecology-themed simulation. BCG runs an online case with a chatbot interviewer. Bain and the strategy arms of the Big Four use structured aptitude and situational screens. The industry describes this as modernisation. What actually happened is more specific and more consequential: the firms moved the highest-volume rejection point in recruiting away from a human reading a CV and toward an automated instrument — and then kept the case interview anyway.
Understanding that sequence is the whole game for a candidate. The digital assessment is where most applications end. The case interview is where offers are decided. They test different things, they reward different preparation, and conflating them is the most expensive mistake in consulting recruiting.
What each instrument is actually built to measure
| Instrument | Format | Primary construct | Function in the funnel |
|---|---|---|---|
| Gamified simulation (e.g. Solve) | Timed scenario, ecosystem or system-management tasks, no business content | Process: how you gather information, revise under feedback, manage constraints | Volume screen before interview |
| Chatbot / online case (e.g. BCG's) | Structured business case with fixed branching, quantitative items | Business judgement and numeracy under time pressure | Volume screen, closer to the real work |
| Aptitude / SJT battery | Numerical, verbal, situational judgement | Reasoning speed and decision preference | Volume screen, cheapest to run |
| Live case interview | 60 minutes, interviewer-led or candidate-led, with a partner-level round | Structuring, communication, comfort with being challenged, client presence | Offer decision |
The two categories are not competing measures of the same thing. A gamified simulation is deliberately content-free so that prior business exposure does not decide the outcome — the design intent is to reduce the advantage held by candidates from target schools with active consulting clubs. A live case is the opposite: it is a deliberate simulation of a client conversation, where the ability to be wrong gracefully in front of a senior person is part of the assessed material.
The evidence question, answered properly
Claims that simulations predict consulting success some precise multiple better than case interviews circulate widely and are not supportable. No firm publishes its criterion-validity study; no independent researcher has access to the performance data required to run one. What does exist is a robust cross-occupational literature, and it points somewhere uncomfortable for the industry's current spend.
For twenty-five years the reference was Schmidt and Hunter's 1998 meta-analysis in Psychological Bulletin. In 2022, Sackett, Zhang, Berry and Lievens published a systematic re-analysis in the Journal of Applied Psychology (107(11), 2040–2068) showing that the older coefficients had been inflated by overcorrection for range restriction. The revised ranking:
| Method | 1998 corrected r | 2022 revised r | Implication for consulting |
|---|---|---|---|
| Structured interview | 0.51 | ~0.42 | The structured case interview is the strongest tool the firms own |
| Job knowledge test | 0.48 | ~0.40 | Supports content-based screens over content-free ones |
| Work sample / simulation | 0.54 | ~0.33 | Simulations are useful, not dominant |
| General mental ability | 0.51 | ~0.31 | Aptitude batteries alone are weak instruments |
| Unstructured interview | 0.38 | ~0.19 | The unstructured partner chat is the weakest stage in the process |
Read against that table, the industry's design choices look defensible in one respect and questionable in another. Defensible: keeping a structured case interview as the decision stage is correct, because it is the best-validated method available. Questionable: placing a content-free simulation as the dominant volume filter means the largest single cut in the funnel is made by an instrument with materially lower validity than the stage it precedes. The screen is cheap and scalable, which is a real argument, but it is an efficiency argument rather than an accuracy one — and firms should say so.
Why the firms did it anyway — four reasons that are not marketing
1. Volume. Application inflation from one-click platforms and AI-assisted CV generation destroyed the CV screen. Any filter that a candidate can produce in thirty seconds at zero cost stops carrying signal. An assessment that costs an hour of attention restores a floor.
2. Consistency. A CV screen distributed across dozens of recruiters in dozens of offices has no rubric reliability. A single instrument applied identically is defensible to a regulator, to a court, and internally.
3. Access. A content-free assessment weakens the advantage of the case-coaching ecosystem that clusters at a handful of universities. This is genuine and measurable in principle, though the preparation market adapts within about two recruiting cycles.
4. Signal on a specific trait. The behaviour a gamified simulation observes best is the one classic case prep suppresses: what a candidate does when the correct method is not known in advance. Consulting work is mostly that condition. On this narrow construct, a simulation sees something an interview struggles to see, because in an interview the candidate is performing method fluency.
What this means for how you prepare
The practical error is symmetrical. Candidates from consulting-heavy schools over-prepare frameworks and get eliminated by the digital screen. Candidates from other backgrounds pass the screen, then meet a case interview they have never rehearsed and lose at the final gate. The allocation that follows from the evidence:
- Digital assessment — treat it as a systems and clock problem. Two or three full-length timed run-throughs to remove interface penalty. Under time pressure, prioritise finishing the required tasks over perfecting any one of them; scoring rewards completed process, and abandoned items score nothing.
- Case interview — build structure from the problem, not from a library. Memorised frameworks are now a negative signal because every interviewer has heard them. What scores is a structure derived out loud from the specific question, with the driver you intend to test named first.
- Quantitative discipline is non-negotiable. Practise arithmetic on paper, state units, sanity-check the order of magnitude aloud, and say the assumption before the number. A defensible wrong figure outperforms an unexplained right one at every firm.
- Rehearse being interrupted. A partner round tests whether you can revise under challenge without either collapsing or defending a dead position. This is trainable only with a live partner, not with written cases.
- Prepare the fit interview as rigorously as the case. It is a structured behavioural interview — the highest-validity method in the table — and most candidates treat it as small talk. Six situations, each with a decision, a constraint, a measured result, and a reflection.
Where the whole system is still weak
Three open problems, stated plainly because the sector does not.
Contestability. A candidate cut by a scored simulation usually receives no specific reason. Under the EU AI Act, employment-related assessment systems are classified as high-risk with transparency and human-oversight obligations, and firms operating in Europe will have to be able to explain outcomes at a level of detail most current tooling does not produce.
Construct drift. Once a preparation industry forms around any instrument, part of what it measures becomes exposure to that industry. This is exactly the bias the instrument was introduced to remove, arriving through a different door on a two-year lag.
No public validity evidence. Firms assert predictive improvement and publish none. Until someone publishes a criterion-validity study with a defined performance outcome and a control comparison, the honest statement is that the digital screen is well-designed and unproven at the level of the individual firm.
A note on incremental validity
One technical point decides whether a firm should run three screens or one. Validity coefficients do not add. Two instruments that measure overlapping constructs — an aptitude battery and a content-free simulation both loading heavily on reasoning speed — produce far less combined prediction than their individual figures suggest, while an instrument that measures something genuinely different adds disproportionately. That is the strongest published argument for pairing a structured interview with a work sample rather than stacking two cognitive screens, and it is the argument a candidate can use to interpret a process: a firm that runs one screen and one deep structured stage has read the literature, and a firm that runs four sequential timed tests is managing volume.
The cycle, as a calendar
Most candidates lose consulting recruiting on timing rather than on ability. The funnel is seasonal, the digital screen has a short window, and by the time a candidate feels ready the cohort has closed. The shape of a European cycle, in the terms that matter for planning:
| Window | What is happening | What a candidate should already have done |
|---|---|---|
| Six to nine months before start | Applications open for internships and full-time intake; referral windows are widest | CV artefacts fixed; one referral conversation held; assessment practice begun |
| Two to four weeks after applying | Digital assessment invitation, with a short expiry | At least two full timed run-throughs already completed on the specific format |
| Following two to six weeks | First-round cases, usually two interviews with managers or engagement leaders | Fifteen to twenty live cases done out loud with a partner, not read |
| Final round | Partner cases plus an unstructured conversation on motivation and judgement | Fit stories rehearsed to the same standard as the cases |
| Offer window | Short decision deadlines, deliberately | Comparison criteria decided in advance: practice, staffing model, mobility |
Two planning consequences follow. First, the digital assessment cannot be prepared reactively — the invitation window is often shorter than a serious preparation cycle, so practice has to precede the application. Second, live case volume is the single strongest driver of first-round outcomes among candidates who reach that stage, and live volume takes calendar weeks that cannot be compressed.
What a candidate should ask, and what the answer reveals
Consulting recruiting is unusual in that the firm expects to be interrogated. Six questions worth spending your slot on, and the signal each one carries:
- How is staffing decided in the first two years? A market model where consultants bid for projects builds different careers from a model where a staffing office allocates. Neither is wrong; one of them suits you.
- What share of the practice's work is implementation rather than strategy? This is the most reliable predictor of what your weeks will actually contain, and it is rarely visible in recruiting material.
- How many people in this office made manager from the analyst class three years ago? A specific number, answered without hesitation, indicates a firm that tracks development. Vagueness is also an answer.
- What is the feedback cadence, and who writes it? Consulting progression is evaluation-driven. If nobody can describe the mechanism, it is informal, and informal systems favour whoever is already visible.
- How does the firm use the digital assessment in the final decision? Whether the score travels into the interview stage or is discarded after screening is material, and recruiters generally answer honestly.
- What happens to people who leave after three years? The strength of the alumni route is part of the compensation, and firms with strong exits state them without prompting.
The purpose of these questions is not to impress. It is to collect the four or five facts that determine whether an offer is worth accepting, at the only moment in the relationship when the firm has an incentive to answer precisely.
The one-line verdict
The digital screen decides whether you are seen; the structured case decides whether you are hired. Prepare them as two separate disciplines, in that order, and allocate your hours to the stage that is actually in front of you rather than the one that is more enjoyable to practise. The industry's own instruments say the interview carries the most predictive weight, which is unexpectedly good news: the highest-stakes stage of consulting recruiting is also the most systematically preparable one, provided you rehearse it out loud, with a partner, against a clock, long before the invitation arrives. Nothing in this dossier rewards talent as much as it rewards sequencing.
Method and limits
All validity coefficients are cross-occupational population estimates from two peer-reviewed meta-analyses: Schmidt and Hunter (1998), Psychological Bulletin 124(2), 262–274, and Sackett, Zhang, Berry and Lievens (2022), Journal of Applied Psychology 107(11), 2040–2068. They are not firm-specific and should not be read as such. Descriptions of Solve, the BCG online case and aptitude batteries reflect the firms' own public careers documentation and widely reported candidate experience; no firm discloses scoring weights, cut scores or pass rates, and none are asserted here. An earlier version of this article claimed a specific multiple by which simulations outperform case interviews at predicting consulting success; that claim has been removed because no published study supports it. Statements about preparation-market adaptation, construct drift and the relative strength of funnel stages are analytical judgements, labelled as such, not findings.
