Monday, August 3

Two 5.6%s: Six rulers, one graduating class

 


One week in April 2026, a new graduate who had just sent out her résumé could read two contradictory headlines on the same screen.

One said hiring plans for the class of 2026 were up 5.6% from last year — grads finally catching a break. The other was a Bloomberg cover story: this class is walking into the toughest entry-level market in years.

Neither is clickbait. Neither is made up. Behind each stands a legitimate data source.

Set them side by side for a day and you reach a natural conclusion: one of them has to be lying, or at least one of them quietly picked the number that suited it.

There's another possibility.

Neither is lying. The two are simply not measuring the same thing. One measures how many people employers plan to hire; the other measures where job seekers actually ended up. The first is a plan, the second an outcome. The first is a survey some companies filled out in February; the second is where a whole generation stands in April. A bathroom scale and a thermometer never contradict each other.

The same graduating class is a single negative. And in 2026 there are at least six developing solutions for it.


Six rulers

The employer survey. The national association of colleges and employers asks 185 employers a simple question: how many new graduates do you plan to hire next year? Fielded mid-February to mid-March 2026. Not probability sampling — members and non-members answer if they feel like it. It measures intent.

A Chinese job platform's posting database. It counts new postings published on its own site; if the title or description hits the keyword, it counts. The report says so itself: it only reflects what happens on the platform.

PwC's barometer and Stanford's payroll study — the real pair here. PwC is built on a billion online job postings. Stanford uses administrative records from the payroll system: twenty-five thousand firms, 4.6 million workers. The first counts postings that went up. The second counts people who actually exist on a payroll.

Each has to assign every occupation an "AI exposure" score, and those two exposures are not the same thing. PwC's is ability-level: break the occupation into abilities one by one, then assess whether AI could in theory cover them. Stanford's is task-level, and folds in actual usage logs. One measures what's possible in theory; the other, what's already happening. The graduations on these two rulers cannot be converted into each other.

Handshake, connecting posting volume on one end to student and recruiter surveys on the other. It counts postings and also asks people whether they're afraid.

The New York Fed's monthly series. It doesn't ask firms, doesn't count postings, doesn't ask about attitudes. It measures one thing: after searching everywhere, how many people still haven't landed.

Six teams, each standing in front of its own ruler and nothing else. The people running the survey can't see the payroll. The people counting postings can't see beyond the platform. The people measuring outcomes don't know how many firms meant to hire.

Anyone standing in front of any one of these rulers will honestly read out the direction they can see.


Changing rulers

Start with the smallest switch, so small it hardly looks like a problem: that +5.6% is reported as a median.

In the fall 2025 round, the same organization, the same employers, originally reported the mean — +1.6%. Which is where many stories got the line "a rebound from 1.6% to 5.6%," a steep recovery curve.

Recalculate that fall batch as a median and what you get is −2.4%.

This needs no third party to interpret. It's written in the data source's own asterisked footnote. That steep recovery curve has a mean at one end and a median at the other. They aren't even on the same ruler.

And even aligning the statistic, the problem only moves — because +5.6% measures a plan, not hiring that has happened.

That same spring, the sixth ruler gives: March 2026, unemployment for recent graduates 5.6%; for all workers 4.2%.

The same number landing on the same class. One says employers plan to hire 5.6% more. The other says 5.6% of those who searched still haven't landed. There is no arithmetic relationship between them; the coincidence is pure. But it's enough to show one thing: a number, on its own, carries no direction. The direction comes from the ruler.

Two more lenses follow.

Flow and stock. Postings are flow; people on payroll are stock. "Entry-level jobs are increasing" can be true and false at once depending on which you measure: more postings and fewer people employed, or the reverse.

Relative and absolute. Stanford's most-cited number is 16% — a relative decline against low-exposure occupations. In absolute terms the same data gives −6%. Both are correct; one asks "how much worse than everyone else," the other "how much less than before." Slip from relative to absolute and the picture in the reader's head changes completely. For how slippery this is: a paper citing that study wrote it down as 13%.

There's a quieter slip too. That 16% comes with a 95% confidence interval — but the original doesn't publish the interval's upper and lower values. It only draws a shaded band. What was citable was never a point; it was a band. By the third retelling the band is gone and only a number is left.


Three layers

The tempting thought at this point: if it's all measurement basis, every contradiction dissolves the same way.

It doesn't.

Layer one: all correct, nobody is wrong.

Employer hiring intentions are recovering. One analysis found firms with the highest AI investment intensity grew total headcount 10.2% and entry-level headcount 12% in the two years after adoption. And this class really is having a hard time. All three measure different things — intent, internal firm stock, individual outcome — and hold at once. That analysis also wrote its own caveat: correlation isn't causation; intensive AI adopters may already have been larger, more technology-intensive and faster-growing beforehand. Cite the number, cite the warning with it.

Layer two: this one really is an error.

The claim circulating on the Chinese internet — "AI campus-hiring postings up twelvefold" — comes from a job platform's report for January–February 2026. But that figure describes experienced mid-to-senior hiring in the new-economy sector: AI's share of new postings rose from 2.29% to 26.23%. The extremely low base is the main source of the multiple. The same platform has a separate campus-hiring figure of an entirely different magnitude.

That report is about experienced hiring from beginning to end. On another page, where it writes that postings requiring under one year of experience fell about 20% year over year, the source itself adds a parenthesis: experienced hires only.

The qualifier was written by the data source. The one who dropped it was whoever repeated it.

There's a thinner layer still: same platform, same indicator — stretch the window from two months to four and twelvefold becomes 8.7-fold.

Layer three: this one is a real fight.

PwC and two other researchers — Lambert, of Warwick and the LSE, and Schindler, of the Ellison Institute in Oxford — used the same underlying data vendors. The same job-postings database, the same hiring records.

PwC read out: entry-level jobs have been "seniorized."

Lambert and Schindler read out something else: what raised the step wasn't mainly AI — it was that nobody is in the office.

Their material is 243 million hiring records and 407 million job postings across the US, UK, Canada and Australia. Estimated separately, generative AI exposure and work-from-home exposure each predict about a five-percentage-point decline in the junior share of new hires by 2025. Two suspects, and each one fits. But put both into the same model and let them compete: the AI coefficient collapses, often becoming statistically indistinguishable from zero, while the work-from-home term holds.

Here a passage is still missing, and leaving it out would be unfair.

Stanford considered the entanglement. They pulled computer occupations out of the sample entirely, and computer-related firms too — the conclusion held. They split occupations into teleworkable and non-teleworkable and looked at each; the more exposed group grew more slowly in both.

But that second test rests on a premise worth spelling out. Telework and AI exposure are already highly correlated — so correlated that once the sample was split the cells ran short, and the two lowest quintiles had to be pooled just to assemble comparable groups. In other words, the test of "looking at them separately" was itself run under the condition that they don't separate very well. And that is precisely the starting point of the later working paper.

It was also exactly when discussing the non-teleworkable group that they left one word in their own concluding sentence. The original says the results for that group indicate their findings are not driven by outsourcing or work-from-home disruptions — at least not solely.

Those words are themselves half a concession to the later criticism. They were written in a 2025 paper, when Lambert and Schindler's did not yet exist.

One more point, and Stanford wrote it into their own text: in the test extending the sample backward, their two exposure measures give different pictures. Under one, the most exposed quintile had already begun growing more slowly from around 2020; under the other, that doesn't appear. They wrote that sentence on their own initiative. Nobody picked it out for them.

So the difference isn't about who was careless. It's two identification strategies: Stanford's is to separate and set aside — remove the suspect, or split into groups and see whether what's left holds. Lambert and Schindler's is confrontation — both suspects stay in the model and you see whose coefficient survives.

Both are legitimate, and they reach different conclusions.

Smoothing it over as "they weren't actually measuring the same thing" would be tidier, but false. This isn't a difference of units. It's a head-on conflict between two causal readings of the same data.

Saying "statistics can lie" is easy. The hard part is pointing at one specific case and saying which kind of error it is. In the first layer nobody was wrong. In the second the retelling was wrong. In the third, two serious research teams genuinely disagree.


What's left

Not the truth. A sentence far narrower than the original, but one that holds:

PwC's analysis shows that in the United States, within the most AI-exposed tier of entry-level jobs, the ones it classifies as "seniorized" grew 35% in postings between 2019 and 2025, while the rest shrank 10%.

Count the hats that sentence wears: what a posting says is not who got hired; the most exposed tier is not all entry-level jobs; the United States is not the world.

In that same report, PwC added a note to its own chart: this chart is not saying AI caused these effects; the structural characteristics of the most exposed tier, and other shocks, may also be at work. The people who produced the number narrowed the causal opening themselves first.

The much-quoted "sevenfold" deserves the same treatment. It compares the most exposed entry-level jobs against the least exposed ones — not against ordinary jobs. And that chart's sample is only Canada, Singapore, the UK and the US. The sevenfold is real; the range it governs is far smaller than the sevenfold that got passed around.

The camps, laid out:

Three pieces of evidence say it is AI. PwC — postings, ability-level exposure. Stanford — payroll, task-level exposure. And the newest: a study of 62 million résumés across 285,000 US firms finding that AI-adopting firms cut junior hiring 9–10% within six quarters, driven entirely by reduced hiring rather than layoffs, with senior hiring unaffected.

But be careful with "three pieces." What they share is only a common direction, not stackable evidence. The two that genuinely corroborate each other are payroll and résumés. PwC's measures posting text and cannot be converted into the other two — it stands on the same side, but it isn't a third vote.

On the other side, the strongest is Lambert and Schindler's, and its strength is that it's the only one here doing causal identification head-on. Three further clues point the same way: in Handshake's data, entry-level tech postings fell about 15% and healthcare about 12%, with no sign that more exposed categories fell harder; in employers' own accounts, AI ranks fifth among reasons for hiring less; and one research institute notes the hires rate fell 0.8 percentage points, with a frozen market as the main cause.

On that employer-account point the claim has to stay narrow: only 19 firms answered. Nineteen firms can't hold up a statistic. It's enough only to say that even among the dozen or so cutting back on purpose, those putting AI ahead of budget and business demand were a minority. Narrow, and therefore solid.

There's also a study formalizing the chain "the step breaks → senior talent runs dry" into a dynamic game-theoretic model. Its reasoning is complete, but it's reasoning, not observation.

All of it is locked to one variable: which grade of evidence you accept. Postings basis or payroll basis; employers' accounts or administrative data. Loosen that lock and the whole chain recalculates.


One number that needs no exposure measure

Every argument so far has revolved around which occupations are more exposed. One set of numbers doesn't touch that question at all.

An annual report on new graduates shows: among those who worked while in school, 82% landed a job. Among those who didn't, 41%.

By that report's account, this is the same graduating class split in two by whether they'd stepped on the first rung — and the gap is double.

It can't prove AI did anything. It shows one thing: the first step itself has weight.

As for what happens when that step gets higher, two accounts are alive at once. A radiologist wrote on a public blog that the ability to read complex scans isn't independent of volume on easy ones, and that once AI absorbs the easy ones, residents inherit the hard cases but lose the substrate that calibrates judgment. A paralegal said in a public class that whoever learns to use AI early will leapfrog everyone. Both are public posts, neither an interview, neither identity independently verified. They aren't evidence — only two accounts alive at the same time. And voices like these turned up in only three trades: legal drafting, radiology, junior programming. Three trades, not every trade.

Has there been a precedent? We looked.

America's emergency shipbuilding in the Second World War: the gap surfaced within one to two years of expansion, and in-plant training averaging about six months turned trainees into machinists. It worked — but the boundary conditions were hard. Britain's apprenticeship system went through several steep declines taking roughly one working generation; the modern apprenticeship reversed the direction but never restored the scale. Japan's successor internships in traditional crafts: 1,785 applied over three years and 63 received placements, against an apprenticeship that runs eight to ten years. For aircraft maintenance and printing, no completed precedent of successful repair was found.

So history offers neither "an inevitable break" nor "it will fill in on its own." It offers a dividing line: skills that can be decomposed and standardized can be replenished in months; judgment that only grows from long presence on the ground takes a recovery period close to its original formation period.

And that line falls precisely on a question with no answer yet — the kind of judgment that was supposed to grow slowly on the first rung: if the height of that rung has changed, where does it grow now?

History offers no precedent here. It tells us only that once questions like this appear, developing them takes a working generation.


Back to you

Six rulers, two directions. Plans recovering, outcomes poor. Postings rising, headcount falling. A 16% relative decline, a 6% absolute one. Not one of these numbers is false.

For China, mechanism only, not figures — because within what's publicly available there is no cross-tabulation matching that administrative payroll data by age, occupation and exposure. The official basis since 2024 splits into 16–24 and 25–29, excluding enrolled students. That change of basis is itself what shows how hard cross-country comparison is: even on who counts as young, the two sides aren't using the same ruler.

One thing can be set side by side, though — not a number, but the same sentence.

One side's postings database read out "seniorized." In the other side's platform report, "stripping out the junior tier" is the phrase they chose themselves: of new postings in the first two months of 2026, those requiring three or more years of experience made up 73.34%.

Two markets, two sets of statistics, two different phrases, describing the same shape. They can be set side by side because both are about what the hiring side wrote down as requirements — the same ruler. Their unemployment figures cannot, because those are two different rulers.

Another batch of reports is coming. It will disguise itself again as an objective report requiring no choice from you.

The version you believe — which ruler did you read it from?

No comments:

Post a Comment

Sixty People Watch Which Door You Walk Toward

  A drill with no missile March 16, 2022. Newport News, Virginia. The aircraft carrier USS George Washington sits in dry dock. Below dec...