Field Notes

synthesis

Male variability, education, and life outcomes

Greater male variability shapes the entire distribution of male life outcomes but it is the bottom tail that carries the heaviest social cost. Boys are overrepresented among the lowest performers in school, among the least likely to attend and finish college, among the long-term unemployed, and among suicide decedents. The same biological pattern that produces male overrepresentation among elite mathematicians and CEOs also produces disproportionate male failure.

The chain from variability through education to life outcomes is real at every link, but its strength varies, and the strongest links are not the ones that get the most attention.

How variability reaches the classroom

The chain starts from Greater male variability: male variance ratios on cognitive tests typically fall between 1.05 and 1.15, and under the approximately normal distribution that test scores follow, even a variance ratio of 1.08 puts meaningfully more boys than girls at both the top and bottom of the ability distribution. The inference is disputed at the level of magnitude and mechanism — the same variance ratio can correspond to radically different tail ratios under different underlying distributions, and the pattern varies by subject and country — but the bottom-tail overrepresentation in test scores that this synthesis builds on is not itself in dispute. Some cold facts about math and gender works the pattern through one subject in detail, where boys hold 61 percent of the top decile of SAT mathematics scores and 56 percent of the bottom decile.

The grade penalty

Boys receive worse school grades than girls across subjects and grade levels. Voyer and Voyer (2014) meta-analysed 502 effect sizes from 369 samples and found a mean female advantage in teacher-assigned marks of d = 0.225. The effect is largest in language courses (d = 0.374) and smallest in mathematics (d = 0.069). The female grade advantage is stable across publication years going back decades; the popular narrative of a recent “boy crisis” obscures a persistent pattern that predates the policy changes often blamed for it.

Why boys receive lower grades is better understood than the popular commentary suggests. Cornwell, Mustard, and Van Parys (2013), using longitudinal data from the Early Childhood Longitudinal Study, found that boys’ lower ratings on “Approaches to Learning” — a composite of attentiveness, task persistence, eagerness to learn, organisational skills, and flexibility — accounted for essentially all of the gender grade gap. When test scores and these non-cognitive skill ratings were controlled simultaneously, the gender grading gap disappeared for white students and was substantially reduced for Black and Hispanic students. Boys were not being graded down for being boys; they were being graded down for behaviour that schools systematically reward and that boys are developmentally less equipped to display.

Teacher bias exists but its role is modest. Terrier (2020) found a 5.2% bias against boys in math grading in French middle schools, using blind versus non-blind score comparisons. Lavy and Sand (2018) found similar effects in Israel with long-term consequences for advanced course selection. But Contreras (2023) found that individual teachers’ grading bias was not persistent across their classes, suggesting the gap is explained by student behaviour rather than stable teacher prejudice.

The grading mechanism matters because it determines what kind of intervention would help. If the gap is bias, the solution is teacher training. If the gap is behaviour, the solutions are developmental (redshirting boys, reducing the behavioural demands of early grades) or pharmacological (the ADHD diagnosis pathway). The evidence favours the behavioural account.

The ADHD relative-age effect

Boys are diagnosed with ADHD at roughly twice the rate of girls, but part of that gap is attributable to relative age. A Finnish nationwide register study (Sayal et al. 2017) found that the youngest children in a school year were 26–31% more likely to receive an ADHD diagnosis than the oldest children in the same grade. The effect was strongest for diagnoses made before age 10 and has been increasing over time. The same Harvard study of over 400,000 children that LaCorte cites in What if boys were never the problem reports a roughly one-third increase in ADHD diagnosis risk for the youngest children in a class, with the only difference being a few months of age.

That question is now largely resolved, and ADHD relative-age effect carries the answer. The excess is predominantly overdiagnosis of the relatively young rather than underdiagnosis of the relatively old: the additional cases concentrate among mild presentations, and the same gradient appears at almost identical magnitude for intellectual disability and depression, which no ADHD-specific case-finding story explains. Overdiagnosis of ADHD in children and adolescents supplies the wider evidence base for that direction, synthesising 334 studies and locating the harm at the mild-symptom margin.

Two qualifications matter for this synthesis. The effect is similar in size for girls and boys, so it does not explain the two-to-one sex gap in diagnosis; it establishes only that the diagnostic process is sensitive to developmental position in a cohort. And part of the gradient may be genuine difficulty rather than misread immaturity, since being youngest is associated with worse grades and worse peer relationships. On that reading the pharmacological pathway is not simply mislabelling age-normal boys — it is labelling children who are, in part, actually struggling because of where their birthday falls.

What does not work

Several proposed fixes for the boys’ education gap have been tested and found wanting.

Male teachers. Nearly 90% of elementary school teachers are women, and the intuitive fix — recruit more men so boys have role models — has no detectable effect on boys’ outcomes. Helbig (2012) analysed data from 146,315 elementary students across 21 countries and found that boys perform no better in reading or mathematics when taught by a male teacher. In some countries, girls actually benefited from female teachers.

Redshirting. Richard Reeves (2022) proposes starting boys in school a year later to let their prefrontal cortex development catch up. The research on redshirting is thin. Economists who have studied it find small benefits that tend to fade, while the cost of an extra year of childcare and a one-year delay in entering the workforce are real.

Boy-friendly teaching methods. Sommers’s War on Boys states the popular version of this proposal — boy-targeted reading material, action-friendly writing assignments, an end to zero-tolerance discipline, and restored recess — on the premise that schools have become actively hostile to boys. A four-year UK Department for Education trial across dozens of schools tested various boy-targeted teaching approaches and could not demonstrate effectiveness. The techniques that engaged boys turned out to engage girls as well. The evidence above favours a developmental-mismatch account over the hostility framing, which matters for the remedy: a mismatch is addressed by changing what schools demand and when, not by segregating instruction by sex.

Title IX-style federal intervention. When the college degree gap favoured men by 13 points in 1972, Congress passed Title IX, universities were investigated, and a decades-long national effort pushed girls into classrooms, sports, and STEM programmes. When the gap reversed and grew larger in the opposite direction, no comparable national effort materialised. LaCorte (2026) attributes this to timing — “help girls” became background noise, and suggesting boys need help now sounds like undoing an old victory — but the policy vacuum also reflects that the tested interventions do not work.

The college gap and its meaning

In 1970, men earned 57% of US bachelor’s degrees. Enrolment equalised in 1982 and continued shifting; men now earn approximately 42%. The ratio is nearly the mirror image of 1970.

The conventional story attributes this to boys’ school failure, but the timing complicates that account. The female grade advantage is older and steadier than the college-enrolment flip. What changed is not male performance but women’s entry into higher education at scale, an expected consequence of reduced discrimination and changing norms. The college gap is partly a story of female success, not only male failure.

That said, the downstream consequences of not completing college are increasingly severe, and they fall disproportionately on men. A four-year degree roughly doubles lifetime earnings. Men without a degree have been exiting the labour force in large numbers, while men with degrees continue working at historical rates. Marriage markets have stratified along educational lines: the men excluded from marriage are predominantly those at the bottom of the education and earnings distribution.

LaCorte (2026) presents the education gap as a “precursor” to suicide and other adult crises. The correlation is real but the causal inference is overstated.

Hughes, Liu, and Qin (2025) meta-analysed 53 studies from 23 countries and identified unemployment (RR 1.84), divorce or separation (RR 2.27), low income (RR 2.69), and low education (prevalence among decedents: 40.5%) as concentrated risk factors for male suicide. But Lorant, Kapadia, Perelman, et al. (2021), analysing over 102,000 suicides across 392 million person-years in 12 European populations, tested the causal hypothesis directly using an instrumental-variable design that exploited changes in compulsory schooling laws. They found no evidence that education itself reduces suicide risk. The education–suicide association appears to be confounded by shared factors such as mental health vulnerability.

The same underlying traits — poor impulse control, low conscientiousness, mental health difficulties — may drive both the education gap and the suicide gap without either causing the other. LaCorte’s “precursor” claim is correlationally defensible but causally misleading: the classroom gap signals trouble, but closing it would not necessarily close the suicide gap unless the intervention addressed the shared underlying factors.

The gender-equality paradox

One pattern that complicates any simple account of male disadvantage as socially constructed is the gender-equality paradox: countries with higher levels of gender equality, as measured by indices such as the World Economic Forum’s Global Gender Gap Index, often show larger sex differences in STEM participation, not smaller ones (Stoet and Geary 2018).

The paradox is partly contested on measurement grounds (Richardson, Reiches, Bruch, et al. 2020). But Herlitz, Honig, Hedebrant, and Asperholm (2024), in a systematic review of 54 articles and new analyses of 27 meta-analyses and large-scale studies, found that more psychological sex differences are larger, rather than smaller, in countries with better living conditions. Economic indicators such as GDP were the most sensitive predictors of sex-difference magnitude. The authors concluded that “the magnitude of most psychological sex differences will remain unchanged or become more pronounced with improvements in living conditions.”

This limb of the argument has weakened considerably since it was written, and it should now be read with the correction attached. Stoet and Geary’s headline STEM measure required a corrigendum: it was not the share of women among STEM graduates, as the paper stated, but a constructed propensity ratio. More seriously, Ilmarinen and Lönnqvist, in Deconstructing the gender-equality paradox, decomposed the difference-score correlation the whole literature relies on and found the strong results driven by men’s and women’s country means correlating at .93 to .97 — a residual of two nearly identical series rather than evidence of divergence. Is the gender-equality paradox a measurement artifact sets out both grounds.

What that leaves is weaker than this synthesis originally claimed. The paradox cannot currently be used to make the pure social-construction account harder to sustain, because the statistic the claim rests on does not measure what it was read as measuring. Whether any cross-country pattern survives proper decomposition is open, and Herlitz and colleagues’ review is cited here as support in terms that may not survive reading it directly. The honest position is that this synthesis has one fewer limb than it had, not that the opposite conclusion now holds.

Farrell’s power paradox and the two male crises

Farrell’s The Myth of Male Power observed that visible male power at the top — CEOs, presidents, generals — conceals average male powerlessness below. Men dominate the most dangerous occupations, die younger, are incarcerated at higher rates, and are subject to a “masculinity tax” in sentencing. The Disposable Male argues the same disposability from an evolutionary rather than a political premise, which matters because the two accounts imply different remedies.

The power paradox has a mirror image: visible male failure at the bottom can conceal the ordinary male middle. Most men are not in crisis in any domain. They graduate, work, marry, and do not die by suicide. The male-tail phenomenon concentrates trouble at the edges of the distribution, and the men at the bottom are genuinely in trouble — but they are not the average man.

LaCorte (2026) usefully distinguishes two different male populations with different problems: boys at the top who could have attended college but chose trades or entrepreneurship instead, a decision that is sometimes a rational market move given rising college costs and declining returns for some majors; and boys at the bottom whose difficulties compound from early school struggles into adult hardship. The first group resolves its own problem; the second does not.

Open questions

This synthesis raised three questions as their own notes. One has since been answered and became a concept: ADHD relative-age effect now states what the gradient represents rather than asking. Two remain open:

Two verification tasks remain:

  • Independently verify the four-year UK boy-friendly-teaching trial that LaCorte (2026) cites. The claim that it demonstrated null effects is taken from the video rather than from the original study.
  • Check Farrell’s 1993 statistical claims (workplace-death percentages, sentencing-disparity figures, life-expectancy gap) against current data.

Built on 16 sources (5 archived here, 11 external).

Working out connections…