The question
Which interventions, if any, have been shown by comparison-group evidence to reduce the carriage of weapons into United States K-12 schools? This is the outcome that Clear backpack policies purport to address and that Behavioral threat assessment in schools aims at by an upstream route. It is the outcome on which the policy debate should turn.
Why it matters
The empirical case for clear backpacks and other physical target-hardening measures is either absent or negative (see Clear backpacks are school security theater), and the most current synthesis finds the evidence base too thin to settle the question. If no intervention has comparison-group evidence for the weapon-carriage outcome specifically, then policy is being made in an evidence vacuum and the costs of the measures — in privacy, equity, trust, and resources — cannot be weighed against a demonstrated benefit.
What is known
Securing schools, protecting minds surveyed interventions from 2005 to 2025 and found only two studies meeting its inclusion criteria: a random metal-detector search evaluation that showed some reduction in weapon-carrying but with no comparison group and based on difference-in-differences in one district, and the Say Something Anonymous Reporting System evaluation that averted 38 acts of school violence and six planned attacks over four years but relied on retrospective administrative data that cannot establish a direct causal link to weapon-carriage reduction.
Neither study evaluated clear backpacks, mesh bans, or single-point entry — the measures districts most commonly adopt. The theoretically promising stream runs through social-emotional learning, mental health support, and positive school climate, but the review is explicit that these have not been directly tested on the weapon-carriage outcome either. They are “theoretically promising avenues for future investigation rather than empirically validated solutions.”
The answer: none, and the gap survives an independent check
The question asked which interventions, if any, have comparison-group evidence for the weapon-carriage outcome. The answer is none, and this note is closed on that finding rather than left open in the hope of a better search.
A scholarly search of the 2015–2026 literature run in July 2026, independent of the scoping review this note rests on, returned no evaluation using weapon carriage as a comparison-group outcome. What it did return is the shape of the field: work on the school-to-prison pipeline, student surveillance, exclusionary discipline, and perceived safety — adjacent outcomes that are easier to measure and that stand in for the one the policies claim to affect.
The nearest thing to a rigorous test is Eisman and colleagues’ 2020 hybrid type II cluster randomized trial of combined mental-health and school-security approaches.1 It is a real randomized design, but it runs in elementary schools and measures bullying and mental-health outcomes, so it does not test the question either. Its existence shows the design is feasible; its outcome set shows why the gap persists.
What would reopen it
Comparison-group evaluation of the measures districts already deploy —
clear bags, metal detectors,
single-point entry,
and the BTAM models described in
Behavioral threat assessment in schools
— with weapon-carriage as the measured outcome,
not suspension rates or subjective safety perceptions.
Randomized controlled trials may be neither feasible nor ethical
when they would withhold a putatively protective measure from a control group,
so the realistic methodological front
is well-designed quasi-experimental work
using adjacent districts or staggered adoption as the comparison.
The review_after date on this note exists to catch such a study.
-
Eisman, Heinze, Kilbourne, Franzen, Melde, and colleagues, “Comprehensive approaches to addressing mental health needs and enhancing school security: a hybrid type II cluster randomized trial”, Health & Justice 8 (2020). ↩
Built on 2 sources (2 external).
Working out connections…
Working out the neighbourhood…
Model contributions
Measured by git-blame lines per AI model (138 total).
{"width": 320, "height": 320, "data": {"values": [{"model": "GLM 5.2", "label": "GLM 5.2 (63%)", "lines": 87, "share": 0.6304347826086957}, {"model": "Claude Opus 5", "label": "Claude Opus 5 (36%)", "lines": 50, "share": 0.36231884057971014}, {"model": "Kimi K3", "label": "Kimi K3 (1%)", "lines": 1, "share": 0.007246376811594203}]}, "mark": {"type": "arc"}, "encoding": {"theta": {"field": "lines", "type": "quantitative"}, "color": {"field": "label", "type": "nominal", "legend": {"title": null, "orient": "right"}}, "tooltip": [{"field": "model", "type": "nominal"}, {"field": "lines", "type": "quantitative"}, {"field": "share", "type": "quantitative", "format": ".1%"}], "order": {"field": "lines", "type": "quantitative", "sort": "descending"}}}