Field Notes

synthesis

Rättssäkerhet in Swedish criminal cases

Swedish criminal procedure declares a demanding proof standard, but the standard is not calibrated against a population whose true guilt is independently known. Courts decide the labels later used to claim that courts are reliable. Successful appeals and exonerations reveal only errors that the same correction system managed to recognize.

The critical problem is therefore not merely that some innocent people are eventually exonerated. It is that the system has no method for measuring how many incorrect factual narratives remain legally final, while its procedures contain mechanisms that can turn an early allegation into a self-reinforcing conviction.

Rättssäkerhet carries three senses, and the one this claim uses is worth naming. Neither the EU-law register, where it means legal certainty, nor the formal register, where it means predictability and the structural conditions for it, says anything about whether a court got the facts right: a tribunal applying clear rules faultlessly to a false factual narrative satisfies both. The sense this note argues in is the one Justitiekanslern used in Rättssäkerheten i brottmål — confidence that nobody is convicted unless guilt has been proved beyond reasonable doubt — which is a claim about outcomes. This note then goes one step past even that, to whether the system can measure how often it fails the test at all.

The strongest Swedish evidence for that claim comes from the state’s own Felaktigt dömda project, the Bergwall commission, Moa Lidén’s Confirmation bias in criminal cases, and causal evidence about Swedish lay judges and political influence. Thérèse Juel’s Fällda för sexövergrepp and Dömda add case narratives and help locate the human consequences, but their claims require the source criticism stated in their source cards.

“Beyond reasonable doubt” is not scientific validation

Scientific proof normally depends on observable outcomes, replication, error measurement, and methods that can be challenged against known or independently testable states of the world. A criminal trial usually lacks those conditions. It reconstructs a past event from incomplete traces under institutional time and resource constraints.

“Beyond reasonable doubt” is a normative decision threshold, not an empirically calibrated confidence level. The phrase cannot tell us whether a judge’s felt certainty corresponds to one error in a hundred, one in ten, or no stable frequency at all.

The recurring claim that the standard means “about 98 percent certain” has a recognizable source in Swedish legal writing, but not the status often implied by casual repetition. Kvalitetssäkring av bevisprövningen i brottmål records 98 percent as an attempted translation and then says it is impossible to verify whether judgments are factually correct at that rate.

Probabilistic interpretations of beyond reasonable doubt draws the implication carefully. If each conviction really represented a calibrated 0.98 probability of guilt, one innocent defendant in fifty would be the expected result. The harder objection is that no calibrated model produces the 0.98. The number can therefore convert an unvalidated feeling of certainty into a scientific-looking quantity without measuring either the individual case or the system’s error rate.

The familiar p < 0.05 convention makes the defect more visible. As the ASA statement on p-values stresses, even that formally calculated statistic does not mean that a hypothesis is 95 percent likely to be true, and scientific decisions should not rest on the threshold alone. The Swedish 98-percent gloss lacks even the specified model and repeatable calculation that give a p-value its limited meaning. Yet the legal decision can permanently harm the accused and end meaningful investigation of a different perpetrator.

Rättssäkerheten i brottmål recognized this problem in practical terms. It warned that a judge may confuse personal belief in guilt with satisfaction of an objective standard, and urged courts to test alternative hypotheses rather than rely on a total impression. The report’s own small later samples did not establish system-wide catastrophic failure. That is a genuine limitation on prevalence claims, not a refutation of the mechanisms it identified.

What the first JK project actually found

The 2006 project examined eleven serious convictions followed by a new trial and full acquittal. Eight concerned sexual offences against children. This was a selected set of known failures, so it cannot estimate the rate of error. It can reveal how errors survived.

The report found defects in every investigated case. The recurring pattern was not one bad witness or one careless judge, but a chain:

Stage Repeated failure Why later review may not cure it
Allegation Suggestion, contagion, treatment influence, or vague time and place The original conditions disappear and the account hardens through repetition
Investigation Investigative confirmation bias, leading questions, missing alternative suspects, and incomplete files Prosecutors and courts see an apparently coherent record built by the first hypothesis
Expertise Treating clinicians used as court experts and opinions extending beyond the science Technical language gives a credibility judgment borrowed authority
Defence Missing witnesses, weak counter-expertise, and inadequate challenge The trial record falsely appears to contain the best available opposition
Judgment Unvalidated credibility cues and unsourced “general experience” Appellate review inherits both the record and the earlier framing
Reopening The convicted person carries the burden and depends on prosecution-side investigation Finality and confirmation bias make the error hard to reconstruct

The report said that the original evidence evaluation was wrong in all eleven cases, with reservations about two. After renewed investigation and trial, it found no remaining suspicion against any of the eight defendants in the sexual-offence cases. In three cases, grounds eventually accepted for reopening had already appeared in an earlier rejected petition.

That last finding matters. A system that eventually reopens a case can still have rejected the same grounds for years. The evidence was not weaker on the earlier petition; institutional confidence was higher, and outside work had not yet changed the frame.

Bergwall shows the failure chain at full scale

The Bergwall commission documents the same chain operating for years in Sweden’s most scrutinized murder cases. That chain is Coercive interrogation and false confessions, and a documented US case reached the same structural result through different pressure tactics. Sture Bergwall, then known as Thomas Quick, was convicted of eight murders between 1994 and 2001 on confessions he later retracted; every conviction was set aside by 2013.

The commission’s findings read as a checklist of the mechanisms this synthesis describes. Investigators adopted the suspect’s own theory that wrong answers were “deliberate deviations,” so disconfirmation stopped functioning. Interviewing conveyed investigative findings to the confessor, so his account converged on the facts it should have been tested against. A stable core of interrogator, prosecutor, therapist, and advising psychologist developed what the report calls groupthink. Court experts had already worked for the prosecution. Therapy aimed at recovering repressed memories, conducted under heavy benzodiazepine medication, fed directly into the murder files. And the trials lacked opposition entirely: with one late exception, even the defence argued for conviction.

The episode also sharpens the reopening critique. The convictions fell only after journalists and a new lawyer reconstructed the record from outside, not because any internal control detected the pattern. A system whose corrections arrive mainly through external investigation cannot cite its correction record as evidence of internal reliability.

The correction channel runs through journalism

Bergwall is not an isolated pattern. Sweden’s other prominent contested cases show the same division of labour: institutions produce and defend the narrative, and journalists perform the error detection.

Kevinfallet shows the mechanism operating with no adversarial forum at all. Two brothers aged five and seven were designated killers of four-year-old Kevin Hjalmarsson in 1998 entirely inside the investigation, since children below criminal responsibility face no trial. The sealed interrogations, opened by SVT and Dagens Nyheter in 2017, contained heavily leading questioning of small children, one session in which the older boy said 61 times that he could not or did not want to continue, no recorded confession, and unresolved alibi evidence. The prosecutor cleared both men in 2018, nineteen years after the file was closed.

Da Costa-fallet shows a finding no procedure can reach. The 1988 acquittal of two doctors carried reasons declaring it proven beyond reasonable doubt that they had dismembered the victim, an offence already time-barred and never adjudicated. An acquittal cannot be appealed by the acquitted, and resning targets convictions, so a career-destroying factual finding sat permanently outside every correction mechanism. The 2026 ex gratia payments closed the matter financially without any procedure ever testing the finding itself.

Kaj Linna carried the reopening burden in practice. Convicted of the Kalamark robbery murder in 2005 on one witness’s account, Linna was acquitted in 2017 after the podcast Spår re-interviewed that witness and recorded him changing his story on decisive points. The material that satisfied Högsta domstolen was produced by two podcasters, twelve years into the sentence, doing witness work no institution was funded to repeat. Samir Sabri, Joy Rahman, and Esa Teittinen were likewise acquitted only after reopened proceedings that outside advocates and journalists helped force.

Mordet i Västermalmsgallerian shows the reopening gate operating in the present tense. A convicted man’s brother confesses to the killing; journalists find an undisclosed cut in the courtroom video and recover the probable weapon where police had searched and found nothing; and the reopening prosecutor declines examination, reasoning that even matching DNA would not warrant resning. The gate answered the physical evidence in advance, before anyone looked at it.

These cases differ in era, offence, and outcome, but they agree on the structural point this synthesis makes: the institutions that produce a guilt narrative also control every channel through which it can be revised, and in practice revision has depended on actors outside the system doing unfunded investigative work.

Administrative coercion runs the pattern without the safeguards

The same failure chain operates outside criminal procedure, where the formal safeguards are weaker still. Adam och övergreppen documents Helsingborg’s social services falsely designating a seven-year-old a sexual abuser and his father a perpetrator. Criminal-law institutions worked as intended: police, prosecutors, and courts all dismissed the accusations. The compulsory care, years of family separation, and perpetrator-framed therapy happened anyway, through LVU proceedings in which the deciding courts saw only the dossier the accusing agency assembled. The agency concealed contradicting information, including the child’s own statements, and the external investigation ordered after Uppdrag granskning’s 2025 broadcast found that no evidence had ever existed, describing the handling as deficient in rättssäkerhet.

The structural lesson extends this synthesis beyond the criminal courts. An administrative body that investigates, formulates the hypothesis, selects the record, and then litigates on that record holds the same narrative monopoly as a criminal investigation under tunnel vision, while its targets have weaker rights to counsel, disclosure, and adversarial testing. Guilt-equivalent designations can therefore be imposed, and years of coercion executed, below the evidentiary thresholds this note criticizes as uncalibrated even in criminal cases. Here too the correction arrived through journalists reading several thousand pages of case files.

False accusations cannot be dismissed by counting classifications

False accusations as an epistemic risk separates deliberate fabrication from sincere error, memory contamination, mistaken identification, and prosecutorial overinterpretation. All can produce a false criminal accusation in the operative sense that an innocent person becomes the object of a guilt narrative.

Studies reporting a small percentage of police reports as “confirmed false” count only accusations that investigators could and chose to classify that way. They cannot observe false accusations that remain plausible, become convictions, or are never independently resolvable. The opposite mistake is also unacceptable: an acquittal is not proof that the complainant lied.

The honest conclusion is methodological. There is no reliable denominator for all true and false accusations, especially when the allegation cannot be compared with independently auditable traces whose provenance and error processes can themselves be tested. An additional witness is evidence, not ground truth. Under Witness reports are not ground truth, separate people can share perceptual error, post-event information, suggestive interviewing, or a common investigative frame. Risk must be studied through known failure mechanisms, auditability, and the cost of an uncorrected error, not through confident percentages built from institutional labels.

Sexual-offence cases expose the proof problem

Sexual offences often present the hardest legitimate adjudicative problem: the event may occur in private; physical evidence may be absent or equivocal; and both a true complainant and a falsely accused defendant may have little corroboration available. The need to prosecute real violence does not solve that epistemic problem.

The first JK report found particularly serious failures in its eight sexual-offence cases: deficient child interviews, outside influence, expert overreach, inadequate testing of accounts, and judgments relying on the supposedly self-experienced quality of a narrative. Fällda för sexövergrepp describes the same mechanisms across ten additional contested cases, including allegations whose dates shifted after an alibi appeared.

Evidence patterns in rape and unlawful-threat cases found that one selected evidence pattern was followed by conviction in 90 percent of sampled rape cases and 18 percent of sampled unlawful-threat cases. The authors’ candidate explanations were a lower operative proof threshold or an overvaluation of the support evidence in rape cases; they claimed nothing about political pressure. In the ensuing exchange, Wegerstad argued that the selected situations were not truly comparable and that a bivariate comparison of hand-selected judgments cannot separate offence type from confounding differences, while Dahlman answered the chance objection with statistical significance, which leaves the confounding objection standing. The discrepancy remains a striking descriptive signal whose cause the design cannot identify.

Political salience can still enter through institutions. Politics in the courtroom found that randomly assigned Left Party lay judges increased convictions in cases with female victims by about 14 percentage points. That result does not identify every case as a sexual offence, but it provides causal evidence that a politically feminist party affiliation affected verdicts for a fact pattern directly relevant to gendered criminal cases.

Swedish fact-finding combines discretion and weak calibration

Free evaluation of evidence in Sweden permits courts to consider relevant material without a general exclusionary code and gives judges broad responsibility for weight. Freedom from rigid evidentiary rules can prevent technical acquittals. It can also hide method inside experience.

The first JK report warned that free evaluation may become too free when courts use speculative trauma theories, demeanor, story coherence, or “general experience” without an empirical foundation. The central defect is not discretion itself. It is discretion without a disclosed, testable method for translating evidence into the proof threshold.

Scandinavian legal realism belongs in the background, not as a monocausal explanation. Its influence on Swedish legal method, procedural doctrine, and legislative source orientation is documented, and its rights skepticism helps explain why uncalibrated credibility assessment met no principled jurisprudential resistance: a culture trained to hear rights-based objection as metaphysics had no native vocabulary for demanding calibrated proof as a matter of the accused’s right. The more direct route to a “vibes-based” outcome remains unvalidated credibility assessment operating under broad evidentiary discretion, Realist foundations of Swedish rättssäkerhet deficits argues that connection across three channels, and Did legal realism weaken Swedish rights protection keeps the stronger causal claim open.

Political lay judges are not a peer jury

Swedish ordinary criminal trials do include lay participation, but Swedish lay judges and political influence makes the jury label misleading. Parties nominate lay judges; political assemblies elect them; and they vote with the professional judge.

The causal Gothenburg study found that Sweden Democrat lay judges increased convictions for defendants with Arabic-sounding names and that Left Party lay judges increased convictions in cases with female victims. The 2026 Riksrevision audit of Swedish lay judges is reviewing composition, suitability, training, and bias after cases involving removal and retrial.

The problem is not that every lay judge follows a party line. It is that political selection is built into the adjudicating panel and can measurably change legal outcomes.

Anonymous witnesses narrow the defence’s testing surface

Since January 2025, Anonyma vittnen allows courts to receive testimony whose source the defendant cannot identify, for serious crimes where a concrete threat to the witness exists and lesser protections fail. The threat problem the reform answers is real, particularly in gang-related prosecutions.

The rättssäkerhet cost is structural rather than hypothetical. Credibility testing in Swedish courts already rests on unvalidated discretionary judgment; anonymity removes the defence’s ability to investigate the witness’s relationship to the parties, motive to lie, or prior reliability. Lagrådet and Advokatsamfundet criticized the proposal in unusually strong terms on exactly this ground. The safeguard the law offers, particularly careful evaluation of anonymous testimony, is the same uncalibrated evaluation whose weakness this synthesis documents.

Remand makes the process coercive before conviction

Swedish remand detention and restrictions adds a separate failure mode. Nine months is now a defeasible statutory cap on continuous remand before charge, not a maximum total period; a court can permit an extension for exceptional reasons. Restrictions intended to prevent obstruction can produce isolation from other people and information.

The CPT has repeatedly criticized Swedish remand restrictions. In 2026, JO inspections of remand isolation found that a majority of detainees at four inspected facilities were isolated, several for an extended period. The official Färre i häkte och minskad isolering inquiry had already concluded that restrictions were used more widely than justified and that the surrounding culture required fundamental change. The harm is not made humane by describing detention as administrative or evidence-protective. Prolonged isolation before adjudication can impair health, defence preparation, and resistance to an investigator’s narrative.

Lidén’s experiments identify a second effect: judges who had ordered detention later rated the prosecution evidence as stronger and were more likely to convict. The coercive measure may therefore pressure the accused while anchoring the decision-maker. This is the detention-stage form of Investigative confirmation bias, and it is the direct warrant for separating the judge who detains from the judge who tries.

Harriette Broman remand isolation is a concrete corrected case. Broman spent 556 days under restrictions, alone for 23 hours a day, before full acquittal on appeal. The duration was about eighteen months rather than two years, but the example supports the broader claim: Swedish procedure can impose prolonged solitary confinement on a person whom the final adjudication does not find guilty.

The system’s physical capacity has since deteriorated. JO inspections of remand detainees in police arrest found that lack of prison-service places left remand detainees in police arrest for more than two weeks, sometimes alone for 23 hours a day without meaningful activity. Double occupancy in Swedish remand prisons found single cells routinely used for two people, including six-square-metre rooms, with ventilation, equipment, privacy, and health concerns.

The 2026 More flexible remand and prison enforcement proposal would abolish the statutory starting right to a single room and legalize temporary police-arrest placement for capacity reasons. That proposal is not yet law, but it shows that the policy response has shifted from the 2016 goal of less detention and isolation toward legal flexibility for overcrowding.

The encryption claim requires precision. Swedish police may compel biometric unlocking, but a suspect is not legally required to state a PIN or password. No reviewed source proves a general policy of prolonging detention to force encryption keys. A particular case could still show that interaction; it would require detention records linking continued custody to device access or alleged obstruction.

Duress credentials and coercive extraction raises a harder value question. A duress wipe can create serious legal exposure after seizure or a preservation duty. Domestic evidence law does not reach the further question: whether a person is morally obliged to cooperate under isolation that international standards treat as prohibited or inhuman. Supporting a user-controlled safety feature is also distinct from a business advising a particular person to destroy evidence in an active case.

Device extraction creates evidence by selection

The right to silence that shapes an interrogation is itself a distinctly Swedish construction: Swedish right to silence and förklaringsbörda rests on RB 35:4’s evidence-weighing rule, older and facially more permissive than England’s 1994 reform, which Högsta domstolen’s own förklaringsbörda doctrine has narrowed in practice, while RB 23:12 categorically bars the kind of interrogation deception that current US law permits. Swedish encryption and passcode disclosure law extends the same privilege to a seized device’s passcode.

Mobile-device extraction and evidentiary selection connects privacy directly to adjudicative reliability. Cellebrite is not reserved for terrorism; public Swedish investigation files show UFED use as ordinary criminal-forensics infrastructure.

The extraction may contain thousands or millions of artifacts. The court receives a selected set. Authentic messages and timestamps can therefore support a false narrative when the date range, surrounding conversation, alternative account, or tool limitation is omitted. Cognitive and human factors in digital forensics shows why technical expertise does not remove this human selection problem.

Swedish law permits important seizure and copying decisions by investigators, prosecutors, and in urgent situations police, with judicial review often available after the fact rather than required before every extraction. That is a more accurate criticism than saying there is no law or oversight. The weakness is broad access combined with limited ex-ante adjudication and an evidentiary pipeline whose filters may be invisible at trial.

Bulk encrypted-chat evidence extends the same problem across borders. Swedish prosecutions have relied heavily on EncroChat material obtained by French authorities through methods protected by French defence secrecy. Swedish courts admit the material under free evaluation of evidence, and Högsta domstolen declined to review the question after the prosecution service argued that practice was already uniform. The result is conviction-grade evidence whose collection method, error rate, and selection filters cannot be examined by the defence, the trial court, or any Swedish institution. Whatever the guilt of particular defendants, that combination places the reliability question outside the adversarial process entirely. Whether any Swedish court has nonetheless discounted weight or demanded disclosure when the defence objected is the open question How do Swedish courts handle EncroChat reliability challenges.

Reopening is structurally backward

Sweden has no separate criminal-cases review commission equivalent to the Norwegian or British model. The 2013 reforms created clearer duties to reopen an investigation and some access to counsel, but the system still routes factual reinvestigation through prosecutors and police after their side obtained the conviction.

The first JK report called this structure backward. The convicted person must first produce enough new material to activate institutions that may already be committed to the old narrative. Finality then works as an evidentiary presumption even though finality is not evidence of factual accuracy.

The corrected cases in this note are subject to a selection effect that cuts both ways. Every acquittal after resning is, in part, the system eventually working, and counting such cases measures the Correction channel, not the error rate. The epistemically dark set is the denials. Billy Butt marks its boundary: convicted of nine rapes in 1993, he attached to a 2006 application a letter from some of his own accusers attesting his innocence, and Högsta domstolen still denied resning, with two of five justices dissenting in favor of reopening. Whether Butt is innocent is unknowable from outside, and that is the point. If recanting complainants and a split Supreme Court do not clear the threshold, the denied set can contain wrongful convictions that no achievable evidence would ever reopen, and the system records each denial as confirmation that finality was correct.

The international frame does not dilute the claim

Two misreadings of this synthesis are worth closing off. The first treats the documented mechanisms as a Swedish pathology: they are not. Confirmation-biased investigation, confession centricity, expert overreach, non-disclosure, and journalism-dependent correction are documented in every comparable democracy, including every Nordic country. The second concludes that Sweden is therefore ordinary and the critique overblown: that misreads the comparison in the opposite direction. Norway responded to its scandal sequence by building an independent reopening commission that publishes everything; Sweden responded to its own with reports and no institutional change. Hidden miscarriage risk and correction-channel opacity across democracies works through the evidence: the failure mechanisms are ubiquitous, but Sweden’s inability to see them is near-maximal among its ranked peers, and its low visible exoneration rate is evidence about its gate, not about its accuracy.

A skeptical reform program

No reform can make historical guilt scientifically observable. The attainable goal is to make uncertainty visible, reduce correlated error, and make adverse narratives reproducible and contestable.

The evidence reviewed here supports:

  1. an independent reopening and investigation body with power to obtain unused material and commission new expertise
  2. separation between judges who order detention and judges who determine guilt
  3. recorded, preserved, non-leading interviews with disclosure of the full interaction
  4. a requirement to answer each material defence hypothesis in written reasons
  5. state-funded independent expertise where the prosecution relies on technical or behavioral claims
  6. disclosure of complete digital-forensics search protocols, exclusions, limitations, and contextual material
  7. offence-specific statistics on evidence patterns, appeals, reopening applications, and dissent
  8. selection of lay adjudicators without party nomination, or a professional-only fact-finding model
  9. strict and progressively heavier review of prolonged remand and isolation
  10. institutional error reviews that study all contributing decisions rather than assign blame to the last visible actor

Privacy is an epistemic safeguard

The connection to Case for privacy and security is not merely that innocent people have secrets. Large private archives let the state search backward for fragments that fit an accusation. The more complete the archive, the more possible stories can be assembled from true but decontextualized facts.

Privacy and security reduce this attack surface. They protect third parties, privileged relationships, intimate experimentation, and the unrecorded context that cannot follow an extracted artifact into a courtroom. After lawful seizure, auditability and full contextual disclosure become the parallel safeguard.

The presumption of innocence is therefore not only a courtroom instruction. It has an architectural implication: institutions should not accumulate, extract, and selectively narrate more private data than they can reliably interpret.

Open research

  • Trace the post-2009 implementation of both JK projects’ recommendations, especially independent reopening and reasoned evidence analysis.
  • Obtain detention applications and decisions in cases where encrypted devices remained inaccessible.
  • Review the National Audit Office lay-judge report after October 2026.
  • Build a Swedish corpus linking full preliminary-investigation files, digital-forensics reports, judgments, appeals, and reopening decisions.
  • Compare sexual-offence evidence patterns before and after major statutory and doctrinal changes without using conviction as ground truth.
  • Read prop. 2024/25:20 and the statute text on Anonyma vittnen, then follow early applications and appellate treatment.
  • Ground the redlinked reopened cases, Samir Sabri, Joy Rahman, and Esa Teittinen, in primary judgments and the journalism that drove each resning.
  • Survey Kalla fakta, Uppdrag granskning, and P3 Dokumentär archives for further contested-conviction and administrative-coercion cases worth entity notes.

Built on 25 sources (20 archived here, 5 external).

Working out connections…