Privacy is not a claim that a person has something shameful to hide. It is the ability to decide who can learn intimate facts about one's life, relationships, movements, finances, health, politics, and vulnerabilities.
Security is the practical ability to keep those facts, accounts, devices, and services confidential, available, and accurate. Privacy without security leaks. Security without privacy can become a tightly controlled surveillance system.
The two are connected but not interchangeable. Privacy threat modeling should always ask both: who could obtain or alter this information, and who should not have been collecting it in the first place?
The core claim of this synthesis is that privacy limits the conversion of information into Data as coercive power. Across commercial breaches, spyware campaigns, government databases, health systems, dating platforms, and forensic extraction, the recurring pattern is the same: concentrated sensitive data becomes leverage when weak controls, covert access, or exploitative sharing expose it.
Why privacy is a right rather than a preference
Privacy protects autonomy, dignity, intimacy, association, bodily integrity, confidentiality, and the conditions for a self-authored life. People need confidential space to form relationships, deliberate, experiment, change their minds, seek care, and manage the boundaries between family, work, politics, and intimacy.
The rights framing is not rhetorical decoration on a preference. Article 8 of the European Convention on Human Rights protects private and family life, home, and correspondence, and permits interference only where it is in accordance with law, pursues a legitimate aim, and is necessary in a democratic society.1 The three-limb test is doing real work: "in accordance with law" has been read to require foreseeable, accessible rules, which is why several bulk-collection regimes failed on legality before anyone reached proportionality. See Big Brother Watch v United Kingdom. Articles 7 and 8 of the CFR go further, separating respect for private life from the protection of personal data. That separation is unusual and consequential: it makes data protection a right in its own right rather than a derivative of secrecy, so a person retains a claim over information that is already known to others.
This matters for the argument that follows, because most of the harms below involve data that was never secret. A home address is known to the postal service, a neighbour, and a landlord. What changes when a data broker assembles it alongside a workplace, a car registration, a set of relatives, and a nightly location pattern is not secrecy but availability: the cost of selecting a person, locating them, and making a threat credible falls to near zero. The useful frame is contextual integrity — information flows carry norms attached to the context that produced them, and a flow can be a violation even when every individual fact in it was already disclosed somewhere.
"Nothing to hide" fails on three grounds, and it is worth separating them because they have different evidentiary status. First, it misdescribes the interest: the person deciding what is shameful is not the subject but whoever holds the file, and that party's standards change with the government, the employer, and the decade. Second, it assumes the record is accurate; misattributed data produces consequences that are extremely expensive to reverse, which is the mechanism examined in False accusations as an epistemic risk. Third, it treats privacy as an individual good that an individual may waive. It is also a collective condition: a journalist's source, an abuse survivor's shelter address, and a clinician's patient list are all protected by norms that no single person can maintain alone.
Note What this note does not argue
It does not argue that any particular legal instrument is well drafted, and it does not rank privacy against competing interests in a given case. It establishes that privacy is load-bearing, which is the premise the vault's proportionality arguments start from.
Chilling effects are the empirically contested part of the rights case, and the honest position is narrower than the advocacy version. Survey and quasi-experimental work after 2013 found measurable declines in searches and page views on privacy-sensitive topics, and self-reported reticence among journalists and lawyers. The effect sizes are modest and the identification is imperfect. What survives scrutiny is not "surveillance silences everyone" but something more specific and more relevant to policy: the people whose participation is most conditional — the newly diagnosed, the undocumented, the politically exposed, the recently separated — are the ones who withdraw first.
Why privacy is a safety issue
The distance between a data record and a physical injury is shorter than the abstraction of "privacy" suggests. A location trace is a route to a person. A therapy record is a lever against them. An identity document is a key to their accounts. Treating disclosure as a reputational nuisance understates the mechanism by several steps.
The clearest case is intimate partner violence and stalking. Separation is the period of highest lethality, and it is also the period when the abuser's information advantage decays: they no longer share a home, a router, or a calendar. Rebuilding that advantage is exactly what stalkerware, people-search sites, vehicle registration lookups, and family-tracking features sell. The safety literature treats address confidentiality as a primary protective measure precisely because location is the input that converts intent into capability.
The same conversion runs through journalism and civil society. A source's protection is not the reporter's discretion alone; it is the metadata around the contact, the device the reporter carries, and the retention policies of every intermediary in between. Targeted implants collapse all of it at once, which is why the El Faro case below is not a story about one newsroom's misfortune but about the failure of the entire compartmentalisation model on which source protection rests.
A third pathway runs through refugees, defectors, and people with protected status. Here the exposed population is defined by the very fact of appearing in the dataset: to be listed is to be identifiable as someone a state wants to find. There is no mitigation available to the individual after the fact, because the information is not about their behaviour but about their category. The Afghan relocation leak is the reference case, and it is also the cleanest demonstration that a breach's severity is a property of the dataset's subject population rather than of its size.
Privacy supports equal participation
Exposure is not distributed evenly, and this is the part of the argument most often lost when privacy is discussed as a consumer preference. The cost of being known is a function of how far a person's life diverges from whatever the surrounding institutions treat as unremarkable. For most people in most weeks that cost is near zero, which is precisely why the aggregate framing is misleading: the average tells you nothing about the distribution's tail, and the tail is where the policy question lives.
Sexual orientation is the standard illustration. In a jurisdiction where same-sex conduct is criminalised, the mere fact of holding a particular application is prosecutable evidence, so the advertising identifier is not metadata — it is the offence. In a jurisdiction where it is lawful, the same disclosure can still cost a job, a lease, a family. Health status behaves the same way, which is the reason the Vastaamo extortion worked at all: the leverage was not the content of any individual session note but the fact of attendance.
Immigration status, criminal record, debt, and religious affiliation follow the same shape. So does the intersection with gender: address confidentiality matters more to women leaving violent relationships than to the median user, and a "share your location with friends" default that is harmless for most is dangerous for them. A design that optimises for the median therefore does not produce a small harm evenly spread; it produces a large harm concentrated on a small population who cannot opt out of being in it.
Sweden makes the tension explicit rather than hiding it. The offentlighetsprincipen gives broad public access to official records, and the personnummer gives every resident a single stable key that binds those records together. The combination is a deliberate constitutional choice with real democratic value, and it also means that Swedish population-scale linkage is trivially available to any commercial actor holding a publishing certificate. Protected identity exists, but it is administratively heavy, incomplete in practice, and available only after a documented threat.
Privacy is market infrastructure
Privacy is usually argued as a civil liberties question, which cedes the economic ground unnecessarily. Confidentiality is a precondition for ordinary commerce: a negotiating position, a supplier price, an unannounced product, a merger discussion, and a payroll file are all commercially sensitive for the same reason a medical record is personally sensitive. The party who learns them early extracts value from the party who generated them.
Consumer-side, the mechanism is price discrimination and risk sorting. Perfect information about willingness to pay transfers the entire consumer surplus to the seller; perfect information about individual risk dissolves the pooling that makes insurance function. Neither is a hypothetical extreme, and neither requires anyone to behave maliciously. They are what sufficiently good prediction does by default.
The structural failure is an externality, and it is worth stating precisely because it explains why exhortation does not work. A firm holding personal data internalises the fine and the brand damage. It does not internalise the harm distributed across the people in the file:
where p is the probability of a breach, F the expected fine, B the brand cost, and hi the harm to each of n data subjects. Security investment is set against the left-hand side. Any regime that wants the right-hand side to govern must either move F up until it approximates the sum, or reduce n by making collection itself expensive. This is the entire case for Data minimization as a design rule rather than a compliance slogan: it is the only term in the expression a firm can reduce without spending money.
Equifax is the pure form of the failure, because the data subjects were never customers. They had no contract to exit, no product to stop buying, and no way to withhold the data in the first place. The brand term B was therefore close to zero by construction, and the market signal that is supposed to discipline the firm never existed.2 The 2019 FTC settlement was announced at up to $700M across the FTC, the CFPB and 50 states. Against roughly 147 million affected people that is under $5 each before legal costs — an order of magnitude below any plausible hi, which is the point. 2019-07-22-ftc-equifax-settlement.pdfPDF
A map from data to harm
The cases below repeat one structure. Data is collected for a plausible purpose, concentrated for administrative convenience, exposed by a breach, a covert access, or a commercial sharing arrangement, and then used as leverage. The step that decides severity is not the exposure but the concentration, because concentration is what makes the exposure worth attempting and what determines how many people a single failure reaches.
| Data class | What exposure yields immediately | Downstream harm | Reversible |
|---|---|---|---|
| Identifiers | Impersonation, account recovery, credit application | Fraud liability, benefits fraud, arrest on a misattributed record | No — the identifier is practically permanent |
| Location traces | Home, workplace, routine, companions | Stalking, targeted violence, inference of religion or health from destinations | Partly — routines can be changed at cost |
| Health and therapy records | Diagnosis, treatment, the fact of attendance | Extortion, employment and insurance loss, family rupture | No |
| Sexual orientation and practice | Status inferred from app presence alone | Outing, dismissal, prosecution where criminalised | No |
| Financial history | Income, debt, dependants, vulnerability | Price discrimination, risk sorting, precisely targeted fraud | Partly |
| Biometrics and genome | Identity binding, kinship structure | Lifelong re-identification, exposure of relatives who never consented | No |
| Communications content | Beliefs, associations, plans, sources | Prosecution, blackmail, burned sources | No |
| Full device extraction | Everything above, from one seizure | All of the above, compounded and admissible | No |
Case studies
Ten cases, selected because each isolates a different step in Figure 1 rather than because each was large. Sizes are given where they are documented, but the argument does not run on them: the Afghan relocation leak involved a few thousand records and is the most severe entry in the list.
Ashley Madison
In 2015 a group calling itself the Impact Team published the customer database of a site marketed for extramarital affairs, covering roughly 32 million accounts. The record was thin — email addresses, partial payment data, stated preferences — and none of it was medically or financially sensitive in the usual taxonomy. Its leverage came entirely from context. Extortion letters referencing the dump were still circulating years later, addressed to people whose only exposure was an address in a file, including addresses that had never been verified as belonging to a user. Several suicides were reported in the months after publication. The case establishes that sensitivity is contextual rather than categorical, and that a dataset's harm continues to accrue long after the breach ceases to be news.
Equifax
In 2017 an unpatched Apache Struts vulnerability, disclosed and fixed months earlier, gave attackers access to the records of roughly 147 million people: names, dates of birth, social security numbers, and in many cases driver's licence numbers. Almost none of them had chosen to be in the file. Credit bureaus collect from lenders, not from subjects, which removes the exit that consumer markets rely on for discipline. The technical failure was mundane; the structural failure is that the firm's incentive to prevent it was decoupled from the harm it caused. This case is the empirical anchor for the externality argument above.
Vastaamo
The Finnish psychotherapy provider Vastaamo held session notes in a database that was, on the investigators' account, exposed with weak protection over an extended period. In 2020 the attacker demanded payment from the company and then, when it refused, contacted patients individually and demanded a few hundred euros each under threat of publishing their therapy notes. Roughly 33,000 people were affected; a subset of the notes was published. Aleksanteri Kivimäki was convicted in 2023 on more than 20,000 counts and sentenced to six years and three months. Vastaamo is the single most instructive case in this list, because the extortion was retail rather than corporate: the attacker discovered that leverage over individuals is easier to monetise than leverage over the institution that failed them.
Tip Design rule the Vastaamo case yields
Where the fact of attendance is itself the sensitive datum, no amount of content encryption helps. The control that would have mattered is not storing the association between an identity and a service in a form any single compromise can read — which is a data model decision made years before the incident, not a security product bought after it.
SpyFone
SpyFone sold covert monitoring software that harvested location, messages, photos, and browsing from a target device while hiding its own presence. The FTC banned the company and its chief executive from the surveillance business in 2021, citing both the covert design and the failure to secure what it collected — the product exposed the harvested data to third parties as well. The case matters because it names the category honestly: the customer is not the data subject, the concealment is the feature, and the market is people who want an information advantage over someone who lives with them.
Grindr
Grindr appears twice. In 2018 it was found to have transmitted user HIV status to two analytics vendors as part of ordinary product telemetry. In 2021 the Norwegian data protection authority fined it roughly 65 million kroner for sharing user data with advertising partners without a valid legal basis, on the reasoning that being a Grindr user is itself information about sexual orientation and therefore special-category data under GDPR Article 9. That reasoning is the durable part. It establishes that an inference drawn from application presence carries the protection of the category inferred, which generalises far beyond this app.
Pegasus against El Faro
Between 2020 and 2021, forensic analysis identified Pegasus infections on the devices of more than twenty journalists and staff at the Salvadoran outlet El Faro, with some phones reinfected repeatedly over months. A mercenary implant against a newsroom defeats source protection in a way no institutional policy can compensate for: it does not intercept a channel, it takes the endpoint, which means the reporter's notes, the source's identity, and the unpublished draft are all available at once. This is the case that establishes why endpoint integrity is a precondition for the rest of the stack, and why arguments about Memory Integrity Enforcement and comparable mitigations are not vendor marketing.
Afghan relocation data leak
A UK Ministry of Defence spreadsheet containing details of Afghans who had applied for relocation on the basis of work with British forces was mishandled in 2022 and subsequently circulated. The affected population was defined by having assisted a foreign military — that is, by being precisely the category the returning government was searching for. A related 2021 incident had already exposed hundreds of applicants' addresses by sending an email with recipients in the visible field. The government obtained an unusually broad injunction that delayed public reporting for a long period, which added a second harm: the people in the file could not learn that they were in it. Record count is a poor proxy for severity; subject population is the variable that matters.
OPM
The 2015 compromise of the US Office of Personnel Management exposed background-investigation records for approximately 21.5 million people, plus fingerprints for around 5.6 million. The SF-86 form is not an employment record; it is a structured inventory of the applicant's foreign contacts, debts, drug history, mental health treatment, and relationships — assembled expressly to identify the levers a foreign service might use. The state built a counterintelligence targeting database about its own cleared personnel and then lost it. Fingerprints, unlike passwords, cannot be rotated.
VTech
The 2015 breach of the children's tablet maker VTech exposed profiles for around 6.4 million children, including names, dates of birth, genders, and in some cases photographs and chat logs between children and parents. It is included here for one reason: the subjects could not consent, cannot meaningfully be notified, and will carry the record for a lifetime measured from a point before they could read it. Whatever the correct policy is for adolescent data — a question the vault treats separately under Age assurance — it cannot be derived from a consent model that the youngest subjects were never in a position to enter.
23andMe
In 2023 attackers used credentials leaked from unrelated services to log into a small number of 23andMe accounts and then, through the opt-in relatives feature, harvested profile data for roughly 6.9 million people who had never been compromised themselves. Curated lists targeting people of Ashkenazi Jewish and Chinese descent were subsequently offered for sale. The company entered bankruptcy in 2025 and the database became an asset in the proceedings. See 23andMe data breach for the timeline. Three distinct lessons compound here: genetic data implicates relatives who never transacted with the company; a social feature can be a bulk export path without any vulnerability in the classical sense; and a privacy promise is only as durable as the corporate entity that made it.
Why lawful extraction is still an ethical question
Everything above concerns unlawful or unconsented access. The harder case is access that is lawful, authorised, and still disproportionate — which is where mobile forensic extraction sits. A modern phone extraction is not the seizure of a document. It is a copy of the location history, message archive, photo library, health data, and deleted-but-recoverable remnants of one person, and typically of everyone who corresponded with them.
The legal authority is usually real. A warrant, a consent form, or a border-search power exists, and the tooling — Cellebrite and its competitors — is bought under public procurement. The gap opens between the authority's stated scope and the extraction's actual scope. An investigator authorised to look for one conversation receives an image containing everything, and the filtering that is supposed to constrain what reaches the file is implemented in a vendor tool whose behaviour is not usually disclosed to the defence.
"likely to generate in the minds of the persons concerned the feeling that their private lives are the subject of constant surveillance"
Court of Justice of the European Union, Digital Rights Ireland, C-293/12 and C-594/12, on generalised retention of traffic data
That sentence was written about retention, but it describes extraction more accurately, because retention produces a record someone might consult and extraction produces a copy someone already has. The vault's open question here is narrow and answerable, which is why it is a question note rather than a complaint: if the filter configuration determines what evidence exists, its disclosure status is a fair-trial issue and not a procurement detail.
Danger A hard boundary, not a preference
Consent to a device search given at a roadside, a border, or a police station is not a reliable basis for anything in this analysis. The person is not free to leave, is rarely told the extraction's scope, and generally cannot revoke it once the image exists. Where the vault relies on a consented extraction as evidence of anything, that reliance must be stated explicitly at the point of use.
Objections
Four objections are strong enough to state in their own terms rather than in a weakened form.
"The data was anonymised." Anonymisation is a property of a dataset in a context, not an intrinsic attribute, and the context includes every other dataset an adversary can obtain. The recurring result across two decades of re-identification work is that a handful of quasi-identifiers — a birth date, a postcode, a sex, or four timestamped location points — suffices to isolate most individuals in a population-scale release.3 The four-location-points figure comes from de Montjoye et al. (2013) on mobility traces, and the birth-date-postcode-sex result from Sweeney (2000). Both are frequently overstated in secondary sources: they establish uniqueness in the released dataset, which is a necessary but not sufficient condition for re-identifying a named person. The correct response is not to abandon the concept but to stop treating it as a binary that a processing step confers.
"People consented." Consent is doing work it cannot bear at this scale. It is granted once, in advance, to a party the person cannot audit, over uses that did not exist at the time, with the alternative being exclusion from a service that has no substitute. The vault's position is that consent is a legitimate basis for a narrow, contemporaneous, revocable decision, and a fiction for anything structural. Where a design depends on consent to be defensible, that is a signal the design is wrong.
"Privacy protects offenders." It does, in the same sense that every procedural protection does. The relevant question is the exchange rate, and it is answerable rather than rhetorical: how many additional detections does a given measure produce, and how many false positives, at what base rate. This is where most access proposals fail arithmetically rather than ethically.
Warning Base rates decide the argument, and they are usually omitted
A classifier with 99% sensitivity and 99% specificity, applied to a population where one message in a million is the target, returns roughly one true positive for every ten thousand false ones. Any proposal for population-scale scanning that does not state the assumed base rate and the resulting positive predictive value has not been evaluated; it has been described. See Probabilistic interpretations of beyond reasonable doubt and the ASA statement on p-values for the underlying reasoning.
"Security requires access." This is the strongest objection and the one with the most institutional weight behind it. The response is not that lawful access is never justified but that the mechanism proposed determines the answer. Targeted access against an identified device, under judicial control, is compatible with the analysis above. Access implemented as a standing capability against everyone's devices is not, because it changes the property being protected from "this person's data is confidential" to "this person's data is confidential until a process says otherwise" — and every case in the list above is an instance of a process saying otherwise. The architectural version of this argument is developed in Apple Private Cloud Compute and, for the verification side, in Reproducible builds.
Regulatory setting
The instruments are not the argument, but the argument has to land somewhere. In the EU, GDPR supplies the baseline — lawful basis, purpose limitation, minimisation, and the special-category regime that made the Grindr reasoning available. The ePrivacy Directive still governs terminal-equipment access, which is why cookie consent sits in a different instrument from the rest. The Digital Services Act adds transparency and risk-assessment duties for large platforms without altering the data-protection baseline. The proposed regulation on child sexual abuse material is the live pressure point, because a detection order against interpersonal messaging is exactly the standing capability the previous section distinguishes from targeted access.
The UK Online Safety Act runs a parallel track with a different structure: duties of care enforced by a regulator, backed by a technology-notice power whose interaction with end-to-end encryption remains unresolved in practice. The United States has no general statute and instead a sectoral patchwork — HIPAA for covered health entities, FCRA for consumer reporting, plus a growing set of state laws — which is why the Equifax and SpyFone actions were brought on unfairness and deception theories rather than on a privacy right.
Question Open on the frontier
Whether Swedish mobile-extraction filter configurations are disclosed to the defence is unresolved and materially affects the extraction section above. Tracked at Are Swedish mobile-extraction filters disclosed at trial?.
Three checks are outstanding on this note. They are listed as work rather than as caveats, because each is executable and each would change a specific sentence:
- Done: Verify the Vastaamo conviction count and sentence against the district court judgment rather than press reporting
- Done: Replace the secondary summary of the Norwegian Grindr decision with the Datatilsynet decision text
- To do: Confirm the current status of the EU child sexual abuse regulation proposal before the
review_afterdate; the detection-order provisions have been redrafted at least twice - To do: Check whether the Afghan relocation injunction has been fully discharged, and cite the judgment rather than reporting about it
- To do: Recount the
sourcesinventory: two entries appear to point at the same archived PDF under different filenames