Your 23andMe report shows 1% of what's in your raw data
23andMe's standard reports cover a narrow slice of your genome. The raw data file contains ~650,000 genotyped SNPs — and you can read the rest legally, for free.
You paid around €100 for a 23andMe kit, spit in a tube, and got back a clean PDF full of ancestry pie charts and a handful of "health predisposition" reports. Useful, but you're seeing a curated sliver. The real file — the raw data — sits behind a download button most people never click.
Here's what's in it, why it matters, and how to read it.
What the raw data actually is
When 23andMe processes your spit, they use a genotyping chip that reads your DNA at roughly 650,000 specific positions (version v5 — older kits read fewer). For each position, they record whether you have zero, one, or two copies of the reference allele.
That file is a ~15 MB tab-separated text document that looks like:
rsid chromosome position genotype
rs548049170 1 69869 TT
rs3131972 1 752721 GG
rs4988235 2 136608646 AA
rs1801133 1 11856378 AG
Each line = one SNP (single nucleotide polymorphism). Every one of those 650,000 lines has published research behind it somewhere.
What 23andMe reports cover
23andMe's health reports — the ones you pay for — touch roughly 150 variants across ~50 conditions. That's about 0.02% of your raw data. They stick to variants with FDA-level evidence because they're regulated as a medical device.
Everything else — caffeine metabolism, warfarin sensitivity, lactose tolerance, folate cycling, nicotine response, statin response, drug-gene interactions for 300+ medications — sits in the raw file, untouched.
How to download it
- Log in at you.23andme.com
- Profile icon → Browse Raw Data
- Click Download in the top right
- Complete the security check
- You'll get an email with a link (~1 hour later). Download the zip.
The file inside is called something like genome_YourName_v5_Full_20230810.txt. That's the entire genotype record — it doesn't expire and doesn't need a subscription to read.
What you can actually learn from it
With the right cross-referencing tools, the raw file reveals:
- Drug metabolism. How fast you process antidepressants (CYP2C19), blood thinners (VKORC1, CYP2C9), opioids (CYP2D6), clopidogrel, statins. Critical for anyone on multiple medications.
- Disease susceptibility. Hereditary haemochromatosis (HFE), Factor V Leiden, age-related macular degeneration risk, type 2 diabetes polygenic risk.
- Nutrient metabolism. Folate (MTHFR), B12 (MTRR), choline (PEMT), vitamin D binding (GC), omega-3 conversion (FADS).
- Fitness genetics. Sprinter vs endurance (ACTN3), mitochondrial training response (PPARGC1A), exercise-induced fat loss (FTO).
- Sleep and circadian rhythm. Clock gene variants (CLOCK, BMAL1, PER3), caffeine sensitivity (CYP1A2, ADORA2A).
- Carrier status for ~400 recessive conditions, many not in 23andMe's default reports.
Privacy considerations
Your raw data is the most sensitive digital artefact you own. It can:
- Identify you uniquely (better than any password)
- Identify your relatives out to cousins
- Expose predispositions that insurers or employers could misuse
Before uploading it anywhere:
- Check the receiver's privacy policy. Look for "we do not share or sell genetic data" — explicitly.
- Prefer services that delete raw files after processing (24h lifecycle is a good benchmark).
- Avoid services that require a paid tier just to delete your data.
- Know your local laws (in the EU, GDPR gives you unconditional right to erasure).
Tools that read it
Historically Promethease was the go-to (now owned by MyHeritage). There's also Whale DNA (that's us) — we focus on drug-gene interactions (PharmGKB), clinical variants (ClinVar), and GWAS polygenic signals, and deliver a personalised action plan rather than a raw database dump.
Whatever tool you choose, the key point is this: the data is already yours. You paid for it. Don't leave 99% of its value locked behind a report you've already read three times.
Ready to see what's hiding in your file? Upload it to Whale DNA →