# Frequently Asked Questions

<a id="how_it_works"></a>
## How does it work?

codegen.eu turns your raw genetic data into a report you can explore - free to start. Upload your provider's file and we match your genotypes against a curated database of variants and **4,000+ topics** (diseases and traits, grouped into 15 major categories), score each one against the general population, and let you search it like documentation. The most relevant findings are surfaced first.

Not sure where to begin? Open the **dashboard** to see your strongest signals versus the population, your established (GWAS) associations, and fitness & trait highlights. From there, search any disease, gene or rs-ID, or browse by category from the left rail.

<a id="login"></a>
## How do I log back in?

We do not keep named accounts. To return to your report, simply **re-upload the same raw genetic data file** - your DNA is your login. We store only an anonymous, one-way hash of a noise-only subset of your data; when the same file is uploaded again the hash matches and your saved **favorites** come back.

If you chose **Erase all data** at logout, your report data was deleted - but your **paid access is kept**, so you never lose what you paid for. Upload the same file again and your paid status is restored and the report regenerated. (Comments are **anonymous and never linked to you or your hash**, so they are not affected - see the Privacy Policy.) See the [Privacy Policy](#privacy) for exactly what we keep (almost nothing) and why.

<a id="dna_hash"></a>
## How can I access my report if you don't store my DNA?

Your DNA file *is* your login. We compute a one-way **hash** from a small, deliberately uninformative subset of your data - it carries no health, trait, or ancestry signal and cannot be reversed into your genome. When you upload the same file again the hash matches and your saved favorites return. We never keep the file itself, your name, or your email. The full design is in the *DNA Hash Architecture* annex of the [Privacy Policy](#privacy).

<a id="impact_score"></a>
## What is the impact score of a genotype?

Every finding shows an **impact score** number on a roughly **0–4+** scale, with a small **colored status dot** beside it marking the **risk band**:

- **Green - Good**: a beneficial or protective genotype.
- **Blue - Info**: informational, little or neutral effect.
- **Amber - Warning**: elevated impact score (≈ 2–3).
- **Red - Bad**: high impact score (above 3).

The impact score estimates **how meaningful a given genotype is for you**. It is set during curation by our deep-research pipeline and reflects a combination of signals:

- **Strength and recency of the evidence** - how many independent studies describe the variant, and how recent they are.
- **The variant's predicted biological effect** - whether it changes a protein, sits in a conserved or functionally important position, and how pathogenicity and effect-prediction tools rate it.
- **How constrained the gene is** - variants in genes that rarely tolerate change tend to matter more.
- **How much your genotype departs from the common one** - a typical genotype carries little signal; an unusual one carries more.

As a rule of thumb, read carefully anything scored above ~3, and pay attention to the [genotype frequency](#frequency) - rare variants are often the most worth understanding.

<a id="frequency"></a>
## What is the genotype frequency?

The frequency chip (e.g. `2.1%`) is how common **your exact genotype** is in the general population. A low frequency means a rarer variant, which is often more relevant to you simply because fewer people carry it. Frequency on its own does **not** mean higher risk - it is shown separately from the impact score so the two are never confused.

<a id="minor_allele"></a>
## What are the "minor allele" and MAF?

At most positions two alleles occur in the population. The **minor allele** is the less common one; the **major allele** is the more common. **MAF** (minor allele frequency) is how common that minor allele is across the general population, on a scale of 0 to 0.5 - a low MAF (e.g. `0.02`) means the variant is rare, so carrying it is more distinctive. MAF describes the *allele* across everyone, which is different from the genotype [frequency](#frequency) chip - how common *your exact pair of alleles* is.

<a id="genotype_call"></a>
## What does my genotype, e.g. (A;A), mean?

A genotype is the **two alleles** you carry at a position - one from each parent. `(A;A)` means both copies carry the same allele (**homozygous**); `(A;G)` means the two copies differ (**heterozygous**). A variant's spec sheet lists your genotype next to the reference and risk alleles, so you can see how many risk copies you carry - one of the things that feeds the [impact score](#impact_score).

<a id="variant_groups"></a>
## What is a variant group?

Most findings describe a single SNP. A **variant group** bundles several SNPs that act together - for example a haplotype, or a set of variants in the same gene whose combined pattern is what matters. It is read and scored just like a single genotype, with its own impact score and references, but it summarizes a small set of positions rather than one. Because a variant group has no single public identifier, its card does not link out to external databases. Variant groups appear in search and topic results alongside ordinary genotypes.

<a id="private_snp"></a>
## What is a "Private" SNP?

A few genotypes are marked **Private**. Their identifier starts with `i` (for example `i123456`) instead of `rs…`: these are **proprietary markers used by 23andMe** for variants it genotypes directly, and they have **no public dbSNP rsID**. We still analyze them and show their impact score, gene, and summary - but because there is no public identifier, we can't link out to dbSNP, PubMed, or ClinVar for them, and any mapping to a public variant is uncertain. That is what the **Private** chip on a card signals.

<a id="topic_score"></a>
## What is a topic score, and what does the distribution graph show?

A **topic** is a disease or trait (for example *LDL cholesterol* or *type 2 diabetes*) that many variants contribute to. Its **topic score** is a **polygenic z-score**: we add up the effect of every trait-associated variant you carry - each weighted by how strong and well-replicated its GWAS evidence is - and express the total in **standard deviations (σ)** from the tested-population average. `0σ` is exactly average, `+1.5σ` is well above, `−1.0σ` is below.

The score is **signed by direction**, not by good-or-bad. A **positive** score means your genetics tend to **raise** the trait; a **negative** score means they **lower** it. For a **disease**, raising means *higher inherited risk* and lowering means *lower risk*. For a **measurement** (cholesterol, height, blood pressure), raising or lowering is only a direction - whether it is good or bad depends on the trait, so we do not label it.

Your **percentile** is your rank against the tested population: the share of people who score below you. *84th percentile* means your score is higher than about 84% of the tested cohort. We use the **empirical rank** against real scored genomes, not a theoretical curve.

On a topic page, the **distribution graph** draws the population's score distribution, with a marker at **your** score (e.g. `You · +1.42σ · 84th pct`). The further your marker sits from the center, the more your genetics push the trait in that direction. This is **relative inherited predisposition, not a diagnosis or a probability of disease** - see the [limitations](#topic_limitations).

<a id="established"></a>
## What does the "Established · GWAS" badge mean?

Topics come in two kinds. **Normal** topics gather everything we can reasonably associate with a trait. **Established (GWAS)** topics are stricter: only high-confidence, genome-wide-significant associations contribute, and their [topic score](#topic_score) is a proper **polygenic z-score** built from harmonized, meta-analyzed GWAS effect sizes. The badge marks an established topic; its score is computed from that high-confidence subset only.

<a id="established_associations"></a>
## What are my "established associations"?

The dashboard panels list the [Established (GWAS)](#established) topics your genome contributes to, split by **direction**: **Raises - above average** are topics your genetics push *up* (a positive [z-score](#topic_score)), and **Lowers - below average** are the ones they push *down*. Each row shows the topic and your score in σ, with an arrow for the direction; open one to see the contributing variants and the evidence. For diseases, "raises" means higher inherited risk and "lowers" means lower; for measurements it is simply a direction. Because only replicated GWAS signals feed these lists, they are deliberately conservative - a short, high-confidence place to start.

<a id="topic_limitations"></a>
## How reliable is a topic score? (limitations)

A topic score is a **research-grade signal, not a clinical test**. Read it with these limits in mind:

- **Not a validated risk.** It is built from GWAS Catalog top hits, which are ascertained (their effect sizes are slightly inflated, only partly corrected) and are not genome-wide LD-clumped, so correlated variants can be counted more than once. It is a *relative* signal within our cohort, never a diagnosis or a probability of disease.
- **Ancestry.** Most GWAS effects come from European-ancestry studies, and percentiles are relative to our own tested cohort. Centering removes systematic drift but not which variants carry across ancestries.
- **Coverage.** Many topics rest on just **one** of your variants; these are marked *low confidence*. The more variants contribute, the more stable the score.
- **The extremes are capped.** Scores are clamped to ±5σ, so the very top and bottom of the range can tie.
- **One input among many.** Genetics is one factor; environment and lifestyle matter at least as much.

**Methodology - key references.** Our scoring follows established polygenic-score practice:

1. Choi SW, Mak TS, O'Reilly PF. *A guide to performing Polygenic Risk Score analyses.* Nat Protoc 2020.
2. *Comparison of inverse-variance weighting and the weighted sum of z-scores* for meta-analysis (PMC5287121).
3. *Polygenic risk scores: from research tools to clinical instruments.* Genome Medicine 2020.
4. *Leveraging effect-size distributions to improve polygenic risk scores derived from GWAS summary statistics.* PLOS Comput Biol.

<a id="polygenic_profile"></a>
## What is the "strongest signals" graph on the dashboard?

There is no single overall risk number - inherited risk is the sum of many small-effect variants, so a single score would be misleading. Instead the dashboard shows your **polygenic profile**: the ten places where your DNA deviates most from the tested population, whatever the underlying trait.

Each row plots the typical population range (the 25–75th percentile band) and a bar running from the population median to **you**, tinted to mark whether that signal is protective or elevated. The rows are deliberately unlabeled by condition - read the set as a whole, as a rough polygenic proxy for your overall genetic advantage/disadvantage (how far your genome sits from average), not a verdict on any one disease. The bands come from population score distributions computed per rank.

<a id="raw_data"></a>
## What raw data format is supported, and how do I download it from my provider?

Your file must list, for each SNP, the **rsID, chromosome, position, and genotype**. **Whole-genome (WGS) sequencing files are not supported** - only SNP genotyping (array) data. Upload it as a compressed archive (recommended; `.zip` or `.gz`) or a plain text file (`.txt` or `.csv`). Each SNP is one line; columns may be separated by spaces, tabs, commas, or colons:

```
# rsid       chromosome  position  genotype
rs4477212    1           82154     AA
rs3094315    1           752566    AA
rs3131972    1           752721    GG
```

Here is how to download your raw SNP data from each supported provider:

- **23andMe** - *Profile → Resources → Browse Raw Genotyping Data → Download* (you receive a zipped `.txt`; it is prepared and emailed to you, which can take up to a week). [Instructions ↗](https://customercare.23andme.com/hc/en-us/articles/212196868-Accessing-Your-Raw-Genetic-Data)
- **AncestryDNA** - *Settings → Download or delete → Download DNA Data*; tick the acknowledgment and confirm via the link in the email Ancestry sends. You get a tab-delimited `.txt`. [Instructions ↗](https://support.ancestry.com/s/article/Downloading-DNA-Data)
- **MyHeritage** - *DNA → Manage DNA kits → ⋯ (three dots) → Download*; confirm via the emailed link (valid 24 hours) and re-enter your password. You get a CSV inside a ZIP. Use a computer or Android device (not iPhone/iPad). [Instructions ↗](https://www.myheritage.com/help/en/articles/12851869-how-do-i-download-my-raw-dna-data-file-from-myheritage)
- **FamilyTreeDNA** - *Results & Tools → Autosomal DNA → Download Raw Data*, and choose **"Build 37 Concatenated"**. Two-factor authentication must be enabled. The file is a CSV compressed as `.gz` (GZIP). [Instructions ↗](https://help.familytreedna.com/hc/en-us/articles/14860944283407-Downloading-Your-Family-Finder-Data)
- **LivingDNA** - *Your name → Profiles → select your profile → Download*, tick the consent box, then **Download autosomal data** (a `.txt` file). Only available if you took an actual Living DNA test. [Instructions ↗](https://support.livingdna.com/hc/en-us/articles/360011384960-How-do-I-download-my-raw-data)
- **Genes for Good** - a research study at the University of Michigan. When your results are ready you receive an **email with a download link and a 6-digit access code**; on a computer, enter the code and download the **23andMe-format text file** (not the whole-genome VCF). [Instructions ↗](https://genesforgood.sph.umich.edu/faq/results)
- **WeGene** - the interface is Chinese-only: *个人中心 (Personal Center) → 原始数据 ("Raw Data")* to download a 23andMe-compatible SNP text file. The separate whole-genome (WGS) download is not a SNP file and is not supported here. [Help center ↗](https://www.wegene.com/help)
- **24Genetics** - there is no self-service download; email **[info@24genetics.com](mailto:info@24genetics.com)** to request your raw data (sent free of charge). Choose the SNP array text/CSV file - the whole-genome product is far larger and is not supported. [FAQ ↗](https://24genetics.com/faqs-frequently-asked-questions/)

<a id="orientation"></a>
## Why do some alleles differ between my raw file and the report?

For a few SNPs the allele letters look swapped compared with your raw file. This is **strand orientation**: DNA is double-stranded, and different providers and reference builds sometimes report a variant from opposite strands (for example, a `C/T` on one strand is `G/A` on the other). We canonicalize every genotype to a single consistent orientation so that the same person's data produces the same result regardless of which provider's file they upload. The **Orientation** field in a variant's spec sheet tells you which strand applies.

<a id="accuracy"></a>
## How accurate is the information?

Findings are assembled by frontier AI models from public scientific sources - peer-reviewed literature plus public databases such as **dbSNP**, the **GWAS Catalog**, **ClinVar**, and **OMIM** - and scored against the general population. This is automated processing and **it contains errors**.

The service is for **informational and educational purposes only** and is **not medical advice**. Treat every statement as a starting point that needs independent verification, and always consult a qualified physician for diagnosis and before making any health decision. See our [Terms of Service](#terms) for the full disclosure on AI-generated content.

<a id="incomplete"></a>
## Why is a variant - or my whole report - missing things?

Coverage is deliberately incomplete. Many SNPs have no usable literature, your raw file only contains the positions your provider genotyped, and the pipeline **suppresses findings where the risk of being wrong is too high**. So a missing variant or topic does not mean "no risk" - it means we did not have enough confidence to say anything. Treat the report as a starting point, never as a complete picture.

<a id="print_report"></a>
## Can I print or save my report?

A full report can run to thousands of pages, so there is no single "print everything" button. You can, however, print or save any individual page - a topic, the dashboard, or a search - using your browser's print / **Save as PDF** command (`Ctrl/Cmd + P`).
