Reference
Where the Data Lives
Twelve public databases and tools that underpin modern genetics research. This site draws on most of them directly — from variant registration to population frequencies to clinical interpretation to literature discovery — plus a few more worth knowing about for your own digging.
Every rsID, allele frequency, and clinical significance rating you see on this site originates in one of these databases. They are free, continuously maintained by research institutions and government bodies, and together form the backbone of how modern genetics is read and published.
A free literature-mapping tool that builds citation networks around a paper or search term — surfacing related, similar, and citing work far beyond what a plain keyword search returns. Used here to widen the study search around a variant beyond what PubMed and Semantic Scholar turn up directly.
The authoritative registry for every known SNP. Run by the National Center for Biotechnology Information, it assigns each variant its rsID — the universal naming system used across all genetics research globally. The source of allele frequencies, consequence annotations (synonymous, missense, stop-gained), and chromosomal coordinates.
A community-maintained wiki of clinical SNP interpretations. Each rsID page carries a magnitude score (0–10) summarising the strength of associated evidence, links to supporting studies, and plain-language summaries of what the variant is thought to do. Less formal than ClinVar, but often more readable — and frequently updated faster.
A gene-centric portal from the Weizmann Institute of Science. Each card aggregates data from over 150 sources into one page — known variants, protein structure, expression patterns by tissue, disease associations, drug interactions, and pathway memberships. A good first stop when you want the full picture around a specific gene rather than a specific variant.
The Genome Aggregation Database, produced by the Broad Institute. Contains whole-genome and whole-exome sequencing data from over 125,000 individuals across diverse ancestry groups — making it the gold standard for population-level allele frequency data. Particularly valuable for understanding how rare or common a variant is in specific ancestral populations, and whether it's ever been observed in a healthy cohort.
Online Mendelian Inheritance in Man — the authoritative catalogue of human genes and genetic disorders, continuously updated since 1966. Each entry links a gene or locus to its known phenotypic effects, with evidence classifications and references to the original discovery literature. The bedrock resource for understanding gene–disease relationships.
NCBI's database of clinically interpreted variants, aggregating submissions from diagnostic laboratories, research groups, and clinicians worldwide. Each variant receives a clinical significance rating — pathogenic, likely pathogenic, benign, likely benign, or uncertain — based on the cumulative weight of evidence submitted. Essential for understanding whether a specific variant has a formally recognised disease connection.
A general academic search engine, not genetics-specific — but still one of the widest nets you can cast over the published literature for a variant or gene. Not used programmatically on this site (it has no public API and rate-limits automated access hard), but still worth knowing about for manual literature digging.
Leiden Open Variation Database — a federated network of gene-specific variant databases, each independently curated by a research group or diagnostic lab specialising in that gene. This shared instance is the index of which genes have their own LOVD database and where to find it.
A variant search and annotation engine that aggregates dbSNP, ClinVar, gnomAD, and dozens of other sources into a single lookup per variant, plus ACMG classification guidance. Useful as a one-stop cross-reference once you already have an rsID from elsewhere on this site.
EMBL-EBI's catalogue of published genome-wide association studies — every SNP that's been statistically linked to a trait or disease across a study population, with the effect size and the paper it came from. The place to check whether a variant has ever turned up in a large population study, beyond single-gene reports.
Run by the Medical College of Wisconsin, RGD is built around the rat as a model organism but every gene page cross-links its orthologs across species — human, mouse, zebrafish, and more — via the RGD Orthologs and Alliance Orthologs sections. Also supports keyword search across genes, variants, QTLs, and phenotypes, plus curated disease portals. Note: orthology data is gene-level only — it shows which species carry the equivalent gene, not whether a specific variant is conserved across them.