Where MyGeneLog's Data Comes From: Inside the GWAS Catalog

Where MyGeneLog's Data Comes From: Inside the GWAS Catalog

By MyGeneLog Team · September 5, 2026 · 45 views

Share:

Every variant page on this site ends with a source citation. This post is about what that source actually is, and how MyGeneLog decides what makes it onto the site at all.

The GWAS Catalog

The GWAS Catalog is a free, publicly-funded database run jointly by EMBL-EBI and the NHGRI (part of the U.S. National Institutes of Health). It catalogs findings from genome-wide association studies — large research studies that scan the genome for statistical links between specific variants and specific traits — going back to 2005. It's the standard reference the genomics research community itself uses, not a consumer product.

What has to be true before a variant is added

Not every entry in that catalog makes it onto MyGeneLog. A candidate variant has to clear several bars first: the association has to meet genome-wide statistical significance, the finding has to point to a single, cleanly two-allele position (not a messy multi-allele one), the underlying paper has to be identifiable by its PubMed ID, and the described trait has to be specific enough to actually explain something to a reader — not a vague research-category label. Anything that doesn't clear all of these is simply left out, not published in a rougher form.

Two ways new variants get added

Some variants are added by hand, after a closer look at the specific literature behind them. Others are collected automatically, several times a day, directly from the GWAS Catalog's own data, filtered through the same quality checks described above. Either way, every page cites its actual source — never "MyGeneLog research."

Why this matters

None of this is proprietary science. It's published, peer-reviewed research, made freely available by publicly funded institutions, that we've translated into plain language. The public API exists so other people can build on the same open dataset, the same way we built on the GWAS Catalog's.

Share:

Frequently asked questions

Is the GWAS Catalog free to use?

Yes — it is a publicly funded, open-access database run by EMBL-EBI and the NHGRI (part of the U.S. NIH), free for anyone to use.

Does every GWAS Catalog finding make it onto MyGeneLog?

No. A candidate has to meet genome-wide statistical significance, resolve to a clean two-allele position, cite a real PubMed ID, and describe a specific, meaningful trait — anything that fails any of those checks is left out.

Can I use MyGeneLog's data in my own project?

Yes — see the public API docs. It is free and open, with no signup required.

APOE and Alzheimer's Risk: What the Research Actually Says
← Previous
APOE and Alzheimer's Risk: What the Research Actually Says
Genetics Conferences Worth Knowing About
Next →
Genetics Conferences Worth Knowing About