The core idea
The resource records variation observed in a defined set of genomes. Its value comes from reliable measurement and contextual comparison, not from defining a single ideal Indian genome.
1. A dated original research resource
The peer-reviewed IndiGenomes paper by Jain, Bhoyar and colleagues appeared online in Nucleic Acids Research on 23 October 2020 and in the journal's 8 January 2021 issue. Its DOI is 10.1093/nar/gkaa923. The researchers analysed 1,029 Indian volunteers and developed a resource of observed genomic variation. The online date and issue date refer to the same paper, not two separate discoveries.
A database paper is original research when collecting, processing and validating data creates a new basis for investigation. Here the contribution is not a claim that one sequence represents everyone. It gives other researchers a comparison set, documented methods and frequency information. A resource becomes scientifically useful when another team can understand what was measured and what population the measurements describe.
Sources: IndiGenomes original paper: Nucleic Acids Research, 2020/2021 (full PDF) ↗ · CSIR-IGIB: IndiGenomes reference resource ↗
2. A variant is a difference, not a verdict
The four DNA bases are commonly represented by A, C, G and T. Sequencing determines their order. A reference genome supplies a coordinate system for comparison, rather like page and line numbers in a shared edition of a book. A variant is a difference relative to the stated reference. The reference version is not automatically healthier, more original or more desirable.
An allele is one version at a genomic location. In a simple autosomal (non-sex-chromosome) example, a person carries two copies, one inherited from each parent. Two different versions make the genotype heterozygous; matching versions make it homozygous at that location. These terms describe sequence relationships. They do not by themselves identify disease, ability, character or membership of a social group.
Sources: NHGRI: human genomic variation ↗ · NHGRI glossary: allele ↗
3. From short reads to a variant catalogue
Sequencing often produces many short reads rather than one uninterrupted chromosome. Software aligns reads with a reference and examines disagreements. It must distinguish genuine biological differences from sequencing errors, poor-quality bases and ambiguous alignment in repeated regions. Variant calling is the inference step that identifies likely differences; annotation adds information such as location within or near a gene.
IndiGenomes used whole-genome sequencing at approximately 30-fold average coverage. Coverage here concerns repeated sequence observations, not the number of volunteers. An average also hides unevenness: some positions receive many observations and others few. Quality filters and reproducible processing therefore matter even when the overall read count is large. An annotated variant remains a candidate for interpretation, not an experimentally proven biological effect.
Sources: IndiGenomes original paper: Nucleic Acids Research, 2020/2021 (full PDF) ↗ · NHGRI: sequencing genomes, Adam Phillippy colloquium, 2024 ↗ · NHGRI: human genomic variation ↗
4. Worked example: copies are not people
Take an illustrative autosomal location in ten people with complete usable data. There are 20 allele copies. Suppose two people each have one copy of variant V and one person has two copies. The allele count is 1 + 1 + 2 = 4. Allele frequency is 4 ÷ 20 = 20%. The fraction of people carrying at least one copy is instead 3 ÷ 10 = 30%.
Now suppose one of the seven people without V has no reliable measurement at this location. The counted copies fall to 18, so the observed frequency becomes 4 ÷ 18, about 22.2%. We have not discovered another V copy; the denominator changed. Database fields often separate allele count from the number of successfully assessed copies for exactly this reason. Missing data should never silently become a reference genotype.
Sources: NHGRI glossary: allele ↗ · CSIR-IGIB: IndiGenomes reference resource ↗
5. Worked example: compare overlapping catalogues
The paper reports 55,898,122 variants. Of these, 37,249,254 were also found in the compared gnomAD set and 21,485,966 in the compared 1000 Genomes set; 20,853,355 were shared across all three. Adding the first two overlaps counts the shared portion twice. The number seen in either comparison is therefore 37,249,254 + 21,485,966 − 20,853,355 = 37,881,865.
Subtracting from the catalogue total gives 18,016,257 absent from both comparison sets, about 32.23% of the total. This is a statement about specified database versions and methods. It does not mean these variants occur only in Indians, occur in every Indian, or were newly created in these volunteers. New to a comparison catalogue and unique to a population are very different propositions.
An overlap is counted once
| Within the IndiGenomes catalogue | Variants |
|---|---|
| Also in the compared gnomAD set | 37,249,254 |
| Also in the compared 1000 Genomes set | 21,485,966 |
| Shared across all three; subtract once | 20,853,355 |
| In either comparison set | 37,881,865 |
| Absent from both compared sets | 18,016,257 |
Sources: IndiGenomes original paper: Nucleic Acids Research, 2020/2021 (full PDF) ↗
6. What a reference sample can represent
The volunteers were selected to include geographical diversity and described as self-declared healthy. That description does not make the sample a random census of India or establish lifelong absence of illness. A sample can be valuable while still underrepresenting particular places or histories. A frequency should therefore travel with information about sample size, selection, measurement quality and the relevant comparison group.
Imagine a variant appears once among 200 assessed copies. Its observed frequency is 0.5%. Finding no copy in a second sample of 20 copies does not establish absence from that population. Small samples can miss uncommon variants. Increasing sample size improves opportunities to observe variation, while improving recruitment diversity addresses a different problem: who is included in the first place.
Sources: IndiGenomes original paper: Nucleic Acids Research, 2020/2021 (full PDF) ↗ · NHGRI: human genomic variation ↗
7. A reference resource needs responsible use
Genomic information can also concern relatives. Informed consent should explain collection, future use, data sharing and privacy protections in language participants understand. Public summary frequencies and an individual's genomic record have different privacy implications. Scientific openness therefore requires thoughtful governance, not unrestricted copying of every person's information. Learning the mathematics does not require gathering personal genetic data.
A useful reference can help researchers question an apparent rarity, prioritise follow-up or identify gaps in knowledge. Frequency alone cannot diagnose a person or prove that a variant causes a trait. That requires additional clinical, functional and family evidence appropriate to the question. The achievement is an improved foundation for research, accompanied by the responsibility to interpret it carefully.
Sources: NHGRI: the informed consent process ↗ · NHGRI: human genomic variation ↗ · CSIR-IGIB: IndiGenomes reference resource ↗
PUT IT INTO PRACTICE
Practice: build a fictional variant record
- Create a ten-person fictional table with genotypes AA, AV or VV at one autosomal location. Use invented labels, not real people.
- Count V copies and successfully observed copies separately. Calculate allele frequency and the proportion of carriers.
- For catalogue overlaps of 70 and 50 variants with 30 shared, calculate the union. If your catalogue contains 120 variants, calculate how many are absent from both comparisons.
- Add a short note naming the reference, sample and missing-data rule. Explain one conclusion your table cannot support.
Check your understanding
What is the answer to the catalogue exercise?
The union is 70 + 50 − 30 = 90. Of 120 variants, 30 are absent from both comparisons. The overlap must be subtracted once.
Why is allele frequency different from the percentage of carriers?
An individual can contribute one or two copies of a variant. Allele frequency counts copies; carrier proportion counts people.
Does 30-fold coverage mean 30 people were sequenced?
No. It describes average repeated sequence coverage, while the number of participants is a separate quantity.
Does a missing database entry prove a variant is unique to India?
No. Absence may reflect sampling, versions or detection methods. The claim must remain limited to the compared datasets.
Why is the reference allele not a definition of normal health?
A reference is a comparison convention. Biological significance requires evidence beyond whether a sequence matches that convention.
What does informed consent contribute to data quality and trust?
It clarifies what participants understand and authorise, supports voluntary participation and establishes appropriate uses of sensitive information.
