Genetics & Molecular

2001

Human Genome Project: First Draft Human Genome Sequence

The public consortium's draft, which NHGRI describes as covering about 90 percent of the genome, was made freely available, giving clinicians and researchers a free reference for finding disease genes and reading individual variation.

A DNA sequencing chromatogram
Public domain (Wikimedia Commons)

Key people

Francis Collins
Director of the National Human Genome Research Institute; led the public consortium
Craig Venter
Led Celera Genomics' private sequencing effort and its 2001 Science paper
John Sulston
Cambridge biologist, co-author of the 2001 draft and 2002 Nobel laureate
Eric Lander
Whitehead Institute genome center scientist and first author of the 2001 Nature paper

Source

Nature, 2001 (opens in a new tab)

The Human Genome Project officially began on October 1, 1990, and came to involve 20 universities and research centers in the United States, United Kingdom, France, Germany, Japan and China. Its goal was to sequence the roughly three billion base pairs of the human genome, and it was projected to take 15 years and cost $3 billion. A private rival, Craig Venter's team at Celera Genomics, pursued the sequence with a whole-genome shotgun strategy; its 2001 paper described a 2.91-billion-base-pair consensus of the euchromatic portion of the genome, the part that holds most genes, assembled partly with the public consortium's data.

The International Human Genome Sequencing Consortium, led by Francis Collins as director of the National Human Genome Research Institute, announced on February 12, 2001, that its draft sequence and initial analysis would appear in Nature, where the paper was published on February 15, with Eric Lander of the Whitehead Institute's genome center as first author and John Sulston among the British co-authors. Celera's paper appeared in Science the next day. Under the Bermuda Principles, agreed in 1996, the publicly funded centers placed human sequence in the public domain within 24 hours of generating it. The consortium had announced a working draft at a White House ceremony with President Clinton on June 26, 2000.

The 2001 analysis estimated about 35,000 genes, a figure later revised to about 20,000; Celera counted 26,588 well-supported protein-coding transcripts and about 12,000 more computationally predicted genes with weaker support. Celera also reported that almost half the genes lay scattered in low-GC regions separated by large tracts of apparently noncoding sequence. For clinical genetics, the draft became a reference for positional cloning: when a family linkage study pointed to a chromosomal region, researchers could look up the genes in that interval.

On April 14, 2003, the project announced its completion, more than two years ahead of schedule; the finished sequence covered 92 percent of the genome with fewer than 400 gaps, against more than 150,000 in the working draft. The reference underpinned the dense SNP arrays used in genome-wide association studies of common diseases. Clinical sequencing still depends on aligning a patient's DNA against a reference genome built on the public consortium's work.

The genome project also established data-sharing norms that shaped subsequent large-scale biology initiatives. NHGRI credits two meetings in Bermuda with setting the rules for rapid release of sequence data. Collins directed NHGRI from 1993 to 2008 and later led the NIH for 12 years, the longest tenure of any NIH director. Sulston shared the 2002 Nobel Prize in Physiology or Medicine for mapping a cell lineage in the worm C. elegans and showing that specific cells undergo programmed death during normal development.

Keep exploring

All 526 moments in the history of medicine. This one is in chapter 7, Genes and pandemics