Genetics & Molecular

2003

Human Genome Project Completion

On 14 April 2003 the public consortium declared the human genome essentially finished, 50 years after Watson and Crick described the double helix. The 2004 Nature paper reported 2.85 billion nucleotides covering about 99% of the euchromatic genome, with about one error per 100,000 bases.

Logo of the Human Genome Project
U.S. Department of Energy, Human Genome Project / Public domain (Wikimedia Commons)

Key people

Francis Collins
Director, National Human Genome Research Institute; led U.S. effort
John Sulston
Director of the Sanger Centre at Hinxton from 1992 to 2000; led adoption of the Bermuda Principles with Robert Waterston
Craig Venter
Led Celera Genomics competing private sequencing effort
Eric Lander
Head of the Whitehead Institute/MIT Center for Genome Research; first author of the 2001 draft sequence paper

Source

Nature. 2004;431(7011):931-945. (opens in a new tab)

When James Watson and Francis Crick described the double helix in April 1953, the idea of reading the entire sequence of three billion base pairs was a thought experiment. Fifty years later, almost to the month, it had been done. On 14 April 2003 the International Human Genome Sequencing Consortium announced that the human genome was essentially complete, two years ahead of the schedule set when the project formally launched in 1990.

The publicly funded consortium coordinated sequencing centers in six countries: the United States, the United Kingdom, France, Germany, Japan, and China. Francis Collins, director of the National Human Genome Research Institute in Bethesda, was the consortium's de facto leader, and John Sulston directed the Sanger Centre at Hinxton, near Cambridge, from its founding in 1992 until 2000. Craig Venter's company Celera Genomics ran a competing private effort using whole-genome shotgun sequencing, a method the public consortium criticized on technical grounds, and both groups published draft sequences in February 2001. The finished sequence published in Nature in 2004 contained 2.85 billion nucleotides with only 341 gaps, covered about 99% of the euchromatic genome, and had an error rate of about one per 100,000 bases.

The decision to release data publicly and immediately, formalized in the Bermuda Principles agreed at a meeting in Bermuda in February 1996, distinguished the consortium's approach from Celera's and set a precedent for genomic data sharing. The principles called for automatic, rapid release of sequence assemblies of 2,000 bases or more, and the participants pledged not to seek patents, so researchers could use the sequence long before it was published. The project cost $2.7 billion in 1991 dollars, less than the $3 billion Congress had been told to expect.

The reference sequence made genome-wide association studies possible. By the late 2000s, researchers had used the reference genome to identify hundreds of loci associated with common diseases, including type 2 diabetes, coronary artery disease, inflammatory bowel disease, and multiple cancers. Turning those statistical associations into mechanisms and drug targets proved much slower, but the catalog of variants became a standard tool in human genetics.

In clinical medicine, the completed genome enabled the development of targeted sequencing panels for inherited cancer risk, pharmacogenomic testing for drug metabolism variants, and eventually the whole-genome sequencing pipelines used in neonatal intensive care and rare disease diagnosis. The 2003 sequence still left about 8% of the genome unread, and many of the remaining gaps lay in duplicated segments; on March 31, 2022, the Telomere-to-Telomere consortium announced that it had filled the remaining gaps.

Keep exploring

All 526 moments in the history of medicine. This one is in chapter 7, Genes and pandemics