Why Sequence a Whole Genome?
We have seen that all the genetic information of an organism lies in the sequence of bases in its DNA. If two people differ in their traits, then somewhere their DNA sequences must differ too. This simple idea sparked an enormous question: could we read out the entire base sequence of a human being?
That is exactly what the Human Genome Project (HGP) set out to do — launched in 1990 and completed in 2003, a 13-year effort. It was rightly called a mega project, because the human genome holds about 3 × 10⁹ base pairs. To picture the scale: at the early cost of about 3 US dollars per base pair, the price tag came to roughly 9 billion dollars, and printing the sequence of a single cell (1000 letters per page, 1000 pages per book) would fill about 3300 books.
Handling that flood of data needed fast computers, which gave birth to a whole new field — Bioinformatics.
Goals of the HGP
The project set itself six clear goals:
- Identify all the genes in human DNA — about 20,000–25,000 of them.
- Determine the sequence of the 3 billion base pairs that make up our DNA.
- Store this information in databases.
- Improve the tools for analysing the data.
- Transfer the related technologies to other sectors, such as industry.
- Address the ethical, legal and social issues (ELSI) raised by the project.
It was coordinated by the U.S. Department of Energy and the National Institutes of Health, with the Wellcome Trust (U.K.) as a major partner and contributions from Japan, France, Germany, China and others.
[NEET Tip] The two numbers examiners love: HGP ran 1990–2003 (13 years), and the genome holds about 3 × 10⁹ bp. The sixth goal, ELSI, is the one students most often forget.
How Was It Done? — Two Approaches
Reading 3 billion bases needed clever strategy. Two complementary approaches were used:
- Expressed Sequence Tags (ESTs): focus only on the genes that are actually expressed as RNA — a shortcut to the functional genes.
- Sequence Annotation: the "blind" approach — sequence the whole genome, coding and non-coding alike, and then work out what each region does.
The practical workflow was: isolate total DNA → break it into small random fragments → clone these in suitable hosts (bacteria and yeast) using special vectors called BAC (bacterial artificial chromosomes) and YAC (yeast artificial chromosomes) → sequence the amplified fragments with automated sequencers based on the method of Frederick Sanger.
The overlapping pieces were then aligned by specialised computer programs and assigned to each chromosome. The very last chromosome to be finished was chromosome 1, completed in May 2006.
Salient Features of the Human Genome
The project handed us some striking facts about ourselves:
| Feature | Value |
|---|---|
| Total genome size | 3164.7 million bp |
| Average gene size | about 3000 bases |
| Largest gene | dystrophin (~2.4 million bases) |
| Total number of genes | about 30,000 (far fewer than expected) |
| Bases identical across all humans | 99.9% |
| Genome that codes for protein | less than 2% |
A few more headline points:
- Functions are still unknown for over 50% of the genes discovered.
- Repeated sequences make up a very large portion of the genome — they have no direct coding role but tell us about chromosome structure and evolution.
- Chromosome 1 has the most genes (2968); the Y chromosome has the fewest (231).
- About 1.4 million sites of single-base differences — SNPs (single nucleotide polymorphisms, said "snips") — were mapped, a goldmine for tracing disease genes and human history.
Applications & Future Challenges
The real payoff of the HGP is not the sequence itself but what we do with it. Knowing how DNA varies between people opens new ways to diagnose, treat and one day prevent thousands of disorders.
Before the HGP, researchers studied one or a few genes at a time. With the full sequence and high-throughput technologies, we can now ask questions on a whole-genome scale — studying every gene in a tissue or tumour, or how thousands of genes and proteins work together in networks.
Many non-human model organisms were also sequenced — bacteria, yeast, Caenorhabditis elegans, Drosophila, rice and Arabidopsis — because comparing genomes teaches us about biology, health care, agriculture and the environment.
The challenge ahead is turning raw sequence into meaning, a task that will occupy scientists for decades.
Memory Capsule — Section 12
- HGP: launched 1990, completed 2003 (13 years); genome ≈ 3 × 10⁹ bp; gave rise to Bioinformatics.
- Goals: identify ~20,000–25,000 genes · sequence the 3 billion bp · build databases · improve tools · transfer technology · address ELSI.
- Methods: ESTs (expressed genes) and Sequence Annotation (whole genome); cloning in BAC/YAC; sequencing by Sanger method.
- Features: 3164.7 million bp · avg gene ~3000 bp · ~30,000 genes · 99.9% bases identical · <2% codes for protein.
- Chromosome 1 = most genes (2968); Y = fewest (231); ~1.4 million SNPs mapped.
Solved Examples — Section 12
Q1. When was the Human Genome Project started and completed, and how long did it run?
Answer: It was launched in 1990 and completed in 2003 — a 13-year project. It was coordinated mainly by the U.S. Department of Energy and the National Institutes of Health.
Q2. Roughly how many genes were estimated in the human genome, and how did this compare with earlier guesses?
Answer: About 30,000 genes — much lower than earlier estimates of 80,000 to 1,40,000. The goals also speak of identifying about 20,000–25,000 genes, so the human gene count turned out surprisingly small.
Q3. Name the two main methodological approaches used in the HGP and state the key difference.
Answer: Expressed Sequence Tags (ESTs) targeted only the genes that are expressed as RNA. Sequence Annotation sequenced the entire genome (coding plus non-coding) and assigned functions afterwards. ESTs focus on genes; annotation maps everything.
Q4. Which chromosome carries the most genes and which carries the fewest?
Answer: Chromosome 1 has the most genes (2968), while the Y chromosome has the fewest (231).
Q5. What fraction of the human genome actually codes for proteins, and what makes up much of the rest?
Answer: Less than 2% of the genome codes for proteins. A very large portion is made of repeated (repetitive) sequences, which have no direct coding role but inform us about chromosome structure and evolution.
Q6. What are SNPs, and why are they valuable?
Answer: SNPs (single nucleotide polymorphisms) are sites where a single DNA base differs between people; about 1.4 million were identified. They help locate disease-associated sequences on chromosomes and trace human evolutionary history.