The Genetic Code
The genetic code is the set of rules by which the sequence of nucleotides in mRNA is read as a sequence of amino acids in a protein.
- The code is a triplet code - a set of three nucleotides (a codon) specifies one amino acid. With four bases, there are 4 x 4 x 4 = 64 codons. George Gamow first proposed (on theoretical grounds) that the code should be a triplet; Nirenberg, Matthaei and Har Gobind Khorana deciphered it (Khorana chemically synthesised RNAs of known sequence).
- Of the 64 codons, 61 code for amino acids and 3 are stop (termination) codons - UAA, UAG and UGA. AUG is the start (initiator) codon and also codes for methionine.
Salient features
- Degenerate (redundant) - most amino acids are coded by more than one codon.
- Unambiguous and specific - one codon codes for only one amino acid.
- Nearly universal - a codon means the same amino acid from bacteria to humans (this is what lets bacteria make human proteins).
- Non-overlapping and comma-less - read continuously in a fixed reading frame, three bases at a time.
Mutations and the code
- A point mutation changes a single base pair - as in sickle-cell anaemia (the sixth codon GAG becomes GUG).
- Insertion or deletion of bases (not in multiples of three) causes a frameshift mutation, altering every codon downstream.
One-liners: triplet code, 64 codons; 61 sense + 3 stop (UAA, UAG, UGA); AUG = start = methionine; code is degenerate, unambiguous, nearly universal, non-overlapping; Gamow proposed triplet, Khorana/Nirenberg deciphered.
tRNA and Translation
tRNA - the adaptor molecule
Francis Crick postulated an adaptor molecule that reads the codon on one side and carries the correct amino acid on the other - this is the tRNA. Each tRNA has an anticodon loop (bases complementary to the codon) and an amino-acid acceptor end; it is drawn as a clover-leaf. There is an initiator tRNA, but no tRNA for the stop codons.
Translation
Translation is the polymerisation of amino acids into a polypeptide in the order dictated by the mRNA; the amino acids are joined by peptide bonds.
- Charging (aminoacylation) - each amino acid is activated and joined to its specific tRNA.
- The ribosome is the site of protein synthesis; its two subunits come together on the mRNA. The rRNA of the ribosome is a ribozyme - it catalyses the formation of the peptide bond.
- The mRNA has untranslated regions (UTRs) at both the 5' and 3' ends that are not translated.
- The three steps are initiation (ribosome assembles at the start codon), elongation (aminoacyl-tRNAs add amino acids one by one) and termination (a stop codon releases the finished polypeptide).
One-liners: tRNA = adaptor (Crick); has anticodon + amino-acid end; no tRNA for stop codons; ribosome = site, rRNA is the ribozyme (peptide-bond catalyst); steps = initiation, elongation, termination; amino acids joined by peptide bonds.
Regulation of Gene Expression - the lac Operon
In bacteria, gene expression is regulated mainly at the level of transcription. The lac operon (worked out by Francois Jacob and Jacques Monod) is the model example. An operon is a unit made of a promoter, an operator and a set of structural genes.
- The lac operon has three structural genes: z (codes beta-galactosidase, which hydrolyses lactose into glucose and galactose), y (codes permease, which lets lactose into the cell) and a (codes transacetylase). They are transcribed together as one polycistronic mRNA.
- A separate regulatory gene i produces the repressor protein, which binds the operator and blocks RNA polymerase, keeping the operon switched off.
- Lactose acts as the inducer: when present, it (as allolactose) binds the repressor and inactivates it, freeing the operator so that RNA polymerase can transcribe the z, y and a genes - the operon is switched on. The lac operon is therefore an inducible operon under negative regulation.

One-liners: lac operon = Jacob & Monod; structural genes z (beta-galactosidase), y (permease), a (transacetylase); i gene -> repressor -> binds operator -> OFF; lactose = inducer (inactivates repressor -> ON); inducible, negative regulation, polycistronic.
The Human Genome Project and DNA Fingerprinting
Human Genome Project (HGP)
The HGP (1990-2003) aimed to sequence all the DNA in the human genome. Salient features:
- The human genome has about 3164.7 million (3.3 billion) base pairs and roughly 30,000 genes - far fewer than expected.
- Chromosome 1 has the most genes (about 2968) and the Y chromosome the fewest (about 231).
- Less than 2% of the genome codes for proteins; much of the rest is repetitive DNA.
- Methods included the Expressed Sequence Tags (ESTs) approach (focusing on the expressed genes) and sequence annotation (sequencing the whole genome, then assigning functions). Bioinformatics handles the huge data.
DNA fingerprinting
DNA fingerprinting identifies individuals from the variation in repetitive DNA sequences. It relies on polymorphism in stretches called VNTRs (variable number of tandem repeats), a type of satellite DNA. The technique (developed by Alec Jeffreys; in India by Lalji Singh) involves: isolating DNA -> cutting it with restriction endonucleases -> separating fragments by electrophoresis -> blotting onto a membrane -> hybridising with a labelled VNTR probe -> detecting the bands by autoradiography. It is used in forensics, paternity testing and studies of diversity.
One-liners: HGP - ~3164.7 million bp, ~30,000 genes, chromosome 1 most / Y fewest, <2% codes protein; ESTs = expressed genes; DNA fingerprinting uses VNTR polymorphism (satellite DNA); Jeffreys (India: Lalji Singh).