For over a century, we have been inclined to believe that the genetic alphabet of all known life on Earth has four letters: T, G, A, and C. These four nitrogenous bases are recognized as thymine, guanine, adenine, and cytosine respectively. In order to create base pairs, C pairs with G and T pairs with A. This is the four letter alphabet that has been accepted and utilized since 1893 after German biochemists Albrecht Kossel and Albert Neumann revealed these four bases in nucleic acid molecules (large biological compounds that carry genetic information). However, researchers at the University of San Diego have challenged this precedent. They have demonstrated how a natural enzyme is able to accurately read and copy an eight-letter genetic sequence.
The study focuses on an essential enzyme for transcription called RNA polymerase—specifically from the bacteria known as E. coli. In transcription, which is the first step of gene expression, RNA polymerase reads a DNA sequence containing T, G, A, and C and transcribes it into an RNA sequence with U (uracil), G, A, and C, in which U substitutes T. However, this new discovery shows how RNA polymerase is able to read and transcribe synthetic bases. Scientists manufactured four artificial letters designed to replicate T, G, A, and C, then added them alongside the four natural bases. These new letters are P, Z, B, and S, where P links with Z, and B pairs with S: identical to how the four original nitrogenous bases form pairs. By examining snapshots taken from high resolution cryo-electron microscopes, which can remarkably zoom into a scale smaller than the width of an atom, the researchers at San Diego observed how the E. coli enzyme utilized the same biochemical processes and signals to recognize and transcribe the synthetic bases. Furthermore, in a related study published by Proceedings of the National Academy of Science with the same researchers, scientists concluded that RNA polymerase can spot and identify synthetic base pairs without hydrogen bonds, which are crucial when linking the four natural nitrogenous bases.
So what does this mean for the future of biotechnology in terms of the developing pool of the genetic alphabet? Well, increasing the genetic alphabet artificially leaves plenty of room for error as one letter can simply pair with the wrong counterpart, causing an error in protein synthesis. On the other hand, the large library of genetic letters also broadens what can be created. When in the wrong hands, what gets built with that expanded alphabet can be unpredictably dangerous to an immune system’s limited resilience.
However, it’s possible the advantages of an expanded genetic library will prevail over these setbacks. For instance, the implementation of man-made bases in genetic DNA can help identify liver cancer cells. Artificial letters can be designed to express special structural details that can link specifically to a liver cancer cell’s genetic letters. If pairs form with the cancer cell’s genetic bases and are observed through high-resolution screening, a patient can be diagnosed earlier when it has a higher chance of being cured. Once these cells are located, a small dose of chemotherapy can easily be employed to target the cancerous liver cells and destroy them. Additionally, cancer screening tools are highly expensive, whereas synthetic bases can be rapidly produced in a chemistry lab for significantly cheaper, making manufacturing costs more economical and treatment more affordable for patients. This easy diagnosis of liver cancer can act as a blueprint for other fatal diseases such as Huntington’s disease, Alzheimer’s disease, cystic fibrosis, and many others.





































































































