The researchers asked whether a multi-subunit cellular RNA polymerase could selectively read DNA containing the four synthetic Hachimoji letters P, Z, B and S, write their partners into RNA, and continue transcription after a synthetic base pair was added.
The research question and why it matters
The researchers asked whether a multi-subunit cellular RNA polymerase could selectively read DNA containing the four synthetic Hachimoji letters P, Z, B and S, write their partners into RNA, and continue transcription after a synthetic base pair was added.
Researchers have spent decades making nonstandard base pairs that preserve DNA's overall shape while changing its hydrogen-bonding chemistry. A 2019 Science paper introduced the Hachimoji eight-letter DNA and RNA system and showed information storage, folding and transcription with the simpler single-subunit T7 polymerase. A 2023 study from members of the current team explained how E. coli RNA polymerase handled the six-letter version containing B:S. The new paper adds the remaining P:Z pair, tests all four synthetic letters together at the pairing level and provides structural explanations for both useful recognition and a major Z:G mismatch.
What researchers found
For each synthetic template letter, the matching synthetic RNA building block produced a clear one-letter extension within 15 seconds. Under the reported single-turnover conditions, incorporation rates for the P:Z directions were about twofold lower than for a natural G:C comparison, and chase assays showed that polymerase could continue making a longer RNA after crossing the synthetic site. The original Z base also showed a prominent mismatch with guanine. Raising the base's pKa from about 7.8 to above 10 in the Z* analogue substantially reduced that mismatch while preserving pairing with P. Across four cryo-EM structures, the synthetic pairs adopted familiar Watson-Crick-like geometry inside the enzyme, while different structures captured the polymerase's trigger loop in open and closed states.
Key results from the tested systems
four natural plus four synthetic
The Hachimoji system combines A, T/U, G and C with P, Z, B and S.
template-substrate combinations screened
Each of four synthetic template letters was tested against four natural and four synthetic incoming RNA building blocks.
P:Z incorporation versus natural G:C
The reported single-turnover rates remained in the same broad range under the study's laboratory conditions.
resolution of four cryo-EM structures
The maps showed the synthetic pairs fitting the polymerase active site with Watson-Crick-like geometry.
How the research worked
The team assembled purified E. coli RNA polymerase on short laboratory-made DNA/RNA scaffolds, placing one of four synthetic letters at the next template position. They tested all eight possible incoming RNA triphosphates against each synthetic template base, creating 32 pairing conditions. Radioactively labeled transcription products were separated on denaturing gels to track correct additions and mismatches over time. Single-turnover measurements compared incorporation rates with a natural G:C pair, while chase experiments tested whether the enzyme could move past a synthetic site and keep extending RNA. Because the original Z base could lose a proton and pair incorrectly with guanine, the researchers also tested a modified Z* analogue designed to reduce that error. Finally, single-particle cryo-EM captured four polymerase complexes containing P:Z or P:Z* at 2.42 to 2.75 angstrom resolution, and the atomic models and maps were deposited in public structural databases.
How to interpret this design
The design determines what kind of conclusion the evidence can support. Direct measurement strengthens the reported observation, while generalization beyond the tested subjects, material, place or conditions requires additional evidence.
The evidence comes from a controlled physical or chemical system. That control helps establish what happened under the tested conditions, while scale-up, durability, manufacturing and real-world performance remain separate questions.
What strengthens or limits the finding?
The paper combines orthogonal biochemical and structural methods, tests every pairing of four synthetic template bases with eight possible incoming substrates, reports repeated assays, resolves four deposited cryo-EM structures and makes source data available. The conclusion is nevertheless limited to a reconstituted E. coli transcription system; the researchers did not build a self-sustaining eight-letter cell, test heritable stability or demonstrate full protein production from the expanded alphabet.
The result is meaningfully informative, but identifiable limitations could alter the size, reach or causal interpretation of the finding.
Funding and disclosure context
The recorded funding source is: National Institutes of Health grants R01 GM102362, R01 GM148476 and R01 AI135146; National Science Foundation grants MCB-2419300 and MCB-1939086; cryo-EM data collection at the Stanford-SLAC Cryo-EM Center supported by National Institute of General Medical Sciences grant 1R24GM154186. The recorded conflict information is: The authors declared no competing interests. Funding or a disclosed relationship does not by itself invalidate a result, but it is relevant when judging design choices, analysis and the need for independent replication.
What it means
The experiment closes an important mechanical gap in expanded-genetic-alphabet research. Earlier work showed that eight-letter DNA and RNA can store information and that simpler polymerases can process synthetic pairs. This study shows that the more complex, multi-subunit enzyme used by bacteria can also recognize all four synthetic letters and continue transcription. That could help researchers design chemically richer RNA aptamers, molecular sensors or other engineered nucleic acids. Those applications remain development goals, not products demonstrated by this paper.
Deeper analysis
Transcription is one checkpoint, not an artificial organism
DNA becomes biologically useful through a chain of operations: replication preserves it, transcription makes RNA, and translation can turn selected RNA messages into proteins. This study strengthens the transcription link using a cellular-style bacterial enzyme. It does not assemble the whole chain inside a living cell, so phrases such as 'new life' or 'rewritten organism' would go far beyond the evidence.
Shape compatibility explains why the natural enzyme can cope
The four synthetic letters rearrange hydrogen-bond donors and acceptors while retaining the broad paired shape that polymerases expect. The cryo-EM maps show P:Z and P:Z* occupying the active site in a geometry close to natural pairs and engaging the enzyme's moving trigger loop. The result supports a design principle: molecular machinery can sometimes accept new chemistry when the geometry of its checkpoints is preserved.
The fidelity problem was chemically diagnosable
Z's nitro group lowers its pKa enough that some molecules lose a proton near physiological pH and resemble a partner for guanine. Replacing that group with a carboxamide raised the pKa above 10 and sharply reduced the visible mismatch. That makes the study more than a compatibility demonstration; it connects an error mechanism to a rational chemical correction, although it does not establish cell-scale mutation rates.
Structural detail and system performance answer different questions
High-resolution structures can show why a selected molecular state is plausible, while gels and kinetics show whether copying actually proceeds. Using both makes the case stronger. Neither method, however, substitutes for long-template, whole-cell and multigeneration tests. The next step is integration: proving that the expanded letters remain accurate and useful when all biological constraints operate at once.
What it does NOT prove
- It does not show that a living E. coli cell carried, replicated or inherited an eight-letter genome. The enzyme and short nucleic-acid scaffolds were assembled and tested outside cells.
- It does not demonstrate complete gene expression from eight-letter DNA. Transcription into RNA was tested, but translation of full Hachimoji messages into proteins was not achieved.
- It does not establish error-free copying. The unmodified Z base mispaired with guanine, and other smaller mismatches appeared in forced single-substrate tests.
- It does not show that synthetic letters are safe, stable or controllable in an organism or ecosystem.
- It does not demonstrate a medicine, diagnostic or industrial product. Proposed aptamer and biotechnology uses require further design, cell-based testing and validation.
Important limitations
- The experiments used purified bacterial polymerase and short synthetic scaffolds, leaving out cellular nucleotide transport, metabolism, DNA repair, replication, toxicity and evolutionary pressure.
- Several gel-based transcription assays were performed in two or three independent replicates, a modest evidence base for estimating rare error frequencies.
- The mismatch tests sometimes supplied one substrate without its correct competitor. The authors note that this can exaggerate errors, but the setup does not directly measure fidelity in a complete cellular nucleotide pool.
- The paper reports strong selectivity and relative rates under specific concentrations and reaction conditions; those values may change with sequence context, temperature, nucleotide balance or another polymerase.
- Four cryo-EM structures show compatible geometry at selected stages and sites. They do not capture every step of long, repeated transcription through many synthetic bases.
- Z* reduced prominent mismatches, but the experiments did not establish a genome-scale mutation rate or long-term chemical stability for an eight-letter system.
- The team did not test protein translation from the full eight-letter alphabet or show a new biological function produced by the transcripts.
- The work was peer reviewed and openly accessible, and the authors declared no competing interests, but independent laboratories have not yet replicated this exact eight-letter multi-subunit transcription system.
How this fits with previous research
Researchers have spent decades making nonstandard base pairs that preserve DNA's overall shape while changing its hydrogen-bonding chemistry. A 2019 Science paper introduced the Hachimoji eight-letter DNA and RNA system and showed information storage, folding and transcription with the simpler single-subunit T7 polymerase. A 2023 study from members of the current team explained how E. coli RNA polymerase handled the six-letter version containing B:S. The new paper adds the remaining P:Z pair, tests all four synthetic letters together at the pairing level and provides structural explanations for both useful recognition and a major Z:G mismatch.
Questions still unanswered
- Can living cells import or manufacture all four synthetic nucleotide triphosphates and use them without disrupting ordinary metabolism?
- What are the true error rates when long templates, many synthetic sites and all competing substrates are present together?
- Can replication, transcription and translation be integrated into a stable eight-letter information system inside a cell?
- Will engineered polymerases improve speed and fidelity beyond what natural E. coli RNA polymerase achieved?
- Can expanded RNA molecules deliver reproducible sensing, catalytic or therapeutic functions that four-letter nucleic acids cannot?
- What containment and biosafety controls would be needed before any organism-level application?
Relevant U.S. government resources
These resources serve different purposes. A registry can verify what researchers planned, a repository can locate government-funded work, and an agency page can supply authoritative background. None automatically proves that this paper's conclusion is correct.
OSTI.GOV research search ↗
DOE's research repository is used to locate related national-laboratory reports, accepted manuscripts and funding-linked technical work.
Reuse note: Facts and discoveries are summarized here in original language. We link to government material instead of copying it wholesale, and we do not reuse agency logos, photographs, charts or third-party material unless the specific reuse rights are verified.
A bacterial enzyme transcribed an eight-letter genetic alphabet in the lab
This review was developed from the source record below and, when separately available, the primary paper or government report. The summary and analysis on this page are original editorial writing.
- Source organization
- University of California San Diego
- Source type
- University
- Authors
- Qingrong Li, Hyo-Joong Kim, Yan Liu, Juntaek Oh, Peini Hou, Shuichi Hoshika, Grigore Pintilie, Sriram Aiyer, Jenny Chong, Dmitry Lyumkis, Steven A. Benner and Dong Wang
- Journal / report
- Nature Communications
- Publication date
- September 2, 2026
- DOI
- 10.1038/s41467-026-76668-0
- PMID
- Not available
- Institution
- University of California San Diego-led collaboration with the Foundation for Applied Molecular Evolution, SLAC National Accelerator Laboratory and Stanford University, the Salk Institute for Biological Studies, and Scripps Research
- Funding
- National Institutes of Health grants R01 GM102362, R01 GM148476 and R01 AI135146; National Science Foundation grants MCB-2419300 and MCB-1939086; cryo-EM data collection at the Stanford-SLAC Cryo-EM Center supported by National Institute of General Medical Sciences grant 1R24GM154186
- Conflicts
- The authors declared no competing interests
- Open access
- Yes
- Reuse approach
- Facts summarized in original language from the UC San Diego report, open peer-reviewed paper, structural-data records and prior research; no source wording, figures, tables, photographs, molecular renderings or code reproduced.
AI-assisted editorial process: AI tools helped organize sources and draft this review. The linked research records—not AI output—are the evidence. Publication standards and corrections are publisher-directed. Read our AI transparency policy.