Compositions and methods for the in vivo synthesis of non-natural polypeptides
In vivo synthesis of non-natural polypeptides using unnatural DNA and tRNA pairs addresses limitations of existing methods by enabling the incorporation of multiple unnatural amino acids, improving enzymatic activity and solubility.
Patent Information
- Application Number
- JP2025175088
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-03-12
- Filing Date
- 2025-10-17
- Publication Date
- 2026-03-05
AI Technical Summary
Current methods for synthesizing unnatural proteins or polypeptides are limited by the inability to introduce multiple unnatural amino acids and often result in reduced enzymatic activity, solubility, or yield, and lack a suitable post-translational modification environment.
In vivo synthesis of non-natural polypeptides using unnatural DNA molecules encoding unnatural base pairs, mRNA molecules with non-canonical codons, and tRNA molecules with complementary anticodons to incorporate two or more unnatural amino acids into the polypeptide.
Facilitates the site-specific incorporation of multiple unnatural amino acids into polypeptides, enhancing enzymatic activity and solubility, and providing a complete post-translational modification environment.
Smart Images

Figure 2026036689000031 
Figure 2026036689000032 
Figure 2026036689000033
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 62 / 913,664, filed October 10, 2019, and U.S. Provisional Application No. 62 / 988,882, filed March 12, 2020, each of which is incorporated by reference in its entirety.
[0002] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy was created on October 6, 2010, is titled "36271-809_601_SL.txt", and is 21 kilobytes in size.
[0003] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH This invention was made with U.S. government support under Grant No. GM118178 awarded by the National Institutes of Health. The government has certain rights in this invention.
[0004] INCORPORATION BY REFERENCE All publications, patents, and patent applications mentioned herein are incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification supersedes and / or takes precedence over any such conflicting material. [Background technology]
[0005] The natural genetic code consists of 64 codons, enabled by four letters of the genetic alphabet. Three codons are used as stop codons, and the remaining 61 sense codons are recognized by transfer RNAs (tRNAs) that are charged by cognate aminoacyl-tRNA synthetases (also referred to herein simply as tRNA synthetases) carrying one of the 20 amino acids that give rise to proteins. While canonical amino acids enable significant biological diversity, there are many chemical functions and associated reactivities that they do not provide. The ability to expand the genetic code to include unnatural or non-canonical amino acids (ncAAs) could confer desired functions or activities to proteins and dramatically facilitate many known and emerging applications of proteins, such as therapeutic drug development. Current methods for synthesizing unnatural proteins or polypeptides containing unnatural amino acids have limitations. Notably, most methods only allow for the introduction of a single unnatural amino acid or several copies of a single unnatural amino acid into an unnatural polypeptide. Similarly, non-natural polypeptides synthesized by currently available methods often have reduced enzymatic activity, solubility, or yield. Summary of the Invention [Problem to be solved by the invention]
[0006] One alternative solution to address these limitations is to synthesize non-natural polypeptides in cell-free or in vitro expression systems. However, such expression systems are inadequate to provide a post-translational modification environment in which the redox properties of the non-natural polypeptide and other post-translational modifications of the synthesized non-natural polypeptide can be fully realized. Thus, there remains a need for compositions and methods for the in vivo synthesis of non-natural polypeptides containing non-natural amino acids. [Means for solving the problem]
[0007] Described herein are compositions, methods, cells (non-engineered and engineered), semisynthetic organisms (SSOs), reagents, genetic material, plasmids, and kits for the in vivo synthesis of non-naturally occurring polypeptides or proteins, wherein each non-naturally occurring polypeptide or protein comprises two or more non-naturally occurring amino acids that are decoded by the cell.
[0008]
[0003] Described herein is an in vivo method for synthesizing an unnatural polypeptide, the method comprising: providing at least one unnatural deoxyribonucleic acid (DNA) molecule comprising at least four unnatural base pairs; transcribing the at least one unnatural DNA molecule to provide a messenger RNA (mRNA) molecule comprising at least two unnatural codons; transcribing the at least one unnatural DNA molecule to provide at least two transfer RNA (tRNA) molecules each comprising at least one unnatural anticodon, wherein the at least two unnatural base pairs in the corresponding DNAs are in a sequence context such that the unnatural codon of the mRNA molecule is complementary to each unnatural anticodon of the tRNA molecule; and synthesizing the unnatural polypeptide by translating the unnatural mRNA molecule utilizing at least two unnatural tRNA molecules, wherein each unnatural anticodon directs the site-specific incorporation of an unnatural amino acid into the unnatural polypeptide. In some embodiments, the at least two unnatural base pairs comprise a base pair selected from dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1, or dNaM-dTAT1.
[0009] In some embodiments, a method of synthesizing an unnatural polypeptide includes: providing at least one unnatural deoxyribonucleic acid (DNA) molecule comprising at least four unnatural base pairs, wherein the at least one unnatural DNA molecule encodes (i) a messenger RNA (mRNA) molecule comprising at least a first and a second unnatural codon, and (ii) at least a first and a second transfer RNA (tRNA) molecule, wherein the first tRNA molecule comprises a first unnatural anticodon and the second tRNA molecule comprises a second unnatural anticodon, and wherein the at least four unnatural base pairs in the at least one DNA molecule encode a first unnatural codon and a second unnatural codon in the mRNA molecule. and a second unnatural codon in a sequence context such that they are complementary to the first and second unnatural anticodons, respectively; transcribing at least one unnatural DNA molecule to provide an mRNA; transcribing the at least one unnatural DNA molecule to provide at least first and second tRNA molecules; and synthesizing an unnatural polypeptide by translating the unnatural mRNA molecule utilizing the at least first and second unnatural tRNA molecules, wherein each of the at least first and second unnatural anticodons directs the site-specific incorporation of an unnatural amino acid into the unnatural polypeptide.
[0010] In some embodiments, the methods include at least two unnatural codons, each comprising a first unnatural nucleotide located at the first, second, or third position of the codon, and optionally the first unnatural nucleotide located at the second or third position of the codon. In some examples, the methods include at least two unnatural codons, each comprising the nucleic acid sequence NNX or NXN, and the unnatural anticodon comprising the nucleic acid sequence XNN, YNN, NXN, or NYN, thereby forming unnatural codon-anticodon pairs comprising NNX-XNN, NNX-YNN, or NXN-NYN, where N is any naturally occurring nucleotide, X is a first unnatural nucleotide, and Y is a second unnatural nucleotide different from the first unnatural nucleotide, and XY forms an unnatural base pair (UBP) in DNA.
[0011] In some embodiments, a UBP is formed between the codon sequence of an mRNA and the anticodon sequence of a tRNA to facilitate translation of the mRNA into a non-naturally occurring polypeptide. The tRNA-anticodon UBP, in some instances, comprises a codon sequence (e.g., UUX) comprising three consecutive nucleic acids read from 5' to 3' of the mRNA, and an anticodon sequence (e.g., YAA or XAA) comprising three consecutive nucleic acids read from 5' to 3' of the tRNA. In some embodiments, when the mRNA codon is UUX, the tRNA anticodon is YAA or XAA. In some embodiments, when the mRNA codon is UGX, the tRNA anticodon is YCA or XCA. In some embodiments, when the mRNA codon is CGX, the tRNA anticodon is YCG or XCG. In some embodiments, when the mRNA codon is AGX, the tRNA anticodon is YCU or XCU. In some embodiments, when the mRNA codon is GAX, the tRNA anticodon is YUC or XUC. In some embodiments, when the mRNA codon is CAX, the tRNA anticodon is YUG or XUG. In some embodiments, when the mRNA codon is GXU, the tRNA anticodon is AYC. In some embodiments, when the mRNA codon is CXU, the tRNA anticodon is AYG. In some embodiments, when the mRNA codon is GXG, the tRNA anticodon is CYC. In some embodiments, when the mRNA codon is AXG, the tRNA anticodon is CYU. In some embodiments, when the mRNA codon is GXC, the tRNA anticodon is GYC. In some embodiments, when the mRNA codon is AXC, the tRNA anticodon is GYU. In some embodiments, when the mRNA codon is GXA, the tRNA anticodon is UYC. In some embodiments, when the mRNA codon is CXC, the tRNA anticodon is GYG. In some embodiments, when the mRNA codon is UXC, the tRNA anticodon is GYA. In some embodiments, when the mRNA codon is AUX, the tRNA anticodon is YAU or XAU. In some embodiments, when the mRNA codon is CUX, the tRNA anticodon is XAG or YAG.In some embodiments, when the mRNA codon is UUX, the tRNA anticodon is XAA or YAA. In some embodiments, when the mRNA codon is GUX, the tRNA anticodon is XAC or YAC. In some embodiments, when the mRNA codon is UAX, the tRNA anticodon is XUA or YUA. In some embodiments, when the mRNA codon is GGX, the tRNA anticodon is XCC or YCC.
[0012] In some embodiments, at least one unnatural DNA molecule is transcribed into messenger RNA (mRNA) containing an unnatural base described herein (e.g., d5SICS, dNaM, dTPT3, dMTMO, dCNMO, dTAT1). Exemplary mRNA codons are encoded by exemplary regions of unnatural DNA containing three consecutive deoxyribonucleotides (NNN), including TTX, TGX, CGX, AGX, GAX, CAX, GXT, CXT, GXG, AXG, GXC, AXC, GXA, CXC, TXC, ATX, CTX, TTX, GTX, TAX, or GGX, where X is an unnatural base attached to a 2' deoxyribosyl moiety. Exemplary mRNA codons resulting from transcription of exemplary unnatural DNAs include three consecutive ribonucleotides (NNN) each containing UUX, UGX, CGX, AGX, GAX, CAX, GXU, CXU, GXG, AXG, GXC, AXC, GXA, CXC, UXC, AUX, CUX, UUX, GUX, UAX, or GGX, where X is an unnatural base attached to a ribosyl moiety. In some embodiments, the unnatural base is in the first position of the codon sequence (XNN). In some embodiments, the unnatural base is in the second (or middle) position of the codon sequence (NXN). In some embodiments, the unnatural base is in the third (last) position of the codon sequence (NNX).
[0013] In some embodiments, the method involves a codon comprising at least one G and an anticodon comprising at least one C. In some examples, the method involves a codon comprising X and Y. wherein X and Y are independently selected from the group consisting of: (i) 2-thiouracil, 2'-deoxyuridine, 4-thiouracil, uracil-5-yl, hypoxanthine-9-yl(I), 5-halouracil; 5-propynyl-uracil, 6-azo-uracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, uracil-5-oxaacetic acid methyl ester, uracil-5-oxaacetic acid, 5-methyl (ii) 5-hydroxyuracil, 5-hydroxy-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, uracil-5-oxaacetic acid, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil or dihydrouracil; hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytosine arabinoside, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azocytosine, azacytosine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1,4]benzoxazine-2(3H) -one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one), or pyridoindole cytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one);(iii) 2-aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2'-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted adenines, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-azaadenine; (iv) 2-methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaadenine, 7-methyladenine, 7-deazaadenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted adenines, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-azaadenine; (v) hypoxanthine, xanthine, 1-methylinosine, queosine, beta-D-galactosylqueosine, inosine, beta-D-mannosylqueosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine, or 2-pyridone. In some embodiments, X and Y are independently selected from the group consisting of: [ka] In some examples, X is [ka] is.
[0014] In some embodiments, Y is [ka] is.
[0015] In some embodiments, the methods described herein comprise an unnatural codon-anticodon pair NNX-XNN, where NNX-XNN is selected from the group consisting of UUX-XAA, UGX-XCA, CGX-XCG, AGX-XCU, GAX-XUC, CAX-XUG, AUX-XAU, CUX-XAG, GUX-XAC, UAX-XUA, and GGX-XCC. In some embodiments, the methods described herein comprise an unnatural codon-anticodon pair NNX-YNN, where NNX-YNN is selected from the group consisting of UUX-YAA, UGX-YCA, CGX-YCG, AGX-YCU, GAX-YUC, CAX-YUG, AUX-YAU, CUX-YAG, GUX-YAC, UAX-YUA, and GGX-YCC. In some examples, the methods described herein involve the unnatural codon-anticodon pair NXN-NYN, where NXN-NYN is GXU- The anticodons are selected from the group consisting of AYC, CXU-AYG, GXG-CYC, AXG-CYU, GXC-GYC, AXC-GYU, GXA-UYC, CXC-GYG, and UXC-GYA. In some embodiments, the methods described herein include at least two unnatural tRNA molecules, each containing a different unnatural anticodon. In some examples, the at least two unnatural tRNA molecules include a pyrrolysyl-tRNA from Methanosarcina and a tyrosyl-tRNA from Methanocaldococcus jannaschii, or a derivative thereof. In some embodiments, the methods include charging the at least two unnatural tRNA molecules with an aminoacyl-tRNA synthetase. In some examples, the tRNA synthetase is selected from the group consisting of chimeric PylRS (chPylRS) and M. jannaschii AzFRS (MjpAzFRS). In some embodiments, the methods described herein include charging the at least two unnatural tRNA molecules with at least two different tRNA synthetases. In some examples, the at least two different tRNA synthetases include a chimeric PylRS (chPylRS) and a M. jannaschii AzFRS (MjpAzFRS).
[0016] In some embodiments, methods for in vivo synthesis of non-natural polypeptides are described herein. In some embodiments, the non-natural polypeptides comprise two, three, or more non-natural amino acids. In some instances, the non-natural polypeptides comprise at least two non-natural amino acids that are the same. In some embodiments, the non-natural polypeptides comprise at least two different non-natural amino acids. In some instances, the non-natural amino acids are: lysine analogs; aromatic side chains; azide groups; alkyne groups; or aldehyde or ketone groups. In some examples, the unnatural amino acid does not contain an aromatic side chain. In some embodiments, the unnatural amino acid is N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propargylethoxy-carbonyl-L-lysine (PraK), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azidophenylalanine (pAzF), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-Amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L -phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)- 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine.
[0017] In some embodiments, the methods for in vivo synthesis of non-naturally occurring polypeptides described herein include at least one non-naturally occurring DNA molecule in the form of a plasmid. In some instances, the at least one non-naturally occurring DNA molecule is integrated into the genome of a cell. In some embodiments, the at least one non-naturally occurring DNA molecule encodes a non-naturally occurring polypeptide. In some embodiments, the methods described herein involve in vivo replication and transcription of the non-natural DNA molecule and in vivo translation of the transcribed mRNA molecule in a cellular organism. In some embodiments, the cellular organism is a microorganism. In some embodiments, the cellular organism is a prokaryote. In some embodiments, the cellular organism is a bacterium. In some instances, the cellular organism is a gram-positive bacterium. In some embodiments, the cellular organism is a gram-negative bacterium. In some instances, the cellular organism is Escherichia coli. In some embodiments, the cellular organism comprises a nucleoside triphosphate transporter. In some instances, the nucleoside triphosphate transporter comprises the amino acid sequence of PtNTT2. In some embodiments, the nucleoside triphosphate transporter comprises a truncated amino acid sequence of PtNTT2. In some alternatives, the truncated amino acid sequence of PtNTT2 is at least 80% identical to PtNTT2 encoded by SEQ ID NO:1. In some embodiments, the cellular organism comprises at least one non-natural DNA molecule. In some embodiments, the at least one non-natural DNA molecule comprises at least one plasmid. In some embodiments, the at least one non-natural DNA molecule is integrated into the genome of the cell. In some examples, the at least one non-natural DNA molecule encodes a non-natural polypeptide. In some examples, the methods described in the present disclosure may be in vitro methods that include synthesizing a non-natural polypeptide in a cell-free system.
[0018] In some embodiments, methods for in vivo synthesis of unnatural polypeptides are described herein, wherein the unnatural polypeptide comprises an unnatural sugar moiety. In some embodiments, the unnatural base pair comprises at least one unnatural nucleotide comprising an unnatural sugar moiety. In some embodiments, the unnatural sugar moiety is selected from the group consisting of OH, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NHF; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n -NH2 and -O(CH2) n ON[(CH2) n CH3)]2, where n and m are from 1 to about 10); and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S); modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving group, reporter group, intercalator, group for improving the pharmacokinetic properties of an oligonucleotide, or group for improving the pharmacodynamic properties of an oligonucleotide, and any combination thereof.
[0019] In some embodiments, described herein are cells for the in vivo synthesis of unnatural polypeptides, the cells comprising at least two different unnatural codon-anticodon pairs, wherein each unnatural codon-anticodon pair comprises an unnatural codon from a unnatural messenger RNA (mRNA) and an unnatural anticodon from a unnatural transfer ribonucleic acid (tRNA), the unnatural codon comprising a first unnatural nucleotide and the unnatural anticodon comprising a second unnatural nucleotide; and at least two different unnatural amino acids are each covalently linked to a corresponding unnatural tRNA. In some examples, the cells further comprise at least one unnatural DNA molecule comprising at least four unnatural base pairs (UBPs). In some embodiments, described herein are cells for the in vivo synthesis of unnatural polypeptides, the cells comprising at least one unnatural DNA molecule comprising at least four unnatural base pairs, wherein the at least one unnatural DNA molecule comprises: (i) a messenger RNA (mRNA) encoding the unnatural polypeptide and comprising at least a first and a second unnatural codon; The cell further comprises an mRNA molecule and at least the first and second tRNA molecules. In some embodiments, the cell encodes an mRNA (RNA) molecule, and (ii) at least first and second transfer RNA (tRNA) molecules, wherein the first tRNA molecule comprises a first unnatural anticodon and the second tRNA molecule comprises a second unnatural anticodon, and wherein at least four unnatural base pairs in the at least one DNA molecule are in a sequence context such that the first and second unnatural codons of the mRNA molecule are complementary to the first and second unnatural anticodons, respectively. In some embodiments, the cell further comprises an mRNA molecule and at least the first and second tRNA molecules. In some embodiments of the cell, at least the first and second tRNA molecules are covalently linked to an unnatural amino acid. In some embodiments, the cell further comprises an unnatural polypeptide.
[0020] In some embodiments, the first non-natural nucleotide is located at the second or third position of the non-natural codon and complementarily base pairs with the second non-natural nucleotide of the non-natural anticodon. In some examples, the first non-natural nucleotide and the second non-natural nucleotide are [ka] and optionally, the second base is different from the first base. In some embodiments, the cell further comprises at least one unnatural DNA molecule comprising at least four unnatural base pairs (UBPs). In some examples, the at least four unnatural base pairs are independently selected from the group consisting of dCNMO / dTPT3, dNaM / dTPT3, dCNMO / dTAT1, or dNaM / dTAT1. In some examples, the at least one unnatural DNA molecule comprises at least one plasmid. In some embodiments, the at least one unnatural DNA molecule is integrated into the genome of the cell. In some embodiments, the at least one unnatural DNA molecule encodes a non-natural polypeptide. In some embodiments, the cell described herein expresses a nucleoside triphosphate transporter. In some alternatives, the nucleoside triphosphate transporter comprises the amino acid sequence of PtNTT2. In some examples, the nucleoside triphosphate transporter comprises a truncated amino acid sequence of PtNTT2, and optionally, the truncated amino acid sequence of PtNTT2 is at least 80% identical to PtNTT2 encoded by SEQ ID NO: 1. In some embodiments, the cell expresses at least two tRNA synthetases. In some embodiments, the at least two tRNA synthetases are chimeric PylRS (chPylRS) and M. jannaschii AzFRS (MjpAzFRS). In some embodiments, the cell comprises an unnatural nucleotide comprising an unnatural sugar moiety. In some instances, the unnatural sugar moiety is selected from the group consisting of: modifications at the 2' position: OH, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NHF; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH, 2'-O(CH)OCH, where alkyl, alkenyl and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n -NH2 and -O(CH2) n ON[(CH2) nCH3)]2, where n and m are from 1 to about 10); and / or modifications at the 5'-position: 5'-vinyl, 5'-methyl (R or S); modifications at the 4'-position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving group, reporter group, intercalator, group for improving the pharmacokinetic properties of an oligonucleotide, or group for improving the pharmacodynamic properties of an oligonucleotide, and any combination thereof. In some embodiments, the cell comprises at least one unnatural nucleotide base that is recognized by RNA polymerase during transcription. In some embodiments, the cells described herein translate at least one unnatural polypeptide comprising at least two unnatural amino acids.In some examples, the at least two unnatural amino acids are N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propargylethoxy-carbonyl-L-lysine (PraK), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azidophenylalanine (pAzF), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxonona acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L- Phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine The amino acid sequence is independently selected from the group consisting of 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine. In some examples, the cells described herein are isolated cells. In some alternatives, the cells described herein are prokaryotic. In some examples, the cells described herein comprise cell lines.
[0021] Various aspects of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description and accompanying drawings that set forth illustrative embodiments, in which the principles of the disclosure are utilized. [Brief explanation of the drawings]
[0022] [Figure 1]
[0023] Figure 1 illustrates a workflow using unnatural base pairing (UBP) for site-specific incorporation of non-canonical amino acids (ncAAs) into non-natural polypeptides or proteins using unnatural X-Y base pairs. The incorporation of three ncAAs into a non-natural polypeptide or protein is shown as an example only; any number of ncAAs can be incorporated. [Figure 2] FIG. 1 depicts exemplary unnatural nucleotide base pairs (UBPs). [Figure 3] FIG. 1 is a diagram depicting the deoxyriboX analog. The deoxyribose and phosphate have been omitted for clarity. [Figure 4] Figures 4A-B are diagrams illustrating ribonucleotide analogs. Figure 4A depicts a ribonucleotide X analog, with the ribose and phosphate omitted for clarity. Figure 4B depicts a ribonucleotide Y analog, with the ribose and phosphate omitted for clarity. [Figure 5A]
[0023] Figure 5A illustrates exemplary unnatural amino acids. Figure 5A is taken from Figure 2 of Young et al., "Beyond the canonical 20 amino acids: expanding the genetic lexicon," J. of Biological Chemistry 285(15):11039-11044 (2010). [Figure 5B] Figure 5A illustrates exemplary unnatural amino acids. Figure 5B is an exemplary unnatural amino acid lysine derivative. [Figure 5C] Figure 5A illustrates exemplary unnatural amino acids. Figure 5C is an exemplary unnatural amino acid phenylalanine derivative. [Figure 5D]
[0023] Figures 5D-5G illustrate exemplary unnatural amino acids. These unnatural amino acids (UAAs) are genetically encoded in proteins (Figure 5D - UAA#1-42; Figure 5E - UAA#43-89; Figure 5F - UAA#90-128; Figure 5G - UAA#129-167). Figures 5D-5G are taken from Table 1 in Dumas et al., Chemical Science 2015, 6, 50-69. [Figure 5E]
[0023] Figures 5D-5G illustrate exemplary unnatural amino acids. These unnatural amino acids (UAAs) are genetically encoded in proteins (Figure 5D - UAA#1-42; Figure 5E - UAA#43-89; Figure 5F - UAA#90-128; Figure 5G - UAA#129-167). Figures 5D-5G are taken from Table 1 in Dumas et al., Chemical Science 2015, 6, 50-69. [Figure 5F]
[0023] Figures 5D-5G illustrate exemplary unnatural amino acids. These unnatural amino acids (UAAs) are genetically encoded in proteins (Figure 5D - UAA#1-42; Figure 5E - UAA#43-89; Figure 5F - UAA#90-128; Figure 5G - UAA#129-167). Figures 5D-5G are taken from Table 1 in Dumas et al., Chemical Science 2015, 6, 50-69. [Figure 5G]
[0023] Figures 5D-5G illustrate exemplary unnatural amino acids. These unnatural amino acids (UAAs) are genetically encoded in proteins (Figure 5D - UAA#1-42; Figure 5E - UAA#43-89; Figure 5F - UAA#90-128; Figure 5G - UAA#129-167). Figures 5D-5G are taken from Table 1 in Dumas et al., Chemical Science 2015, 6, 50-69. [Figure 6-1]Figures 6A–D illustrate protein production in nonclonal SSO using unnatural codons and anticodons. The unnatural codons and anticodons are written in their DNA coding sequences. Figure 6A shows the chemical structure of the dNaM-dTPT3 UBP. Figure 6B shows the chemical structures of the ncAAs, AzK, PrK, and pAzF. Figure 6C shows a schematic diagram of the gene cassette used to express sfGFP151 (NNN) and M. mazei tRNAPyl (NNN), where NNN refers to any designated codon or anticodon. Figure 6D shows normalized fluorescence (au, arbitrary units) from nonclonal SSO cultures at the endpoint of protein expression (i.e., t = 180 min after the addition of aTc) using the designated codons and anticodons with and without AzK in the medium. Each replicate culture originated from a different batch of competent SSO starter cells transformed with a plasmid carrying the UBP (n = 3 biological replicates). Averages are shown for individual data points. One representative cropped Western blot of purified sfGFP subjected to SPAAC with TAMRA-PEG4-DBCO from SSO cultures is shown above each codon and anticodon (α-GFP channel only). The inset in Figure 6D is a scatter plot of the mean endpoint fluorescence (from Figure 6D) in the presence of AzK versus the mean quantified relative protein shift induced by SPAAC (n = 3 biological replicates). The top seven codons selected for further analysis are boxed. [Figure 6-2] Continued from Figure 6-1. [Figure 7]Figures 7A-B illustrate the analysis of protein production and codon orthogonality in clonal SSOs. Unnatural codons and anticodons are written in their DNA coding sequences. Figure 7A shows normalized fluorescence from clonal SSOs at the endpoint of protein expression (i.e., t = 180 min after the addition of aTc) for the top seven codons and anticodons (left) with and without AzK and four other selected codons (right). Each duplicate culture was grown from an individual SSO colony (left: n = 3, right: n = [5, 4, 3, 3]; biological repeats). Averages are shown for individual data points. One representative cropped Western blot of purified sfGFP from an SSO culture subjected to SPAAC with TAMRA-PEG4-DBCO is shown (α-GFP channel only). Figure 7B shows normalized fluorescence from clonal SSO cultures at the endpoint of expression for the AXC, GXT, and AGX codons and the GYT, AYC, and XCT anticodons. All pairwise combinations were tested, both with and without AzK in the medium, and with and without the ribonucleoside triphosphates NaMTP and TPT3TP in the medium. Each culture was grown from a single colony, and the mean ± standard deviation is shown (black text; n = 3; biological replicates). [Figure 8-1]Figures 8A-F illustrate the simultaneous decoding of two unnatural codons. The unnatural codon and unnatural anticodon are written in their DNA coding sequences. Figure 8A is a schematic diagram of the gene cassette containing sfGFP190, 200 (GXT, AXC), M. mazei tRNAPyl (AYC), and M. jannaschii tRNApAzF (GYT). Figures 8B-C are time plots of normalized fluorescence during sfGFP expression in the presence of the indicated ncAAs. IPTG was added at t = -60 min, and aTc was added at t = 0. Each replicate expression was performed on cultures grown from individual SSO colonies (n = 3; biological replicates). Average and individual data points are shown. Figure 8B illustrates the expression of the cassette in Figure 8A as well as a control clonal SSO showing expression of a cassette containing only the appropriate tRNA and a single codon. Figure 8C illustrates the clonal expression of cassettes containing sfGFP190, 200 (TAA, TAG), M. mazei tRNAPyl (TTA), and M. jannaschii tRNApAzF (CTA), as well as a control cassette containing the appropriate suppressor tRNA and a single stop codon, as indicated. Figure 8D shows a pseudocolored Western blot of a TAMRA fluorescence scan of purified sfGFP from the SSO of Figures 8B-C, with and without conjugation to TAMRA-PEG4-DBCO via SPAAC. Images are cropped from the same blot (UBP construct and stop codon suppressor) but positioned to align unshifted bands for ease of comparison of electrophoretic migration. Figure 8E shows a time-course plot of normalized fluorescence during clonal expression of the dual-codon / tRNA cassette from Figures 8B-C with the addition of PrK and pAzF. Average and individual data points are shown (n=3; biological replicates). Figure 8F shows a pseudocolored Western blot of TAMRA fluorescence scans of α-GFP and purified sfGFP from the SSO in Figure 8E, with and without conjugation to TAMRA-PEG4-DBCO by SPAAC and to TAMRA-PEG4-azide by CuAAC. [Figure 8-2] Continued from Figure 8-1. [Figure 9]Figures 9A-C illustrate the simultaneous decoding of three unnatural codons. The unnatural codons and anticodons are written in their DNA coding sequences. Figure 9A is a schematic diagram of a gene cassette containing sfGFP151, 190, 200 (AXC, GXT, AGX), M. mazei tRNAPyl (XCT), M. jannaschii tRNApAzF (GYT), and E. coli tRNASer (AYC). Figure 9B is a time plot of normalized fluorescence during sfGFP expression in the absence or presence of AzK and / or pAzF. IPTG was added at t = -60 min, and aTc was added at t = 0. Each replicate expression was performed on cultures grown from individual SSO colonies (n = 3; biological replicates). Average and individual data points are shown. Figure 9C shows a representative resolved mass spectrum from HRMS analysis of intact sfGFP purified from the SSO in Figure 9B. Peak labels indicate the molecular weight and quantification of each peak relative to other related species. Standard single-letter amino acid codes used are shown. The mean ± standard deviation is shown for each of these species (n = 3). [Figure 10] Figure 1 illustrates the initial screening of unnatural codons in non-clonal SSOs. The unnatural codon and unnatural anticodon are written in their DNA coding sequence. Paired strip chart of normalized fluorescence from SSO cells at the endpoint of protein expression (i.e., t = 180 min after addition of aTc) for selected codon / anticodon pairs with a UBP in either the first, second, or third position of the codon. Plus / minus indicates the addition of 20 mM AzK to the medium. Each replicate is derived from a different batch of competent SSO starter cells (n = 3; biological replicates). [Figure 11]Figures 11A-B illustrate Western blots and fluorescence scans for non-clonal SSO expression. The non-natural codons and non-natural anticodons are written in their DNA coding sequences. Figure 11A is a pseudocolored Western blot of a TAMRA fluorescence scan of α-GFP and purified sfGFP from the culture in Figure 6D conjugated to TAMRA-PEG4-DBCO by SPAAC. The plus / minus sign indicates whether SPAAC was performed. Three experiments were performed (designated 1, 2, and 3; biological replicates). Three experiments for each set (NXN / NYN and NNX / XNN) were processed in parallel. Figure 11B shows quantification of the relative shift (i.e., the signal of the shifted band divided by the total signal of the shifted and unshifted bands) in the Western blot (Figure 11A) for the indicated codon / anticodon pair. The plus / minus sign indicates whether SPAAC was performed. Mean ± standard deviation and individual data points are shown (n = 3). [Figure 12] Figures 12A-B illustrate Western blots and fluorescence scans for cloned SSO expression. Unnatural codons and anticodons are written in their DNA coding sequence. Figure 12A, False-colored Western blot of TAMRA fluorescence scan of α-GFP and purified sfGFP from the culture in Figure 7A conjugated to TAMRA-PEG4-DBCO by SPAAC. The indicated (cropped) region migrated between standard protein markers of 32 kDa and 25 kDa. Figure 12B, Quantification of relative shifts in the Western blot (Figure 12A) for the indicated codons. Mean ± standard deviation and individual data points are shown (n=3 except for CXC n=5 and GXG n=4). [Figure 13]Figure 7 illustrates clonal SSO expression in the absence of TPT3TP. Unnatural codons and anticodons are written in their DNA coding sequences. Normalized fluorescence from clonal SSOs at the end point of protein expression (i.e., t = 180 min after the addition of aTc) for the top four self-pairing codons / anticodons. Each replicate expression was performed in cultures grown from individual colonies as performed in Figure 7A (n = 3; biological replicates). Mean ± standard deviation shown for fluorescence and quantified Western blot protein shift (i.e., relative shift; gel not shown), as well as individual data points for fluorescence. [Figure 14] Figure 1 illustrates controls for dual-codon expression. The unnatural codon and unnatural anticodon are written in their DNA coding sequence. Time plot of normalized fluorescence during sfGFP expression of the indicated genotypes with and without the indicated ncAA in the medium. IPTG was added at t = -60 min, and aTc was added at t = 0. Each replicate expression was performed on cultures grown from individual colonies (n = 3; biological replicates). Average and individual data points are shown. [Figure 15] Figures 15A-B illustrate HRMS analysis of proteins from dual-codon expression. HRMS analysis (n=3; biological replicates) of intact sfGFP purified from an SSO expressing sfGFP151, 190, 200 (GXT, AXC), tRNAPyl (AYC), and tRNApAzF (GYT) with AzK and pAzF in the medium as shown in Figure 8B. Standard single-letter amino acid code was used. Figure 15A shows the analyzed spectrum along with annotation of associated peaks and their relative abundance to each other. Figure 15B shows peak assignments and interpretation. [Figure 16]Figures 16A-B illustrate HRMS analysis of proteins from triple-codon expression. HRMS analysis of intact sfGFP purified from an SSO expressing sfGFP151, 190, 200 (AXC, GXT, AGX), tRNAPyl (XCT), tRNApAzF (GYT), and tRNASer (AYC) with AzK and pAzF in the medium as shown in Figure 9B (n=3 biological replicates). Standard single-letter amino acid code was used. Figure 16A shows the analyzed spectrum along with annotation of associated peaks and their relative abundance to each other. Figure 16B shows peak assignments and interpretation. DETAILED DESCRIPTION OF THE INVENTION
[0023] Specific Terms Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not limiting of the claimed subject matter. In this application, the use of the singular includes the plural unless expressly stated otherwise. It should be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless expressly stated otherwise. Furthermore, the use of the term "including," as well as other forms such as "include," "includes," and "included," is not limiting.
[0024] As used herein, ranges and amounts may be expressed as "about" a particular value or range. About includes the exact amount. Thus, "about 5 μL" means "about 5 μL" or "5 μL." In general, the term "about" includes amounts that are expected to be within experimental error.
[0025] As used herein, phrases such as "under conditions suitable to provide" or "under conditions sufficient to produce" in the context of synthetic methods refer to conditions such as time, temperature, solvent, reaction conditions, etc., that are within the ordinary skill of an experimenter to vary, which provide a useful amount or yield of reaction product. The term "reaction product" refers to reaction conditions such as reactant concentrations, etc. The desired reaction product need not be the only reaction product, nor need the starting materials be completely consumed, provided that the desired reaction product may be isolated or otherwise further used.
[0026] "Chemically feasible" means a bonding arrangement or compound that does not violate the generally understood rules of organic structure. For example, it is understood that structures within the definitions of the claims that contain a pentavalent carbon atom that does not occur in nature in certain circumstances are not within the scope of the claims. The structures disclosed herein, in all their embodiments, are intended to include only "chemically feasible" structures; for example, any recited structure that is not chemically feasible in a structure shown with variable atoms or groups is not intended to be disclosed or claimed herein.
[0027] An "analog" of a chemical structure, as the term is used herein, refers to a chemical structure that may not be readily synthetically derived from the parent structure, but that retains substantial similarity to the parent structure. In some embodiments, a nucleotide analog is a non-natural nucleotide. In some embodiments, a nucleoside analog is a non-natural nucleoside. Related chemical structures that are readily synthetically derived from the parent chemical structure are referred to as "derivatives."
[0028] Thus, as used herein, the term polynucleotide refers to DNA, RNA, DNA- or RNA-like polymers, such as peptide nucleic acids (PNAs), locked nucleic acids (LNAs), phosphorothioates, unnatural bases, etc., which are well known in the art. Polynucleotides can be synthesized on an automated synthesizer, for example, using phosphoramidite chemistry or other chemical approaches adapted for synthesizer use.
[0029] DNA includes, but is not limited to, cDNA and genomic DNA. DNA can be bound to other biomolecules, including, but not limited to, RNA and peptides, by covalent or non-covalent means. RNA includes coding RNA, such as messenger RNA (mRNA). In some embodiments, the RNA is rRNA, RNAi, snoRNA, microRNA, siRNA, snRNA, exRNA, piRNA, long ncRNA, or any combination or hybrid thereof. In some examples, the RNA is a component of a ribozyme. DNA and RNA may be in any form, including, but not limited to, linear, circular, supercoiled, single-stranded, and double-stranded.
[0030] Peptide nucleic acids (PNAs) are synthetic DNA / RNA analogs in which a peptide-like backbone replaces the sugar-phosphate backbone of DNA or RNA. PNA oligomers exhibit higher binding strength and greater specificity in binding to complementary DNA, making PNA / DNA base mismatches more destabilizing than similar mismatches in DNA / DNA duplexes. This binding strength and specificity also applies to PNA / RNA duplexes. PNAs are not readily recognized by nucleases or proteases, making them resistant to enzymatic degradation. PNAs are also stable over a wide pH range. Nielsen PE, Egholm M, Berg RH, Buchardt O (December 1991). "Sequence-selective recognition of DNA by strand displacement with thymine-substituted polyamide," Science 254(5037):1497-500. doi:10.1126 / science.1962210. PMID 1962210; and Egholm M, Buchardt O, Christensen L, Behrens C, Freier SM, Driver DA, Berg RH, Kim SK, Norden B, and Nielsen PE (1993), "PNA Hybridizes to Complementary Oligonucleotides Obeying Watson-Crick Hydrogen Bonding Rules." "Complementary Oligonucleotides Obeying the Watson-Crick Hydrogen Bonding Rules" Nature 365(6446):566-8. doi:10.1038 / 365566a0. See also PMID 7692304.
[0031] Locked nucleic acids (LNA) are modified RNA nucleotides in which the ribose moiety of the LNA nucleotide is modified with an additional bridge connecting the 2' oxygen and 4' carbon. This bridge "locks" the ribose in the 3'-endo (N) conformation, which is commonly found in A-form duplexes. LNA nucleotides can be mixed with DNA or RNA residues in oligonucleotides as needed. Such oligomers can be chemically synthesized or commercially available. The locked ribose conformation enhances base stacking and backbone preorganization. See, e.g., Kaur, H; Arora, A; Wengel, J; Maiti, S (2006), "Thermodynamic, Counterion, and Hydration Effects for the Incorporation of Locked Nucleic Acid Nucleotides into DNA Duplexes," Biochemistry 45(23):7347-7355.doi:10.1021 / bi060307w. PMID 16752924; Owczarzy R.; You Y., Groth CL, Tataurov AV (2011), "Stability and mismatch discrimination of locked nucleic acid-DNA duplexes", Biochem. 50(43):9352-9367. doi:10.1021 / bi200904e. PMC 3201676. PMID 21928795; Alexei A. Koshkin; Sanjay K. Singh, Poul Nielsen, Vivek K. Rajwanshi, Ravindra Kumar, Michael Meldgaard, Carl Erik Olsen, Jesper Wengel (1998), "LNA (Locked Nucleic Acids): Synthesis of the adenine, cytosine, guanine, 5-methylcytosine, thymine and uracil bicyclonucleoside monomers, oligomerization, and unprecedented nucleic acid recognition," Tetrahedron 54(14):3607-30. doi:10.1016 / S0040-4020(98)00094-5; and Satoshi Obika; Daishu Nanbu, Yoshiyuki Hari, Ken-ichiro Morio, Yasuko In, Toshimasa Ishida, Takeshi Imanishi (1997), "2'-O,. Synthesis of 2'-O,4'-C-methylene uridine and -cytidine. Novel bicyclic nucleosides with fixed C3'-endo sugar puckering. See "Bicyclic nucleosides having a fixed C3'-endo sugar puckering," Tetrahedron Lett. 38(50):8735-8. doi:10.1016 / S0040-4039(97)10322-7.
[0032] Molecular beacons or molecular beacon probes are oligonucleotide hybridization probes that can detect the presence of specific nucleic acid sequences in a homogeneous solution. Molecular beacons are hairpin-shaped molecules with an internally quenched fluorophore that can detect the presence of a target nucleic acid sequence. Fluorescence is restored upon binding to the array. See, e.g., Tyagi S, Kramer FR (1996), "Molecular beacons: probes that fluoresce upon hybridization," Nat Biotechnol. 14(3):303-8. PMID 9630890; Tapp I, Malmberg L, Rennel See E, Wik M, Syvanen AC (April 2000), "Homogeneous scoring of single-nucleotide polymorphisms: comparison of the 5'-nuclease TaqMan assay and Molecular Beacon probes," Biotechniques 28(4):732-8. PMID 10769752; and Akimitsu Okamoto (2011), "ECHO probes: a fluorescence-controlled concept for practical nucleic acid sensing," Chem. Soc. Rev. 40:5815-5828.
[0033] In some embodiments, a nucleobase is generally the heterocyclic base moiety of a nucleoside. A nucleobase may be naturally occurring, modified, or may not have similarity to a natural base, for example, synthesized by organic synthesis. In certain embodiments, a nucleobase comprises any atom or group of atoms that can interact with a base of another nucleic acid, with or without the use of hydrogen bonds. In certain embodiments, a non-natural nucleobase is not derived from a natural nucleobase. It should be noted that a non-natural nucleobase does not necessarily possess basic properties, but is referred to as a nucleobase for simplicity. In some embodiments, when referring to a nucleobase, "(d)" indicates that the nucleobase may be attached to deoxyribose or ribose.
[0034] In some embodiments, a nucleoside is a compound comprising a nucleobase portion and a sugar portion.Nucleosides include, but are not limited to, naturally occurring nucleosides (found in DNA and RNA), abasic nucleosides, modified nucleosides, and nucleosides with pseudobase and / or sugar groups.Nucleosides include nucleosides containing any of a variety of substituents.Nucleosides can be glycosidic compounds formed by glycosidic linkage between a nucleobase and a reducing group of a sugar.
[0035] In some embodiments, the non-natural mRNA codons and non-natural tRNA anticodons described in this disclosure can be written in their DNA coding sequences, for example, a non-natural tRNA anticodon can be written as GYU or GYT.
[0036] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0037] Compositions and methods for the in vivo synthesis of non-natural polypeptides Disclosed herein are compositions and methods for the in vivo synthesis of non-natural polypeptides with an expanded genetic alphabet. In some instances, the compositions and methods described herein include a non-natural nucleic acid molecule encoding a non-natural polypeptide, wherein the non-natural polypeptide comprises a non-natural amino acid. In some instances, the non-natural polypeptide comprises at least two non-natural amino acids. In some instances, the non-natural polypeptide comprises at least three non-natural amino acids. In some instances, the non-natural polypeptide comprises two non-natural amino acids. In some instances, the non-natural polypeptide comprises three non-natural amino acids. In some instances, the at least two non-natural amino acids incorporated into the non-natural polypeptide can be the same or different non-natural amino acids. In some instances, the non-natural amino acids are incorporated into the non-natural polypeptide site-specifically. In some instances, A non-naturally occurring polypeptide is a non-naturally occurring protein.
[0038] In some instances, the compositions and methods described herein include semisynthetic biology (SSO). In some instances, the method includes incorporating at least one unnatural base pair (UBP) into at least one unnatural nucleic acid molecule. In some embodiments, the method includes incorporating one UBP into at least one unnatural nucleic acid molecule. In some embodiments, the method includes incorporating two UBPs into at least one unnatural nucleic acid molecule. In some embodiments, the method includes incorporating three UBPs into at least one unnatural nucleic acid molecule. A UBP base pair is formed by pairing between unnatural nucleobases of two unnatural nucleosides. In some embodiments, the unnatural nucleic acid molecule is an unnatural DNA molecule.
[0039] In some embodiments, the at least one unnatural nucleic acid molecule is or comprises one molecule (e.g., a plasmid or a chromosome). In some embodiments, the at least one unnatural nucleic acid molecule is or comprises two molecules (e.g., two plasmids, two chromosomes, or a chromosome and a plasmid). In some embodiments, the at least one unnatural nucleic acid molecule is or comprises three molecules (e.g., three plasmids, two plasmids and a chromosome, a plasmid and two chromosomes, or three chromosomes). Examples of chromosomes include genomic chromosomes into which UBPs have been incorporated and artificial chromosomes (e.g., bacterial artificial chromosomes) that comprise UBPs. In some embodiments, at least one unnatural DNA molecule that comprises at least four unnatural base pairs is used, and when the at least one unnatural DNA molecule is two or more molecules, the at least four unnatural base pairs may be distributed in any feasible manner among the two or more molecules (e.g., one in the first and three in the second, two in the first and two in the second, etc.).
[0040] In some examples, at least one unnatural nucleic acid molecule, optionally including a UBP, is transcribed to provide a messenger RNA molecule that includes at least one unnatural codon with at least one unnatural nucleotide. In some embodiments, transcribing refers to producing one or more RNA molecules that are complementary to a portion of a DNA molecule. In some cases, the unnatural nucleotide occupies the first, second, or third codon position of the unnatural codon, e.g., the second or third codon position. In some cases, two unnatural nucleotides occupy the first and second, first and third, second and third, or first and third codon positions of the unnatural codon. In some cases, three unnatural nucleotides occupy all three codon positions of the unnatural codon. In some cases, the mRNA with the unnatural nucleotides includes at least two unnatural codons (in some embodiments, the phrase "at least two unnatural codons" is interchangeable with "at least a first and a second unnatural codon"). In some cases, the mRNA with the unnatural nucleotides includes two unnatural codons. In some cases, the mRNA with the unnatural nucleotide contains three unnatural codons.
[0041] In some embodiments, a non-naturally occurring nucleic acid molecule, optionally comprising a UBP, is transcribed to provide at least one tRNA molecule, wherein the tRNA molecule comprises a non-naturally occurring anticodon with at least one non-naturally occurring nucleotide. Optionally, the non-naturally occurring nucleotide occupies the first, second, or third anticodon position of the non-naturally occurring anticodon. Optionally, two non-naturally occurring nucleotides occupy the first and second, first and third, second and third, or first and third anticodon positions of the non-naturally occurring anticodon. Optionally, three non-naturally occurring nucleotides occupy all three anticodon positions of the non-naturally occurring anticodon. Optionally, a non-naturally occurring nucleic acid molecule, optionally comprising a UBP, is transcribed to provide at least two tRNAs comprising at least two non-naturally occurring anticodons. Optionally, the at least two non-naturally occurring anticodons can be the same or different. In some instances, a non-naturally occurring nucleic acid molecule, optionally comprising a UBP, is transcribed to give two tRNAs containing non-naturally occurring anticodons, which may be the same or different. In some instances, a non-naturally occurring nucleic acid molecule, optionally comprising a UBP, is transcribed to give three tRNAs containing three non-naturally occurring anticodons, which may be the same or different.
[0042] In some embodiments, at least one unnatural codon encoded by the mRNA can be complementary to at least one unnatural anticodon of the tRNA to form an unnatural codon-anticodon pair. In some embodiments, the compositions and methods described herein include synthesizing an unnatural polypeptide with one, two, three, or more unnatural codon-anticodon pairs. In some embodiments, the compositions and methods described herein include synthesizing an unnatural polypeptide with two unnatural codon-anticodon pairs. In some embodiments, the compositions and methods described herein include synthesizing an unnatural polypeptide with three unnatural codon-anticodon pairs.
[0043] In some cases, the compositions and methods described herein include synthesizing an unnatural polypeptide with one, two, three, or more unnatural amino acids using one, two, three, or more unnatural codon-anticodon pairs. In some cases, the compositions and methods described herein include synthesizing an unnatural polypeptide with two unnatural amino acids using two unnatural codon-anticodon pairs. In some cases, the compositions and methods described herein include synthesizing an unnatural polypeptide with three unnatural amino acids using three unnatural codon-anticodon pairs.
[0044] In some examples, the unnatural codon comprises the nucleic acid sequence XNN, NXN, NNX, XXN, XNX, NXX, or XXX, and the unnatural anticodon comprises the nucleic acid sequence XNN, YNN, NXN, NYN, NNX, NNY, NXX, NYY, XNX, YNY, XXN, YYN, or YYY, thereby forming an unnatural codon-anticodon pair. In some cases, the unnatural codon-anticodon pair comprises NNX-XNN, NNX-YNN, or NXN-NYN, where N is any naturally occurring nucleotide, X is a first unnatural nucleotide, and Y is a second unnatural nucleotide. In some embodiments, the naturally occurring nucleotide includes nucleotides having a standard base, such as adenine, thymine, uracil, guanine, or cytosine, as well as nucleotides having naturally occurring modified bases such as pseudouridine and 5-methylcytosine. In some embodiments, the unnatural codon-anticodon pair comprises at least one G in the codon and at least one C in the anticodon. In some embodiments, the unnatural codon-anticodon pair includes at least one G or C in the codon and at least one complementary C or G in the anticodon.X and Y are each independently selected from the group consisting of: (i) 2-thiouracil, 2'-deoxyuridine, 4-thiouracil, uracil-5-yl, hypoxanthine-9-yl(I), 5-halouracil; 5-propynyl-uracil, 6-azo-uracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, uracil-5-oxaacetic acid methyl ester, uracil-5-oxaacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-Methoxyuracil, uracil-5-oxaacetic acid, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, 5-hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytosine arabinoside, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azocytosine, azacytosine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][. 1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one), pyridoindole cytidine (H-pyrimido [3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one), 2-aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2'-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted Adenine, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine, 6-azaadenine, 2-methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thioguanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, 8-halo, 8-aza amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted guanines, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine, 6-azaguanine, hypoxanthine, xanthine, 1-methylinosine, queosine, beta-D-galactosylqueosine, inosine, beta-D-mannosylqueosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone. In some embodiments, X and Y are independently selected from the group consisting of: [ka]
[0045] In some cases, the unnatural codon-anticodon pair comprises NNX-XNN, where NNX-XNN is selected from the group consisting of AAX-XUU, AUX-XAU, ACX-XGU, AGX-XCU, UAX-XUA, UUX-XAA, UCX-XGA, UGX-XCA, CAX-XUG, CUX-XAG, CCX-XGG, CGX-XCG, GAX-XUC, GUX-XAC, GCX-XGC, and GGX-XCC. In some cases, the unnatural codon-anticodon pair comprises NNX-YNN, where NNX-YNN is selected from the group consisting of AAX-YUU, AUX-YAU, ACX-YGU, AGX-YCU, UAX-YUA, UUX-YAA, UCX-YGA, UGX-YCA, CAX-YUG, CUX-YAG, CCX-YGG, CGX-YCG, GAX-YUC, GUX-YAC, GCX-YGC, and GGX. In some embodiments, the unnatural codon-anticodon pair comprises NXN-NXN, where NXN-NXN is selected from the group consisting of AXA-UXU, AXU-AXU, AXC-GXU, AXG-CXU, UXA-UXA, UXU-AXA, UXC-GXA, UXG-CXA, CXA-UXG, CXU-AXG, CXC-GXG, CXG-CXG, GXA-UXC, GXU-AXC, GXC-GXC, and GXG-CXC. In some examples, the unnatural codon-anticodon pair comprises NXN-NYN, where NXN-NYN is selected from the group consisting of AXA-UYU, AXU-AYU, AXC-GYU, AXG-CYU, UXA-UYA, UXU-AYA, UXC-GYA, UXG-CYA, CXA-UYG, CXU-AYG, CXC-GYG, CXG-CYG, GXA-UYC, GXU-AYC, GXC-GYC, and GXG-CYC.
[0046] In some embodiments, the unnatural codon-anticodon pair comprises XNN-NNX, where XNN-NNX is selected from the group consisting of XAA-UUX, XAU-AUX, XAC-AGX, XAG-CUX, XUA-UAX, XUU-AAX, XUC-GAX, XUG-CAX, XCA-UGX, XCU-AGX, XCC-GGX, XCG-CGX, XGA-UCX, XGU-ACX, XGC-GCX, and XGG-CCX. In some embodiments, the unnatural codon-anticodon pair comprises XNN-NNY, where XNN-NNY is selected from the group consisting of XAA-UUY, XAU-AUY, XAC-AGY, XAG-CUY, XUA-UAY, XUU-AAY, XUC-GAY, XUG-CAY, XCA-UGY, XCU-AGY, XCC-GGY, XCG-CGY, XGA-UCY, XGU-ACY, XGC-GCY, and XGG-CCY.
[0047] In some embodiments, the unnatural codon-anticodon pair comprises XXN-NXX, where XXN-NXX is selected from the group consisting of XXA-UXX, XXU-AXX, XXC-GXX, and XXG-CXX. In some embodiments, the unnatural codon-anticodon pair comprises XXN-NYY, where XXN-NYY is selected from the group consisting of XXA-UYY, XXU-AYY, XXC-GYY, and XXG-CYY. In some alternatives, the unnatural codon-anticodon pair comprises XNX-XNX, where XNX-XNX is selected from the group consisting of XAX-XUX, XUX-XAX, XCX-XGX, and XGX-XCX. In some embodiments, the unnatural codon-anticodon pair comprises XNX-YNY, where XNX-YNY is selected from the group consisting of XAX-YUY, XUX-YAY, XCX-YGY, and XGX-YCY. In some cases, the unnatural codon-anticodon pair comprises NXX-XXN, where NXX-XXN is selected from the group consisting of AXX-XXU, UXX-XXA, CXX-XXG, and GXX-XXC. In some examples, the unnatural codon-anticodon pair comprises NXX-YYN, where NXX-YYN is selected from the group consisting of AXX-YYU, UXX-YYA, CXX-YYG, and GXX-YYC. In some cases, the unnatural codon-anticodon pair comprises XXX-XXX or XXX-YYY.
[0048] In an exemplary workflow 100 (FIG. 1) of a method for generating unnatural polypeptides with an expanded genetic alphabet (FIG. 2), a D gene encoding a protein 102 and a tRNA 103 each containing complementary unnatural nucleobases (X, Y) are generated. NA 101 is transcribed 104 to generate tRNA 106 and mRNA 107. X is a first unnatural nucleotide and Y is a second unnatural nucleotide. After charging the tRNA with the unnatural amino acid 105, the mRNA 107 is translated 108 to produce a protein 110 that contains one or more unnatural amino acids 109. The methods and compositions described herein allow for the site-specific incorporation of unnatural amino acids, in some instances with high fidelity and yield. Also described herein are semisynthetic organisms that contain an expanded genetic alphabet, methods for using semisynthetic organisms to generate protein products, including those that contain at least one unnatural amino acid residue.
[0049] The selection of unnatural nucleobases allows for the optimization of one or more steps in the methods described herein. For example, nucleobases are selected for high efficiency replication, transcription, and / or translation. In some instances, one or more unnatural nucleobase pairs are utilized for the methods described herein. For example, a first set of nucleobases containing deoxyribonucleotides is used for DNA replication (e.g., a first nucleobase and a second nucleobase configured to form a first base pair), and a second set of nucleobases (e.g., a third nucleobase and a fourth nucleobase configured to bind to ribose and form a second base pair) is used for transcription / translation. Complementary pairing between the first set of nucleobases and the second set of nucleobases, in some instances, allows transcription of a gene to produce a tRNA or protein from a DNA template containing nucleobases from the first set. Complementary pairing (second base pairs) between the second set of nucleobases, in some instances, allows translation by matching tRNA and mRNA containing unnatural nucleic acids. In some cases, the nucleobases in the first set are bound to a deoxyribose moiety. In some cases, the nucleobases in the first set are bound to a ribose moiety. In some cases, the nucleobases in both sets are unique. In some cases, at least one nucleobase is the same in both sets. In some cases, the first nucleobase and the third nucleobase are the same. In some embodiments, the first base pair and the second base pair are not the same. In some cases, the first base pair, the second base pair and the third base pair are not the same.
[0050] In some embodiments, the yield of a non-naturally occurring polypeptide or protein synthesized by the compositions and methods disclosed herein is higher than the yield of the same non-naturally occurring polypeptide or protein synthesized by other methods. In some instances, the yield of a non-naturally occurring polypeptide or protein synthesized by the compositions and methods disclosed herein is at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than the yield of the same non-naturally occurring polypeptide or protein synthesized by other methods. Examples of other methods include methods that utilize amber codon suppression.
[0051] In some instances, the solubility of a non-natural polypeptide or protein synthesized by the compositions and methods disclosed herein is higher compared to the solubility of the same non-natural polypeptide or protein synthesized by other methods. In some instances, the solubility of a non-natural polypeptide or protein synthesized by the compositions and methods disclosed herein is at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than the same non-natural polypeptide or protein synthesized by other methods. In some instances, the biological activity of a non-natural protein synthesized by the compositions and methods disclosed herein is higher compared to the biological activity of the same non-natural protein synthesized by other methods. In some instances, the biological activity of a non-natural protein synthesized by the compositions and methods disclosed herein is at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than the biological activity of the same non-natural protein synthesized by other methods.
[0052] In some embodiments, the compositions and methods for in vivo synthesis of non-naturally occurring polypeptides described herein utilize or include semi-synthetic organisms (SSOs). In some embodiments, the SSOs undergo clonal expansion during synthesis of the non-naturally occurring polypeptide. In some instances, the SSOs do not undergo clonal expansion during synthesis of the non-naturally occurring polypeptide. In some instances, the SSOs can be arrested at any stage of the cell cycle during synthesis of the non-naturally occurring polypeptide. In some embodiments, the compositions and methods described herein In some cases, the non-naturally occurring polypeptides can be synthesized in vitro. In some cases, the compositions and methods described herein can include a cell-free system for synthesizing the non-naturally occurring polypeptides.
[0053] nucleic acid molecule In some embodiments, the nucleic acid (e.g., also referred to herein as a nucleic acid molecule of interest) is from any source or composition, e.g., DNA, cDNA, gDNA (genomic DNA), RNA, siRNA (short inhibitory RNA), RNAi, tRNA, mRNA, or rRNA (ribosomal RNA), and is in any form (e.g., linear, circular, supercoiled, single-stranded, double-stranded, etc.). In some embodiments, the nucleic acid comprises a nucleotide, nucleoside, or polynucleotide. In some cases, the nucleic acid includes natural and non-natural nucleic acids. In some cases, the nucleic acid also includes non-natural nucleic acids, such as DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-native backbones, etc.). It is understood that the term "nucleic acid" does not refer to or infer a specific length of a polynucleotide chain, and thus, polynucleotides and oligonucleotides are also included within its definition. Exemplary natural nucleotides include, but are not limited to, ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, GMP, dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Exemplary natural deoxyribonucleotides include dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Exemplary natural ribonucleotides include ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, and GMP. For natural RNA, the uracil base is uridine. Nucleic acid can be vector, plasmid, phagemid, autonomously replicating sequence (ARS), centromere, artificial chromosome, yeast artificial chromosome (e.g., YAC), or other nucleic acid that can replicate in host cell or replicates in host cell.In some cases, non-natural nucleic acid is nucleic acid analog.In additional cases, non-natural nucleic acid is derived from extracellular source.In other cases, non-natural nucleic acid can be available in the intracellular space of organisms provided herein, for example, genetically modified organisms.In some embodiments, a non-natural nucleotide is not a naturally occurring nucleotide. In some embodiments, a nucleotide that does not contain a natural base comprises a non-natural nucleobase.
[0054] unnatural nucleic acids Nucleotide analogs or non-natural nucleotides include nucleotides containing some type of modification in either the base, sugar, or phosphate moiety. In some embodiments, the modification includes a chemical modification. In some cases, the modification occurs in the 3'OH or 5'OH group, the backbone, the sugar moiety, or the nucleotide base. In some examples, the modification optionally includes a non-naturally occurring linker molecule and / or an inter- or intra-strand crosslink. In one aspect, the modified nucleic acid includes one or more modifications of the 3'OH or 5'OH group, the backbone, the sugar moiety, or the nucleotide base, and / or the addition of a non-naturally occurring linker molecule. In one aspect, the modified backbone includes a backbone other than a phosphodiester backbone. In one aspect, the modified sugar includes a sugar other than deoxyribose (in modified DNA) or other than ribose (in modified RNA). In one embodiment, the modified base comprises a base other than adenine, guanine, cytosine or thymine (in modified DNA) or a base other than adenine, guanine, cytosine or uracil (in modified RNA).
[0055] In some embodiments, the nucleic acid comprises at least one modified base. In some examples, the nucleic acid comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or and further modified bases. In some cases, modifications to the base moiety include natural and synthetic modifications of A, C, G, and T / U, as well as different purine or pyrimidine bases. In some embodiments, the modifications are modified forms of adenine, guanine, cytosine, or thymine (in modified DNA), or modified forms of adenine, guanine, cytosine, or uracil (modified RNA).
[0056] Modified bases of non-natural nucleic acids include, but are not limited to, uracil-5-yl, hypoxanthine-9-yl (I), 2-aminoadenin-9-yl, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and Included are cytosine, 6-azouracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, particularly 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine. Certain unnatural nucleic acids, such as 5-substituted pyrimidines, 6-azapyrimidines and N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, 2-aminopropyladenine, 5-propynyluracil, 5-propynylcytosine, 5-methylcytosine, those that increase the stability of duplex formation, universal nucleic acids, hydrophobic nucleic acids, promiscuous nucleic acids, size-expanded nucleic acids, fluorinated nucleic acids, 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, include 2-aminopropyladenine, 5-propynyluracil, and 5-methylcytosine. cytosine, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyl (-C≡C-CH3)uracil, 5-propynylcytosine, other alkynyl derivatives of pyrimidine nucleic acids, 6-azouracil, 6-azocytosine, 6-azothymine5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl, other 5-substituted uracils and cytosines, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine cytidines, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, tricyclic pyrimidines, phenoxazine cytidines ([5,4-b][1,4]benzoxazin-2(3H)-ones), phenothiazine cytidines (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-ones), G-clamps, phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-ones), 2H-pyrimido[4,5-b]indol-2-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one), pyridoindole cytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one), purine or pyrimidine bases replaced by other heterocycles, 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, 2-pyridone, azacytosine, 5-bromocytosine, ...tidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, bromocytidine, Bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil and 5-iodouracil, 2-amino-adenine, 6-thio-guanine, 2-thio, O-thymine, 4-thio-thymine, 5-propynyl-uracil, 4-thio-uracil, N4-ethylcytosine, 7-deazaguanine, 7-deaza-8-azaguanine, 5-hydroxycytosine, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, and the like, as described in U.S. Patent Nos. 3,687,808; 4,845,205; 4,910,300; 4,948,882; 5,093,232; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,18 No. 7; No. 5,459,255; No. 5,484,908; No. 5,502,177; No. 5,525,711; No. 5,552,540 No. 5,587,469; No. 5,594,121; No. 5,596,091; No. 5,614,617; No. 5,645,985 5,681,941; 5,750,692; 5,763,588; 5,830,653 and 6,005,096; WO 99 / 62923; Kandimalla et al. (2001) Bioorg. Med. Chem. 9:807-813; The Concise Encyclopedia of Polymer Science and Engineering, Kroschwitz, JI (ed.), John Wiley & Sons, 1990, pp. 858-859; Englisch et al., Angewandte Chemie, International Edition, 1991, Vol. 30, p. 613; and Sanghvi, Chapter 15, Antisense Research and Applications, Crooke and Lebleu (eds.), CRC Press, 1993, pp. 273-288. Additional base modifications can be found, for example, in U.S. Pat. No. 3,687,808; Englisch et al., Angewandte Chemie, International Edition, 1991, Vol. 30, p. 613. In some instances, the non-natural nucleic acid comprises the nucleobase of Figure 3. In some instances, the non-natural nucleic acid comprises the nucleobase of Figure 4A. In some instances, the non-natural nucleic acid comprises the nucleobase of Figure 4B.
[0057] Non-natural nucleic acids containing various heterocyclic bases and various sugar moieties (and sugar analogs) are available in the art, and in some cases, nucleic acids contain one or several heterocyclic bases other than the five main base components of naturally occurring nucleic acids. For example, heterocyclic bases in some cases include uracil-5-yl, cytosin-5-yl, adenin-7-yl, adenin-8-yl, guanin-7-yl, guanin-8-yl, 4-aminopyrrolo[2.3-d]pyrimidin-5-yl, 2-amino-4-oxopyrrolo[2,3-d]pyrimidin-5-yl, 2-amino-4-oxopyrrolo[2.3-d]pyrimidin-3-yl groups, where the purine is linked to the sugar moiety of the nucleic acid through the 9-position, the pyrimidine through the 1-position, the pyrrolopyrimidine through the 7-position, and the pyrazolopyrimidine through the 1-position.
[0058] In some embodiments, modified bases of non-natural nucleic acids are represented below, where the wavy line or R identifies the point of attachment to the deoxyribose or ribose:
[0059] [ka] [ka] [ka] [ka]
[0060] In some embodiments, the nucleotide analogue is also modified at the phosphate moiety. Modified phosphate moieties include, but are not limited to, those with modifications in the linkage between two nucleotides, such as phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates, including 3'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramidates, including 3'-aminophosphoramidates and aminoalkylphosphoramidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, and boranophosphates. The linkage of these phosphates or modified phosphates between two nucleotides can be via a 3'-5' linkage or a 2'-5' linkage. It is understood that the linkage includes reverse orientations such as 3'-5' to 5'-3' or 2'-5' to 5'-2'. Various salts, mixed salts and free acid forms are also included. Numerous U.S. patents teach how to make and use nucleotides containing modified phosphates, including but not limited to: 3,687,808; 4,469,863; 4,476,301; 5,023,243; 5,177,196; 5,188,897; 5,264,423; 5,276,019; 5,278,302; 5,286,717; 5,286,718; 5,286,720; 5,286,730; 5,286,740; 5,286,750; 5,286,760; 5,286,770; 5,286,780; 5,286,718; 5,286,720; 5,286,740; 5,286,750; 5,286,740; 5,286,750; 5,286,750; 5,286,760; 5,286,760; 5,286,718; 5,286,720; 5,286,74 ... ,321,131; No. 5,399,676; No. 5,405,939; No. 5,453,496; No. 5,455,233; No. 5,466,677; No. 5,476,925; No. 5,519,126; No. 5,536,821; No. 5,541,306; No. 5,550,111; No. 5,563,253; No. 5,571,799; No. 5,587,361; and No. 5,625,050.
[0061] In some embodiments, the non-naturally occurring nucleic acids are 2',3'-dideoxy-2',3'-didehydro-nucleosides (PCT / US2002 / 006460), 5'-substituted DNA and RNA derivatives (PCT / US2011 / 033961; Saha et al., J. Org Chem., 1995, Vol. 60, pp. 788-789; Wang et al., Bioorganic & Medicinal Chemistry Letters, 1999, Vol. 9, pp. 885-890; and Mikhailov et al., Nucleosides & Nucleotides, 1991, Vol. 10(Is. 1-3), pp. 339-343; Leonid et al., 1995, Vol. 14(Is. 3-5), pp. 901-905; and Eppacher et al., Helvetica Chimica Acta, 2004, Vol. 87, pp. 3004-3020; PCT / JP2000 / 004720; PCT / JP2003 / 002342; PCT / JP2004 / 013216; PCT / JP2005 / 020435; PCT / JP2006 / 315479; PCT / JP2006 / 324484; PCT / JP2009 / 056718; PCT / JP2010 / 067560), or 5'-substituted monomers made as monophosphates with modified bases (Wang et al., Nucleosides Nucleotides & Nucleic Acids, 2004, Vol. 23(1 & 2), pp. 317-337).
[0062] In some embodiments, non-natural nucleic acids contain modifications at the 5' and 2' positions of the sugar ring (PCT / US94 / 02993), such as 5'-CH2-substituted 2'-O-protected nucleosides (Wu et al., Helvetica Chimica Acta, 2000, Vol. 83, pp. 1127-1143 and Wu et al., Bioconjugate Chem. 1999, Vol. 10, pp. 921-924). In some cases, non-natural nucleic acids contain amide-linked nucleoside dimers prepared for incorporation into oligonucleotides, where the 3'-linked nucleoside (5' to 3') in the dimer contains 2'-OCH3 and 5'-(S)-CH3 (Mesmaeker et al., Synlett, 1997, pp. 1287-1290). Non-natural nucleic acids can include 2'-substituted 5'-CH2 (or O) modified nucleosides (PCT / US92 / 01020). Non-natural nucleic acids can include 5'-methylene phosphonate DNA and RNA monomers and dimers (Bohringer et al., Tet. Lett., 1993, Vol. 34, pp. 2723-2726; Collingwood et al., Synlett, 1995, Vol. 7, pp. 703-705; and Hutter et al., Helvetica Chimica Acta, 2002, Vol. 85, pp. 2777-2806). Non-natural nucleic acids can include 5'-phosphonate monomers with 2'-substitutions (U.S. Patent Application Publication No. 2006 / 0074035) and other modified 5'-phosphonate monomers (WO 1997 / 35869). Non-naturally occurring nucleic acids can include 5'-modified methylene phosphonate monomers (EP 614907 and EP 629633). Non-naturally occurring nucleic acids can include 5'- or 6'-phosphonate ribonucleic acids containing hydroxyl groups at the 5' and / or 6' positions. Nucleoside analogs (Chen et al., Phosphorus, Sulfur and Silicon, 2002, Vol. 777, pp. 1783-1786; Jung et al., Bioorg. Med. Chem., 2000, Vol. 8, pp. 2501-2509; Gallier et al., Eur. J. Org. Chem., 2007, pp. 925-933; and Hampton et al., J. Med. Chem., 1976, Vol. 19(8), pp. 1029-1033) can be included. Non-naturally occurring nucleic acids can include 5'-phosphonate deoxyribonucleoside monomers and dimers having a 5'-phosphate group (Nawrot et al., Oligonucleotides, 2006, Vol. 16(1), pp. 68-82). Non-natural nucleic acids can include nucleosides having a 6'-phosphonate group, where the 5' and / or 6' positions are unsubstituted or contain a thio-tert-butyl group (SC(CH3)3) (and its analogs); a methyleneamino group (CH2NH2) (and its analogs) or a cyano group (CN) (and its analogs) (Fairhurst et al., Synlett, 2001, Vol. 4, pp. 467-472; Kappler et al., J. Med. Chem., 1986, Vol. 29, pp. 1030-1038; Ka ppler et al., J. Med. Chem., 1982, Vol. 25, pp. 1179-1184; Vrudhula et al., J. Med. Chem., 1987, Vol. 30, pp. 888-894; Hampton et al., J. Med. Chem., 1976, Vol. 19, pp. 1371-1377; Geze et al., J. Am. Chem. Soc., 1983, Vol. 105 (No. 26), pp. 7638-7640; and Hampton et al., J. Am. Chem. Soc., 1973, Vol. 95 (No. 13), pp. 4404-4414).
[0063] In some embodiments, non-natural nucleic acids also include modifications to the sugar moiety. In some cases, the nucleic acid contains one or more nucleosides with modified sugar groups. Such sugar-modified nucleosides may be conferred enhanced nuclease stability, increased binding affinity, or some other advantageous biological property. In certain embodiments, the nucleic acid includes a chemically modified ribofuranose ring moiety. Examples of chemically modified ribofuranose rings include, but are not limited to, the addition of substituents (5' and / or 2' substituents; bridging of two ring atoms to form bicyclic nucleic acids (BNAs); S, N(R), or C(R1)(R2) (R = H, C1-C) at the oxygen atom of the ribosyl ring. 12 Examples of chemically modified sugars can be found in WO 2008 / 101157, U.S. Patent Application Publication No. 2005 / 0130923, and WO 2007 / 134181.
[0064] In some instances, modified nucleic acids contain modified sugars or sugar analogs. Thus, in addition to ribose and deoxyribose, the sugar moiety can be a pentose, deoxypentose, hexose, deoxyhexose, glucose, arabinose, xylose, lyxose, or a cyclopentyl group of a sugar "analog." The sugar can be in the form of a pyranosyl or furanosyl. The sugar moiety can be a furanoside of ribose, deoxyribose, arabinose, or 2'-O-alkylribose, and the sugar can be linked to the respective heterocyclic base in either the [alpha] or [beta] anomeric configuration. Sugar modifications include, but are not limited to, 2'-alkoxy-RNA analogs, 2'-amino-RNA analogs, 2'-fluoro-DNA, and 2'-alkoxy- or amino-RNA / DNA chimeras. For example, sugar modifications can include 2'-O-methyl-uridine or 2'-O-methyl-cytidine. Sugar modifications include 2'-O-alkyl-substituted deoxyribonucleosides and 2'-O-ethylene glycol-like ribonucleosides. The preparation of these sugars or sugar analogs and the respective "nucleosides" in which such sugars or sugar analogs are linked to heterocyclic bases (nucleobases) is known. Sugar modifications can also be made using and combined with other modifications.
[0065] Modifications to the sugar moiety include natural modifications of ribose and deoxyribose, as well as non-natural modifications. Sugar modifications include, but are not limited to, modifications at the 2-position: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C6 10 Alkyl or C2-C 10 It may be alkenyl or alkynyl. 2' sugar modifications include, but are not limited to, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2)n NH2, -O(CH2) n CH3, -O(CH2) n ONH2 and -O(CH2) n ON[(CH2) n CH3)]2 (wherein n and m are from 1 to about 10).
[0066] Other modifications at the 2' position include, but are not limited to, C1 to C 10Examples of suitable sugars include lower alkyl, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl, O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving groups, reporter groups, intercalators, groups for improving the pharmacokinetic or pharmacodynamic properties of oligonucleotides, and other substituents with similar properties. Similar modifications can also be made at other positions on the sugar, particularly the 3' position of the sugar in the 3'-terminal nucleotide or in 2'-5'-linked oligonucleotides, and the 5' position of the 5'-terminal nucleotide. Modified sugars also include those containing modifications at the bridging ring oxygens, such as CH2 and S. Nucleotide sugar analogs can also have sugar mimetics, such as cyclobutyl moieties, in place of the pentofuranosyl sugar.U.S. Patent Nos. 4,981,957; 5,118,800; 5,319,080; 5,359,044; 5,393,878; 5,446,137; 5,466,786; 5,514,785; 5,519,134; 5,567,811; 5,576,42 No. 7; No. 5,591,722; No. 5,597,909; No. 5,610,300; No. 5,627,053; No. 5,639,873; No. 5 , 646,265; 5,658,873; 5,670,633; 4,845,205; 5,130,302; 5,134,06 There are numerous U.S. patents that teach the preparation of such modified sugar structures and detail and describe the range of base modifications, such as U.S. Patents Nos. 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121; 5,596,091; 5,614,617; 5,681,941; and 5,700,920, each of which is incorporated herein by reference in its entirety.
[0067] Examples of nucleic acids with modified sugar moieties include, but are not limited to, nucleic acids containing 5'-vinyl, 5'-methyl (R or S), 4'-S, 2'-F, 2'-OCH3, and 2'-O(CH2)2OCH3 substituents. Substituents at the 2' position include allyl, amino, azido, thio, O-allyl, O-(C1-C2). 1O alkyl), OCF3, O(CH2)2SCH3, O(CH2)2-ON(R m )(R n ) and O-CH2-C(=O)-N(R m )(R n ) (wherein each R m and R n are independently H or substituted or unsubstituted C1-C 10 It may also be selected from the group consisting of alkyl.
[0068] In certain embodiments, the nucleic acids described herein comprise one or more bicyclic nucleic acids. In certain such embodiments, the bicyclic nucleic acids comprise a bridge between the 4' and 2' ribosyl ring atoms. In certain embodiments, the nucleic acids provided herein comprise one or more bicyclic nucleic acids. The acid comprises one or more bicyclic nucleic acids, wherein the bridge comprises a 4'-2' bicyclic nucleic acid. Examples of such 4'-2' bicyclic nucleic acids include, but are not limited to, those of the formula: 4'-(CH2)-O-2' (LNA); 4'-(CH2)-S-2'; 4'-(CH2)2-O-2' (ENA); 4'-CH(CH3)-O-2' and 4'-CH(CHOCH3)-O-2', and analogs thereof (see U.S. Pat. No. 7,399,845); 4'-C(CH3)(C H3)-O-2' and analogs thereof (see WO 2009 / 006478, WO 2008 / 150729, U.S. Patent Application Publication No. 2004 / 0171570, U.S. Patent No. 7,427,672, Chattopadhyaya et al., J. Org. Chem., Vol. 209, No. 74, pp. 118-134, and WO 2008 / 154401).See, for example, Singh et al., Chem. Commun., 1998, Vol. 4, pp. 455-456; Koshkin et al., Tetrahedron, 1998, Vol. 54, pp. 3607-3630; Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 2000, Vol. 97, pp. 5633-5638; Kumar et al., Bioorg. Med. Chem. Lett., 1998, Vol. 8, pp. 2219-2222; Singh et al., J. Org. Chem., 1998, Vol. 63, pp. 10035-10039; Srivastava et al., J. Am. Chem. Soc., 2007, Vol. 129 (No. 26), pp. 8362-8379; Elayadi et al., Curr. Opinion Invens. Drugs, 2001, Vol. 2, pp. 558-561; Braasch et al., Chem. Biol., 2001, Vol. 8, pp. 1-7; Oram et al., Curr. Opinion Mol. Ther., 2001, Vol. 3, pp. 239-243; U.S. Patent Nos. 4,849,513; 5,015,733; 5,118,800; 5,118,802; 7,053,207; 6,268,490; 6,770,748; 6,794,499; 7,034,133; 6,525,191; 6,670,461; and 7,399,845; WO 2004 / 106356, WO 1994 / 14226, WO 2005 / 02 1570, WO 2007 / 090071 and WO 2007 / 134181; U.S. Patent Application Publication Nos. 2004 / 0171570, 2007 / 0287831 and 2008 / 0039618; U.S. Provisional Application Nos. 60 / 989,574, 61 / 026,995, 61 / 026,998, 61 / 056,564, 61 / 086,231, 61 / 097,787 and 61 / 099,844; and International Application Nos. PCT / US2008 / 064591, PCT See also US2008 / 066154, PCT US2008 / 068922 and PCT / DK98 / 00393.
[0069] In certain embodiments, nucleic acids include linked nucleic acids. Nucleic acids can be linked together using any internucleic acid linkage. Two major classes of internucleic acid linkage groups are defined by the presence or absence of a phosphorus atom. Representative phosphorus-containing internucleic acid linkages include, but are not limited to, phosphodiesters, phosphotriesters, methylphosphonates, phosphoramidates, and phosphorothioates (P=S). Representative non-phosphorus-containing internucleic acid linkages include, but are not limited to, methylenemethylimino (-CH2-N(CH3)-O-CH2-), thiodiesters (-OC(O)-S-), thionocarbamate (-OC(O)(NH)-S-); siloxanes (-O-Si(H)2-O-); and N,N * -dimethylhydrazine (-CH2-N(CH3)-N(CH3)). In certain embodiments, internucleic acid linkages having chiral atoms can be prepared as racemic mixtures or as separate enantiomers, such as alkylphosphonates and phosphorothioates. Non-natural nucleic acids can contain a single modification. Non-natural nucleic acids can contain multiple modifications within one moiety or between different moieties.
[0070] Backbone phosphate modifications to nucleic acids include, but are not limited to, methyl Phosphonates, phosphorothioates, phosphoramidates (bridged or unbridged), phosphotriesters, phosphorodithioates, phosphodithioates, and boranophosphates can be used in any combination. Other non-phosphate linkages can also be used.
[0071] In some embodiments, backbone modifications (e.g., methylphosphonate, phosphorothioate, phosphoramidate, and phosphorodithioate internucleotide linkages) can confer immunomodulatory activity to the modified nucleic acids and / or enhance their stability in vivo.
[0072] In some examples, the phosphorus derivative (or modified phosphate group) is attached to the sugar or sugar analog moiety and can be a monophosphate, diphosphate, triphosphate, alkylphosphonate, phosphorothioate, phosphorodithioate, phosphoramidate, etc. Exemplary polynucleotides containing modified phosphate or non-phosphate linkages are described in Peyrottes et al., 1996, Nucleic Acids Res. 24:1841-1848; Chaturvedi et al., 1996, Nucleic Acids Res. 24:2318-2323; and Schultz et al. (1996) Nucleic Acids Res. 24:2966-2973; Matteucci, 1997, "Oligonucleotide Analogs: an Overview" in Oligonucleotides as Therapeutic Agents, (Chadwick and Cardew, eds.) John Wiley and Sons, New York, NY; Zon, 1993, "Oligonucleoside Phosphorothioates" in Protocols for Oligonucleotides and Analogs, Synthesis and Properties, Humana Press, pp. 165-190; Miller et al., 1971, JACS 93:6657-6665; Jager et al., 1988, Biochem. 27:7247-7246; Nelson et al., 1997, JOC 62:7278-7287; U.S. Patent No. 5,453,496; and Micklefield, 2001, Curr. Med. Chem. 8:1157-1179.
[0073] In some cases, backbone modifications include replacing phosphodiester linkages with alternative moieties, such as anionic, neutral, or cationic groups. Examples of such modifications include anionic internucleoside linkages; N3'-P5' phosphoramidate modifications; boranophosphate DNA; prooligonucleotides; neutral internucleoside linkages, such as methylphosphonate; amide-linked DNA; methylene (methylimino) linkages; formacetal and thioformacetal linkages; backbones containing sulfonyl groups; morpholino oligos; peptide nucleic acids (PNAs); and positively charged deoxyribonucleic acid guanidine (DNG) oligos (Micklefield, 2001, Current Medicinal Chemistry 8:1157-1179). Modified nucleic acids can include chimeric or mixed backbones containing one or more modifications, such as a combination of phosphate linkages, such as a combination of phosphodiester and phosphorothioate linkages.
[0074] Alternatives to phosphate include, for example, short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages. These include those with morpholino linkages (formed in part from the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl backbones; These include cetyl and thioformacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, amide backbones, and others with mixed N, O, S, and CH2 moieties. Numerous U.S. patents disclose how to make and use these types of phosphate replacements, including, but not limited to, U.S. Patent Nos. 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; 5,434,257; 5,466,677; 5,47 0,967; 5,489,677; 5,541,307; 5,561,225; 5,596,086; 5,602,240; 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439.It is also understood that in nucleotide substitutes, both the sugar and phosphate moieties of nucleotide can be replaced, for example, by amide-type linkage (aminoethylglycine) (PNA). U.S. Patent No. 5,539,082; U.S. Patent No. 5,714,331; and U.S. Patent No. 5,719,262 teach the preparation and use of PNA molecules, each of which is incorporated herein by reference.See also Nielsen et al., Science, 1991, vol. 254, pp. 1497-1500.For example, to enhance cellular uptake, other types of molecules can be linked (conjugated) to nucleotide or nucleotide analogue.Conjugate can be chemically linked to nucleotide or nucleotide analogue.Such conjugates include, but are not limited to, lipid moieties such as cholesterol moieties (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, Vol. 86, pp. 6553-6556), cholic acid (Manoharan et al., Bioorg. Med. Chem. Let., 1994, Vol. 4, pp. 1053-1060), thioethers such as hexyl-S-tritylthiol (Manoharan et al., Ann. KY. Acad. Sci., 1992, Vol. 660, pp. 306-309; Manoharan et al., Bioorg. Med. Chem. Let., 1993, Vol. 3, pp. 2765-2770), thiocholesterol (Oberhauser et al., Nucl. Acids Res., 1992, Vol. 20, pp. 533-538), aliphatic chains, for example, dodecanediol or undecyl residues (Saison-Behmoaras et al., EM5OJ, 1991, Vol. 10, pp. 1111-1118; Kabanov et al., FEBS. Lett., 1990, vol. 259, pp. 327-330; Svinarchuk et al., Biochimie, 1993, vol. 75, pp. 49-54), phospholipids, such as di-hexadecyl-rac-glycerol or triethylammonium l-di-O-hexadecyl-rac-glycero-SH-phosphonate (Manoharan et al., Tetrahedron Lett., 1995, vol. 36, pp. 3651-3654; Shea et al., Nucl. Acids Res., 1990, vol. 18, pp. 3777-3783), polyamines or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides, 1995, vol. 14, pp. 969-973), or adamantane acetic acid (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651-3654), a palmityl moiety (Mishra et al., Biochem. Biophys. Acta, 1995, 1264, 229-237), or an octadecylamine or hexylamino-carbonyl-oxycholesterol moiety (Crooke et al., J. Pharmacol. Exp. Ther., 1996, 277, 923-937). Numerous United States patents teach the preparation of such conjugates, including, but not limited to, U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; and 5,578,717. , No. 5,580,731; No. 5,591,584; No. 5,109,124; No. 5,118,802; No. 5,138,045; No. 5,414,077 No. 5,486,603; No. 5,512,439; No. 5,578,718; No. 5,608,046; No. 4,587,044; No. 4,605,735 No. 4,667,025; No. 4,762,779; No. 4,789,737; No. 4,824,941; No. 4,835,263; No. 4,876,33 No. 5; No. 4,904,582; No. 4,958,013; No. 5,082,830; No. 5,112,963; No. 5,214,136; No. 5,245,02 No. 2; No. 5,254,469; No. 5,258,506; No. 5,262,536; No. 5,272,250; No. 5,292,873; No. 5,317,0 No. 98; No. 5,371,241, No. 5,391,723; No. 5,416,203, No. 5,451,463; No. 5,510,475; No. 5,512,6 Nos. 67; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.
[0075] Described herein are nucleobases used in compositions and methods for the replication, transcription, translation, and incorporation of unnatural amino acids into proteins. In some embodiments, the nucleobases described herein have the structure: [ka] wherein each X is independently carbon or nitrogen; R2 is optional and, when present, is independently hydrogen, alkyl, alkenyl, alkynyl, methoxy, methanethiol, methaneseleno, halogen, cyano, or azido group; each Y is independently sulfur, oxygen, selenium, or a secondary amine; each E is independently oxygen, sulfur, or selenium; the wavy line indicates the point of attachment to a ribosyl, deoxyribosyl, or dideoxyribosyl moiety, or an analog thereof, which is in free form, optionally attached to a monophosphate, diphosphate, or triphosphate group, including an α-thiotriphosphate, β-thiotriphosphate, or γ-thiotriphosphate group, or contained in RNA or DNA, or an RNA or DNA analog. In some embodiments, R2 is lower alkyl (e.g., C1-C6), hydrogen, or halogen. In some embodiments of the nucleobases described herein, R2 is fluoro. In some embodiments of the nucleobases described herein, X is carbon. In some embodiments of the nucleobases described herein, E is sulfur. In some embodiments of the nucleobases described herein, Y is sulfur. In some embodiments of the nucleobases described herein, the nucleobase has the structure: [ka] In some embodiments of the nucleobases described herein, E is sulfur and Y is sulfur. In some embodiments of the nucleobases described herein, the wavy line indicates the point of attachment to the ribosyl or deoxyribosyl moiety. In some embodiments of the nucleobases described herein, the wavy line indicates the point of attachment to the ribosyl or deoxyribosyl moiety that is connected to the triphosphate group. In some embodiments of the nucleobases described herein, the nucleobase is a component of a nucleic acid polymer. In some embodiments of the nucleobases described herein, the nucleobase is a component of a tRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of an anticodon in a tRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of an mRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of a codon in an mRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of RNA or DNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of a codon in DNA. In some embodiments of the nucleobases described herein, the nucleobase forms a nucleobase pair with another complementary nucleobase.
[0076] Nucleic acid base pairing In some embodiments, an unnatural nucleotide forms a base pair (unnatural base pair; UBP) with another unnatural nucleotide during or after incorporation into DNA or RNA. In some embodiments, a stably incorporated unnatural nucleic acid is a unnatural nucleic acid that can base pair with another nucleic acid, e.g., a natural or unnatural nucleic acid. In some embodiments, a stably incorporated unnatural nucleic acid is a unnatural nucleic acid that can base pair with another unnatural nucleic acid (unnatural nucleobase pair (UBP)). For example, a first unnatural nucleic acid can base pair with a second unnatural nucleic acid. For example, one pair of unnatural nucleoside triphosphates that can base pair during or after incorporation into a nucleic acid includes the triphosphate of (d)5SICS ((d)5SICSTP) and the triphosphate of (d)NaM ((d)NaMTP). Other examples include, but are not limited to, the triphosphate of (d)CNMO ((d)CNMOTP) and the triphosphate of (d)TPT3 ((d)TPT3TP). Such unnatural nucleotides can have a ribose or deoxyribose sugar moiety (indicated by "(d)"). For example, one pair of unnatural nucleoside triphosphates that can base pair when incorporated into a nucleic acid includes the triphosphate of TAT1 (TAT1TP) and the triphosphate of NaM (NaMTP). In some embodiments, one pair of unnatural nucleoside triphosphates that can base pair when incorporated into a nucleic acid includes the triphosphate of dCNMO (dCNMOTP) and the triphosphate of TAT1 (TAT1TP). In some embodiments, one pair of unnatural nucleoside triphosphates that can base pair when incorporated into a nucleic acid includes the triphosphate of dTPT3 (dTPT3TP) and the triphosphate of NaM (NaMTP). In some embodiments, unnatural nucleic acids do not substantially base pair with natural nucleic acids (A, T, G, C). In some embodiments, a stably integrated non-naturally occurring nucleic acid can base pair with a naturally occurring nucleic acid.
[0077] In some embodiments, a stably incorporated non-natural (deoxy)ribonucleotide is a non-natural (deoxy)ribonucleotide that can form a UBP but does not substantially base pair with any of the respective natural (deoxy)ribonucleotides. In some embodiments, a stably incorporated non-natural (deoxy)ribonucleotide is a non-natural (deoxy)ribonucleotide that can form a UBP but does not substantially base pair with one or more natural nucleic acids. For example, a stably incorporated non-natural nucleic acid is substantially unable to base pair with A, T, and C, but can base pair with G. For example, a stably incorporated non-natural nucleic acid is substantially unable to base pair with A, T, and G, but can base pair with C. For example, a stably incorporated non-natural nucleic acid is substantially unable to base pair with C, G, and A, but can base pair with T. For example, a stably incorporated non-natural nucleic acid is substantially unable to base pair with C, G, and T, but can base pair with A. For example, a stably integrated non-naturally occurring nucleic acid is substantially unable to base pair with A and T, but can base pair with C and G. For example, a stably integrated non-naturally occurring nucleic acid is substantially unable to base pair with A and C, but can base pair with T and G. For example, a stably integrated non-naturally occurring nucleic acid is substantially unable to base pair with A and G, but can base pair with C and T. For example, a stably integrated non-naturally occurring nucleic acid is substantially unable to base pair with C and T, but can base pair with A and G. For example, a stably integrated non-naturally occurring nucleic acid is substantially unable to base pair with C and G, but can base pair with T and G. For example, a stably integrated non-naturally occurring nucleic acid is substantially unable to base pair with T and G, but can base pair with A and G. For example, a stably integrated non-naturally occurring nucleic acid is substantially unable to base pair with G, but can base pair with A, T, and C.For example, a stably integrated non-naturally occurring nucleic acid may be substantially unable to base pair with A, but may be able to base pair with G, T, and C. For example, a stably integrated non-naturally occurring nucleic acid may be substantially unable to base pair with T, but may be able to base pair with G, A, and C. For example, a stably integrated non-naturally occurring nucleic acid may be substantially unable to base pair with C, but may be able to base pair with G, T, and A.
[0078] Exemplary unnatural nucleotides capable of forming unnatural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, 5SICS, d5SICS, NaM, dNaM, dTPT3, dMTMO, dCNMO, TAT1, and combinations thereof. In some embodiments, unnatural nucleotide base pairs include, but are not limited to, [ka] Examples include:
[0079] Engineered Organisms In some embodiments, the methods and plasmids disclosed herein are further used to generate engineered organisms, such as organisms that can incorporate and replicate unnatural nucleotides or unnatural nucleic acid base pairs (UBPs) and use nucleic acids containing unnatural nucleotides to transcribe mRNAs and tRNAs used to translate unnatural polypeptides or proteins containing at least one unnatural amino acid residue. In some cases, the unnatural amino acid residue is site-specifically incorporated into the unnatural polypeptide or protein. In some examples, the organism is a non-human semisynthetic organism (SSO). In some examples, the organism is a semisynthetic organism (SSO). In some examples, the SSO is a cell. In some examples, the in vivo method involves a semisynthetic organism (SSO). In some examples, the semisynthetic organism comprises a microorganism. In some examples, the organism comprises a bacterium. In some examples, the organism comprises a gram-negative bacterium. In some examples, the organism comprises a gram-positive bacterium. In some examples, the organism comprises E. coli. Such modified organisms may variously include additional components such as DNA repair machinery, modified polymerases, nucleotide transporters, or other components. In some examples, the SSO includes E. coli strain YZ3. In some examples, the SSO includes E. coli strain ML1 or ML2, e.g., Ledbetter et al., J. Am. Chem.Soc. 2018, Vol. 140 (No. 2), p. 758, Figure 1 (B-D). Optionally, the SSO is a cell line. Optionally, the cell line is an immortalized cell line. In some examples, the cell line comprises primary cells. In some examples, the cell line comprises stem cells. In some examples, the SSO is an organoid.
[0080] In some cases, the cells used are genetically transformed with an expression cassette encoding a heterologous protein, such as a nucleoside triphosphate transporter that can transport unnatural nucleoside triphosphates into cells, and optionally a CRISPR / Cas9 system for eliminating DNA that lacks unnatural nucleotides (e.g., E. coli strain YZ3, ML1 or ML2).In some cases, the cells further comprise enhanced activity for the uptake of unnatural nucleic acids.In some cases, the cells further comprise enhanced activity for the transfer of unnatural nucleic acids.
[0081] In some embodiments, Cas9 and a suitable guide RNA (sgRNA) are encoded on separate plasmids. In some examples, Cas9 and the sgRNA are encoded on the same plasmid. In some cases, Cas9, a nucleic acid molecule encoding the sgRNA, or a nucleic acid molecule comprising unnatural nucleotides are located on one or more plasmids. In some examples, Cas9 is encoded on a first plasmid, and the sgRNA and the nucleic acid molecule comprising unnatural nucleotides are encoded on a second plasmid. In some examples, Cas9, the sgRNA, and the nucleic acid molecule comprising unnatural nucleotides are encoded on the same plasmid. In some examples, the nucleic acid molecule comprises two or more unnatural nucleotides. In some examples, Cas9 is integrated into the genome of a host organism, and the sgRNA is encoded on a plasmid or in the genome of the organism.
[0082] In some examples, a first plasmid encoding Cas9 and an sgRNA and a second plasmid encoding a nucleic acid molecule comprising non-natural nucleotides are introduced into an engineered microorganism. In some examples, a first plasmid encoding Cas9 and a second plasmid encoding an sgRNA and a nucleic acid molecule comprising non-natural nucleotides are introduced into an engineered microorganism. In some examples, a plasmid encoding Cas9, an sgRNA, and a nucleic acid molecule comprising non-natural nucleotides is introduced into an engineered microorganism. In some examples, the nucleic acid molecule comprises two or more non-natural nucleotides. Contains nucleotides.
[0083] In some embodiments, living cells are generated that incorporate at least one non-natural nucleic acid molecule that comprises at least one unnatural base pair (UBP) within its DNA (plasmid or genome). Optionally, the at least one non-natural nucleic acid molecule comprises one, two, three, four, or more UBPs. In some examples, the at least one non-natural nucleic acid molecule is a plasmid. Optionally, the at least one non-natural nucleic acid molecule is integrated into the genome of the cell. In some embodiments, the at least one non-natural nucleic acid molecule encodes a non-natural polypeptide or protein. Optionally, the at least one non-natural nucleic acid molecule is transcribed to provide an unnatural codon for the mRNA and an unnatural anticodon for the tRNA. In some embodiments, the at least one non-natural nucleic acid molecule is a non-natural DNA molecule.
[0084] In some instances, the unnatural base pair comprises a pair of unnatural mutually base-pairing nucleotides that can form an unnatural base pair under in vivo conditions when the unnatural mutually base-pairing nucleotides, as their respective triphosphates, are taken up into a cell by the action of a nucleotide triphosphate transporter. The cell can be genetically transformed with an expression cassette encoding a nucleotide triphosphate transporter so that the nucleotide triphosphate transporter is expressed and available for transport of the unnatural nucleotide into the cell. The cell can be a prokaryotic or eukaryotic cell, and the pair of unnatural mutually base-pairing nucleotides, as their respective triphosphates, can be the triphosphate of dTPT3 (dTP3TP) and the triphosphate of dNaM (dNaMTP) or the triphosphate of dCNMO (dCNMOTP).
[0085] In some embodiments, the cell is a cell that is genetically transformed with a nucleic acid, e.g., an expression cassette encoding a nucleotide triphosphate transporter that can transport such unnatural nucleotides into the cell. The cell can include a heterologous nucleoside triphosphate transporter, where the heterologous nucleoside triphosphate transporter can transport natural and unnatural nucleoside triphosphates into the cell.
[0086] In some cases, the methods described herein also include contacting the genetically transformed cells with the respective triphosphates in the presence of potassium phosphate and / or a phosphatase or nucleotidase inhibitor. During or after such contact, the cells can be placed in a life-sustaining medium suitable for cell growth and replication. The cells can be maintained in the life-sustaining medium so that the respective triphosphate forms of the unnatural nucleotides are incorporated into nucleic acids within the cells through at least one replication cycle of the cells. The pair of unnatural mutually base-pairing nucleotides can include, as their respective triphosphates, a triphosphate of dTPT3 (dTPT3TP) and a triphosphate of dCNMO or dNaM (dCNOM or dNaMTP), the cell can be Escherichia coli, and dTPT3TP and dNaMTP can be imported into the E. coli by the transporter PtNTT2, and an E. coli polymerase, such as Pol III or Pol II, can use the unnatural triphosphate to replicate DNA containing the UBP, thereby incorporating the unnatural nucleotide and / or unnatural base pair into the nucleic acid of a cell within the cellular environment. Furthermore, ribonucleotides, such as NaMTP and TAT1TP, 5FMTP and TPT3TP, are imported into E. coli by the transporter PtNTT2 in some examples. In some examples, the PtNTT2 for importing a ribonucleotide is a truncated PtNTT2, wherein the truncated PtNTT2 has an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90% identical to the amino acid sequence of an untruncated PtNTT2. An example of PtNTT2 (NCBI accession number EEC49227.1, GI:217409295) has the following amino acid sequence (SEQ ID NO: 1): 1 MRPYPTIALI SVFLSAATRI SATSSHQASA LPVKKGTHVP 41 DSPKLSKLYI MAKTKSVSSS FDPPRGGSTV APTTPLATGG 81 ALRKVRQAVF PIYGNQEVTK FLLIGSIKFF IILALTLTRD 121 TKDTLIVTQC GAEAIAFLKI YGVLPAATAF IALYSKMSNA 161 MGKKMLFYST CIPFFTFFGL FDVFIYPNAE RLHPSLEAVQ 201 AILPGGAASG GMAVLAKIAT HWTSALFYVM AEIYSSVSVG 241 LLFWQFANDV VNVDQAKRFY PLFAQMSGLA PVLAGQYVVR 281 FASKAVNFEA SMHRLTAAVT FAGIMICIFY QLSSSYVERT 321 ESAKPAADNE QSIKPKKKKP KMSMVESGKF LASSQYLRLI 361 AMLVLGYGLS INFTEIMWKS LVKKQYPDPL DYQRFMGNFS 401 SAVGLSTCIV IFFGVHVIRL LGWKVGALAT PGIMAILALP 441 FFACILLGLD SAVINGS FGTIQSLLSK TSKYALFDPT 481 TQMAYIPLDD ESKVKGKAAI DVLGSRIGKS GGSLIQQGLV 521 FVFGNIINAA PVVGVVYYSV LVAWMSAAGR LSGLFQAQTE 561 MDCADKMEAK TNKEK
[0087] Described herein are compositions and methods that include the use of three or more unnatural base-pairing nucleotides. Such base-pairing nucleotides, in some cases, enter cells by the use of nucleotide transporters or by standard nucleic acid transformation methods known in the art (e.g., electroporation, chemical transformation, or other methods). In some cases, the base-pairing unnatural nucleotides enter cells as part of a polynucleotide, such as a plasmid. One or more base-pairing unnatural nucleotides that enter cells as part of a polynucleotide (RNA or DNA) do not themselves need to replicate in vivo. For example, a double-stranded DNA plasmid or other nucleic acid containing a first unnatural deoxyribonucleotide and a second unnatural deoxyribonucleotide with bases configured to form a first unnatural base pair is electroporated into cells. The cell culture medium is treated with a third unnatural deoxyribonucleotide and a fourth unnatural deoxyribonucleotide, each having a base configured to form a second unnatural base pair with the other, where the base of the first unnatural deoxyribonucleotide and the base of the third unnatural deoxyribonucleotide form a second unnatural base pair, and the base of the second unnatural deoxyribonucleotide and the base of the fourth unnatural deoxyribonucleotide form a third unnatural base pair. In some examples, in vivo replication of the initially transformed double-stranded DNA plasmid subsequently results in a replicated plasmid containing the third unnatural deoxyribonucleotide and the fourth unnatural deoxyribonucleotide. Alternatively, or in combination, ribonucleotide variants of the third unnatural deoxyribonucleotide and the fourth unnatural deoxyribonucleotide are added to the cell culture medium. In some examples, these ribonucleotides are incorporated into RNA, such as mRNA or tRNA. In some examples, the first, second, third, and fourth deoxynucleotides contain different bases. In some examples, the first, third, and fourth deoxynucleotides contain different bases. In some examples, the first and third deoxynucleotides contain the same base.
[0088] By practicing the methods of the present disclosure, one skilled in the art can obtain a population of living, growing cells that have at least one unnatural nucleotide and / or at least one unnatural base pair (UBP) in at least one nucleic acid maintained within at least some of the individual cells, where the at least one nucleic acid is stably propagated within the cells, and the cells exhibit a transcriptional regulation of the one or more unnatural nucleotides when contacted with (e.g., grown in the presence of) the unnatural nucleotides in a life-sustaining medium suitable for the growth and replication of the organism. A suitable nucleotide triphosphate transporter is expressed to provide cellular uptake of the phosphate form.
[0089] After transport into the cell by the nucleotide triphosphate transporter, the unnatural base-pairing nucleotide is incorporated into intracellular nucleic acids by the cellular machinery, for example, the cell's own DNA and / or RNA polymerase, heterologous polymerases, or polymerases evolved using directed evolution (Chen T, Romesberg FE, FEBS Lett. 2014 January 21; 588 (2): 219-29; Betz K et al., J Am Chem Soc. 2013 December 11; 135 (49): 18637-43). Unnatural nucleotides can be incorporated into cellular nucleic acids, such as genomic DNA, genomic RNA, mRNA, tRNA, structural RNA, microRNA, and self-replicating nucleic acids (e.g., plasmids, viruses, or vectors).
[0090] In some cases, genetically engineered cells are generated by introducing nucleic acids, such as heterologous nucleic acids, into cells. In some instances, the nucleic acids introduced into the cells are in the form of a plasmid. In some instances, the nucleic acids introduced into the cells are integrated into the genome of the cells. Any cell described herein is a host cell and can include an expression vector. In one embodiment, the host cell is a prokaryotic cell. In another embodiment, the host cell is Escherichia coli. In some embodiments, the cell contains one or more heterologous polynucleotides. Nucleic acid reagents can be introduced into microorganisms using various techniques. Non-limiting examples of methods used to introduce heterologous nucleic acids into various organisms include transformation, transfection, transduction, electroporation, ultrasound-mediated transformation, conjugation, particle bombardment, etc. In some instances, the addition of carrier molecules (e.g., bis-benzimidazolyl compounds, see, e.g., U.S. Pat. No. 5,595,899) can increase DNA uptake in cells typically considered difficult to transform by conventional methods. Conventional methods of transformation are readily available to those skilled in the art and can be found in Maniatis, T., E. Fritsch and J. Sambrook (1982) Molecular Cloning: a Laboratory Manual; Cold Spring Harbor Laboratory, Cold Spring Harbor, NY.
[0091] In some instances, genetic transformation can be achieved by using, but not limited to, direct introduction of expression cassettes in plasmids, viral vectors, viral nucleic acids, phage nucleic acids, phage, cosmids and artificial chromosomes, or by introducing genetic material into cells or carriers such as cationic liposomes.Such methods are available in the art and can be easily adapted for use in the methods described herein.Transfer vectors can be any nucleotide constructs (such as plasmids) used to deliver genes to cells, or as part of a general strategy for delivering genes, for example, as part of recombinant retroviruses or adenoviruses (Ram et al., Cancer Res. 53:83-88, (1993)). Suitable means for transfection, including viral vectors, chemical transfectants, or physico-mechanical methods such as electroporation and direct diffusion of DNA, are described, for example, in Wolff, JA et al., Science, Vol. 247, pp. 1465-1468, (1990); and Wolff, JA, Nature, Vol. 352, pp. 815-818, (1991).
[0092] For example, DNA encoding a nucleoside triphosphate transporter or polymerase expression cassette, and / or vector, can be expressed by, but not limited to, calcium-mediated transformation, electroporation, microinjection, lipofection, biolistics, or other methods. The vector can be introduced into the cell by any method, including transfection.
[0093] In some cases, the cell contains an unnatural nucleoside triphosphate incorporated into one or more nucleic acids within the cell. For example, the cell is a living cell that can incorporate at least one unnatural nucleotide into DNA or RNA maintained within the cell. The cell can also incorporate at least one unnatural base pair (UBP) comprising a pair of unnatural mutually base-pairing nucleotides into nucleic acids within the cell under in vivo conditions, where the unnatural mutually base-pairing nucleotides, e.g., their respective triphosphates, are taken up into the cell by the action of a nucleoside triphosphate transporter and the gene is present (e.g., introduced) into the cell by genetic transformation. For example, upon incorporation into nucleic acids maintained within the cell, dTPT3TP and dCNMOTP can form a stable unnatural base pair that is stably propagated by the DNA replication machinery of the organism when grown in a life-sustaining medium containing, for example, dTPT3TP and dCNMOTP.
[0094] In some cases, the cells can replicate nucleic acids containing unnatural nucleotides. Such methods can include genetically transforming the cells with an expression cassette encoding a nucleoside triphosphate transporter capable of transporting one or more unnatural nucleotides into the cells as their respective triphosphates under in vivo conditions. Alternatively, cells previously genetically transformed with an expression cassette capable of expressing the encoded nucleoside triphosphate transporter can be utilized. The method can also include contacting or exposing the genetically transformed cells to potassium phosphate and the respective triphosphate forms of at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides capable of forming an unnatural base pair (UBP)) in a life-sustaining medium suitable for cell growth and replication, and maintaining the transformed cells in the life-sustaining medium in the presence of the respective triphosphate forms of at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides capable of forming an unnatural base pair (UBP)) under in vivo conditions through at least one replication cycle of the cells.
[0095] In some embodiments, the cells comprise stably integrated unnatural nucleic acids. Some embodiments include cells (e.g., E. coli) that stably incorporate nucleotides other than A, G, T, and C into nucleic acids maintained within the cells. For example, the nucleotides other than A, G, T, and C are d5SICS, dCNMO, dNaM, and / or dTPT3, which, upon incorporation into the nucleic acid of the cells, can form stable unnatural base pairs within the nucleic acid. In one aspect, the unnatural nucleotides and unnatural base pairs are stably propagated by the organism's replication apparatus when an organism transformed with a gene for a triphosphate transporter is grown in a life-sustaining medium containing potassium phosphate and the triphosphate forms of d5SICS, dNaM, dCNMO, and / or dTPT3.
[0096] In some cases, the cell comprises an expanded genetic alphabet. The cell can comprise a stably integrated unnatural nucleic acid. In some embodiments, a cell with an expanded genetic alphabet comprises a unnatural nucleic acid containing an unnatural nucleotide that can pair with another unnatural nucleotide. In some embodiments, a cell with an expanded genetic alphabet comprises a unnatural nucleic acid that hydrogen bonds to another nucleic acid. In some embodiments, a cell with an expanded genetic alphabet comprises a unnatural nucleic acid that is not hydrogen bonded to another nucleic acid with which it is base-paired. In some embodiments, a cell with an expanded genetic alphabet comprises a unnatural nucleic acid containing an unnatural nucleotide that has a nucleobase that base-pairs with another unnatural nucleotide through hydrophobic and / or packing interactions. In some embodiments, a cell with an expanded genetic alphabet comprises a unnatural nucleic acid that has a nucleobase that base-pairs with another unnatural nucleotide through non-hydrogen bonding interactions. , including unnatural nucleic acids that base pair with another nucleic acid. A cell with an expanded genetic alphabet is a cell that can copy a homologous nucleic acid to form a nucleic acid that includes an unnatural nucleic acid. A cell with an expanded genetic alphabet is a cell that contains an unnatural nucleic acid that is base-paired with another unnatural nucleic acid (unnatural nucleobase pair (UBP)).
[0097] In some embodiments, cells form unnatural DNA base pairs (UBPs) from imported unnatural nucleotides under in vivo conditions. In some embodiments, the activity of potassium phosphate and / or phosphatase and / or nucleotidase inhibitors can facilitate the transport of unnatural nucleotides. The method includes the use of cells expressing a heterologous nucleoside triphosphate transporter. When such cells are contacted with one or more nucleoside triphosphates, the nucleoside triphosphates are transported into the cells. The cells are in the presence of potassium phosphate and / or phosphatase and nucleotidase inhibitors. Unnatural nucleoside triphosphates can be incorporated into nucleic acids within the cells by the cells' natural machinery (i.e., polymerases), e.g., by base pairing with each other to form unnatural base pairs within the cells' nucleic acids. In some embodiments, UBPs are formed between DNA and RNA nucleotides with unnatural bases.
[0098] In some embodiments, the UBP is incorporated into a cell or population of cells when exposed to the non-natural triphosphate. In some embodiments, the UBP is substantially always incorporated into a cell or population of cells when exposed to the non-natural triphosphate.
[0099] In some embodiments, inducing expression of a heterologous gene, e.g., a nucleoside triphosphate transporter (NTT), in a cell can result in slower cell growth and increased uptake of one or more unnatural triphosphates compared to the growth and uptake of the unnatural triphosphate in a cell without inducing expression of the heterologous gene. Uptake variously includes transport of nucleotides into the cell, such as by diffusion, osmosis, or via the action of a transporter. In some embodiments, inducing expression of a heterologous gene, e.g., an NTT, in a cell can result in increased cell growth and increased uptake of unnatural nucleic acids compared to the growth and uptake of the cell without inducing expression of the heterologous gene.
[0100] In some embodiments, the UBP is incorporated during logarithmic growth phase. In some embodiments, the UBP is incorporated during non-logarithmic growth phase. In some embodiments, the UBP is incorporated during substantially linear growth phase. In some embodiments, the UBP is stably incorporated into a cell or population of cells after growth over a period of time. For example, the UBP can be stably incorporated into a cell or population of cells after growth over at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or 50 or more replications. For example, the UBP can be stably incorporated into a cell or population of cells after growth for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours of growth. For example, the UBP can be stably incorporated into a cell or population of cells after growth for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or 31 days of growth. For example, the UBP can be stably integrated into a cell or population of cells after propagation for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months of growth. can be stably integrated into a cell or population of cells after propagation over 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 50 years of growth.
[0101] In some embodiments, the cell further utilizes an RNA polymerase to generate an mRNA containing one or more unnatural nucleotides. In some instances, the cell further utilizes a polymerase to generate a tRNA containing an anticodon comprising one or more unnatural nucleotides. In some instances, the tRNA is charged with an unnatural amino acid. In some instances, the unnatural anticodon of the tRNA pairs with the unnatural codon of the mRNA during translation such that an unnatural polypeptide or protein containing at least one unnatural amino acid is synthesized.
[0102] Natural and Unnatural Amino Acids As used herein, amino acid residue may refer to a molecule that contains both an amino group and a carboxyl group.Suitable amino acids include, but are not limited to, both the D-isomer and the L-isomer of naturally occurring amino acids, and non-naturally occurring amino acids prepared by organic synthesis or any other method.As used herein, the term amino acid includes, but is not limited to, α-amino acids, natural amino acids, unnatural amino acids, and amino acid analogs. The term "α-amino acid" may refer to a molecule that contains both an amino group and a carboxyl group attached to a carbon, referred to as the α-carbon. For example: [ka]
[0103] The term "β-amino acid" can refer to a molecule that contains both an amino group and a carboxyl group in the β configuration.
[0104] A "naturally occurring amino acid" can refer to any one of the 20 amino acids commonly found in peptides synthesized in nature, and are known by the single letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V.
[0105] The following table provides a summary of the properties of natural amino acids.
[0106] [Table 1]
[0107] "Hydrophobic amino acids" include small and large hydrophobic amino acids. "Small hydrophobic amino acids" may be glycine, alanine, proline, and analogs thereof. "Large hydrophobic amino acids" may be valine, leucine, isoleucine, phenylalanine, methionine, tryptophan, and analogs thereof. "Polar amino acids" may be serine, threonine, asparagine, glutamine, cysteine, tyrosine, and analogs thereof. "Charged amino acids" may be lysine, arginine, histidine, aspartic acid, glutamate, and analogs thereof.
[0108] An "amino acid analog" can be a molecule that is structurally similar to an amino acid and can be substituted for an amino acid in the formation of a peptidomimetic macrocycle. Amino acid analogs include, but are not limited to, beta-amino acids and amino acids in which the amino or carboxy group has been replaced with a similarly reactive group (e.g., replacement of a primary amine with a secondary or tertiary amine, or replacement of the carboxy group with an ester).
[0109] A "non-standard amino acid (ncAA)" or "unnatural amino acid" can be an amino acid that is not one of the 20 amino acids commonly found in peptides synthesized in nature and is known by the one-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. In some cases, non-natural amino acids are a subset of non-standard amino acids.
[0110] The amino acid analogue may comprise a β-amino acid analogue. Examples of β-amino acid analogues include, but are not limited to, cyclic β-amino acid analogues; β-alanine; (R)-β-phenylalanine; (R)-1,2,3,4-tetrahydro-isoquinoline-3-acetic acid; (R)-3-amino-4-(1-naphthyl)-butyric acid; (R)-3-amino-4-(2,4-dichlorophenyl)butyric acid; (R)-3-amino-4-(2-chlorophenyl)-butyric acid; (R)-3-amino-4-(2-cyanophenyl)-butyric acid; (R)-3-amino-4-(2-fluorophenyl)-butyric acid; (R)-3-Amino-4-(2-furyl)-butyric acid;(R)-3-Amino-4-(2-methylphenyl)-butyric acid;(R)-3-Amino-4-(2-naphthyl)-butyric acid;(R)-3-Amino-4-(2-thienyl)-butyric acid;(R)-3-Amino-4-(2-trifluoromethylphenyl)-butyric acid;(R)-3-Amino-4-(3,4-dichlorophenyl)butyric acid;(R)-3-Amino-4-(3,4-difluorophenyl)butyric acid;(R)-3-Amino-4-(3-benzothienyl)-butyric acid;(R)-3-Amino-4-(3-chlorophenyl)butyric acid )-butyric acid;(R)-3-amino-4-(3-cyanophenyl)-butyric acid;(R)-3-amino-4-(3-fluorophenyl)-butyric acid;(R)-3-amino-4-(3-methylphenyl)-butyric acid;(R)-3-amino-4-(3-pyridyl)-butyric acid;(R)-3-amino-4-(3-thienyl)-butyric acid;(R)-3-amino-4-(3-trifluoromethylphenyl)-butyric acid;(R)-3-amino-4-(4-bromophenyl)-butyric acid;(R)-3-amino-4-(4-chlorophenyl)-butyric acid;(R)-3-amino-4-(4-sia (R)-3-amino-4-(4-fluorophenyl)-butyric acid;(R)-3-amino-4-(4-iodophenyl)-butyric acid;(R)-3-amino-4-(4-methylphenyl)-butyric acid;(R)-3-amino-4-(4-nitrophenyl)-butyric acid;(R)-3-amino-4-(4-pyridyl)-butyric acid;(R)-3-amino-4-(4-trifluoromethylphenyl)-butyric acid;(R)-3-amino-4-pentafluoro-phenylbutyric acid;(R)-3-amino-5-hexenoic acid;(R)-3-amino-5-hexynoic acid(R)-3-Amino-5-phenylpentanoic acid;(R)-3-Amino-6-phenyl-5-hexenoic acid;(S)-1,2,3,4-Tetrahydro-isoquinoline-3-acetic acid;(S)-3-Amino-4-(1-naphthyl)-butyric acid;(S)-3-Amino-4-(2,4-dichlorophenyl)butyric acid;(S)-3-Amino-4-(2-chlorophenyl)-butyric acid;(S)-3-Amino-4-(2-cyanophenyl)-butyric acid;(S)-3-Amino-4-(2-fluorophenyl)-butyric acid;(S)-3-Amino-4-(2-furyl)-butyric acid;(S)-3-Amino (S)-3-amino-4-(2-methylphenyl)-butyric acid;(S)-3-amino-4-(2-naphthyl)-butyric acid;(S)-3-amino-4-(2-thienyl)-butyric acid;(S)-3-amino-4-(2-trifluoromethylphenyl)-butyric acid;(S)-3-amino-4-(3,4-dichlorophenyl)butyric acid;(S)-3-amino-4-(3,4-difluorophenyl)butyric acid;(S)-3-amino-4-(3-benzothienyl)-butyric acid;(S)-3-amino-4-(3-chlorophenyl)-butyric acid;(S)-3-amino-4-(3-cyanophenyl)-butyric acid -3-Amino-4-(3-fluorophenyl)-butyric acid;(S)-3-Amino-4-(3-methylphenyl)-butyric acid;(S)-3-Amino-4-(3-pyridyl)-butyric acid;(S)-3-Amino-4-(3-thienyl)-butyric acid;(S)-3-Amino-4-(3-trifluoromethylphenyl)-butyric acid;(S)-3-Amino-4-(4-bromophenyl)-butyric acid;(S)-3-Amino-4-(4-chlorophenyl)butyric acid;(S)-3-Amino-4-(4-cyanophenyl)-butyric acid;(S)-3-Amino-4-(4-fluorophenyl)butyric acid;(S)- 3-Amino-4-(4-iodophenyl)-butyric acid;(S)-3-Amino-4-(4-methylphenyl)-butyric acid;(S)-3-Amino-4-(4-nitrophenyl)-butyric acid;(S)-3-Amino-4-(4-pyridyl)-butyric acid;(S)-3-Amino-4-(4-trifluoromethylphenyl)-butyric acid;(S)-3-Amino-4-pentafluoro-phenylbutyric acid;(S)-3-Amino-5-hexenoic acid;(S)-3-Amino-5-hexynoic acid;(S)-3-Amino-5-phenylpentanoic acid;(S)-3-Amino-6-phenyl-5-hexenoic acid1,2,5,6-Tetrahydropyridine-3-carboxylic acid; 1,2,5,6-Tetrahydropyridine; 3-Amino-3-(2-chlorophenyl)-propionic acid;3-Amino-3-(2-thienyl)-propionic acid;3-Amino-3-(3-bromophenyl)-propionic acid;3-Amino-3-(4-chlorophenyl)-propionic acid;3-Amino-3-(4-methoxyphenyl)-propionic acid;3-Amino-4,4,4-trifluorobutyric acid;3-Aminoadipic acid;D-β-phenyladipic acid Lanine;β-Leucine;L-β-Homoalanine;L-β-Homoaspartic acid γ-benzyl ester;L-β-Homoglutamic acid δ-benzyl ester;L-β-Homoisoleucine;L-β-Homoleucin;L-β-Homomethionine;L-β-Homophenylalanine;L-β-Homoproline;L-β-Homotryptophan;L-β-Homovaline;L-Nω-Benzyloxycarbonyl-β-Homolysine;Nω-L- β-Homoarginine;O-benzyl-L-β-homohydroxyproline;O-benzyl-L-β-homoserine;O-benzyl-L-β-homothreonine;O-benzyl-L-β-homotyrosine;γ-Trityl-L-β-homoasparagine;(R)-β-Phenylalanine;L-β-Homoaspartic acid γ-t-butyl ester;L-β-Homoglutamic acid δ-t-butyl ester;L-Nω-β-Homolysine;Nδ -trityl-L-β-homoglutamine; Nω-2,2,4,6,7-pentamethyl-dihydrobenzofuran-5-sulfonyl-L-β-homoarginine; Ot-butyl-L-β-homohydroxyproline; Ot-butyl-L-β-homoserine; Ot-butyl-L-β-homothreonine; Ot-butyl-L-β-homotyrosine; 2-aminocyclopentanecarboxylic acid; and 2-aminocyclohexanecarboxylic acid.
[0111] Amino acid analogs may include analogs of alanine, valine, glycine, or leucine. Examples of amino acid analogs of alanine, valine, glycine, and leucine include, but are not limited to, α-methoxyglycine, α-allyl-L-alanine, α-aminoisobutyric acid, α-methyl-leucine, β-(1-naphthyl)-D-alanine, β-(1-naphthyl)-L-alanine, β-(2-naphthyl)-D-alanine, β-(2-naphthyl)-L-alanine, β-(2-pyridyl)-D-alanine, β-(2-pyridyl)-L-alanine. β-(2-Thienyl)-D-alanine;β-(2-Thienyl)-L-alanine;β-(3-Benzothienyl)-D-alanine;β-(3-Benzothienyl)-L-alanine;β-(3-Pyridyl)-D-alanine;β-(3-Pyridyl)-L-alanine;β-(4-Pyridyl)-D-alanine;β-(4-Pyridyl)-L-alanine;β-Chloro-L-alanine;β-Cyano-L-alanine;β-Cyclohexyl-D-alanine;β-Cyclohexyl-L- Alanine; β-Cyclopenten-1-yl-alanine; β-Cyclopentyl-alanine; β-Cyclopropyl-L-Ala-OH; Dicyclohexylammonium salt; β-t-Butyl-D-alanine; β-t-Butyl-L-alanine; γ-Aminobutyric acid; L-α,β-Diaminopropionic acid; 2,4-Dinitrophenylglycine; 2,5-Dihydro-D-phenylglycine; 2-Amino-4,4,4-trifluorobutyric acid; 2-Fluoro-phenylglycine; 3-Amino 4,4,4-Trifluorobutyric acid; 3-Fluorovaline; 4,4,4-Trifluorovaline; 4,5-Dehydro-L-leu-OH; Dicyclohexylammonium salt; 4-Fluoro-D-phenylglycine; 4-Fluoro-L-phenylglycine; 4-Hydroxy-D-phenylglycine; 5,5,5-Trifluoroleucine; 6-Aminohexanoic acid; Cyclopentyl-D-Gly-OH; Dicyclohexylammonium salt; Cyclopentyl-Gly-OH.Dicyclohexylammonium salt; D-α,β-diaminopropionic acid; D-α-aminobutyric acid; D-α-t-butylglycine; D-(2-thienyl)glycine; D-(3-thienyl)glycine; D-2-aminocaproic acid; D-2-indanylglycine; D-allylglycine-dicyclohexylammonium salt; D-cyclohexylglycine; D-norvaline; D-phenylglycine; β-aminobutyric acid; β-aminoisobutyric acid; (2-bromophenyl)glycine; (2-methoxyphenyl)glycine; (2-methylphenyl)glycine; (2-thiazolyl)glycine; (2-thienyl)glycine; 2-amino-3-(dimethylamino)-propionic acid; L-α,β-diaminopropionic acid; L-α-aminobutyric acid; L. -α-t-Butylglycine;L-(3-Thienyl)glycine;L-2-Amino-3-(dimethylamino)-propionic acid;L-2-Aminocaproic acid dicyclohexylammonium salt;L-2-Indanylglycine;L-Allylglycine dicyclohexylammonium salt;L-Cyclohexylglycine;L-Phenylglycine;L-Propargylglycine;L-Norvaline;N-α-Aminomethyl-L-alanine;D-α,γ-Diaminobutyric acid;L-α,γ-Diamino Aminobutyric acid;β-Cyclopropyl-L-alanine;(N-β-(2,4-dinitrophenyl))-L-α,β-diaminopropionic acid;(N-β-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-D-α,β-diaminopropionic acid;(N-β-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-L-α,β-diaminopropionic acid;(N-β-4-methyltrityl)-L-α,β-diamino Propionic acid;(N-β-Allyloxycarbonyl)-L-α,β-diaminopropionic acid;(N-γ-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-D-α,γ-diaminobutyric acid;(N-γ-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-L-α,γ-diaminobutyric acid;(N-γ-4-Methyltrityl)-D-α,γ-diaminobutyric acid;(N-γ-4-Methyltrityl)-L-α,γ-diaminobutyric acid Aminobutyric acid; (N-γ-allyloxycarbonyl)-L-α,γ-diaminobutyric acid; D-α,γ-diaminobutyric acid; 4,5-dehydro-L-leucine; cyclopentyl-D-Gly-OH; cyclopentyl-Gly-OH; D-allylglycine; D-homocyclohexylalanine; L-1-pyrenylalanine; L-2-aminocaproic acid; L-allylglycine; L-homocyclohexylalanine; and N-(2-hydroxy-4-methoxy-Bzl)-Gly-OH.
[0112] The amino acid analogs may include arginine or lysine analogs. Examples of arginine and lysine amino acid analogs include, but are not limited to, citrulline; L-2-amino-3-guanidinopropionic acid; L-2-amino-3-ureidopropionic acid; L-citrulline; Lys(Me)-OH; Lys(N)-OH; Nδ-benzyloxycarbonyl-L-ornithine; Nω-nitro-D-arginine; Nω-nitro-L-arginine; α-methylornithine; 2,6-diaminoheptanedioic acid; L-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxo-cyclohex-1-ylidene)-methyl)-2,6-diaminoheptanedioic acid; )ethyl)-D-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxo-cyclohexen-1-ylidene)ethyl)-L-ornithine; (Nδ-4-methyltrityl)-D-ornithine; (Nδ-4-methyltrityl)-L-ornithine; D-ornithine; L-ornithine; Arg(Me)(Pbf)-OH; Arg(Me)-OH (asymmetric); Arg(Me)-OH (symmetric); Lys(ivDde)-OH; Lys(Me)-OH.HCl; Lys(Me)-OH chloride; Nω-nitro-D-arginine; and Nω-nitro-L-arginine.
[0113] Amino acid analogs may include analogs of aspartic acid or glutamic acid. Examples of amino acid analogs of aspartic acid and glutamic acid include, but are not limited to, α-methyl-D-aspartic acid; α-methyl-glutamic acid; α-methyl-L-aspartic acid; γ-methylene-glutamic acid; (N-γ-ethyl)-L-glutamine; [N-α-(4-aminobenzoyl)]-L-glutamic acid; 2,6-diaminopimelic acid; L-α-aminosuberic acid; D-2-aminoadipic acid; D-α-aminosuberic acid; α-aminopimelic acid; iminodiacetic acid; L-2-aminoadipic acid; threo-β-methyl-aspartic acid; γ-carboxy-D-glutamic acid γ,γ-di-t-butyl ester; γ-carboxy-L-glutamic acid γ,γ-di-t-butyl ester; Glu(OAll)-OH; L-Asu(OtBu)-OH; and pyroglutamic acid.
[0114] Amino acid analogs may include analogs of cysteine and methionine. Cysteine and examples of amino acid analogs of methionine include, but are not limited to, Cys(farnesyl)-OH, Cys(farnesyl)-OMe, α-methyl-methionine, Cys(2-hydroxyethyl)-OH, Cys(3-aminopropyl)-OH, 2-amino-4-(ethylthio)butyric acid, buthionine, buthionine sulfoximine, ethionine, methionine methylsulfonium chloride, selenomethionine, cysteic acid, [2-(4-pyridyl)ethyl]-DL-penicillamine, [2-(4-pyridyl)ethyl]-L-cysteine, 4-methoxybenzyl-D-penicillamine, 4-methoxybenzyl-L-penicillamine, 4-methylbenzyl-D-penicillamine. amine, 4-methylbenzyl-L-penicillamine, benzyl-D-cysteine, benzyl-L-cysteine, benzyl-DL-homocysteine, carbamoyl-L-cysteine, carboxyethyl-L-cysteine, carboxymethyl-L-cysteine, diphenylmethyl-L-cysteine, ethyl-L-cysteine, methyl-L-cysteine, t-butyl-D-cysteine, trityl-L-homocysteine, trityl-D-penicillamine, cystathionine, homocystine, L-homocystine, (2-aminoethyl)-L-cysteine, seleno-L-cystine, cystathionine, Cys(StBu)-OH, and acetamidomethyl-D-penicillamine.
[0115] Amino acid analogs may include analogs of phenylalanine and tyrosine. Examples of amino acid analogs of phenylalanine and tyrosine include β-methyl-phenylalanine, β-hydroxyphenylalanine, α-methyl-3-methoxy-DL-phenylalanine, α-methyl-D-phenylalanine, α-methyl-L-phenylalanine, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, 2,4-dichloro-phenylalanine, 2-(trifluoromethyl)-D-phenylalanine, 2-(trifluoromethyl)-L-phenylalanine, 2-bromo-D-phenylalanine, 2-bromo-L-phenylalanine, 2-chloro-D-phenylalanine, 2-chloro-L-phenylalanine, 2-cyano-D-phenylalanine, 2-cyano-L-phenylalanine, and 2-fluoro-D-phenylalanine. , 2-fluoro-L-phenylalanine, 2-methyl-D-phenylalanine, 2-methyl-L-phenylalanine, 2-nitro-D-phenylalanine, 2-nitro-L-phenylalanine, 2,4,5-trihydroxy-phenylalanine, 3,4,5-trifluoro-D-phenylalanine, 3,4,5-trifluoro-L-phenylalanine, 3,4-dichloro-D-phenylalanine, 3,4-dichloro-L-phenylalanine, 3,4-difluoro-D-phenylalanine, 3,4-difluoro-L-phenylalanine, 3,4-dihydroxy-L-phenylalanine, 3,4-dimethoxy-L-phenylalanine, 3,5,3'-triiodo-L-thyronine, 3,5-diiodo-D-tyrosine, 3,5-diiodo-L-tyrosine, 3,5-Diiodo-L-thyronine, 3-(trifluoromethyl)-D-phenylalanine, 3-(trifluoromethyl)-L-phenylalanine, 3-amino-L-tyrosine, 3-bromo-D-phenylalanine, 3-bromo-L-phenylalanine, 3-chloro-D-phenylalanine, 3-chloro-L-phenylalanine, 3-chloro-L-tyrosine, 3-cyano-D-phenylalanine, 3-cyano-L-phenylalanine, 3-fluoro-D-phenylalanine, 3-fluoro-L-phenylalanine, 3-fluoro-tyrosine, 3-iodo-D-phenylalanine, 3-iodo-L-phenylalanine, 3-iodo-L-tyrosine, 3-methoxy-L-tyrosine, 3-methyl-D-phenylalanine, 3-methyl 4-(trifluoromethyl)-L-phenylalanine, 3-nitro-D-phenylalanine, 3-nitro-L-phenylalanine, 3-nitro-L-tyrosine, 4-(trifluoromethyl)-D-phenylalanine, 4-(trifluoromethyl)-L-phenylalanine, 4-amino-D-phenylalanine, 4-amino-L-phenylalanine, 4-benzoyl-D-phenylalanine, 4-benzoyl-L-phenylalanine, 4-bis(2-chloroethyl)amino-L-phenylalanine, 4-bromo-D-phenylalanine, 4-bromo-L-phenylalanine, 4-chloro-D-phenylalanine, 4-chloro-L-phenylalanine, 4-cyano-D-phenylalanine, 4-cyano-L-phenylalanine, 4-fluoro-D-, These include phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-D-phenylalanine, 4-iodo-L-phenylalanine, homophenylalanine, thyroxine, 3,3-diphenylalanine, thyronine, ethyl-tyrosine, and methyltyrosine.
[0116] Amino acid analogs may include analogs of proline. Examples of amino acid analogs of proline include, but are not limited to, 3,4-dehydro-proline, 4-fluoro-proline, cis-4-hydroxy-proline, thiazolidine-2-carboxylic acid, and trans-4-fluoro-proline.
[0117] Amino acid analogue can include serine and threonine analogue.The example of serine and threonine amino acid analogue includes but is not limited to 3-amino-2-hydroxy-5-methylhexanoic acid, 2-amino-3-hydroxy-4-methylpentanoic acid, 2-amino-3-ethoxybutanoic acid, 2-amino-3-methoxybutanoic acid, 4-amino-3-hydroxy-6-methylheptanoic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-ethoxypropionic acid, 4-amino-3-hydroxybutanoic acid and α-methylserine.
[0118] Amino acid analogs may include analogs of tryptophan. Examples of tryptophan amino acid analogs include, but are not limited to, α-methyl-tryptophan, β-(3-benzothienyl)-D-alanine, β-(3-benzothienyl)-L-alanine, 1-methyl-tryptophan, 4-methyl-tryptophan, 5-benzyloxy-tryptophan, 5-bromo-tryptophan, 5-chloro-tryptophan, 5-fluoro-tryptophan, 5-hydroxytryptophan, 5-hydroxy-L-tryptophan, 5-methoxy-tryptophan, 5-methoxy-L-tryptophan, 5-methyl-tryptophan, 6 -Bromo-tryptophan; 6-chloro-D-tryptophan; 6-chloro-tryptophan; 6-fluoro-tryptophan; 6-methyl-tryptophan; 7-benzyloxy-tryptophan; 7-bromo-tryptophan; 7-methyl-tryptophan; D-1,2,3,4-tetrahydro-norharman-3-carboxylic acid; 6-methoxy-1,2,3,4-tetrahydronorharman-1-carboxylic acid; 7-azatryptophan; L-1,2,3,4-tetrahydro-norharman-3-carboxylic acid; 5-methoxy-2-methyl-tryptophan; and 6-chloro-L-tryptophan.
[0119] The amino acid analog may be racemic. In some cases, the D-isomer of the amino acid analog is used. In some cases, the L-isomer of the amino acid analog is used. In some cases, the amino acid analog contains a chiral center in the R or S configuration. Sometimes, the amino group of the β-amino acid analog is substituted with a protecting group, such as tert-butyloxycarbonyl (BOC group), 9-fluorenylmethyloxycarbonyl (FMOC), tosyl, etc. Sometimes, the carboxylic acid functional group of the β-amino acid analog is protected, for example, as its ester derivative. In some cases, a salt of the amino acid analog is used.
[0120] In some embodiments, the unnatural amino acid is an unnatural amino acid described in Liu CC, Schultz, PG Annu. Rev. Biochem. 2010, 79, 413. In some embodiments, the unnatural amino acid comprises N6(2-azidoethoxy)-carbonyl-L-lysine.
[0121] In some embodiments, an amino acid residue described herein (e.g., in a protein) is mutated to a non-natural amino acid prior to attachment to a conjugate moiety. Mutation to an unnatural amino acid prevents or minimizes the immune system's autoantigen response. As used herein, the term "unnatural amino acid" refers to an amino acid other than the 20 amino acids naturally occurring in proteins. Non-limiting examples of unnatural amino acids include p-acetyl-L-phenylalanine, p-iodo-L-phenylalanine, p-methoxyphenylalanine, O-methyl-L-tyrosine, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, L-3-(2-naphthyl)alanine, 3-methyl-phenylalanine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, tri-O-acetyl-GlcNAcp-serine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-azido-L-phenylalanine. p-Azido-phenylalanine, p-benzoyl-L-phenylalanine, p-boronophenylalanine, O-propargyltyrosine, L-phosphoserine, phosphonoserine, phosphonotyrosine, p-bromophenylalanine, selenocysteine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, N6-(propargyloxy)-carbonyl-L-lysine (Prk), azido-lysine (N6-azidoethoxy-carbonyl-L-lysine, AzK), N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine non-natural analogs of lysine and tyrosine amino acids; non-natural analogs of glutamine amino acids; non-natural analogs of phenylalanine amino acids; non-natural analogs of serine amino acids; non-natural analogs of threonine amino acids; alkyl, aryl, acyl, azido, cyano, halo, hydrazine, hydrazide, hydroxyl, alkenyl, alkynyl, ether, thiol, sulfonyl, seleno, ester, thioacid, borate, boronate, phospho, phosphono, phosphine, heterocyclic, enone, imine, aldehyde, hydroxylamine, keto, or amino substituted amino acids, or combinations thereof; amino acids with photoactivatable crosslinkers; spin-labeled amino acids; fluorescent amino acids; metal-binding amino acids;Metal-containing amino acids; radioactive amino acids; photocaged and / or photoisomerizable amino acids; biotin or biotin analog containing amino acids; keto containing amino acids; polyethylene glycol or polyether containing amino acids; heavy atom substituted amino acids; chemically cleavable or photocleavable amino acids; amino acids with elongated side chains; amino acids containing toxic groups; sugar-substituted amino acids; carbon-bonded sugar-containing amino acids; redox-active amino acids; α-hydroxy-containing acids; aminothioacids; α,α-disubstituted amino acids; β-amino acids; cyclic amino acids other than proline or histidine, and aromatic amino acids other than phenylalanine, tyrosine, or tryptophan;
[0122] In some embodiments, the unnatural amino acid comprises a selectively reactive group or a reactive group for site-selective labeling of a target protein or polypeptide. In some cases, the chemical reaction is a biorthogonal reaction (e.g., a biocompatible and selective reaction). In some cases, the chemical reaction is a Cu(I)-catalyzed or "copper-free" alkyne azide triazole-forming reaction, Staudinger ligation, inverse electron demand Diels-Alder (IEDDA) reaction, "photoclick" chemistry, or a metal-mediated process such as olefin metathesis and Suzuki-Miyaura or Sonogashira cross-coupling. In some embodiments, the unnatural amino acid comprises a photoreactive group that crosslinks upon irradiation, e.g., with UV light. In some embodiments, the unnatural amino acid comprises a photocaged amino acid. In some cases, the unnatural amino acid is a para-substituted, meta-substituted, or ortho-substituted amino acid derivative.
[0123] In some cases, the unnatural amino acid is p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, O-methyl-L-tyrosine, p-methoxyphenylalanine, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, L-3-(2-naphthyl)alanine, 3-methyl-phenylalanine, O-4-allyl-L-tyrosine, 4-propyl-L -tyrosine, tri-O-acetyl-GlcNAcp-serine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, L-phosphoserine, phosphonoserine, phosphonotyrosine, p-bromophenylalanine, p-amino-L-phenylalanine, or isopropyl-L-phenylalanine.
[0124] In some cases, the unnatural amino acid is 3-aminotyrosine, 3-nitrotyrosine, 3,4-dihydroxy-phenylalanine, or 3-iodotyrosine. In some cases, the unnatural amino acid is phenylselenocysteine. In some cases, the unnatural amino acid is a benzophenone, ketone, iodide, methoxy, acetyl, benzoyl, or azide (including a phenylalanine derivative). In some cases, the unnatural amino acid is a benzophenone, ketone, iodide, methoxy, acetyl, benzoyl, or azide-containing lysine derivative. In some cases, the unnatural amino acid comprises an aromatic side chain. In some cases, the unnatural amino acid does not comprise an aromatic side chain. In some cases, the unnatural amino acid comprises an azide group. In some cases, the unnatural amino acid comprises a Michael acceptor group. In some cases, the Michael acceptor group comprises an unsaturated moiety capable of forming a covalent bond via a 1,2-addition reaction. In some cases, the Michael acceptor group comprises an electron-deficient alkene or alkyne. In some cases, the Michael acceptor group includes, but is not limited to, alpha, beta unsaturated: ketone, aldehyde, sulfoxide, sulfone, nitrile, imine, or aromatic. In some cases, the unnatural amino acid is dehydroalanine. In some cases, the unnatural amino acid comprises an aldehyde or ketone group. In some cases, the unnatural amino acid is a lysine derivative comprising an aldehyde or ketone group. In some cases, the unnatural amino acid is a lysine derivative comprising one or more O, N, Se, or S atoms at the beta, gamma, or delta position. In some cases, the unnatural amino acid is a lysine derivative comprising an O, N, Se, or S atom at the gamma position. In some cases, the unnatural amino acid is a lysine derivative in which the epsilon N atom is replaced with an oxygen atom. In some cases, the unnatural amino acid is a lysine derivative that is a non-naturally occurring post-translationally modified lysine.
[0125] In some cases, the unnatural amino acid is an amino acid that includes a side chain and the sixth atom from the alpha position includes a carbonyl group. In some cases, the unnatural amino acid is an amino acid that includes a side chain and the sixth atom from the alpha position includes a carbonyl group and the fifth atom from the alpha position is nitrogen. In some cases, the unnatural amino acid is an amino acid that includes a side chain and the seventh atom from the alpha position is an oxygen atom.
[0126] In some cases, the unnatural amino acid is a serine derivative that includes selenium. In some cases, the unnatural amino acid is selenoserine (2-amino-3-hydroselenopropanoic acid). In some cases, the unnatural amino acid is 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid. In some cases, the unnatural amino acid is 2-amino-3-(phenylselanyl)propanoic acid. In some cases, the unnatural amino acid includes selenium, and oxidation of the selenium results in the formation of an alkene-containing unnatural amino acid.
[0127] In some cases, the unnatural amino acid comprises a cyclooctynyl group. In some cases, the unnatural amino acid comprises a transcycloctenyl group. In some cases, the unnatural amino acid comprises a norbornenyl group. In some cases, the unnatural amino acid comprises a cyclopropenyl group. In some cases, the unnatural amino acid comprises a diazirine group. In some cases, the unnatural amino acid comprises a tetrazine group.
[0128] In some cases, the unnatural amino acid is a lysine derivative in which the side chain nitrogen is carbamylated. In some cases, the unnatural amino acid is a lysine derivative, and the side chain nitrogen is acylated. In some cases, the unnatural amino acid is 2-amino-6-{[(tert-butoxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(tert-butoxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-Boc-N6-methyllysine. In some cases, the unnatural amino acid is N6-acetyllysine. In some cases, the unnatural amino acid is pyrrolysine. In some cases, the unnatural amino acid is N6-trifluoroacetyllysine. In some cases, the unnatural amino acid is 2-amino-6-{[(benzyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(p-iodobenzyloxy)carbonyl]amino}hexanoic acid. In some instances, the unnatural amino acid is 2-amino-6-{[(p-nitrobenzyloxy)carbonyl]amino}hexanoic acid. In some instances, the unnatural amino acid is N6-prolyl lysine. In some instances, the unnatural amino acid is 2-amino-6-{[(cyclopentyloxy)carbonyl]amino}hexanoic acid. In some instances, the unnatural amino acid is N6-(cyclopentanecarbonyl)lysine. In some instances, the unnatural amino acid is N6-(tetrahydrofuran-2-carbonyl)lysine. In some instances, the unnatural amino acid is N6-(3-ethynyltetrahydrofuran-2-carbonyl)lysine. In some instances, the unnatural amino acid is N6-((prop-2-yn-1-yloxy)carbonyl)lysine. In some instances, the unnatural amino acid is 2-amino-6-{[(2-azidocyclopentyloxy)carbonyl]amino}hexanoic acid. In some instances, the unnatural amino acid is N6-((2-azidoethoxy)carbonyl)lysine. In some instances, the unnatural amino acid is 2-amino-6-{[(2-nitrobenzyloxy)carbonyl]amino}hexanoic acid. In some instances, the unnatural amino acid is 2-amino-6-{[(2-cyclooctynyloxy)carbonyl]amino}hexanoic acid.In some instances, the unnatural amino acid is N6-(2-aminobut-3-ynoyl)lysine. In some instances, the unnatural amino acid is 2-amino-6-((2-aminobut-3-ynoyl)oxy)hexanoic acid. In some instances, the unnatural amino acid is N6-(allyloxycarbonyl)lysine. In some instances, the unnatural amino acid is N6-(butenyl-4-oxycarbonyl)lysine. In some instances, the unnatural amino acid is N6-(pentenyl-5-oxycarbonyl)lysine. In some instances, the unnatural amino acid is N6-((but-3-yn-1-yloxy)carbonyl)-lysine. In some instances, the unnatural amino acid is N6-((pent-4-yn-1-yloxy)carbonyl)-lysine. In some instances, the unnatural amino acid is N6-(thiazolidine-4-carbonyl)lysine. In some instances, the unnatural amino acid is 2-amino-8-oxononanoic acid. In some instances, the unnatural amino acid is 2-amino-8-oxooctanoic acid. In some instances, the unnatural amino acid is N6-(2-oxoacetyl)lysine. In some instances, the unnatural amino acid is N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine. In some instances, the unnatural amino acid is N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine. In some instances, the unnatural amino acid is N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine.
[0129] In some instances, the unnatural amino acid is N6-propionyl lysine. In some instances, the unnatural amino acid is N6-butyryl lysine. In some instances, the unnatural amino acid is N6-(but-2-enoyl) lysine. In some instances, the unnatural amino acid is N6-((bicyclo[2.2.1]hept-5-en-2-yloxy)carbonyl) lysine. In some instances, the unnatural amino acid is N6-((spiro[2.3]hex-1-en-5-ylmethoxy)carbonyl) lysine. In some instances, the unnatural amino acid is N6-(((4-(1-(trifluoromethyl)cycloprop-2-en-1-yl)benzyl)oxy)carbonyl) lysine. In some instances, the unnatural amino acid is is N6-((bicyclo[2.2.1]hept-5-en-2-ylmethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is cysteinyl lysine. In some cases, the unnatural amino acid is N6-((1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((3-(3-methyl-3H-diazirin-3-yl)propoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((metanitrobenyloxy)N6-methylcarbonyl)lysine. In some instances, the unnatural amino acid is N6-((bicyclo[6.1.0]non-4-yn-9-ylmethoxy)carbonyl)-lysine. In some instances, the unnatural amino acid is N6-((cyclohept-3-en-1-yloxy)carbonyl)-L-lysine.
[0130] In some instances, the unnatural amino acid is 2-amino-3-(((((benzyloxy)carbonyl)amino)methyl)selanyl)propanoic acid. In some embodiments, the unnatural amino acid is incorporated into the unnatural polypeptide or protein by a recycled amber, opal, or ochre stop codon. In some embodiments, the unnatural amino acid is incorporated into the unnatural polypeptide or protein by a four-base codon. In some embodiments, the unnatural amino acid is incorporated into the protein by a recycled rare-sense codon.
[0131] In some embodiments, the unnatural amino acid is incorporated into the unnatural polypeptide or protein by an unnatural codon that comprises an unnatural nucleotide.
[0132] In some cases, incorporation of unnatural amino acids into proteins is mediated by orthogonal engineered synthetase / tRNA pairs. Such orthogonal pairs include a natural or mutant synthetase that can charge a particular unnatural amino acid to the unnatural tRNA, often while minimizing a) charging of other endogenous or alternative unnatural amino acids to the unnatural tRNA and b) charging of any other (including endogenous) tRNAs. Such orthogonal pairs include a tRNA that can be charged by the synthetase while avoiding charging of other endogenous amino acids by the endogenous synthetase. In some embodiments, such pairs are identified from various organisms, such as bacteria, yeast, archaea, or human sources. In some embodiments, the orthogonal synthetase / tRNA pair includes components from a single organism. In some embodiments, the orthogonal synthetase / tRNA pair includes components from two different organisms. In some embodiments, the orthogonal synthetase / tRNA pair includes components that facilitate translation of different amino acids prior to modification. In some embodiments, the orthogonal synthetase is an engineered alanine synthetase. In some embodiments, the orthogonal synthetase is an engineered arginine synthetase. In some embodiments, the orthogonal synthetase is an engineered asparagine synthetase. In some embodiments, the orthogonal synthetase is an engineered aspartate synthetase. In some embodiments, the orthogonal synthetase is an engineered cysteine synthetase. In some embodiments, the orthogonal synthetase is an engineered glutamine synthetase. In some embodiments, the orthogonal synthetase is an engineered glutamate synthetase. In some embodiments, the orthogonal synthetase is an engineered alanine synthetase. In some embodiments, the orthogonal synthetase is an engineered histidine synthetase. In some embodiments, the orthogonal synthetase is an engineered leucine synthetase. In some embodiments, the orthogonal synthetase is an engineered isoleucine synthetase. In some embodiments, the orthogonal synthetase is an engineered lysine synthetase.In some embodiments, the orthogonal synthetase is an engineered methionine synthetase. In some embodiments, the orthogonal synthetase is a modified phenylalanine synthetase. In some embodiments, the orthogonal synthetase is a modified proline synthetase. In some embodiments, the orthogonal synthetase is a modified serine synthetase. In some embodiments, the orthogonal synthetase is a modified threonine synthetase. In some embodiments, the orthogonal synthetase is a modified tryptophan synthetase. In some embodiments, the orthogonal synthetase is a modified tyrosine synthetase. In some embodiments, the orthogonal synthetase is a modified valine synthetase. In some embodiments, the orthogonal synthetase is a modified phosphoserine synthetase. In some embodiments, the orthogonal tRNA is a modified alanine tRNA. In some embodiments, the orthogonal tRNA is a modified arginine tRNA. In some embodiments, the orthogonal tRNA is a modified asparagine tRNA. In some embodiments, the orthogonal tRNA is a modified aspartate tRNA. In some embodiments, the orthogonal tRNA is a modified cysteine tRNA. In some embodiments, the orthogonal tRNA is a modified glutamine tRNA. In some embodiments, the orthogonal tRNA is a modified glutamate tRNA. In some embodiments, the orthogonal tRNA is a modified alanine glycine. In some embodiments, the orthogonal tRNA is a modified histidine tRNA. In some embodiments, the orthogonal tRNA is a modified leucine tRNA. In some embodiments, the orthogonal tRNA is a modified isoleucine tRNA. In some embodiments, the orthogonal tRNA is a modified lysine tRNA. In some embodiments, the orthogonal tRNA is a modified methionine tRNA. In some embodiments, the orthogonal tRNA is a modified phenylalanine tRNA. In some embodiments, the orthogonal tRNA is a modified proline tRNA. In some embodiments, the orthogonal tRNA is a modified serine tRNA. In some embodiments, the orthogonal tRNA is a modified threonine tRNA.In some embodiments, the orthogonal tRNA is a modified tryptophan tRNA. In some embodiments, the orthogonal tRNA is a modified tyrosine tRNA. In some embodiments, the orthogonal tRNA is a modified valine tRNA. In some embodiments, the orthogonal tRNA is a modified phosphoserine tRNA.
[0133] In some embodiments, unnatural amino acids are incorporated into unnatural polypeptides or proteins by aminoacyl (aaRS or RS)-tRNA synthetase-tRNA pairs. Exemplary aaRS-tRNA pairs include, but are not limited to, the Methanococcus jannaschii (Mj-Tyr) aaRS / tRNA pair, the Methanococcus jannaschii (M. jannaschii) TyrRS mutant pAzFRS (MjpAzFRS), and the E. coli TyrRS (Ec-Tyr) / B. stearothermophilus tRNA pair. CUA vs. E. coli LeuRS (Ec-Leu) / B.stearothermophilus tRNA CUA Examples of such tRNA pairs include the Mj-TyrRS / tRNA pair and the pyrrolysyl-tRNA pair. In some instances, unnatural amino acids are incorporated into unnatural polypeptides or proteins via the Mj-TyrRS / tRNA pair. Exemplary unnatural amino acids (UAA) that can be incorporated by the Mj-TyrRS / tRNA pair include, but are not limited to, para-substituted phenylalanine derivatives such as p-azido-L-phenylalanine (pAzF), N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine, p-aminophenylalanine, and p-methoxyphenylalanine; meta-substituted tyrosine derivatives such as 3-aminotyrosine, 3-nitrotyrosine, 3,4-dihydroxyphenylalanine, and 3-iodotyrosine; phenylselenocysteine; p-boronophenylalanine; and o-nitrobenzyltyrosine.
[0134] In some instances, the unnatural amino acid is Ec-Tyr / tRNA CUA or Ec-Leu / tRNA CUA Ec-Tyr / tRNA pairing is incorporated into non-natural polypeptides or proteins. CUA or Ec-Leu / tRNA CUA Exemplary UAAs that can be incorporated pairwise include, but are not limited to, phenylalanine derivatives containing benzophenone, ketone, iodide, or azide substituents; O-propargyl tyrosine; α-aminocapric acid, O-methyl tyrosine, O-nitrobenzyl cysteine; and 3-(naphthalen-2-ylamino)-2-aminopropanoic acid.
[0135] In some instances, unnatural amino acids are incorporated into unnatural polypeptides or proteins via a pyrrolysyl-tRNA pair. In some instances, the PylRS can be obtained from an archaeal species, such as a methanogenic archaea. In some instances, the PylRS can be obtained from Methanosarcina barkeri, Methanosarcina mazei, or Methanosarcina acetivorans. In some instances, the PylRS can be a chimeric PylRS. Exemplary UAAs that can be incorporated via a pyrrolysyl-tRNA pair include, but are not limited to, amide and carbamate substituted lysines, such as N6-(2-azidoethoxy)-carbonyl-L-lysine (AzK), N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine, These include 2-amino-6-((R)-tetrahydrofuran-2-carboxamido)hexanoic acid, N-ε-D-prolyl-L-lysine and N-ε-cyclopentyloxycarbonyl-L-lysine; N-ε-acryloyl-L-lysine; N-ε-[(1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl]-L-lysine; and N-ε-(1-methylcyclopro-2-enecarboxamido)lysine.
[0136] In some cases, the compositions and methods described herein include using at least two tRNA synthetases to incorporate at least two unnatural amino acids into a non-natural polypeptide or protein. In some cases, the at least two tRNA synthetases can be the same or different. In some cases, the at least two unnatural amino acids can be the same or different. In some cases, the at least two unnatural amino acids incorporated into the non-natural polypeptide are different. In some cases, the at least two different unnatural amino acids can be site-specifically incorporated into the non-natural polypeptide or protein.
[0137] In some examples, unnatural amino acids can be incorporated into the unnatural polypeptides or proteins described herein by the synthetases disclosed in U.S. Patent Nos. 9,988,619 and 9,938,516. Exemplary UAAs that can be incorporated by such synthetases include para-methylazido-L-phenylalanine, aralkyl, heterocyclyl, and heteroaralkyl unnatural amino acids. In some embodiments, such UAAs comprise pyridyl, pyrazinyl, pyrazolyl, triazolyl, oxazolyl, thiazolyl, thiophenyl, or other heterocycles. In some embodiments, such amino acids comprise other chemical groups that can be conjugated to coupling partners, such as azides, tetrazines, or water-soluble moieties. In some embodiments, such synthetases are expressed and used to incorporate UAAs into proteins in vivo. In some embodiments, such synthetases are used to incorporate UAAs into proteins using cell-free translation systems.
[0138] In some instances, unnatural amino acids can be incorporated into the unnatural polypeptides or proteins described herein by naturally occurring synthetases. In some embodiments, the unnatural amino acid is incorporated into a non-natural polypeptide or protein by an organism that is auxotrophic for one or more amino acids. In some embodiments, a synthetase corresponding to the auxotrophic amino acid can charge the corresponding tRNA with the unnatural amino acid. In some embodiments, the unnatural amino acid is selenocysteine or a derivative thereof. In some embodiments, the unnatural amino acid is selenomethionine or a derivative thereof. In some embodiments, the unnatural amino acid is an aromatic amino acid, where the aromatic amino acid comprises an aryl halide, such as iodide. In embodiments, the unnatural amino acid is structurally similar to the auxotrophic amino acid. In some cases, the unnatural amino acid includes an unnatural amino acid shown in Figure 5a.
[0139] In some cases, the unnatural amino acid comprises a lysine or phenylalanine derivative or analog. In some cases, the unnatural amino acid comprises a lysine derivative or lysine analog. In some cases, the unnatural amino acid comprises pyrrolysine (Pyl). In some cases, the unnatural amino acid comprises a phenylalanine derivative or phenylalanine analog. In some cases, the unnatural amino acid is an unnatural amino acid described in Wan et al., "Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool," Biochem Biophys Aceta 1844(6):1059-4070 (2014). In some cases, the unnatural amino acid includes the unnatural amino acids shown in Figure 5B and Figure 5C.
[0140] In some embodiments, the unnatural amino acids include the unnatural amino acids shown in Figures 5D-5G (adapted from Table 1 in Dumas et al., Chemical Science 2015, 6, 50-69).
[0141] In some embodiments, the unnatural amino acids incorporated into the proteins described herein are disclosed in U.S. Patent No. 9,840,493; U.S. Patent No. 9,682,934; U.S. Patent Application Publication No. 2017 / 0260137; U.S. Patent No. 9,938,516; or U.S. Patent Application Publication No. 2018 / 0086734. Exemplary UAAs that can be incorporated by such synthetases include para-methylazido-L-phenylalanine, aralkyl, heterocyclyl, and heteroaralkyl, and lysine-derived unnatural amino acids. In some embodiments, such UAAs comprise pyridyl, pyrazinyl, pyrazolyl, triazolyl, oxazolyl, thiazolyl, thiophenyl, or other heterocycles. In some embodiments, such amino acids comprise azides, tetrazines, or other chemical groups that can be attached to coupling partners such as water-soluble moieties. In some embodiments, the UAAs comprise an azide linked to an aromatic moiety via an alkyl linker. In some embodiments, the alkyl linker is a C1-C10 linker. In some embodiments, the UAA comprises a tetrazine linked to the aromatic moiety via an alkyl linker. In some embodiments, the UAA comprises a tetrazine linked to the aromatic moiety via an amino group. In some embodiments, the UAA comprises a tetrazine linked to the aromatic moiety via an alkylamino group. In some embodiments, the UAA comprises an azide linked to the terminal nitrogen of an amino acid side chain via an alkyl chain (e.g., N6 of a lysine derivative, or N5, N4, or N3 of a derivative containing a shorter alkyl side chain). In some embodiments, the UAA comprises a tetrazine linked to the terminal nitrogen of an amino acid side chain via an alkyl chain. In some embodiments, the UAA comprises an azide or tetrazine linked to an amide via an alkyl linker. In some embodiments, the UAA is an azide- or tetrazine-containing carbamate or an amide of 3-aminoalanine, serine, lysine, or a derivative thereof. In some embodiments, such a UAA is incorporated into a protein in vivo. In some embodiments, such The UAA is incorporated into proteins in a cell-free system.
[0142] cell type In some embodiments, many types of cells / microorganisms are used, for example, for transformation or genetic engineering. In some embodiments, the cells are prokaryotic or eukaryotic cells. In some cases, the cells are microorganisms, such as bacterial cells, fungal cells, yeast, or unicellular protozoa. In other cases, the cells are eukaryotic cells, such as cultured animal, plant, or human cells. In additional cases, the cells are present in an organism, such as a plant or animal.
[0143] In some embodiments, the engineered microorganism is a unicellular organism, capable of dividing and growing frequently. The microorganism may comprise one or more of the following characteristics: aerobic, anaerobic, filamentous, non-filamentous, haploid, diploid, auxotrophic, and / or non-auxotrophic. In certain embodiments, the engineered microorganism is a prokaryotic microorganism (e.g., a bacterium), and in certain embodiments, the engineered microorganism is a non-prokaryotic microorganism. In some embodiments, the engineered microorganism is a eukaryotic microorganism (e.g., a yeast, a fungus, an amoeba). In some embodiments, the engineered microorganism is a fungus. In some embodiments, the engineered organism is a yeast.
[0144] Any suitable yeast may be selected as a host microorganism, engineered microorganism, genetically modified organism, or source of a heterologous or modified polynucleotide. Yeasts include, but are not limited to, Yarrowia yeasts (e.g., Y. lipolytica (formerly classified as Candida lipolytica)), Candida yeasts (e.g., C. revkaufi, C. viswanathii, C. pulcherrima, C. tropicalis, C. utilis), Rhodotorula yeasts (e.g., R. glutinus, R. graminis), Rhodosporidium yeasts (e.g., R. toruloides), Saccharomyces yeasts (e.g., S. cerevisiae, S. bayanus, S. pastorianus, S. carlsbergensis), Cryptococcus yeasts, Trichosporon yeasts (e.g., T. pullans, T. cutaneum), Pichia yeasts (e.g., P. pastoris), and Lipomyces yeasts (e.g., L. starkeyii, L. lipoferus). In some embodiments, suitable yeasts are those of the genus Arachniotus, Aspergillus, Aureobasidium, Auxarthron, Blastomyces, Candida, Chrysosporium, Debaryomyces, Coccidiodes, Cryptococcus, Gymnoascus, Hansenula, Histoplasma, Issatchenkia, Kluyveromyces, Lipomyces, Lssatchenkia, Microsporum, Myxotrichum, Myxozyma, Oidiodendron, Pachysolen, Penicillium, Pichia, Rhodosporidium, Rhodotorula, Saccharomyces, Schizosaccharomyces, Scopulariopsis, Sepedonium, Trichosporon, or Yarrowia.Arachniotus flavoluteus, Aspergillus flavus, Aspergillus fumigatus, Aspergillus niger, and Aureobasidium pullulans, Auxarthron. thaxteri、Blastomyces dermatitidis、Candida albicans、Candida dubliniensis、Candida famata、Candida glabrata、Candida guilliermondii、Candida kefyr、Candida krusei、Can dida lambica、Candida lipolytica、Candida lustitaniae、Candida parapsilosis、Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xestobii, Chrysosporuim keratinophilum, Coccidiodes immitis, Cryptococcus albidus var. diffluens, Cryptococcus laurentii, Cryptococcus neofomans, Debaryomyces hansenii, Gymnoascus dugwayensis, Hansenula anomala, Histoplasma capsulatum, Issatchenkia occidentalis, Isstachenkia orientalis, Kluyveromyces lactis, Kluyveromyces marxianus, Kluyveromyces thermotolerans, Kluyveromyces waltii, Lipomyces lipoferus, Lipomyces starkeyii, Microsporum gypseum, Myxotrichum deflexum, Oidiodendron echinulatum, Pachysolen tannophilis, Penicillium notatum, Pichia anomala, Pichia pastoris, Pichia stipitis, Rhodosporidium toruloides, Rhodotorula glutinus, Rhodotorula graminis, Saccharomyces cerevisiae, Saccharomyces kluyveri, Schizosaccharomyces pombe, Scopulariopsis acremonium, Sepedonium chrysospermum, Trichosporon cutaneum, Trichosporon pullans, Yarrowia lipolytica, or Yarrowia lipolytica (formerly Candida It was classified as lipolytica)In some embodiments, the yeast is a Y. lipolytica strain, including, but not limited to, ATCC20362, ATCC8862, ATCC18944, ATCC20228, ATCC76982, and LGAM S(7)1 strains (Papanikolaou S. and Aggelis G., Bioresort. Technol. 82(1):43-9 (2002)). In certain embodiments, the yeast is a Candida species (i.e., Candida genus) yeast. Any suitable Candida species may be used and / or may be genetically modified for the production of fatty dicarboxylic acids (e.g., octanedioic acid, decanedioic acid, dodecanedioic acid, tetradecanedioic acid, hexadecanedioic acid, octadecanedioic acid, eicosanedioic acid). In some embodiments, suitable Candida species include, but are not limited to, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, and Candida. Candida species include Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xestobii, and any other Candida species yeasts described herein. Non-limiting examples of Candida species strains include, but are not limited to, strains sAA001 (ATCC20336), sAA002 (ATCC20913), sAA003 (ATCC20962), sAA496 (US Patent Application Publication No. 2012 / 0077252), sAA106 (US Patent Application Publication No. 2012 / 0077252), SU-2 (ura3- / ura3-), and H5343 (beta-oxidation blocked; U.S. Patent No. 5,648,247). Any suitable strain of yeast can be utilized as the parent strain for genetic recombination.
[0145] Yeast genera, species, and strains are often so closely related in genetic content that they can be difficult to distinguish, classify, and / or name. In some cases, it can be difficult to distinguish, classify, and / or name strains of C. lipolytica and Y. lipolytica, and in some cases, they may be considered the same organism. In some cases, it can be difficult to distinguish, classify, and / or name various strains of C. tropicalis and C. viswanathii (see, e.g., Arie et al., J. Gen. Appl. Microbiol., 46, 257-262 (2000)). Some C. tropicalis and C. viswanathii strains obtained from ATCC and other commercial or academic sources may be considered equivalent and equally suitable for the embodiments described herein. In some embodiments, some parent strains of C. tropicalis and C. viswanathii are considered to differ only in name.
[0146] Any suitable fungus may be selected as a host microorganism, an engineered microorganism, or a source of a heterologous polynucleotide. Non-limiting examples of fungi include, but are not limited to, Aspergillus fungi (e.g., A. parasiticus, A. nidulans), Thraustochytrium fungi, Schizochytrium fungi, and Rhizopus fungi (e.g., R. arrhizus, R. oryzae, R. nigricans). In some embodiments, the fungus is an A. parasiticus strain, including, but not limited to, the ATCC 24690 strain, and in particular embodiments, the fungus is an A. nidulans strain, including, but not limited to, the ATCC 38163 strain.
[0147] Any suitable prokaryote may be selected as a host microorganism, engineered microorganism, or source of a heterologous polynucleotide. Gram-negative or Gram-positive bacteria may be selected. Exemplary bacteria include, but are not limited to, Bacillus (e.g., B. subtilis, B. megaterium), Acinetobacter, Norcardia, Xanthobacter, Escherichia (e.g., E. coli (e.g., strains DH10B, Stbl2, DH5-alpha, DB3, DB3.1), DB4, DB5, JDP682, and ccdA-over (e.g., Rice National Application No. 09 / 518,188), Streptomyces, Erwinia, Klebsiella, Serratia (e.g., S. marcessans), Pseudomonas (e.g., P. aeruginosa), Salmonella (e.g., S. typhimurium, S. typhi), and Megasphaera (e.g., Megasphaera elsdenii). Fungi also include, but are not limited to, photosynthetic bacteria (e.g., green non-sulfur bacteria (e.g., Choroflexus bacteria (e.g., C. aurantiacus), Chloronema bacteria (e.g., C. gigateum)), green sulfur bacteria (e.g., Chlorobium bacteria (e.g., C. limicola), Pelodictyon bacteria (e.g., P. luteolum), purple sulfur bacteria (e.g., Chromatium bacteria (e.g., C. okenii)), and purple non-sulfur bacteria (e.g., Rhodospirillum bacteria (e.g., R. rubrum), Rhodobacter bacteria (e.g., R. sphaeroides, R. capsulatus), and Rhodomicrobium bacteria (e.g., R. vanellii)).
[0148] Cells of non-microbial origin can be utilized as host microorganisms, engineered microorganisms, or sources of heterologous polynucleotides. Examples of such cells include, but are not limited to, insect cells (e.g., Drosophila (e.g., D. melanogaster), Spodoptera (e.g., S. frugiperda Sf9 or Sf21 cells), and Trichoplusa (e.g., High-Five cells); nematode cells (e.g., C. elegans cells); avian cells; amphibian cells (e.g., Xenopus laevis cells); reptilian cells; mammalian cells (e.g., NIH3T3, 293, CHO, COS, VERO, C127, BHK, Per-C6, Bowes melanoma, and HeLa cells); and plant cells (e.g., Arabidopsis thaliana, Nicotania tabacum, Cuphea acinifolia, Cuphea aequipetala, Cuphea angustifolia, Cuphea appendiculata, Cuphea avigera, Cuphea avigera var. pulcherrima, Cuphea axilliflora, Cuphea bahiensis, Cuphea baillonis, Cuphea brachypoda, Cuphea bustamanta, Cuphea calcarata, Cuphea calophylla, Cuphea calophylla subsp. mesostemon, Cuphea carthagenensis, Cuphea circaeoides, Cuphea confertiflora, Cuphea cordata, Cuphea crassiflora, Cupphea cyanea, Cupphea decandra, Cupphea denticulata, Cuphea disperma, Cupphea epilobiifolia, Cupphea ericoides, Cuphea flava、Cuphea flavisetula、Cuphea fuchsiifolia、Cuphea gaumeri、Cuphea glutinosa、Cuphea heterophylla、Cuphea hookeriana、Cuphea hyssopifolia (Mexican-heather)、Cuphea hyssopoides, Cuphea ignea, Cuphea ingrata, Cuphea jorullensis, Cuphea lanceolata, Cuphea linarioides, Cuphea llavea, Cuphea lophostoma, Cuphea lutea, Cuphea lutescens, Cuphea melanium, Cuphea melvilla, Cuphea micrantha, Cuphea micropetala, Cuphea mimuloides, Cuphea nitidula, Cuphea palustris, Cuphea parsonsia, Cuphea pascuorum, Cuphea paucipetala, Cuphea procumbens, Cuphea pseudosilene, Cuphea pseudovaccinium, Cuphea pulchra, Cuphea racemosa, Cuphea repens, Cuphea salicifolia, Cuphea salvadorensis, Cuphea schumannii, Cuphea sessiliflora, Cuphea sessilifolia, Cuphea setosa, Cuphea spectabilis, Cuphea spermacoce, Cuphea splendida, Cuphea splendida var. viridiflava, Cuphea strigulosa, Cuphea subuligera, Cuphea teleandra, Cuphea thymoides, Cuphea tolucana, Cuphea urens, Cuphea utriculosa, Cuphea viscosissima, Cupphea watsoniana, Cupphea wrightii, Cupphea lanceolata).
[0149] Microorganisms or cells used as host organisms or sources of heterologous polynucleotides are commercially available. The microorganisms and cells described herein, as well as other suitable microorganisms and cells, are also commercially available. Cells are available, for example, from Invitrogen Corporation (Carlsbad, CA), the American Type Culture Collection (Manassas, Virginia), and the Agricultural Research Culture Collection (NRRL; Peoria, Illinois). Host microorganisms and engineered microorganisms can be provided in any suitable form. For example, such microorganisms can be provided in liquid or solid cultures (e.g., agar-based media), which may be primary cultures or may have been passaged one or more times (e.g., diluted and cultured). Microorganisms can also be provided in frozen or dried form (e.g., lyophilized). Microorganisms can be provided at any suitable concentration.
[0150] polymerase A particularly useful function of polymerases is to catalyze the polymerization of nucleic acid chains using existing nucleic acids as templates. Other useful functions are described elsewhere herein. Examples of useful polymerases include DNA polymerases and RNA polymerases.
[0151] The ability to improve the specificity, processivity, or other characteristics of a polymerase for non-naturally occurring nucleic acids is highly desirable in a variety of situations in which the incorporation of non-naturally occurring nucleic acids is desired, including, for example, amplification, sequencing, labeling, detection, cloning, etc.
[0152] In some examples, the disclosure includes polymerases that incorporate non-natural nucleic acids into extending template copies, for example, during DNA amplification. In some embodiments, the polymerase can be modified so that the active site of the polymerase is modified to reduce steric inhibition of non-natural nucleic acids entering the active site. In some embodiments, the polymerase can be modified to provide complementarity with one or more non-natural features of the non-natural nucleic acid. Such polymerases can be expressed or engineered in cells to stably incorporate UBPs into the cells. Thus, the disclosure includes compositions comprising heterologous or recombinant polymerases, and methods of use thereof.
[0153] Polymerases can be modified using methods related to protein engineering. For example, molecular modeling can be performed based on the crystal structure to identify positions in the polymerase where mutations can be made to modify target activity. Residues identified as targets for substitution are described in Bordo et al., J Mol Biol 217:721-729 (1991) and Hayes et al., Proc. The amino acid sequence may be substituted with residues selected using energy minimization modeling, homology modeling, and / or conservative amino acid substitution, as described in Natl Acad Sci, USA 99:15926-15931 (2002).
[0154] Any of a variety of polymerases may be used in the methods or compositions described herein, including, for example, protein-based enzymes isolated from biological systems and functional variants thereof. Reference to a particular polymerase, such as those exemplified below, will be understood to include its functional variants unless otherwise indicated. In some embodiments, the polymerase is a wild-type polymerase. In some embodiments, the polymerase is a modified or mutant polymerase.
[0155] Polymerases with features for improved entry of non-natural nucleic acids into the active site region and for cooperation with non-natural nucleotides in the active site region can also be used, hi some embodiments, the modified polymerase has an altered nucleotide binding site.
[0156] In some embodiments, the modified polymerase has a specificity for non-naturally occurring nucleic acids that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the specificity of the wild-type polymerase for the non-naturally occurring nucleic acids. In some embodiments, the modified or wild-type polymerase has a specificity for non-naturally occurring nucleic acids containing modified sugars that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the specificity of the wild-type polymerase for natural nucleic acids and / or non-naturally occurring nucleic acids without modified sugars. In some embodiments, the modified or wild-type polymerase has a specificity for non-natural nucleic acids containing modified bases that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the specificity of the wild-type polymerase for natural nucleic acids and / or non-natural nucleic acids that do not contain modified bases. In some embodiments, the modified or wild-type polymerase has a specificity for non-natural nucleic acids containing a triphosphate that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the specificity of the wild-type polymerase for nucleic acids containing a triphosphate and / or non-natural nucleic acids that do not contain a triphosphate. For example, a modified or wild-type polymerase can have specificity for non-naturally occurring nucleic acids containing a triphosphate that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the specificity of the wild-type polymerase for non-naturally occurring nucleic acids containing a diphosphate or monophosphate, or no phosphate, or combinations thereof.
[0157] In some embodiments, modified or wild-type polymerases have relaxed specificity for non-naturally occurring nucleic acids. In some embodiments, modified or wild-type polymerases have specificity for non-naturally occurring nucleic acids and for naturally occurring nucleic acids that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the specificity of the wild-type polymerase for naturally occurring nucleic acids. In some embodiments, modified or wild-type polymerases have specificity for non-naturally occurring nucleic acids that contain modified sugars and for naturally occurring nucleic acids that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the specificity of the wild-type polymerase for naturally occurring nucleic acids. In some embodiments, the modified or wild-type polymerase has specificity for non-natural nucleic acids comprising modified bases and specificity for natural nucleic acids that is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the specificity of the wild-type polymerase for natural nucleic acids.
[0158] The lack of exonuclease activity can be a wild-type characteristic or a characteristic conferred by a variant or engineered polymerase, for example, the exo-minus Klenow fragment is a mutant version of the Klenow fragment that lacks 3' to 5' proofreading exonuclease activity.
[0159] The disclosed methods can be used to expand the substrate range of any DNA polymerase that lacks inherent 3' to 5' exonuclease proofreading activity or in which 3' to 5' exonuclease proofreading activity has been abolished (e.g., through mutation). Examples of DNA polymerases include polA, polB (see, e.g., Parrel & Loeb, Nature Struc Biol 2001), polC, polD, polY, polX, and reverse transcriptase (RT), although evolvable, high-fidelity polymerases (PCT / GB2004 / 004643) are preferred. In some embodiments, the modified or wild-type polymerase lacks 3' to 5' proofreading activity. In some embodiments, the modified or wild-type polymerase substantially lacks 3' to 5' proofreading exonuclease activity toward non-naturally occurring nucleic acids. In some embodiments, the modified or wild-type polymerase has 3' to 5' proofreading exonuclease activity. In some embodiments, the modified or wild-type polymerase has 3' to 5' proofreading exonuclease activity toward natural nucleic acids and substantially lacks 3' to 5' proofreading exonuclease activity toward non-naturally occurring nucleic acids.
[0160] In some embodiments, the modified polymerase has a 3' to 5' proofreading exonuclease activity that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the proofreading exonuclease activity of the wild-type polymerase. In some embodiments, the modified polymerase has a 3' to 5' proofreading exonuclease activity towards non-naturally occurring nucleic acids that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the proofreading exonuclease activity of the wild-type polymerase towards natural nucleic acids. In some embodiments, the modified polymerase has 3' to 5' proofreading exonuclease activity towards non-naturally occurring nucleic acids that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the proofreading exonuclease activity of the wild-type polymerase towards natural nucleic acids. In some embodiments, the modified polymerase has 3' to 5' proofreading exonuclease activity towards natural nucleic acids that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the proofreading exonuclease activity of the wild-type polymerase towards natural nucleic acids.
[0161] In some embodiments, polymerases are characterized according to their dissociation rate from nucleic acids. In some embodiments, the polymerase has a relatively low dissociation rate for one or more natural and non-natural nucleic acids. In some embodiments, the polymerase has a relatively high dissociation rate for one or more natural and non-natural nucleic acids. This dissociation rate is an activity of the polymerase that can be adjusted to adjust the reaction rate in the methods described herein.
[0162] In some embodiments, polymerases are characterized according to their fidelity when used with a specific natural and / or non-natural nucleic acid or a collection of natural and / or non-natural nucleic acids.Fidelity generally refers to the accuracy with which a polymerase incorporates correct nucleic acids into the growing nucleic acid strand when making copies of a nucleic acid template.The fidelity of a DNA polymerase can be measured as the ratio of correct to incorrect incorporation of natural and non-natural nucleic acids when natural and non-natural nucleic acids are present, for example, at equal concentrations and compete for strand synthesis at the same site of the polymerase strand-template nucleic acid binary complex. The fidelity of a DNA polymerase can be calculated as the ratio of (kcat / Km) for natural and non-natural nucleic acids and (kcat / Km) for mis-natural and non-natural nucleic acids, where kcat and Km are the Michaelis-Menten parameters in steady-state enzyme kinetics (Fersht, AR (1985) Enzyme Structure and Mechanism, 2nd ed., p350, W.H. Freeman & Co., New York. (Incorporated herein by reference)). In some embodiments, the polymerase, with or without proofreading activity, has a fidelity value of at least about 100, 1000, 10,000, 100,000, or 1 x 106.
[0163] In some embodiments, polymerases from natural sources or variants thereof are screened using assays that detect the incorporation of non-natural nucleic acids with specific structures. In one example, polymerases may be screened for their ability to incorporate non-natural nucleic acids or UBPs (e.g., d5SICSTP, dCNMOTP, dTPT3TP, dNaMTP, dCNMOTP-dTPT3TP, or d5SICSTP-dNaMTP UBPs). Polymerases, e.g., heterologous polymerases, that exhibit modified properties toward non-natural nucleic acids compared to wild-type polymerases may also be used. For example, modified properties may include, for example, Km, kcat, Vmax, polymerase processivity in the presence of non-natural nucleic acids (or naturally occurring nucleotides), average template read length by the polymerase in the presence of non-natural nucleic acids, polymerase specificity toward non-natural nucleic acids, binding rate of non-natural nucleic acids, release rate of products (e.g., pyrophosphate, triphosphate), branching rate, or any combination thereof. In one embodiment, the altered property is a decreased Km for the non-naturally occurring nucleic acid and / or an increased kcat / Km or Vmax / Km for the non-naturally occurring nucleic acid. Similarly, the polymerase optionally has an increased rate of binding of the non-naturally occurring nucleic acid, an increased rate of product release, and / or a decreased rate of branching compared to the wild-type polymerase.
[0164] At the same time, the polymerase can incorporate naturally occurring nucleic acids, e.g., A, C, G, and T, into the extending nucleic acid copy. For example, the polymerase optionally exhibits a specific activity toward naturally occurring nucleic acids that is at least about 5% higher (e.g., 5%, 10%, 25%, 50%, 75%, 100% or more) than the corresponding wild-type polymerase, and a processivity with naturally occurring nucleic acids in the presence of a template that is at least 5% higher (e.g., 5%, 10%, 25%, 50%, 75%, 100% or more) relative to the wild-type polymerase in the presence of naturally occurring nucleic acids. Optionally, the polymerase exhibits a kcat / Km or Vmax / Km for naturally occurring nucleotides that is at least about 5% higher (e.g., about 5%, 10%, 25%, 50%, 75%, or 100% or more) relative to the wild-type polymerase.
[0165] Polymerases used herein that may have the ability to incorporate unnatural nucleic acids of specific structures may also be generated using directed evolution approaches. Nucleic acid synthesis assays may be used to screen polymerase variants with specificity for any of a variety of unnatural nucleic acids. For example, polymerase variants may be screened for the ability to incorporate unnatural nucleoside triphosphates opposite unnatural nucleotides in DNA templates (e.g., dTPT3TP opposite dCNMO, dCNMOTP opposite dTPT3, NaMTP opposite dTPT3, or TAT1TP opposite dCNMO or dNaM). In some embodiments, such assays are in vitro assays, e.g., using recombinant polymerase variants. In some embodiments, such assays are in vivo assays, e.g., expressing polymerase variants in cells. Such directed evolution techniques may be used to screen any suitable polymerase variant for activity against any unnatural nucleic acid described herein. In some cases, the polymerases used herein have the ability to incorporate unnatural ribonucleotides into nucleic acids, such as RNA. For example, NaM or TAT1 ribonucleotides are incorporated into nucleic acids using the polymerases described herein.
[0166] The modified polymerase of the described compositions can optionally be a modified and / or recombinant Φ29-type DNA polymerase. Optionally, the polymerase can be a modified and / or recombinant Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, or L17 polymerase.
[0167] The modified polymerases of the described compositions are optionally engineered and / or recombinant The modified polymerase may be a prokaryotic DNA polymerase, such as DNA polymerase II (Pol II), DNA polymerase III (Pol III), DNA polymerase IV (Pol IV), or DNA polymerase V (Pol V). In some embodiments, the modified polymerase comprises a polymerase that mediates DNA synthesis across a non-directed damaged nucleotide. In some embodiments, genes encoding Pol I, Pol II (polB), Pol IV (dinB), and / or Pol V (umuCD) are constitutively expressed or overexpressed in the engineered cell or SSO. In some embodiments, increased expression or overexpression of Pol II contributes to increased retention of unnatural base pairs (UBPs) in the engineered cell or SSO.
[0168] Nucleic acid polymerases generally useful in the present disclosure include DNA polymerases, RNA polymerases, reverse transcriptases, and mutant or modified versions thereof.DNA polymerases and their properties are described in detail, inter alia, in DNA Replication, 2nd Edition, Kornberg and Baker, W.H. Freeman, New York, NY (1991).Known conventional DNA polymerases useful in the present disclosure include, but are not limited to, Pyrococcus furiosus (Pfu) DNA polymerase (Lundberg et al., 1991, Gene, 108:1, Stratagene), Pyrococcus woesei (Pwo) DNA polymerase (Hinnisdaels et al., 1996, Biotechniques, 20:186-8, Boehringer Mannheim), Thermus Thermophilus (Tth) DNA polymerase (Myers and Gelfand 1991, Biochemistry 30:7661), Bacillus stearothermophilus DNA polymerase (Stenesh and McGowan, 1977, Biochim Biophys Acta 475:32), Thermococcus litoralis (TIi) DNA polymerase (also called Vent™ DNA polymerase, Cariello et al., 1991, Polynucleotides Res, 19:4193, New England Biolabs), 9°Nm™ DNA polymerase (New England Biolabs), Stoffel fragment, Thermo Sequenase™ (Amersham Pharmacia Biotech UK), Therminator™ (New England Biolabs), Thermotoga maritima (Tma) DNA polymerase (Diaz and Sabino, 1998, Braz J Med. Res., 31:1239), Thermus aquaticus (Taq) DNA polymerase (Chien et al., 1976, J. Bacteoriol., 127:1550), DNA polymerase, Pyrococcus kodakaraensis KOD DNA polymerase (Takagi et al., 1997, Appl. Environ. Microbiol., 63:4504), JDF-3 DNA polymerase (thermococcus sp. JDF-3, from patent application WO0132887), Pyrococcus GB-D (PGB-D) DNA polymerase (also called Deep Vent™ DNA polymerase, Juncosa-Ginesta et al., 1994, Biotechniques, 16:820, New England Journal of Medicine). Biolabs), UlTma DNA polymerase (from the thermophilic bacterium Thermotoga maritima; Diaz and Sabino, 1998 Braz J. Med. Res, 31:1239; PE Applied Biosystems), Tgo DNA polymerase (from Thermococcus gorgonarius, Roche from Molecular Biochemicals), E. coli DNA polymerase I (Lecomte and Doubleday, 1983, Polynucleotides Res. 11:7505), T7 DNA polymerase (Nordstrom et al., 1981, J Biol. Chem. 256:3112), and archaeal DP1I / DP2 DNA polymerase II (Cann et al., 1998, Proc. Natl. Acad. Sci. USA 95:14250). Both mesophilic and thermophilic polymerases are contemplated. Thermophilic DNA polymerases include, but are not limited to, ThermoSequenase®, 9°Nm™, Therminator™, Taq, Tne, Tma, Pfu, TfI, Tth, TIi, Stoffel fragment, Vent™ and DeepVent™ DNA polymerase, KOD DNA polymerase, Tgo, JDF-3, and mutants, variants, and derivatives thereof. Polymerases that are 3' exonuclease-deficient mutants are also contemplated. Reverse transcriptases useful in the present disclosure include, but are not limited to, reverse transcriptases from HIV, HTLV-I, HTLV-II, FeLV, FIV, SIV, AMV, MMTV, MoMuLV, and other retroviruses (see Levin, Cell 88:5-8 (1997); Verma, Biochim Biophys Acta. 473:1-38 (1977); Wu et al., CRC Crit Rev Biochem. 3:289-347 (1975)). Further examples of polymerases include, but are not limited to, 9°N™ DNA polymerase, Taq DNA polymerase, Phusion® DNA polymerase, Pfu DNA polymerase, RB69 DNA polymerase, KOD DNA polymerase, and Vent® DNA polymerase. Gardner et al. (2004) "Comparative Kinetics of Nucleotide Analog Incorporation by Vent DNA Polymerase" (J. Biol. Chem., 279(12), 11834-11842; Gardner and Jack "Determinants of nucleotide sugar recognition in an archaeon DNA polymerase" Nucleic Acids Research, 27(12) 2545-2553.) Polymerases isolated from non-thermophilic organisms can be heat-inactivated. An example is DNA polymerases from phages. It is understood that polymerases from any of a variety of sources can be modified to increase or decrease their tolerance to high temperature conditions. In some embodiments, the polymerase can be thermophilic. In some embodiments, thermophilic polymerases can be heat-inactivatable. Thermophilic polymerases are typically useful for high temperature conditions or thermocycling conditions, such as those used in polymerase chain reaction (PCR) techniques.
[0169] In some embodiments, the polymerase is selected from the group consisting of Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, L17, ThermoSequenase®, 9°Nm™, Therminator™ DNA polymerase, Tne, Tma, TfI, Tth, TIi, Stoffel fragment, Vent™ and DeepVent™ DNA polymerase, KOD DNA polymerase, Tgo, JDF-3, Pfu, Taq, T7 DNA polymerase, T7 RNA polymerase, PGB-D, UlTma DNA polymerase, E. coli DNA polymerase I, E. coli DNA polymerase III, archaeal DP1I / DP2. DNA Polymerase II, 9°N™ DNA Polymerase, Taq DNA Polymerase, Phusion® DNA Polymerase, Pfu DNA Polymerase, SP6 These include RNA polymerase, RB69 DNA polymerase, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, and SuperScript® III reverse transcriptase.
[0170] In some embodiments, the polymerase is DNA polymerase I (or Klenow Fragment), Vent polymerase, Phusion® DNA polymerase, KOD DNA polymerase, Taq polymerase, T7 DNA polymerase, T7 RNA polymerase, Therminator™ DNA polymerase, POLB polymerase, SP6 RNA polymerase, E. coli DNA polymerase I, E. coli DNA polymerase III, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, or SuperScript® III reverse transcriptase.
[0171] Nucleotide transporters Nucleotide transporters (NTs) are a group of membrane transport proteins that facilitate the movement of nucleotide substrates across cell membranes and vesicles. In some embodiments, there are two types of NTs: concentrative nucleoside transporters and equilibrative nucleoside transporters. In some cases, NTs also include organic anion transporters (OATs) and organic cation transporters (OCTs). In some cases, nucleotide transporters are nucleoside triphosphate transporters (NTTs).
[0172] In some embodiments, the nucleoside triphosphate transporter (NTT) is derived from a bacterium, a plant, or an alga. In some embodiments, the nucleotide nucleoside triphosphate transporter is TpNTT1, TpNTT2, TpNTT3, TpNTT4, TpNTT5, TpNTT6, TpNTT7, TpNTT8 (T. pseudona), PtNTT1, PtNTT2, PtNTT3, PtNTT4, PtNTT5, PtNTT6 (P. tricornutum), GsNTT (Galdieria sulphuraria), AtNTT1, AtNTT2 (Arabidopsis thaliana), CtNTT1, CtNTT2 (Chlamydia trachomatis), PamNTT1, PamNTT2 (Protochlamydia amoebophila), CcNTT (Caedibacter caryophilus), or RpNTT1 (Rickettsia prowazekii). In some embodiments, NTT is CNT1, CNT2, CNT3, ENT1, ENT2, OAT1, OAT3, or OCT1. In some cases, NTT is PtNTT1, PtNTT2, PtNTT3, PtNTT4, PtNTT5, or PtNTT6.
[0173] In some embodiments, the NTT introduces a non-natural nucleic acid into an organism, e.g., a cell. In some embodiments, the NTT may be modified such that the nucleotide-binding site of the NTT is modified to reduce steric inhibition of the non-natural nucleic acid into the nucleotide-binding site. In some embodiments, the NTT may be modified to provide increased interaction with one or more natural or non-natural features of the non-natural nucleic acid. Such an NTT may be expressed or engineered in a cell to stably introduce a UBP into the cell. Thus, the present disclosure includes compositions comprising heterologous or recombinant NTT and methods of use thereof.
[0174] NTT can be modified using methods related to protein engineering.For example, molecular modeling can be carried out based on crystal structure to identify the position of NTT that can be mutated to change target activity or binding site.The residue identified as the target of substitution can be replaced with the selected residue using energy minimization modeling, homology modeling, and / or conservative amino acid substitution, as described in Bordo et al., J Mol Biol 217:721-729 (1991) and Hayes et al., Proc Natl Acad Sci, USA 99:15926-15931 (2002).
[0175] Any of a variety of NTTs can be used, for example, as a protein-based antibody isolated from a biological system. The enzymes and functional variants thereof may be used in the methods or compositions described herein. Reference to a specific NTT, such as those exemplified below, is understood to include its functional variants unless otherwise indicated. In some embodiments, the NTT is wild-type NTT. In some embodiments, the NTT is a modified or mutant NTT.
[0176] In some embodiments, as used herein, a modified or mutated NTT is an NTT truncated at the N-terminus, C-terminus, or both the N- and C-terminus. In some embodiments, the truncated NTT is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90% identical to the untruncated NTT. In some cases, an NTT as used herein is PtNTT1, PtNTT2, PtNTT3, PtNTT4, PtNTT5, or PtNTT6. In some cases, a PtNTT as used herein is truncated at the N-terminus, C-terminus, or both the N- and C-terminus. In some embodiments, a truncated PtNTT is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90% identical to the untruncated PtNTT. In some instances, the NTT used herein refers to a truncated PtNTT2, wherein the truncated PtNTT2 has an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90% identical to the amino acid sequence of untruncated PtNTT2. An example of untruncated PtNTT2 (NCBI accession number EEC49227.1, GI:217409295) has the amino acid sequence SEQ ID NO:1.
[0177] NTTs with features for improving entry of non-natural nucleic acids into cells and coordinating with non-natural nucleotides in the nucleotide-binding region may also be used. In some embodiments, the modified NTT has a modified nucleotide-binding site. In some embodiments, the modified or wild-type NTT has relaxed specificity for non-natural nucleic acids. For example, the NTT optionally exhibits a specific incorporation activity for non-natural nucleotides that is at least about 0.1% higher (e.g., about 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.1%, 1.2%, 1.5%, 1.8%, 2%, 3%, 4%, 5%, 10%, 25%, 50%, 75%, 100% or more) than the corresponding wild-type NTT. Optionally, NTT exhibits a kcat / Km or Vmax / Km for unnatural nucleotides that is at least about 0.1% higher (e.g., about 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.1%, 1.2%, 1.5%, 1.8%, 2%, 3%, 4%, 5%, 10%, 25%, 50%, 75%, or 100% or more) than wild-type NTT.
[0178] NTTs can be characterized according to their affinity for triphosphates (i.e., Km) and / or incorporation rate (i.e., Vmax). In some embodiments, NTTs have a relative Km or Vmax for one or more natural and unnatural triphosphates. In some embodiments, NTTs have a relatively high Km or Vmax for one or more natural and unnatural triphosphates.
[0179] NTTs from natural sources or variants thereof can be screened using assays that detect the amount of triphosphate (using either mass spectrometry or radioactivity if the triphosphate is appropriately labeled). In one example, NTTs can be screened for their ability to incorporate non-natural triphosphates (e.g., dTPT3TP, dCNMOTP, d5SICSTP, dNaMTP, NaMTP, and / or TPT1TP). NTTs, e.g., heterologous NTTs, that exhibit altered properties toward non-natural nucleic acids compared to wild-type NTTs, can be screened. may be used. For example, the altered property can be, for example, Km, kcat, Vmax in the case of introduction of a triphosphate. In one embodiment, the altered property is a decreased Km in the case of a non-natural triphosphate, and / or an increased kcat / Km or Vmax / Km in the case of a non-natural triphosphate. Similarly, the NTT optionally has an increased rate of binding, an increased rate of intracellular release, and / or an increased rate of cellular entry of the non-natural triphosphate compared to wild-type NTT.
[0180] At the same time, the NTT can introduce natural triphosphates, such as dATP, dCTP, dGTP, dTTP, ATP, CTP, GTP, and / or TTP, into cells. In some cases, the NTT optionally exhibits a specific activity for introducing natural nucleic acids capable of supporting replication and transcription. In some embodiments, the NTT optionally exhibits kcat / Km or Vmax / Km for natural nucleic acids capable of supporting replication and transcription.
[0181] NTTs used herein that may have the ability to introduce unnatural triphosphates of specific structures can also be generated using directed evolution approaches. Nucleic acid synthesis assays can be used to screen NTT variants with specificity for any of a variety of unnatural triphosphates. For example, NTT variants can be screened for their ability to introduce unnatural triphosphates (e.g., d5SICSTP, dNaMTP, dCNMOTP, dTPT3TP, NaMTP, and / or TPT1TP). In some embodiments, such assays are in vitro assays, e.g., using recombinant NTT variants. In some embodiments, such assays are in vivo assays, e.g., expressing NTT variants in cells. Such techniques can be used to screen any suitable NTT variant for activity toward any of the unnatural triphosphates described herein.
[0182] Nucleic Acid Reagents and Tools The nucleotide and / or nucleic acid reagents (or polynucleotides) for use in the methods, cells, or engineered microorganisms described herein contain one or more ORFs, with or without unnatural nucleotides. ORFs can be derived from any suitable source, sometimes genomic DNA, mRNA, reverse-transcribed RNA, or complementary DNA (cDNA), or from a nucleic acid library containing one or more of the foregoing, and can be derived from any organism containing the desired nucleic acid sequence, protein, or activity. Non-limiting examples of organisms from which ORFs can be obtained include, for example, bacteria, yeast, fungi, humans, insects, nematodes, cattle, horses, dogs, cats, rats, or mice. In some embodiments, the nucleotide and / or nucleic acid reagents or other reagents described herein are isolated or purified. ORFs containing unnatural nucleotides can be produced by published in vitro methods. In some cases, the nucleotide or nucleic acid reagents contain unnatural nucleobases.
[0183] The nucleic acid reagent may include a nucleotide sequence adjacent to the ORF that is translated in conjunction with the ORF and encodes an amino acid tag. The nucleotide sequence encoding the tag is located 3' and / or 5' of the ORF in the nucleic acid reagent, thereby encoding the tag at the C-terminus or N-terminus of the protein or peptide encoded by the ORF. Any tag that does not abolish in vitro transcription and / or translation may be utilized and may be appropriately selected by a technician. The tag may facilitate isolation and / or purification of the desired ORF product from the culture or fermentation medium. In some cases, a library of nucleic acid reagents is used with the methods and compositions described herein. For example, a library of at least 100, 1000, 2000, 5000, 10,000, or more than 50,000 unique polynucleotides is present in the library, and each polynucleotide comprises at least one unnatural nucleobase.
[0184] Nucleic acids or nucleic acid reagents, which may or may not contain non-natural nucleotides, may contain certain elements, such as regulatory elements, which are often selected according to the intended use of the nucleic acid. Any of the following elements may be included or excluded from a nucleic acid reagent. For example, a nucleic acid reagent may contain one or more or all of the following nucleotide elements: one or more promoter elements, one or more 5' untranslated regions (5' UTRs), one or more regions into which a target nucleotide sequence may be inserted ("insertion elements"), one or more target nucleotide sequences, one or more 3' untranslated regions (3' UTRs), and one or more selection elements. A nucleic acid reagent may be provided with one or more of such elements, and other elements may be inserted into the nucleic acid before the nucleic acid is introduced into a desired organism. In some embodiments, a provided nucleic acid reagent contains a promoter, a 5' UTR, an optional 3' UTR, and an insertion element into which a target nucleotide sequence is inserted (i.e., cloned) into the nucleic acid reagent. In certain embodiments, the provided nucleic acid reagent comprises a promoter, an insertion element, and an optional 3'UTR, and a 5'UTR / target nucleotide sequence is inserted together with the optional 3'UTR. The elements may be arranged in any order suitable for expression in a selected expression system (e.g., expression in a selected organism, or, for example, expression in a cell-free system), and in some embodiments, the nucleic acid reagent comprises the following elements in a 5'→3' direction: (1) a promoter element, a 5'UTR, and an insertion element; (2) a promoter element, a 5'UTR, and a target nucleotide sequence; (3) a promoter element, a 5'UTR, an insertion element, and a 3'UTR; and (4) a promoter element, a 5'UTR, a target nucleotide sequence, and a 3'UTR. In some embodiments, the UTRs may be optimized to alter or increase transcription or translation of the ORF, either completely natural or containing non-natural nucleotides.
[0185] Nucleic acid reagents, such as expression cassettes and / or expression vectors, may contain various regulatory elements, including promoters, enhancers, translation initiation sequences, transcription termination sequences, and other elements. A "promoter" is generally a sequence of DNA that functions when located in a relatively fixed position relative to the transcription start site. For example, a promoter may be upstream of a nucleotide triphosphate transporter nucleic acid segment. A "promoter" contains core elements required for basic interaction of RNA polymerase and transcription factors and may also contain upstream elements and response elements. An "enhancer" generally refers to a sequence of DNA that functions at a non-fixed distance from the transcription start site and can be either 5' or 3' relative to the transcription unit. Furthermore, enhancers can be present within introns and within the coding sequence itself. They are usually between 10 and 300 nucleotides in length and function in cis. Enhancers function to increase transcription from nearby promoters. Like promoters, enhancers often contain response elements that mediate transcriptional regulation. Enhancers often determine the regulation of expression and can be used to alter or optimize expression of ORFs, including ORFs that are entirely natural or contain non-natural nucleotides.
[0186] As described above, nucleic acid reagents may also contain one or more 5' UTRs and one or more 3' UTRs. For example, expression vectors used in eukaryotic host cells (e.g., yeast, fungi, insects, plants, animals, humans, or eukaryotic cells) and prokaryotic host cells (e.g., viruses, bacteria) may contain sequences that signal transcription termination, which can affect mRNA expression. These regions may be transcribed as polyadenylated segments in the untranslated portion of the mRNA encoding tissue factor protein. The 3' untranslated region also includes a transcription termination site. In some preferred embodiments, the transcription unit contains a polyadenylation region. One advantage of this region is that the transcribed unit is more likely to be processed and transported like mRNA. The identification and use of polyadenylation signals in expression constructs is well established. In some preferred embodiments, a homologous polyadenylation signal is used in transgene constructs. It can be used in various ways.
[0187] A 5' UTR may contain one or more elements endogenous to the nucleotide sequence from which it is derived, and sometimes contains one or more exogenous elements. A 5' UTR may be derived from any suitable nucleic acid, such as genomic DNA, plasmid DNA, RNA, or mRNA, for example, from any suitable organism (e.g., virus, bacteria, yeast, fungus, plant, insect, or mammal). A skilled artisan may select appropriate elements for a 5' UTR based on the selected expression system (e.g., expression in a selected organism, or, for example, expression in a cell-free system). A 5' UTR may contain one or more of the following elements known to skilled artisans: enhancer sequences (e.g., transcription or translation), transcription initiation sites, transcription factor binding sites, translational regulatory sites, translation initiation sites, translation factor binding sites, accessory protein binding sites, feedback regulator binding sites, Pribnow boxes, TATA boxes, -35 elements, E-boxes (helix-loop-helix binding elements), ribosome binding sites, replicons, internal ribosome entry sites (IRES), silencer elements, etc. In some embodiments, the promoter element may be separated such that all 5'UTR elements necessary for proper conditional regulation are contained within the promoter element fragment, or within a functional subsequence of the promoter element fragment.
[0188] The 5'UTR in a nucleic acid reagent may contain a translational enhancer nucleotide sequence. The translational enhancer nucleotide sequence is often located between the promoter and target nucleotide sequence of the nucleic acid reagent. The translational enhancer sequence often binds to ribosomes and may be an 18S rRNA-binding ribonucleotide sequence (i.e., a 40S ribosome-binding sequence) or an internal ribosome entry sequence (IRES). The IRES generally forms an RNA scaffold with a precisely positioned RNA tertiary structure that contacts the 40S ribosomal subunit through a number of specific intermolecular interactions. Examples of ribosomal enhancer sequences are known and can be identified by a skilled artisan (e.g., Mignone et al., Nucleic Acids Research 33:D141-D146 (2005); Paulous et al., Nucleic Acids Research 31:722-733 (2003); Akbergenov et al., Nucleic Acids Research 32:239-247 (2004); Mignone et al., Genome Biology 3(3):reviews0004.1-0001.10 (2002); Gallie, Nucleic Acids Research 30:3401-3411 (2002); Shaloiko et al., DOI:10.1002 / bit.20267; and Gallie et al., Nucleic Acids Research 15:3257-3273 (1987)).
[0189] The translational enhancer sequence may be a eukaryotic sequence such as a Kozak consensus sequence or other sequence (e.g., the Hydrozoa polyp sequence, GenBank Accession No. U07128). The translational enhancer sequence may be a prokaryotic sequence such as a Shine-Dalgarno consensus sequence. In certain embodiments, the translational enhancer sequence is a viral nucleotide sequence. The translational enhancer sequence may be derived from the 5'UTR of a plant virus, such as tobacco mosaic virus (TMV), alfalfa mosaic virus (AMV); tobacco etch virus (ETV); potato virus Y (PVY); turnip mosaic (poty) virus, and pea seed-borne mosaic virus. In certain embodiments, an approximately 67-base omega sequence from TMV is included in the nucleic acid reagent as the translational enhancer sequence (e.g., lacking guanosine nucleotides and containing a 25-nucleotide poly(CAA) central region).
[0190] The 3'UTR contains one or more elements endogenous to the nucleotide sequence from which it is derived. The 3'UTR may contain a sequence encoding a nucleotide ... The 3'UTR often or may not include a polyadenosine tail, and if a polyadenosine tail is present, one or more adenosine moieties may be added to or deleted from it (e.g., about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 adenosine moieties may be added to or deleted from it).
[0191] In some embodiments, modifications of the 5' UTR and / or 3' UTR are used to alter (e.g., increase, add, decrease, or substantially eliminate) promoter activity. Altering promoter activity can then alter the activity (e.g., enzymatic activity) of a peptide, polypeptide, or protein by altering transcription of a nucleotide sequence of interest from an operably linked promoter element containing the modified 5' or 3' UTR. For example, a microorganism may be genetically engineered to express a nucleic acid reagent containing a modified 5' or 3' UTR that can add a new activity (e.g., an activity not normally found in the host organism) or, in certain embodiments, increase expression of an existing activity by increasing transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest (e.g., a homologous or heterologous nucleotide sequence of interest). In some embodiments, a microorganism may be genetically engineered to express a nucleic acid reagent containing a modified 5' or 3' UTR that can, in certain embodiments, decrease expression of an activity by reducing or substantially eliminating transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest.
[0192] Expression of the nucleotide triphosphate transporter from an expression cassette or expression vector can be controlled by any promoter capable of expression in prokaryotic or eukaryotic cells. Promoter elements are typically required for DNA and / or RNA synthesis. Promoter elements often contain a region of DNA that can promote transcription of a particular gene by providing an initiation site for synthesis of the RNA corresponding to that gene. Promoters are generally located near the genes they regulate, upstream of the genes (e.g., 5' of the genes), and in some embodiments, on the same DNA strand as the sense strand of the gene. In some embodiments, promoter elements may be isolated from a gene or organism and inserted in functional association with a polynucleotide sequence to allow for altered and / or regulated expression. Non-native promoters (e.g., promoters not normally associated with a given nucleic acid sequence) used to express nucleic acids are often referred to as heterologous promoters. In certain embodiments, a heterologous promoter and / or 5'UTR can be inserted in functional association with a polynucleotide encoding a polypeptide having a desired activity as described herein. As used herein with respect to a promoter, the terms "operably linked" and "operably associated" refer to the relationship between the coding sequence and the promoter element. A promoter is operably linked or functionally associated with a coding sequence when expression from the coding sequence via transcription is regulated or controlled by the promoter element. The terms "operably linked" and "in functional association" are used interchangeably herein with respect to promoter elements.
[0193] Promoters often interact with RNA polymerase, which is a pre-existing An enzyme that catalyzes the synthesis of nucleic acids using nucleic acid reagents. When the template is a DNA template, an RNA molecule is transcribed before protein synthesis. Enzymes with polymerase activity suitable for use in the methods of the present invention include any polymerase that is active in a selected system using a selected template to synthesize a protein. In some embodiments, a promoter (e.g., a heterologous promoter), also referred to herein as a promoter element, can be operably linked to a nucleotide sequence or an open reading frame (ORF). Transcription from the promoter element can catalyze the synthesis of RNA corresponding to the nucleotide sequence or ORF sequence operably linked to the promoter, which then results in the synthesis of a desired peptide, polypeptide, or protein.
[0194] Promoter elements can be responsive to regulatory control. Promoter elements can also be regulated by selective agents. That is, transcription from a promoter element can be turned on, off, up-regulated, or down-regulated in response to changes in environmental, nutrient, or internal conditions or signals (e.g., heat-inducible promoters, light-regulated promoters, feedback-regulated promoters, hormone-influenced promoters, tissue-specific promoters, oxygen- and pH-influenced promoters, promoters responsive to selective agents (e.g., kanamycin)). Promoters that are influenced by environmental, nutrient, or internal signals are often influenced by signals (direct or indirect) that bind to or near the promoter and increase or decrease expression of the target sequence under certain conditions. As with all methods disclosed herein, the inclusion of native or modified promoters can be used to alter or optimize expression of fully native ORFs (e.g., NTT or aaRS) or ORFs containing non-natural nucleotides (e.g., mRNA or tRNA).
[0195] Non-limiting examples of selective agents or modulators that affect transcription from promoter elements for use in the embodiments described herein include, but are not limited to, (1) nucleic acid segments encoding products that confer resistance to an otherwise toxic compound (e.g., an antibiotic); (2) nucleic acid segments encoding products that are otherwise lacking in the recipient cell (e.g., an essential product, a tRNA gene, an auxotrophic marker); (3) nucleic acid segments encoding products that repress the activity of a gene product; (4) nucleic acid segments encoding products that can be easily identified (e.g., phenotypic markers such as antibiotics (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and cell surface proteins); (5) nucleic acid segments that bind to products that are otherwise deleterious to cell survival and / or function; and (6) nucleic acid segments that otherwise inhibit the activity of any of the nucleic acid segments listed above (e.g., antisense oligonucleotides). (7) nucleic acid segments that bind to a product that modifies a substrate (e.g., a restriction endonuclease); (8) nucleic acid segments that can be used to isolate or identify a desired molecule (e.g., a specific protein binding site); (9) nucleic acid segments that encode specific nucleotide sequences that would not otherwise function (e.g., for PCR amplification of a subpopulation of molecules); (10) nucleic acid segments that, when absent, directly or indirectly confer resistance or sensitivity to a particular compound; (11) nucleic acid segments that encode products that convert toxic or relatively non-toxic compounds into toxic compounds (e.g., herpes simplex thymidine kinase, cytosine deaminase) in recipient cells; (12) nucleic acid segments that inhibit the replication, partitioning, or heritability of the nucleic acid molecule that contains them; (13) nucleic acid segments that encode conditional replication functions, e.g., replication in a particular host or host cell line or under particular environmental conditions (e.g., temperature, nutrient conditions, etc.); and / or (14) nucleic acids that encode one or more mRNAs or tRNAs that contain unnatural nucleotides. In some embodiments, regulatory or selective The agent may be added to alter the existing growth conditions to which the organism is subjected (eg, growth in liquid culture, growth in a fermentor, growth on solid nutrient plates, etc.).
[0196] In some embodiments, modulation of promoter elements can be used to alter (e.g., increase, add, decrease, or substantially eliminate) the activity of a peptide, polypeptide, or protein (e.g., enzymatic activity, etc.). For example, a microorganism may be genetically engineered to express a nucleic acid reagent that can add a new activity (e.g., an activity not normally found in the host organism) or, in certain embodiments, increase the expression of an existing activity by increasing transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest (e.g., a homologous or heterologous nucleotide sequence of interest). In some embodiments, a microorganism may be genetically engineered to express a nucleic acid reagent that can, in certain embodiments, decrease the expression of an activity by reducing or substantially eliminating transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest.
[0197] Nucleic acids encoding heterologous proteins, such as nucleotide triphosphate transporters, may be inserted into or used in any suitable expression system. In some embodiments, nucleic acid reagents are sometimes stably integrated into the chromosome of the host organism, or nucleic acid reagents may, in certain embodiments, be deletions of portions of the host chromosome (e.g., in genetically modified organisms, modifications of the host genome confer the ability to selectively or preferentially maintain desired organisms with genetic modifications). Such nucleic acid reagents (e.g., nucleic acids or genetically modified organisms whose modified genomes confer selectable traits on organisms) can be selected for their ability to direct the production of desired proteins or nucleic acid molecules. Optionally, nucleic acid reagents may be modified so that codons (i) use different tRNAs than those specified in the native sequence to encode the same amino acids, or (ii) encode unusual amino acids, including unconventional or unnatural amino acids (including detectably labeled amino acids).
[0198] Recombinant expression is usefully achieved using an expression cassette, which may be part of a vector such as a plasmid. The vector may include a promoter operably linked to the nucleic acid encoding the nucleotide triphosphate transporter. The vector may also include other elements necessary for transcription and translation, as described herein. The expression cassette, expression vector, and sequences in the cassette or vector may be heterologous to the cell with which the non-natural nucleotide is contacted. For example, the nucleotide triphosphate transporter sequence may be heterologous to the cell.
[0199] Various prokaryotic and eukaryotic expression vectors suitable for carrying, encoding, and / or expressing nucleotide triphosphate transporters can be generated. Examples of such expression vectors include pET, pET3d, pCR2.1, pBAD, pUC, and yeast vectors. The vectors can be used, for example, in various in vivo and in vitro contexts. Non-limiting examples of prokaryotic promoters that can be used include SP6, T7, T5, tac, bla, trp, gal, lac, or maltose promoters. Non-limiting examples of eukaryotic promoters that can be used include constitutive promoters, such as viral promoters, such as CMV, SV40, and RSV promoters, and regulatable promoters, such as inducible or repressible promoters, such as tet promoters, hsp70 promoters, and synthetic promoters regulated by CRE. Examples of vectors for bacterial expression include pGEX-5X-3, and examples of vectors for eukaryotic expression include pCIneo-CMV. Viral vectors that can be used include lentivirus, adenovirus, adeno-associated virus, herpes virus, vaccinia virus, polio virus, AIDS virus, neurotrophic virus, and the like. Retroviral vectors include those related to the genus avian, Sindbis, and other viruses. Any virus family that shares the properties of these viruses and is suitable for use as a vector is also useful. Examples of retroviral vectors that can be used include those described in Verma, American Society for Microbiology, pp. 229-232, Washington, (1985). For example, such retroviral vectors include murine Maloney leukemia virus, MMLV, and other retroviruses that express desired characteristics. Viral vectors typically contain nonstructural early genes, structural late genes, RNA polymerase III transcripts, inverted terminal repeats required for replication and encapsidation, and promoters that control transcription and replication of the viral genome. When designed as a vector, viruses typically have one or more early genes removed, and a gene or gene / promoter cassette is inserted into the viral genome in place of the removed viral nucleic acid.
[0200] Cloning Any convenient cloning strategy known in the art may be used to incorporate elements such as ORFs into a nucleic acid reagent. Elements may be inserted into a template independently of the inserted element using known methods, such as (1) cleaving the template at one or more existing restriction enzyme sites and ligating the desired element, and (2) adding restriction enzyme sites to the template by hybridizing oligonucleotide primers containing one or more appropriate restriction enzyme sites and amplifying by polymerase chain reaction (described in more detail herein). Other cloning strategies utilize one or more insertion sites present in or inserted into the nucleic acid reagent, such as oligonucleotide primer hybridization sites for PCR and others described herein. In some embodiments, a cloning strategy may be combined with genetic manipulation, such as recombination (e.g., recombination of a nucleic acid reagent bearing a nucleic acid sequence of interest into the genome of an organism to be modified, as further described herein). In some embodiments, the cloned ORFs can produce modified or wild-type nucleotide triphosphate transporters and / or polymerases (directly or indirectly) by engineering a microorganism with one or more ORFs of interest, such that the microorganism contains altered nucleotide triphosphate transporter activity or polymerase activity.
[0201] A nucleic acid can be specifically cleaved by contacting the nucleic acid with one or more specific cleaving agents, which often specifically cleave at specific sites and according to specific nucleotide sequences. Examples of enzyme-specific cleaving agents include, but are not limited to, endonucleases (e.g., DNases (e.g., DNase I, II); RNases (e.g., RNase E, F, H, P); Cleavase™ enzyme; Taq DNA polymerase; E. coli DNA polymerase I and eukaryotic structure-specific endonucleases; mouse FEN-1 endonuclease; type I, type II, or type III restriction endonucleases, such as Acc I, Afl III, Alu I, Alw44 I, Apa I, Asn I, Ava I, Ava II, BamH I, Ban II, Bcl I, Bgl I, Bgl II, Bln I, Bsa I, Bsm I, BsmBI, BssH II, BstE II, Cfo I, CIa I, Dde I, Dpn I, Dra I, EcIX I, EcoR I, EcoR II, EcoR V, Hae II, Hae II, Hind II, Hind III, Hpa I, Hpa II, Kpn I, Ksp I, Mlu I, MIuN I, Msp I, Nci I, Nco I, Nde I, Nde II, Nhe I, Not I, Nru I, Nsi I, Pst I, Pvu I, Pvu II, Rsa I, Sac I, Sal I, Sau3A I, Sca I, ScrF I, Sfi I, Sma I, Spe I, Sph I, Ssp I, Stu I, Sty I, Swa I, Taq I, glycosylases (e.g., uracil-DNA glycosylase (UDG), 3-methyladenine DNA glycosylase, 3-methyladenine DNA glycosylase II, pyrimidine hydrate-DNA glycosylase, FaPy-DNA glycosylase, thymine mismatch-DNA glycosylase, hypoxanthine-DNA glycosylase, 5-hydroxymethyluracil DNA glycosylase (HmUDG), 5-hydroxymethylcytosine DNA glycosylase, or 1,N6-etheno-adenine DNA glycosylase); exonucleases (e.g., exonuclease III); ribozymes, and DNAzymes. The sample nucleic acid may be treated with chemical agents or synthesized using modified nucleotides, and the modified nucleic acid may be cleaved. In a non-limiting example, the sample nucleic acid can be treated with (i) an alkylating agent such as methylnitrosourea, which generates several alkylated bases, including N3-methyladenine and N3-methylguanine, which are recognized and cleaved by alkylpurine DNA glycosylase; (ii) sodium bisulfite, which deaminates cytosine residues in DNA to form uracil residues that can be cleaved by uracil N-glycosylase; and (iii) a chemical agent that converts guanine to its oxidized form, 8-hydroxyguanine, which can be cleaved by formamidopyrimidine DNA N-glycosylase. Examples of chemical cleavage processes include, but are not limited to, alkylation (e.g., alkylation of phosphorothioate-modified nucleic acids); acid-labile cleavage of P3'-N5'-phosphoramidate-containing nucleic acids; and osmium tetroxide and piperidine treatment of nucleic acids.
[0202] In some embodiments, the nucleic acid reagent comprises one or more recombinase insertion sites. Recombinase insertion sites are recognition sequences on nucleic acid molecules that participate in integration / recombination reactions by recombination proteins. For example, the recombination site for Cre recombinase is loxP, a 34-base pair sequence consisting of two 13-base pair inverted repeats (which function as recombinase binding sites) flanking an 8-base pair core sequence (e.g., Sauer, Curr. Opin. Biotech. 5:521-527 (1994)). Other examples of recombination sites include the attB, attP, attL, and attR sequences, as well as mutants, fragments, variants, and derivatives thereof, recognized by the recombination protein λInt and by the auxiliary proteins integration host factor (IHF), FIS, and excision enzyme (Xis) (e.g., U.S. Pat. Nos. 5,888,732; 6,143,557; 6,171,861; 6,270,969; 6,277,608; and 6,720,140; U.S. Patent Application Publication Nos. 09 / 517,466 and 09 / 732,914; U.S. Patent Application Publication No. 2002 / 0007051; and Landy, Curr. Opin. Biotech. 3:699-707 (1993)).
[0203] An example of a recombinase cloning nucleic acid is the Gateway® system (Invitrogen, California), which contains at least one recombination site for cloning a desired nucleic acid molecule in vivo or in vitro. In some embodiments, this system utilizes a vector containing at least two different site-specific recombination sites, often based on the bacteriophage lambda system (e.g., att1 and att2), mutated from the wild-type (att0) site. Each mutated site has unique specificity for its cognate partner att site (i.e., its binding partner recombination site) of the same type (e.g., attB1 and attP1, or attL1 and attR1) and does not cross-react with other mutated recombination sites or with the wild-type att0 site. The different site specificities allow for directional cloning or ligation of the desired molecule, thus providing a desired orientation of the cloned molecule. Nucleic acid fragments flanked by recombination sites are ligated using the Gateway® system into a recipient plasmid, sometimes referred to as a Destination Vector. Cloning and subcloning are performed by replacing a selectable marker (e.g., ccdB) adjacent to the att site on the molecule. Desired clones are then selected by transformation of a ccdB-sensitive host strain and positive selection for the marker on the recipient molecule. Similar strategies for negative selection (e.g., the use of toxic genes) can be used in other organisms, such as thymidine kinase (TK) in mammals and insects.
[0204] A nucleic acid reagent may contain one or more origin of replication (ORI) elements. In some embodiments, a template contains two or more ORIs, one that functions efficiently in one organism (e.g., bacteria) and another that functions efficiently in another organism (e.g., a eukaryote such as yeast). In some embodiments, an ORI may function efficiently in one species (e.g., S. cerevisiae, etc.) and another ORI may function efficiently in a different species (e.g., S. pombe, etc.). A nucleic acid reagent may also contain one or more transcriptional regulatory sites.
[0205] A nucleic acid reagent, such as an expression cassette or vector, can contain a nucleic acid sequence encoding a marker product. The marker product is used to determine whether a gene has been delivered to a cell and is expressed after delivery. Examples of marker genes include the E. coli lacZ gene, which encodes β-galactosidase and green fluorescent protein. In some embodiments, the marker can be a selectable marker. When such a selectable marker is successfully transferred to a host cell, the transformed host cell can survive when placed under selective pressure. There are two widely used distinct categories of selection regimes. The first category is based on cellular metabolism and the use of mutant cell lines that lack the ability to grow independently of supplemented media. The second category is dominant selection, which refers to a selection scheme used with any cell type and does not require the use of mutant cell lines. These schemes typically use drugs to inhibit host cell growth. Those cells carrying the novel gene express a protein conferring drug resistance and survive selection. Examples of such dominant selection use the drugs neomycin (Southern et al., J. Molec. Appl. Genet. 1:327 (1982)), mycophenolic acid (Mulligan et al., Science 209:1422 (1980)), or hygromycin (Sugden et al., Mol. Cell. Biol. 5:410-413 (1985)).
[0206] A nucleic acid reagent can contain one or more selection elements (e.g., an element for selecting for the presence of the nucleic acid reagent, but not for the activation of a promoter element that can be selectively regulated). Selection elements are often utilized using known processes to determine whether a nucleic acid reagent is contained in a cell. In some embodiments, a nucleic acid reagent contains two or more selection elements, one element that functions efficiently in one organism and the other element that functions efficiently in another organism. Examples of selection elements include, but are not limited to, (1) nucleic acid segments encoding products that confer resistance to an otherwise toxic compound (e.g., an antibiotic); (2) nucleic acid segments encoding products otherwise lacking in the recipient cell (e.g., an essential product, a tRNA gene, an auxotrophic marker); (3) nucleic acid segments encoding products that suppress the activity of a gene product; (4) nucleic acid segments encoding products that can be easily identified (e.g., phenotypic markers such as antibiotics (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and cell surface proteins); (5) nucleic acid segments that bind products that are otherwise deleterious to cell survival and / or function; (6) nucleic acid segments that otherwise inhibit the activity of any of the nucleic acid segments listed above (e.g., antisense oligonucleotides); (7) nucleic acid segments that bind products that modify substrates (e.g., restriction endonucleases). (8) nucleic acid segments that can be used to isolate or identify a desired molecule (e.g., a specific protein binding site); (9) nucleic acid segments that encode specific nucleotide sequences that may not otherwise function (e.g., for PCR amplification of a subpopulation of molecules); (10) nucleic acid segments that, when absent, directly or indirectly confer resistance or sensitivity to a particular compound; (11) nucleic acid segments that encode products that are toxic or that convert relatively non-toxic compounds into toxic compounds in recipient cells (e.g., herpes simplex thymidine kinase, cytosine deaminase); (12) nucleic acid segments that inhibit the replication, partitioning, or heritability of the nucleic acid molecule that contains them; and / or (13) nucleic acid segments that encode conditional replication functions, such as replication in a particular host or host cell line or under particular environmental conditions (e.g., temperature, nutrient status, etc.).
[0207] Nucleic acid reagents can be in any form useful for in vivo transcription and / or translation. Nucleic acids can be plasmids, such as supercoiled plasmids, yeast artificial chromosomes (e.g., YACs), linear nucleic acids (e.g., linear nucleic acids produced by PCR or restriction digestion), single-stranded, or sometimes double-stranded. Nucleic acid reagents can also be prepared by amplification processes, such as polymerase chain reaction (PCR) or transcription-mediated amplification (TMA). In TMA, two enzymes are used in an isothermal reaction to generate amplification products that are detected by luminescence (e.g., Biochemistry 1996 Jun 25;35(25):8429-38). Standard PCR processes are known (e.g., U.S. Patent Nos. 4,683,202; 4,683,195; 4,965,188; and 5,565,493), and are generally performed in cycles. Each cycle includes heat denaturation (where hybrid nucleic acids dissociate), cooling (where primer oligonucleotides hybridize), and oligonucleotide extension by a polymerase (i.e., Taq polymerase). An example of a PCR cycling process is treating a sample at 95°C for 5 minutes; repeating 45 cycles of 95°C for 1 minute, 59°C for 1 minute and 10 seconds, and 72°C for 1 minute and 30 seconds; then treating the sample at 72°C for 5 minutes. Multiple cycles are often performed using commercially available thermal cyclers. PCR amplification products may be temporarily stored at low temperatures (e.g., 4°C) or frozen (e.g., -20°C) before analysis.
[0208] DNA containing unnatural nucleotides can be generated using cloning strategies similar to those described above. For example, an oligonucleotide containing an unnatural nucleotide at the desired position is synthesized using standard solid-phase synthesis and purified by HPLC. The oligonucleotide is then inserted into a plasmid containing the desired sequence context (i.e., UTR and coding sequence) using a cloning method (such as Golden Gate assembly) using a cloning site such as the BsaI site (although others as described above may also be used).
[0209] Kits and Manufactured Products In certain embodiments, disclosed herein are kits and articles of manufacture for use in one or more of the methods described herein. Such kits include a carrier, package, or container compartmentalized to receive one or more containers, such as vials, tubes, etc., each of which contains one of the distinct elements used in the methods described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. In one embodiment, the containers are formed from a variety of materials, such as glass or plastic.
[0210] In some embodiments, the kit includes suitable packaging material for housing the contents of the kit. In some cases, the packaging material preferably provides a sterile, contaminant-free environment. The packaging materials used herein include, for example, those commonly used in commercially available kits sold for use in nucleic acid sequencing systems. Exemplary packaging materials include, but are not limited to, glass, plastic, paper, foil, and the like, capable of retaining the components described herein within a certain range.
[0211] The packaging material may include a label indicating the specific use of the components. The use of the kit indicated by the label may be one or more of the methods described herein that are appropriate for the particular combination of components present in the kit. For example, the label may indicate that the kit is useful for a method of synthesizing polynucleotides or for a method of sequencing nucleic acids.
[0212] Instructions for use of the packaged reagents or components may also be included in the kit, and typically include specific wording describing reaction parameters such as the relative amounts of kit components and sample to be mixed, maintenance periods for the reagent / sample mixture, temperature, buffer conditions, etc.
[0213] It is understood that not all components required for a particular reaction need be present in a particular kit. Rather, one or more additional components may be provided from other sources. Instructions accompanying the kit may identify the additional components provided and where they can be obtained.
[0214] In some embodiments, kits are provided that are useful for stably incorporating non-native nucleic acids into cellular nucleic acids, for example, using the methods provided by the present disclosure for preparing genetically engineered cells. In one embodiment, the kits described herein include genetically engineered cells and one or more non-native nucleic acids.
[0215] In additional embodiments, the kits described herein provide cells and nucleic acid molecules comprising heterologous genes for introduction into the cells, thereby providing genetically engineered cells, such as expression vectors comprising the nucleic acids of any of the above embodiments described in this paragraph.
[0216] Numbered Embodiments. The present disclosure includes the following non-limiting numbered embodiments: Embodiment 1. A method of synthesizing a non-naturally occurring polypeptide, comprising: a. providing at least one unnatural deoxyribonucleic acid (DNA) molecule comprising at least four unnatural base pairs; b. transcribing at least one non-naturally occurring DNA molecule to provide a messenger RNA (mRNA) molecule comprising at least two non-naturally occurring codons; c. transcribing at least one unnatural DNA molecule to provide at least two transfer RNA (tRNA) molecules each containing at least one unnatural anticodon, wherein at least two unnatural base pairs in the corresponding DNA are in a sequence context such that the unnatural codon of the mRNA molecule is complementary to each unnatural anticodon of the tRNA molecule; and d. synthesizing a non-natural polypeptide by translating a non-natural mRNA molecule utilizing at least two non-natural tRNA molecules, wherein each non-natural anticodon directs the site-specific incorporation of an non-natural amino acid into the non-natural polypeptide. A method comprising: Embodiment 1.1. A method for synthesizing a non-naturally occurring polypeptide, comprising: a. providing at least one unnatural deoxyribonucleic acid (DNA) molecule comprising at least four unnatural base pairs; b. Transcribing at least one unnatural DNA molecule to include at least two unnatural codons providing a messenger RNA (mRNA) molecule comprising: c. transcribing at least one non-natural DNA molecule to provide at least two transfer RNA (tRNA) molecules each containing at least one non-natural anticodon; at least two unnatural base pairs in the corresponding DNA are in a sequence context such that one of the unnatural codons of the mRNA molecule is complementary to the unnatural anticodon of one of the tRNA molecules and at least one of the one or more other unnatural codons is complementary to at least one unnatural anticodon of the other tRNA molecule; and d. synthesizing a non-natural polypeptide by translating a non-natural mRNA molecule utilizing at least two non-natural tRNA molecules, wherein each non-natural anticodon directs the site-specific incorporation of an non-natural amino acid into the non-natural polypeptide. A method comprising: Embodiment 2. A method of synthesizing a non-naturally occurring polypeptide, comprising: a. providing at least one unnatural deoxyribonucleic acid (DNA) molecule comprising at least four unnatural base pairs, wherein the at least one unnatural DNA molecule encodes (i) a messenger RNA (mRNA) molecule comprising at least first and second unnatural codons, and (ii) at least first and second transfer RNA (tRNA) molecules, wherein the first tRNA molecule comprises a first unnatural anticodon and the second tRNA molecule comprises a second unnatural anticodon, and wherein the at least four unnatural base pairs in the at least one DNA molecule are in a sequence context such that the first and second unnatural codons of the mRNA molecule are complementary to the first and second unnatural anticodons, respectively; b. transcribing at least one non-naturally occurring DNA molecule to provide mRNA; c. transcribing at least one non-natural DNA molecule to provide at least first and second tRNA molecules; and d. synthesizing a non-natural polypeptide by translating a non-natural mRNA molecule utilizing at least first and second non-natural tRNA molecules, wherein each of the at least first and second non-natural anticodons directs the site-specific incorporation of an non-natural amino acid into the non-natural polypeptide. A method comprising: Embodiment 3. The method of embodiment 1, 1.1., or 2, wherein the at least two unnatural codons each comprise a first unnatural nucleotide located at the first, second, or third position of the codon, and optionally, the first unnatural nucleotide is located at the second or third position of the codon. Embodiment 4. The method of any one of the preceding embodiments, wherein the at least two unnatural codons each comprise the nucleic acid sequence NNX or NXN, and the unnatural anticodon comprises the nucleic acid sequence XNN, YNN, NXN, or NYN, thereby forming an unnatural codon-anticodon pair comprising NNX-XNN, NNX-YNN, or NXN-NYN, where N is any naturally occurring nucleotide, X is a first unnatural nucleotide, and Y is a second unnatural nucleotide different from the first unnatural nucleotide, and XY or XX forms an unnatural base pair in DNA. Embodiment 4.1. The method of any one of the preceding embodiments, wherein the at least two unnatural codons each comprise the nucleic acid sequence XNN, NXN, NNX, and the unnatural anticodon comprises the nucleic acid sequence NNX, NNY, NXN, NYN, NNX, or NNY, thereby forming an unnatural codon-anticodon pair comprising XNN-NNX, XNN-NNY, NXN-NXN, NXN-NYN, NNX-XNN, or NNX-YNN, where N is any naturally occurring nucleotide, X is a first unnatural nucleotide, and Y is a second unnatural nucleotide different from the first unnatural nucleotide, and wherein X-X or X-Y forms an unnatural base pair in DNA. Embodiment 5. The method of embodiment 4, wherein the codon comprises at least one G or C and the anticodon comprises at least one complementary C or G. Embodiment 6. X and Y are: (i) 2-thiouracil, 2'-deoxyuridine, 4-thiouracil, uracil-5-yl, hypoxanthine-9-yl(I), 5-halouracil; 5-propynyl-uracil, 6-azo-uracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, uracil-5-oxaacetic acid methyl ester, uracil-5-oxaacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, uracil-5-oxaacetic acid, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, or dihydrouracil; (ii) 5-hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytosine arabinoside, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azocytosine, azacytosine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1,4]benzoxazine-2 (3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one), or pyridoindole cytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one); (iii) 2-aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2'-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted adenines, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-azaadenine; (iv) 2-methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thioguanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted guanines, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-azaguanine; and (v) hypoxanthine, xanthine, 1-methylinosine, queosine, beta-D-galactosylqueosine, inosine, beta-D-mannosylqueosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone 6. The method of embodiment 4 or 5, wherein the IL-14 is independently selected from the group consisting of:
[0217] Embodiment 7. The bases comprising each of X and Y are: [ka] 6. The method of embodiment 4 or 5, wherein the IL-14 is independently selected from the group consisting of:
[0218] Embodiment 8. Each X-containing base is [ka] 8. The method of embodiment 7, wherein
[0219] Embodiment 9. Each Y-containing base is [ka] 9. The method of embodiment 7 or 8, wherein
[0220] Embodiment 10. The method of any one of embodiments 4-9, wherein NNX-XNN is selected from the group consisting of UUX-XAA, UGX-XCA, CGX-XCG, AGX-XCU, GAX-XUC, CAX-XUG, AUX-XAU, CUX-XAG, GUX-XAC, UAX-XUA, and GGX-XCC. Embodiment 11. The method of any one of embodiments 4 to 9, wherein NNX-YNN is selected from the group consisting of UUX-YAA, UGX-YCA, CGX-YCG, AGX-YCU, GAX-YUC, CAX-YUG, AUX-YAU, CUX-YAG, GUX-YAC, UAX-YUA, and GGX-YCC. Embodiment 12. The method of any one of embodiments 4 to 9, wherein NXN-NYN is selected from the group consisting of GXU-AYC, CXU-AYG, GXG-CYC, AXG-CYU, GXC-GYC, AXC-GYU, GXA-UYC, CXC-GYG, and UXC-GYA.Embodiment 13. The method of embodiment 12, wherein NXN-NYN is selected from the group consisting of AXG-CYU, GXC-GYC, AXC-GYU, GXA-UYC, CXC-GYG, and UXC-GYA. Embodiment 13.1. The method of any one of embodiments 4.1 to 9, wherein XNN-NNY is selected from the group consisting of XUU-AAY, XUG-CAY, XCG-CGY, XAG-CUY, XGA-UCY, XCA-UGY, XAU-AUY, XCU-AGY, XGU-ACY, XUA-UAY, XUC-GAY, XCC-GGY, XAA-UUY, XAC-GUY, XGC-GCY, XGG-CCY, and XGG-CCY. Embodiment 13.2. The method of any one of embodiments 4.1 to 9, wherein XNN-NNX is selected from the group consisting of XUU-AAX, XUG-CAX, XCG-CGX, XAG-CUX, XGA-UCX, XCA-UGX, XAU-AUX, XCU-AGX, XGU-ACX, XUA-UAX, XUC-GAX, XCC-GGX, XAA-UUX, XAC-GUX, XGC-GCX, XGG-CCX, and XGG-CCX. Embodiment 14. The method of any one of the preceding embodiments, wherein the at least two non-natural tRNA molecules each comprise a different non-natural anticodon. Embodiment 15. The method of embodiment 14, wherein the at least two non-natural tRNA molecules comprise a pyrrolysyl-tRNA from Methanosarcina and a tyrosyl-tRNA from Methanocaldococcus jannaschii, or a derivative thereof. Embodiment 16. The method of any one of embodiments 13, 14, or 15, comprising charging at least two unnatural tRNA molecules with an aminoacyl-tRNA synthetase. Embodiment 17. The method of embodiment 16, wherein the aminoacyl-tRNA synthetase is selected from the group consisting of chimeric PylRS (chPylRS) and M. jannaschii AzFRS (MjpAzFRS). Embodiment 18. The method of embodiment 14 or 15, comprising charging at least two non-natural tRNA molecules with at least two tRNA synthetases. Embodiment 19. The method of embodiment 18, wherein the at least two tRNA synthetases comprise a chimeric PylRS (chPylRS) and a M. jannaschii AzFRS (MjpAzFRS). Embodiment 20. The method of any one of embodiments 1-19, wherein the non-naturally occurring polypeptide comprises two, three, or more non-naturally occurring amino acids. Embodiment 21. The method of any one of embodiments 1-20, wherein the non-natural polypeptide comprises at least two non-natural amino acids that are the same. Embodiment 22. The method of any one of embodiments 1-20, wherein the non-natural polypeptide comprises at least two different non-natural amino acids. Embodiment 23. The unnatural amino acid is Lysine analogues; aromatic side chains; Azide group; an alkyne group; or Aldehyde or ketone group 23. The method of any one of embodiments 1 to 22, comprising: Embodiment 24. The method of any one of embodiments 1-22, wherein the unnatural amino acid does not comprise an aromatic side chain. Embodiment 25. The unnatural amino acid is N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propargylethoxy-carbonyl-L-lysine (PraK), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azidophenylalanine. (pAzF), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenyl Alanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)amino)amino 23. The method of any one of embodiments 1-22, wherein the hydroxybenzoate is selected from N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine. Embodiment 26. The method of any one of the preceding embodiments, wherein at least one non-naturally occurring DNA molecule is in the form of a plasmid. Embodiment 27. The method of any one of embodiments 1 to 26, wherein at least one non-naturally occurring DNA molecule is integrated into the genome of the cell. Embodiment 28. The method of embodiment 26 or 27, wherein at least one non-naturally occurring DNA molecule encodes a non-naturally occurring polypeptide. Embodiment 29. The method of any one of the preceding embodiments, comprising in vivo replication and transcription of the non-naturally occurring DNA molecule in a cellular organism, and in vivo translation of the transcribed mRNA molecule. Embodiment 30. The method of embodiment 29, wherein the cellular organism is a microorganism. Embodiment 31. The method of embodiment 30, wherein the cellular organism is a prokaryote. Embodiment 32. The method of embodiment 31, wherein the cellular organism is a bacterium. Embodiment 33. The method of embodiment 32, wherein the cellular organism is a Gram-positive bacterium. Embodiment 34. The method of embodiment 32, wherein the cellular organism is a Gram-negative bacterium. Embodiment 35. The method of embodiment 34, wherein the cellular organism is E. coli. Embodiment 36. The method of any one of the preceding embodiments, wherein the at least two unnatural base pairs comprise a base pair selected from dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1, or dNaM-dTAT1. Embodiment 37. The method of any one of embodiments 29-36, wherein the cellular organism comprises a nucleoside triphosphate transporter. Embodiment 38. The method of embodiment 37, wherein the nucleoside triphosphate transporter comprises the amino acid sequence of PtNTT2. Embodiment 39. The method of embodiment 38, wherein the nucleoside triphosphate transporter comprises a truncated amino acid sequence of PtNTT2. Embodiment 40. The method of embodiment 39, wherein the truncated amino acid sequence of PtNTT2 is at least 80% identical to PtNTT2 encoded by SEQ ID NO:1. Embodiment 41. The method of any one of embodiments 29-40, wherein the cellular organism comprises at least one non-naturally occurring DNA molecule. Embodiment 42. The method of embodiment 41, wherein the at least one non-naturally occurring DNA molecule comprises at least one plasmid. Embodiment 43. The method of embodiment 42, wherein at least one non-naturally occurring DNA molecule is integrated into the genome of the cell. Embodiment 44. At least one non-naturally occurring DNA molecule encodes a non-naturally occurring polypeptide. 44. The method of embodiment 42 or 43. Embodiment 45. The method of any one of embodiments 1 to 26, which is an in vitro method comprising synthesizing the non-naturally occurring polypeptide in a cell-free system. Embodiment 46. The method of any one of the preceding embodiments, wherein the unnatural base pair comprises at least one unnatural nucleotide comprising an unnatural sugar moiety. Embodiment 47. The unnatural sugar moiety comprises: OH, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH, 2'-O(CH)OCH, where alkyl, alkenyl and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n -NH2 and -O(CH2) n ON[(CH2) n CH3)]2, where n and m are from 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S); Modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving group, reporter group, intercalator, group for improving the pharmacokinetic properties of an oligonucleotide, or group for improving the pharmacodynamic properties of an oligonucleotide, and any combination thereof. 47. The method of embodiment 46, comprising a moiety selected from the group consisting of: Embodiment 48. A cell comprising at least one unnatural DNA molecule comprising at least four unnatural base pairs, wherein the at least one unnatural DNA molecule encodes (i) a messenger RNA (mRNA) molecule encoding an unnatural polypeptide, the messenger RNA (mRNA) molecule comprising at least a first and a second unnatural codon, and (ii) at least a first and a second transfer RNA (tRNA) molecule, wherein the first tRNA molecule comprises a first unnatural anticodon and the second tRNA molecule comprises a second unnatural anticodon, and wherein the at least four unnatural base pairs in the at least one DNA molecule are in a sequence context such that the first and second unnatural codons of the mRNA molecule are complementary to the first and second unnatural anticodons, respectively. Embodiment 49. The cell of embodiment 48, further comprising an mRNA molecule and at least a first and a second tRNA molecule. Embodiment 50. The cell of embodiment 49, wherein at least the first and second tRNA molecules are covalently linked to an unnatural amino acid. Embodiment 51. The cell of embodiment 50, further comprising a non-naturally occurring polypeptide. Embodiment 52. A cell comprising: a. at least two different unnatural codon-anticodon pairs, each unnatural codon-anticodon pair comprising an unnatural codon from an unnatural messenger RNA (mRNA) and an unnatural anticodon from an unnatural transfer ribonucleic acid (tRNA), wherein the unnatural codon comprises a first unnatural nucleotide and the unnatural anticodon comprises a second unnatural nucleotide; and b. at least two different unnatural amino acids, each covalently attached to a corresponding unnatural tRNA Cells containing Embodiment 53. The cell of embodiment 52, further comprising at least one unnatural DNA molecule comprising at least four unnatural base pairs (UBPs). Embodiment 54. The cell of any one of embodiments 48 to 53, wherein the first non-natural nucleotide is positioned at the second or third position of the non-natural codon. Embodiment 54.1. The cell of any one of embodiments 48 to 53, wherein the first non-natural nucleotide is positioned at the first, second, or third position of the non-natural codon. Embodiment 55. The cell of embodiment 54 or 54.1, wherein the first non-natural nucleotide complementarily base pairs with the second non-natural nucleotide of the non-natural anticodon.
[0221] Embodiment 56. The first non-natural nucleotide and the second non-natural nucleotide are: [ka] 56. The cell of any one of embodiments 48-55, wherein the cell comprises a first and a second base each independently selected from the group consisting of:
[0222] Embodiment 57. The cell of any one of embodiments 48 or 50-56, wherein the at least four unnatural base pairs are independently selected from the group consisting of dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1, or dNaM-dTAT1. Embodiment 58. The cell of any one of embodiments 48 or 50-57, wherein the at least one non-naturally occurring DNA molecule comprises at least one plasmid. Embodiment 59. The cell of any one of embodiments 48 or 50-58, wherein at least one non-native DNA molecule is integrated into the genome of the cell. Embodiment 60. The cell of any one of embodiments 50-59, wherein at least one non-naturally occurring DNA molecule encodes a non-naturally occurring polypeptide. Embodiment 61. The cell of any one of embodiments 48 to 60, wherein the cell expresses a nucleoside triphosphate transporter. Embodiment 62. The cell of embodiment 61, wherein the nucleoside triphosphate transporter comprises the amino acid sequence of PtNTT2. Embodiment 63. The method of embodiment 62, wherein the nucleoside triphosphate transporter comprises a truncated amino acid sequence of PtNTT2. Embodiment 64. The method of embodiment 63, wherein the truncated amino acid sequence of PtNTT2 is at least 80% identical to PtNTT2 encoded by SEQ ID NO:1. Embodiment 65. The cell of any one of embodiments 48 to 64, wherein the cell expresses at least two tRNA synthetases. Embodiment 66. The cell of embodiment 65, wherein the at least two tRNA synthetases are a chimeric PylRS (chPylRS) and a M. jannaschii AzFRS (MjpAzFRS). Embodiment 67. The cell of any one of embodiments 48-66, wherein the cell comprises an unnatural nucleotide comprising an unnatural sugar moiety. Embodiment 68. The unnatural sugar moiety comprises: Modifications at the 2' position: OH, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH, 2'-O(CH)OCH, where alkyl, alkenyl and alkynyl are substituted or unsubstituted C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n -NH2 and -O(CH2) n ON[(CH2) n CH3)]2, where n and m are from 1 to about 10; and / or modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S); Modifications at the 4' position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving group, reporter group, intercalator, group for improving the pharmacokinetic properties of an oligonucleotide, or group for improving the pharmacodynamic properties of an oligonucleotide, and any combination thereof. 68. The cell of embodiment 67, selected from the group consisting of: Embodiment 69. The cell of any one of embodiments 48 to 68, wherein at least one unnatural nucleotide base is recognized by an RNA polymerase during transcription. Embodiment 70. The cell of any one of embodiments 48-69, wherein the cell translates at least one unnatural polypeptide comprising at least two unnatural amino acids. Embodiment 71. The at least two unnatural amino acids are N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propargylethoxy-carbonyl-L-lysine (PraK), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azidophenylalanine (pAzF), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino lysine. 8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri- O-Acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl 71. The cell of any one of embodiments 48-70, wherein the N-(((4-azidobenzyl)oxy)carbonyl)-L-lysine is independently selected from the group consisting of N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine, ... and N6-(((4-azidobenzyl)oxy)carbonyl)-L- Embodiment 72. The cell of any one of embodiments 48 to 71, which is isolated. Embodiment 73. The cell of any one of embodiments 48 to 72, which is a prokaryote. Embodiment 74. A cell line comprising the cell of any one of embodiments 48 to 73. [Example]
[0223] Example 1. Initial codon screening As a model system for studying ncAA incorporation, particularly at position Y151, green fluorescent protein and mutants such as sfGFP, which have been shown to tolerate a variety of natural and ncAA substitutions, have been used. A plasmid was constructed containing two dNaM-dTPT3 UBPs, one located within codon 151 of sfGFP and the other located within the Y151 codon of M. mazei tRNA. Pyl The anticodon was positioned to encode the dNaM anticodon (Figure 6C), which was selectively charged with the ncAA N6-((2-azidoethoxy)-carbonyl)-L-lysine (AzK) by PylRS (Figure 6B). Plasmids were constructed to investigate the decoding of six codons, including two unnatural codons in the first position (XTC and XTG; X refers to dNaM), two unnatural codons in the second position (AXC and GXA), and two unnatural codons in the third position (AGX and CAX), as well as opposite-strand context codons (YTC, YTG, AYC, GYA, AGY, and CAY; Y refers to dTPT3).
[0224] Although clonal populations of SSOs can produce higher amounts of pure non-native proteins, likely due to the elimination of misassembled plasmids during in vitro construction, to facilitate initial codon screening, protein expression was first investigated in non-clonal populations of cells, and protein production was analyzed immediately after transformation. A chimeric pyrrolysyl-tRNA synthetase (chPylRS) was used to transform E. coli ML2 (BL21(DE3)lacZYA:PtNTT2(66-575)ΔrecApolB++). IPYEThe cells were grown to early stationary phase in selective medium supplemented with dNaMTP and dTPT3TP, and then transferred to fresh medium. After growth to mid-exponential phase, the cultures were supplemented with NaMTP, TPT3TP, and AzK, and isopropyl-β-D-thiogalactoside (IPTG) was added to encode the T7 RNA polymerase (T7RNAP), chPylRS, and chPylRS. IPYE and tRNA Pyl After 1 h of further growth, anhydrotetracycline (aTc) was added to induce expression of sfGFP, which was monitored by fluorescence.
[0225] The codon in the first position is either a hetero-pairing or a self-pairing anticodon (e.g., XTG in the case of tRNA 2, respectively). Pyl (CAY) or tRNA Pyl Codons with dNaM at the second position showed little fluorescence in the absence of AzK, but in its presence, heteropaired anticodon tRNAs were detected. Pyl (GYT) or tRNA Pyl (TYC)-recoded tRNA Pyl showed significant fluorescence when decoded with the self-pairing anticodon tRNA Pyl (GXT) or tRNA Pyl This was not the case for (TXC). With dTPT3 in the second position, no fluorescence was observed, regardless of whether decoding was attempted with heteropairing or self-pairing tRNAs, regardless of whether AzK was added. The codons CAX and CAY in the third position were not detected with heteropairing or self-pairing tRNAs. Pyl Regardless of whether decoding was attempted with or without AzK, the codons showed high fluorescence in the absence of AzK and, surprisingly, showed less fluorescence upon its addition. This result suggests that the unnatural tRNA at the corresponding third position binds unproductively on the ribosome, blocking readthrough of the unnatural codon by the natural tRNA. In the absence of AzK, AGX and AGY showed almost no fluorescence, and tRNAPyl AGX by (XCT) showed an increase in fluorescence upon addition of AzK.
[0226] Because the codon at the first position did not appear promising, a more comprehensive screen of codons at the second position was performed. Initial analysis showed potential decoding with only NaM in the codon and TPT3 in the anticodon, so the NXN codon and the cognate tRNA Pyl We considered the NYN (NYN) gene. Of the 16 possible codons, we eliminated CXA, CXG, and TXG because the corresponding sequence context was not sufficiently preserved in the DNA of the SSO. Consistent with previous results, the use of codons AXC and GXC in the absence of AzK resulted in little or no fluorescence, whereas in the presence of AzK they resulted in significant fluorescence (Figure 6D). Similarly, for the GXT, CXC, TXC, GXG, GXA, CXT, and AXG codons, the addition of AzK resulted in a significant increase in fluorescence compared to the absence of AzK. The remaining four codons, AXA, AXT, TXA, and TXT, produced little fluorescence, regardless of whether AzK was added, revealing a stringent requirement for at least one GC pair.
[0227] To screen for non-native protein production, sfGFP was purified via a C-terminal Strep II affinity tag and subjected to a strain-promoted azide-alkyne cycloaddition (SPAAC) reaction with dibenzocyclooctyne (DBCO) linked to a rhodamine dye (TAMRA) by four PEG units (DBCO-PEG4-TAMRA). As previously shown, successful conjugation not only labels the ncAA-containing protein with a detectable fluorophore, but also produces a detectable shift in electrophoretic mobility, allowing quantification of AzK-containing protein relative to the total protein produced (i.e., fidelity of ncAA incorporation; Figure 6D). Consistent with previous results, use of the codons GXC and AXC resulted in the production of significant amounts of sfGFP at the AzK residue. Notably, seven additional non-native codons, GXT, CXC, TXC, GXG, GXA, CXT, and AXG, also resulted in significant levels of non-native protein (Figure 6D, Figure 11).
[0228] Finally, a more comprehensive screen of codons at the third position was performed, only then was the self-pairing tRNA Pyl Since only AGX appears to be decoded by (XCT), the cognate self-pairing tRNA Pyl We further investigated codons with dNaM in the third position of the codon (XNN) (Figure 6C). NCX codons were excluded because they result in the NCXA sequence context, which is not well-retained in SSO DNA as described above. Consistent with the initial analysis, in the absence of AzK, these codons generally resulted in more fluorescence than that observed with codons in the second position, but in the presence of AzK, increased variability in fluorescence was observed (Figure 6D). Nevertheless, when proteins were isolated and analyzed as described above, the use of CGX, ATX, CAX, AGX, GAX, TGX, CTX, TTX, GTX, or TAX all resulted in the production of significant levels of unnatural protein (Figure 6D, Figure 11). The codon GGX generated multiple shifted species, and tRNA PylThis suggests that (XCC) decodes one or more natural codons. When the codon AAX was used, no non-native proteins were detected.
[0229] Example 2. Codon characterization in clone SSO To select the most promising codon / anticodon pairs identified in the codon screening above, we compared the fluorescence observed in the presence of AzK and the induced mobility shift in the isolated protein (Figure 6D, inset). Based on this analysis, seven unnatural codon / anticodon pairs, GXC / GYC, GXT / AYC, AXC / GYT, AGX / XCT, CGX / XCG, TTX / XAA, and TGX / XCA, were selected for further characterization. These codon / anticodon pairs were used to identify clone S, eliminating cells transformed with misassembled plasmids or plasmids that had lost their UBP during in vitro construction. Clonal SSOs were obtained by streaking transformants onto solid growth medium containing dNaMTP and dTPT3TP, selecting individual colonies, and confirming plasmid integrity and high UBP retention. High-retaining clones were regrown and induced to produce the proteins described above. Remarkably, the observed fluorescence indicated that each of the seven codon / anticodon pairs produced protein at levels favorably comparable to the amber suppression control. Furthermore, gel shift assays demonstrated that virtually all of the sfGFPs contained ncAAs (Figure 7A, Figure 12). Decoding using the codon / anticodons AGX / XCT, CGX / XCG, TTX / XAA, and TGX / XCA produced sfGFP dependent only on NaMTP in the expression medium with similar AzK content both with and without added TPT3TP (Figure 13).
[0230] The seven unnatural codon / anticodon pairs analyzed above clearly mediated efficient decoding in the ribosome; however, it was possible that other codons from the preliminary non-clonal screen might have demonstrated efficient decoding when analyzed in clone SSO. Therefore, we investigated unnatural protein production in clone SSO with four additional codon / anticodon pairs: TXC / GYA, GXG / CYC, CXC / GYG, and AXT / AYT. Despite high UBP retention (Table 1), AXT showed no fluorescence signal with or without AzK, further supporting the requirement for a GC pair at the second codon position. The fluorescence with added AzK for TXC, CXC, and GXG was comparable to that of the seven initially characterized codons, but was slightly higher in the absence of AzK (Figure 7A). SPAAC gel shift analysis revealed that CXC clearly yielded significantly shifted proteins in the cloned SSO compared to those observed in the preliminary screen with the non-cloned SSO; TXC and GXG likely did too, although the relatively larger error in the data from the preliminary screen prevented quantitative comparison (Figure 7B). The data suggested that for some codons, suboptimal performance in the screen resulted, at least in part, from sequence-dependent differences in in vitro plasmid construction. Nevertheless, the results identified two additional high-fidelity codons, TXC and CXC, suggesting that more viable codons could still be identified.
[0231] To begin evaluating the orthogonality of unnatural codon / anticodon pairs, we selected AXC / GYT, GXT / AYC, and AGX / XCT to test all pairwise combinations of unnatural codons and anticodons for protein production in clone SSO. Upon addition of AzK, significant fluorescence was observed when each unnatural codon was paired with its cognate unnatural anticodon, with virtually no increase over background observed when paired with a noncognate unnatural anticodon (Figure 7B). Thus, AXC / GYT, GXT / AYC, and AGX / XCT were orthogonal in SSO and could be used simultaneously.
[0232] Example 3. Simultaneous decoding of two unnatural codons. To explore simultaneous decoding of multiple codons, the native sfGFP codons at positions 190 and 200 (sfGFP) were replaced with GXT and AXC, respectively. 190、200 The plasmid was first constructed using the tRNA gene (GXT, AXC). Pyl (AYC) and M. jannaschii tRNA pAzF , which is selectively charged with p-azido-L-phenylalanine (pAzF; Figure 6B) by the M. jannaschii TyrRS (MjTyrRS), and its anticodon was recoded to recognize AXC (tRNA pAzF (GYT); Figure 8A). IPYE E. coli ML2 carrying accessory plasmids encoding AzK and MjpAzFRS was transformed with the UBP-containing plasmid, and clones SSO were obtained, grown, and induced to produce sfGFP as described above. Upon presentation of AzK and pAzF, increased cellular fluorescence was observed on the same timescale as expression with the single-codon construct (Figure 8B, Figure S14). 190、200 The level of fluorescence due to expression from (GXT, AXC) is s fGFP 190 (GXT) or sfGFP 200 The amber and ochre controls (sfGFP) were slightly less than half of those observed in the AXC control (AXC), but they were significantly less than those observed in the amber and ochre controls (sfGFP) encoded with the corresponding suppressor tRNAs. 190、200The ncAA concentrations were significantly higher than those observed from the ncAAs (TAA, TAG) (Figure 8C, Figure 14). In both cases, when analyzed by SPAAC gel shift, no unshifted bands were evident, and the mobility of the major band was further delayed compared to that observed with the incorporation of a single ncAA, suggesting that two ncAAs were indeed incorporated (Figure 8D). To confirm that both pAzF and AzK had been incorporated, the purified protein was analyzed using quantitative intact protein mass spectrometry (HRMS ESI-TOF). Consistent with the gel shift assay, this analysis revealed that 91 ± 1.1% of the isolated protein contained both pAzF and AzK, 1.7 ± 0.4% contained a single pAzF, and 7.5 ± 0.78% contained a single AzK (Figure 15). In both cases, the identified impurity masses corresponded to amino acid substitutions consistent with dX to dT mutations, suggesting that the majority of loss of ncAA incorporation fidelity resulted from loss of dNaM or dTPT3 during replication and not due to errors during transcription or translation. UBP retention was based on streptavidin-biotin shift assays. Retention was also determined for tRNAs for which normalization could not be performed. pAzF and tRNA Ser The relative shifts were normalized to the relative shift of the ssDNA template control (i.e., the signal of the shifted band divided by the total signal of the shifted and unshifted bands), except for the . The mean ± standard deviation is shown (Table 1).
[0233] [Table 2-1] [Table 2-2]
[0234] SSO was 16 ± 3.2 μg·ml -1 of purified protein, whereas the amber and ochre suppression controls yielded 6.8 ± 1.1 μg ml -1 However, SSO cultures were noted to grow to lower densities than amber and ochre control cells, resulting in an OD 600When normalized for SSO, the mean was 13 ± 1.6 μg ml -1 The purified protein yielded 2.8±0.28 μg ml -1 SSO brings OD 600 demonstrated that over 4.5-fold more protein was produced by affinity purification using excess Strep-Tactin XT beads to capture sfGFP. Yields were calculated based on the final OD at t = 180 min of expression. 600 The results were normalized to the mean ± standard deviation (Table 2). Therefore, SSO efficiently generates unnatural proteins with two ncAAs.
[0235] [Table 3]
[0236] To characterize protein expression by ncAAs with different functional groups, we used sfGFP. 190、200 (GXT, AXC) were expressed in SSO as described above, but the growth medium contained chPylRS IPYE N, also recognized by 6 -(Propargyloxy)-carbonyl-L-lysine (PrK, Figure 6B) was added instead of AzK. No substantial effect on expression was observed by fluorescence in either the SSO or the amber and ochre controls (Figure 8E). In each case, correct incorporation of PrK and pAzF was verified by SPAAC with TAMRA-PEG4-DBCO followed by copper-catalyzed alkyne-azide cycloaddition (CuAAC) using TAMRA-PEG4-azide, both of which induced observable shifts in electrophoretic mobility. Proteins produced by SSO and the amber and ochre controls exhibit the expected gel shift and TAMRA signal (Figure 8F).
[0237] Example 4. Simultaneous decoding of three unnatural codons Endogenous serine tRNA to explore simultaneous decoding of three orthogonal unnatural codons SerWe used Escherichia coli SerT, which is charged with an endogenous SerRS without anticodon recognition and was previously recoded to decode unnatural codons. IPYE E. coli ML2 carrying an accessory plasmid encoding MjpAzFRS and sfGFP 151、190、200 (AXC, GXT, AGX) and tRNA Pyl (XCT), tRNA pAzF (GYT) and tRNA Ser The cells were transformed with a plasmid expressing (AYC) (Fig. 9A), and clonal SSOs were prepared, grown, and induced to produce protein as described above. Significant fluorescence was observed when AzK and pAzF were added to the medium, similar to the results obtained above for simultaneous decoding of two codons (Fig. 9B, Fig. 14). These cells produced 12.1 ± 1.9 μg ml -1 (7.8±1.1μg ml -1 OD -1 ), which was slightly less than the amount isolated with the decoding of the two unnatural codons (Table 2). To confirm that pAzF, AzK, and Ser were all incorporated, we analyzed the purified protein by quantitative intact protein mass spectrometry (HRMS ESI-TOF) and found that 96 ± 0.63% of the isolated protein contained pAzF, AzK, and Ser, with the major impurity being sfGFP (3.5 ± 0.63%), which contained only AzK and Ser. Protein without Ser incorporation was nearly undetectable (0.20 ± 0.087%), while the mass corresponding to the protein containing only pAzF and Ser was undetectable (Figure 9C, Figure 16). Furthermore, we did not detect any impurities corresponding to multiple insertions of Ser, AzK, or pAzF.
[0238] Example 5. Methods for in vivo expression of non-naturally occurring polypeptides material A complete list of oligonucleotides and plasmids used is in Table 3. Natural ssDNA oligonucleotides and gBlocks were purchased from IDT (San Diego, CA). Sequencing was performed by Genewiz (San Diego, CA). All DNA purification was performed using Zymo Research silica column kits. All cloning enzymes and polymerases were purchased from New England Biolabs (Ipswich, MA). All bioconjugation reagents were purchased from Click Chemistry Tools (Scottsdale, AZ). All unnatural nucleoside triphosphates and nucleoside phosphoramidites used in this study were obtained from commercial sources. sfGFP was synthesized as described in the literature. 200 With the exception of (AGX), all ssDNA dNaM templates were also obtained from commercial sources.
[0239] [Table 4-1] [Table 4-2] [Table 4-3]
[0240] Growth conditions 300 μl of 2xYT (Fisher) supplemented with potassium phosphate (50 mM pH 7) All bacterial experiments were performed in BioScientific medium. Growth was performed in flat-bottom 48-well plates (CELLSTAR, Greiner Bio-One) with shaking (Infors HT Minitron) at 37°C and 200 rpm. Antibiotics were used at the following concentrations (unless otherwise stated): chloramphenicol (5 μg / ml), carbenicillin (100 μg / ml), and zeocin (50 μg / ml). Unnatural nucleoside triphosphates were used at the following concentrations (unless otherwise stated): dNaMTP (150 μM), dTPT3TP (10 μM), NaMTP (250 μM), and TPT3TP (30 μM). UBP medium is defined as the aforementioned 2xYT medium containing dNaMTP and dTPT3TP.
[0241] Plasmid construction Insertion of large inserts (>100 bp), MjpAzFRS, tRNA, or antibiotic resistance cassettes was performed by Gibson assembly of PCR amplicons or gBlocks. Amplicons were treated with DpnI overnight at room temperature before assembly for 1.5 hours at 50°C. Deletions or small insertions (<50 bp; e.g., codon or anticodon mutagenesis, restriction site removal, or Golden Gate target site introduction) were constructed by introducing the desired changes into PCR primer overhangs designed to amplify the entire plasmid. Primers were phosphorylated using T4 PNK prior to PCR, and the resulting PCR amplicons were treated with DpnI overnight at room temperature and recircularized using T4 DNA ligase. After the initial assembly / ligation, plasmids were transformed into electrocompetent XL-10 Gold cells and grown on selective LB Lennox agar (BP Difco). Plasmids were isolated from individual colonies and verified by Sanger sequencing at the time of use. All plasmids used in this study can be found in Table 4. All sfGFP reading frames are in P T7-tetO All tRNAs are regulated by P T7-lacOThe backbone pSYN contains: a replication origin (p15A) bleoR. The backbone pGEX contains: a replication origin (pBR322) ampR. The Golden Gate destination site (dest) consisted of the recognition sequence BsaI-KpnI-BsaI.
[0242] [Table 5-1] [Table 5-2]
[0243] UBP oligo PCR Double-stranded DNA inserts containing UBP-containing sequences were synthesized by PCR with primers (List A) using chemically synthesized dNaM-containing ssDNA oligonucleotides (List B) as templates (OneTaq standard buffer 1x, 0.025 units / µl OneTaq, 0.2 mM dNTPs, 0.1 mM dTPT3TP, 0.1 mM dNaMTP, 1.2 mM MgSO, 1x SYBR Green, 1.0 µM primers, approximately 20 pM template; cycle: 96°C 0:30 min, 96°C 0:30 min, 54°C 0:30 min, 68°C 0:30 min). 4:00 min, fluorescence reading, go to step 2 (<24 times). Position sfGFP 190 and sfGFP 200 The insert fragment for was combined using the same conditions as above, except that both templates were combined by overlap extension at 1 nM. Amplification was monitored, and when the SYBR green trace reached a plateau, the reaction was placed on ice. Products were analyzed via native PAGE (6% acrylamide:bisacrylamide 29:1; SYBR Gold stain in 1x TBE) to verify single amplicons, purified on spin columns (Zymo Research), and quantified using Qubit dsDNA HS (ThermoFisher).
[0244] Golden Gate assembly of SSO expression vectors The UBP-containing inserts were assembled into the pSYN entry vector framework (Table 4) by Golden Gate assembly using a 3:1 molar ratio of each insert to entry vector (Cutsmart buffer 1x, 1 mM ATP, 6.67 units / µl T4 DNA ligase, 0.67 units / µl BsaI-HFv2, 20 ng / µl entry vector DNA; cycle: 37°C 10:00 min, 37°C 5:00 min, 16°C 5:00 min, 22°C 2:00 min, repeat step 2 39 times, 37°C 20:00 min, 55°C 15:00 min, 80°C 30:00 min). For the experiments in Figure 6, BsaI-HF was used. Residual linear DNA and undigested entry vector were first digested with KpnI-HF (0.33 units / μl, 1 hour at 37°C) followed by T5 exonuclease (0.17 units / μl, 30 minutes at 37°C). Products were purified on spin columns and quantified using Qubit dsDNA HS (ThermoFisher).
[0245] Preparation of competent starter cells Strain ML2 (BL21(DE3)lacZYA::PtNTT2(66-575)Δre cA polB ++ ) was transformed with the accessory pGEX plasmids (Table 4) and plated onto LB Lennox agar containing chloramphenicol and carbenicillin. Single colonies were picked and probed with radioactive [α- 32 PtNTT2 activity was verified by [P]dATP incorporation. UBP replication and translation competent cells were incubated at an OD of 0.25–0.30. 600 Cultures were prepared by growing cells in 2xYT medium in baffled culture flasks at 37°C and 250 rpm until 5 min. The cultures were transferred to pre-chilled 50 mL Falcon tubes and gently shaken in an ice-water bath for 2 min. Cells were pelleted by centrifugation (10 min, 3200 rpm), washed in sterile cold water, pelleted, washed again, then pelleted a final time and suspended in 50 μl of 10% glycerol per 10 mL of culture. Cells were used immediately or frozen at -80°C for later use.
[0246] Non-clonal population experiments Freshly prepared competent cells were electroporated (2.5 kV) with approximately 0.4 ng of Golden Gate assembly product and immediately suspended in 950 μl of 2xYT supplemented with potassium phosphate (50 mM pH 7), and 10 μl of the suspension was diluted in 40 μl of UBP medium containing 1.25X dNaMTP and dTPT3TP without Zeocin. After 1 hour of cell recovery at 37°C, 15 μl of cells were suspended in 285 μl of UBP medium containing Zeocin and grown at 37°C with shaking in a 48-well plate. OD 600 Before reaching stationary phase at approximately 1°C, the culture was transferred to ice and stored overnight for protein expression.
[0247] Clone SSO experiment Competent cells were electroporated with Golden Gate assembly products (1–20 ng) and recovered for non-clonal population experiments. Plating was performed by spreading 10 μl of the recovered culture (and its dilutions) onto agar droplets (250 μl of 2xYT 2% agar, 50 mM potassium phosphate) containing chloramphenicol, carbenicillin, zeocin, dNaMTP, and dTPT3TP. Colonies approximately 0.5 mm in diameter were picked and suspended in UBP medium (300 μl) after growth on the plate (12–20 h at 37°C). Before reaching stationary phase with an OD of approximately 1, each culture was transferred to a pre-chilled tube on ice and stored overnight for protein expression. Each culture was prescreened for 1) UBP retention using a streptavidin-biotin shift assay (as described below) and 2) qualitative sfGFP expression by mixing the culture 1:4 with medium already containing the expression components (ribonucleoside triphosphates, ncAAs, IPTG, and anhydrotetracycline). Colonies were discarded if they did not produce any fluorescent signal when the appropriate ncAA was added after 2 hours of incubation at 37°C or overnight at room temperature. Additionally, colonies with <80% UBP retention in sfGFP were discarded. If more than three colonies met these criteria, only the three with the highest UBP retention were selected to limit material costs. The data to the right of the dashed line in Figure 7A were obtained through a slightly modified method. Instead of prescreening colonies as described above, expression was performed on a large number of colonies, but protein analysis was only performed on cultures that showed promising fluorescence during expression. 10 mM AzK was used during expression. Additionally, buffer W2 was used instead of buffer W during protein purification.
[0248] Pre-cloned SSO expression vector For the experiments in Figures 7B, 8, and 9, plasmids from prescreened colonies were isolated (Zymo Research Miniprep) to serve as starting plasmids for (precloned) transformation to facilitate prescreening of colonies. Plasmids were prescreened for fluorescent fluorescence (as described above). Colonies for the data in Figure 7B were prescreened alternatively with and without rNaMTP and rTPT3TP in the presence of AzK to qualitatively generate dark and fluorescent signals, respectively. All precloned plasmids were prescreened for UBP retention (>80%) in sfGFP. Additionally, these plasmids were PCR amplified using the standard OneTaq protocol (Ne...
Claims
1. 1. A method for synthesizing a non-naturally occurring polypeptide, comprising: a. providing at least one unnatural deoxyribonucleic acid (DNA) molecule comprising at least four unnatural base pairs, wherein the at least one unnatural DNA molecule encodes (i) a messenger RNA (mRNA) molecule comprising at least a first and a second unnatural codon, and (ii) at least a first and a second transfer RNA (tRNA) molecule, wherein the first tRNA molecule comprises a first unnatural anticodon and the second tRNA molecule comprises a second unnatural anticodon, and wherein the at least four unnatural base pairs in the at least one DNA molecule are in a sequence context such that the first and second unnatural codons of the mRNA molecule are complementary to the first and second unnatural anticodons, respectively; b. transcribing at least one non-naturally occurring DNA molecule to provide mRNA; c. transcribing at least one non-naturally occurring DNA molecule to provide at least first and second tRNA molecules; d. synthesizing a non-natural polypeptide by translating a non-natural mRNA molecule utilizing at least a first and a second non-natural tRNA molecule, wherein each of the at least a first and a second non-natural anticodon directs the site-specific incorporation of an non-natural amino acid into the non-natural polypeptide; The method comprising:
2. 2. The method of claim 1, wherein the at least two unnatural codons each comprise a first unnatural nucleotide located at the first position, the second position, or the third position of the codon, and optionally the first unnatural nucleotide is located at the second position or the third position of the codon.
3. 3. The method of any one of claims 1-2, wherein the at least two unnatural codons each comprise the nucleic acid sequence NNX or NXN, and the unnatural anticodon comprises the nucleic acid sequence XNN, YNN, NXN, or NYN, thereby forming an unnatural codon-anticodon pair comprising NNX-XNN, NNX-YNN, or NXN-NYN, where N is any naturally occurring nucleotide, X is a first unnatural nucleotide, and Y is a second unnatural nucleotide different from the first unnatural nucleotide, and X-Y forms an unnatural base pair in DNA.
4. 4. The method of claim 3, wherein the codon contains at least one G or C and the anticodon contains at least one complementary C or G.
5. X and Y are: (i) 2-thiouracil, 2'-deoxyuridine, 4-thiouracil, uracil-5-yl, hypoxanthine-9-yl (I), 5-halouracil; 5-propynyl-uracil, 6-azo-uracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, uracil-5-oxaacetic acid methyl ester, uracil-5-oxaacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, uracil-5-oxaacetic acid, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil or dihydrouracil; (ii) 5-hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytosine arabinoside, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azocytosine, azacytosine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1, 4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one), or pyridoindole cytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one); (iii) 2-aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2′-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted adenines, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-azaadenine; (iv) 2-methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thioguanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted guanines, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-azaguanine; and (v) hypoxanthine, xanthine, 1-methylinosine, queosine, beta-D-galactosylqueosine, inosine, beta-D-mannosylqueosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone 5. The method of claim 3 or 4, wherein the hydroxyl group is independently selected from the group consisting of:
6. The bases comprising each of X and Y are: 【Chemistry 1】 6. The method of claim 4 or 5, wherein the hydroxyl group is independently selected from the group consisting of:
7. Each X-containing base is 【Chemistry 2】 The method of claim 6, wherein
8. Each Y-containing base is 【Transformation 3】 The method according to claim 6 or 7, wherein
9. 9. The method of any one of claims 3 to 8, wherein NNX-XNN is selected from the group consisting of UUX-XAA, UGX-XCA, CGX-XCG, AGX-XCU, GAX-XUC, CAX-XUG, AUX-XAU, CUX-XAG, GUX-XAC, UAX-XUA, and GGX-XCC.
10. 9. The method of any one of claims 3 to 8, wherein NNX-YNN is selected from the group consisting of UUX-YAA, UGX-YCA, CGX-YCG, AGX-YCU, GAX-YUC, CAX-YUG, AUX-YAU, CUX-YAG, GUX-YAC, UAX-YUA and GGX-YCC.
11. 9. The method of any one of claims 3 to 8, wherein NXN-NYN is selected from the group consisting of GXU-AYC, CXU-AYG, GXG-CYC, AXG-CYU, GXC-GYC, AXC-GYU, GXA-UYC, CXC-GYG, and UXC-GYA.
12. The method of any one of claims 1 to 11, wherein the at least two non-natural tRNA molecules each comprise a different non-natural anticodon.
13. 13. The method of claim 12, wherein the at least two non-natural tRNA molecules comprise a pyrrolysyl-tRNA from Methanosarcina and a tyrosyl-tRNA from Methanocaldococcus jannaschii, or a derivative thereof.
14. The method of any one of claims 11 to 13, comprising charging at least two non-natural tRNA molecules with an aminoacyl-tRNA synthetase.
15. 15. The method of claim 14, wherein the tRNA synthetase is selected from the group consisting of a chimeric PylRS (chPylRS) and a M. jannaschii AzFRS (MjpAzFRS).
16. 14. The method of claim 12 or 13, comprising charging at least two non-natural tRNA molecules with at least two different tRNA synthetases.
17. 17. The method of claim 16, wherein the at least two different tRNA synthetases comprise a chimeric PylRS (chPylRS) and a M. jannaschii AzFRS (MjpAzFRS).
18. The method of any one of claims 1 to 17, wherein the non-natural polypeptide comprises two, three or more non-natural amino acids.
19. The method of any one of claims 1 to 18, wherein the non-natural polypeptide comprises at least two non-natural amino acids that are the same.
20. The method of any one of claims 1 to 18, wherein the non-natural polypeptide comprises at least two different non-natural amino acids.
21. Unnatural amino acids are Lysine analogues; Aromatic side chains; Azide group; alkyne groups; or Aldehyde or ketone group The method of any one of claims 1 to 20, comprising:
22. 21. The method of any one of claims 1 to 20, wherein the unnatural amino acid does not comprise an aromatic side chain.
23. Unnatural amino acids include N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propargylethoxy-carbonyl-L-lysine (PrK), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azidophenylalanine (pAzF), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxoocta-lysine, and the like. acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl- L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2 21. The method of any one of claims 1 to 20, wherein the amino acid is selected from N6-(((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine.
24. The method of any one of claims 1 to 23, wherein at least one non-naturally occurring DNA molecule is in the form of a plasmid.
25. Claims 1 to 23, wherein at least one non-natural DNA molecule is integrated into the genome of the cell.
10. The method according to any one of the preceding claims.
26. 26. The method of claim 24 or 25, wherein at least one non-naturally occurring DNA molecule encodes a non-naturally occurring polypeptide.
27. 27. The method of any one of claims 1 to 26, comprising the in vivo replication and transcription of the non-naturally occurring DNA molecule in a cellular organism, and the in vivo translation of the transcribed mRNA molecule.
28. 28. The method of claim 27, wherein the cellular organism is a microorganism.
29. 29. The method of claim 28, wherein the cellular organism is a prokaryote.
30. 30. The method of claim 29, wherein the cellular organism is a bacterium.
31. 31. The method of claim 30, wherein the cellular organism is a gram-positive bacterium.
32. 31. The method of claim 30, wherein the cellular organism is a gram-negative bacterium.
33. 33. The method of claim 32, wherein the cellular organism is Escherichia coli.
34. 34. The method of any one of claims 1 to 33, wherein the at least two unnatural base pairs comprise a base pair selected from dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1, or dNaM-dTAT1.
35. The method of any one of claims 27 to 34, wherein the cellular organism comprises a nucleoside triphosphate transporter.
36. 36. The method of claim 35, wherein the nucleoside triphosphate transporter comprises the amino acid sequence of PtNTT2.
37. The method of claim 36, wherein the nucleoside triphosphate transporter comprises a truncated amino acid sequence of PtNTT2, and optionally the truncated amino acid sequence of PtNTT2 is at least 80% identical to PtNTT2 encoded by SEQ ID NO:
1.
38. The method of any one of claims 27 to 37, wherein the cellular organism comprises at least one non-naturally occurring DNA molecule.
39. 39. The method of claim 38, wherein the at least one non-naturally occurring DNA molecule comprises at least one plasmid.
40. 39. The method of claim 38, wherein at least one non-natural DNA molecule is integrated into the genome of the cell.
41. 41. The method of claim 39 or 40, wherein at least one non-naturally occurring DNA molecule encodes a non-naturally occurring polypeptide.
42. The method of any one of claims 1 to 24, which is an in vitro method comprising synthesizing the non-natural polypeptide in a cell-free system.
43. 43. The method of any one of claims 1 to 42, wherein the unnatural base pair comprises at least one unnatural nucleotide comprising an unnatural sugar moiety.
44. The unnatural sugar moiety is: Modifications at the 2' position include: OH, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH 3 , OCN, Cl, Br, CN, CF 3 , OCF 3 , SOCH 3 , S.O. 2 CH 3 , ONO 2 , NO 2 , N 3 or NH 2 F; O-alkyl, S-alkyl or N-alkyl; O-alkenyl, S-alkenyl or N-alkenyl; O-alkynyl, S-alkynyl or N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH 3 , 2'-O(CH 2 ) 2 OCH 3 , wherein alkyl, alkenyl and alkynyl are substituted or unsubstituted C 1 ~C 10 Alkyl, C 2 ~C 10 Alkenyl, C 2 ~C 10 Alkynyl, —O[(CH 2 ) n O] m CH 3 , -O(CH 2 ) n OCH 3 , -O(CH 2 ) n NH 2 , -O(CH 2 ) n CH 3 , -O(CH 2 ) n -NH 2 or -O(CH 2 ) n ON [(CH 2 ) n CH 3 )] 2 wherein n and m are from 1 to about 10; Modifications at the 5' position including: 5'-vinyl, or 5'-methyl (R or S); or a modification at the 4' position, 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, an RNA cleaving group, a reporter group, an intercalator, a group for improving the pharmacokinetic properties of an oligonucleotide, or a group for improving the pharmacodynamic properties of an oligonucleotide; or A group consisting of any combination thereof 44. The method of claim 43, comprising a moiety selected from:
45. 1. A cell comprising at least one unnatural DNA molecule comprising at least four unnatural base pairs, wherein the at least one unnatural DNA molecule encodes: (i) a messenger RNA (mRNA) molecule encoding an unnatural polypeptide and comprising at least a first and a second unnatural codon; and (ii) at least a first and a second transfer RNA (tRNA) molecule, wherein the first tRNA molecule comprises a first unnatural anticodon and the second tRNA molecule comprises a second unnatural anticodon, and wherein the at least four unnatural base pairs in the at least one DNA molecule are in a sequence context such that the first and second unnatural codons of the mRNA molecule are complementary to the first and second unnatural anticodons, respectively.
46. 46. The cell of claim 45, further comprising an mRNA molecule and at least a first and a second tRNA molecule.
47. 47. The cell of claim 46, wherein at least the first and second tRNA molecules are covalently linked to an unnatural amino acid.
48. 48. The cell of claim 47, further comprising a non-naturally occurring polypeptide.
49. A cell, a. at least two different unnatural codon-anticodon pairs, each unnatural codon-anticodon pair comprising an unnatural codon from a unnatural messenger RNA (mRNA) and an unnatural anticodon from a unnatural transfer ribonucleic acid (tRNA), wherein the unnatural codon comprises a first unnatural nucleotide and the unnatural anticodon comprises a second unnatural nucleotide; and b. at least two different non-natural tRNAs, each covalently linked to a corresponding non-natural tRNA; Amino acid The cell comprising:
50. 50. The cell of claim 49, further comprising at least one unnatural DNA molecule comprising at least four unnatural base pairs (UBPs).
51. 51. The cell of any one of claims 45 to 50, wherein the first non-natural nucleotide is located at the second or third position of the non-natural codon.
52. 52. The cell of claim 51, wherein the first unnatural nucleotide complementarily base pairs with the second unnatural nucleotide of the unnatural anticodon.
53. The first non-natural nucleotide and the second non-natural nucleotide are 【Chemistry 4】 53. The cell of any one of claims 45 to 52, wherein the cell comprises a first and a second base independently selected from the group consisting of:
54. 54. The cell of any one of claims 45 or 47 to 53, wherein the at least four unnatural base pairs are independently selected from the group consisting of dCNMO / dTPT3, dNaM / dTPT3, dCNMO / dTAT1, or dNaM / dTAT1.
55. 55. The cell of any one of claims 45 or 47 to 54, wherein the at least one non-naturally occurring DNA molecule comprises at least one plasmid.
56. 55. The cell of any one of claims 45 or 47 to 54, wherein at least one non-naturally occurring DNA molecule is integrated into the genome of the cell.
57. 57. The cell of any one of claims 47 to 56, wherein at least one non-naturally occurring DNA molecule encodes a non-naturally occurring polypeptide.
58. The cell of any one of claims 45 to 57, which expresses a nucleoside triphosphate transporter. Cell.
59. 59. The cell of claim 58, wherein the nucleoside triphosphate transporter comprises the amino acid sequence of PtNTT2.
60. The method of claim 59, wherein the nucleoside triphosphate transporter comprises a truncated amino acid sequence of PtNTT2, and optionally, the truncated amino acid sequence of PtNTT2 is at least 80% identical to PtNTT2 encoded by SEQ ID NO:
1.
61. 61. The cell of any one of claims 45 to 60, which expresses at least two tRNA synthetases.
62. 62. The cell of claim 61, wherein the at least two tRNA synthetases are a chimeric PylRS (chPylRS) and a M. jannaschii AzFRS (MjpAzFRS).
63. 63. The cell of any one of claims 45 to 62, comprising an unnatural nucleotide comprising an unnatural sugar moiety.
64. The unnatural sugar moiety is: Modifications at the 2' position include: OH, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH 3 , OCN, Cl, Br, CN, CF 3 , OCF 3 , SOCH 3 , S.O. 2 CH 3 , ONO 2 , NO 2 , N 3 or NH 2 F; O-alkyl, S-alkyl or N-alkyl; O-alkenyl, S-alkenyl or N-alkenyl; O-alkynyl, S-alkynyl or N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH 3 , 2'-O(CH 2 ) 2 OCH 3 , wherein alkyl, alkenyl and alkynyl are substituted or unsubstituted C 1 ~C 10 Alkyl, C 2 ~C 10 Alkenyl, C 2 ~C 10 Alkynyl, —O[(CH 2 ) n O] m CH 3 , -O(CH 2 ) n OCH 3 , -O(CH 2 ) n NH 2 , -O(CH 2 ) n CH 3 , -O(CH 2 ) n -NH 2 or -O(CH 2 ) n ON [(CH 2 ) n CH 3 )] 2 wherein n and m are from 1 to about 10; Modifications at the 5' position including: 5'-vinyl, 5'-methyl (R or S); or a modification at the 4' position, 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, an RNA cleaving group, a reporter group, an intercalator, a group for improving the pharmacokinetic properties of an oligonucleotide, or a group for improving the pharmacodynamic properties of an oligonucleotide; or Any combination thereof 64. The cell of claim 63, selected from the group consisting of:
65. 65. The cell of any one of claims 45 to 64, wherein at least one unnatural nucleotide base is recognized by an RNA polymerase during transcription.
66. 66. The cell of any one of claims 45-65, wherein the cell translates at least one unnatural polypeptide comprising at least two unnatural amino acids.
67. At least two unnatural amino acids are N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propargylethoxy-carbonyl-L-lysine (PraK), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azidophenylalanine (pAzF), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azido Methyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L- Phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)- 67. The cell of any one of claims 45 to 66, wherein the amino acid residues are independently selected from the group consisting of N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine.
68. The cell of any one of claims 45 to 67, which is isolated.
69. The cell of any one of claims 45 to 68, which is a prokaryote.
70. A cell line comprising a cell according to any one of claims 45 to 69.
Citation Information
Patent Citations
glycoprotein synthesis
JP2006507358A
Method for Site-Specific Protein Incorporation of Keto Amino Acids
JP2006507814A
New method for introducing non-natural amino acid to protein
JP2007097423A
Incorporation of unnatural amino acids into proteins
JP2017532043A
Multi-purpose acylation catalayst and use thereof
WO2007066627A1