Reagents and methods for replication, transcription, and translation in semi-synthetic organisms

JP2025102850A5Pending Publication Date: 2025-11-26THE SCRIPPS RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025051316
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-14
Filing Date
2025-03-26
Publication Date
2025-11-26

AI Technical Summary

Technical Problem

The limitations of the natural amino acids restrict the functional diversity of proteins, and existing methods for expanding the genetic code face challenges such as competition with endogenous release factors and inefficient codon reassignment, particularly in eukaryotes, limiting the incorporation of non-standard amino acids.

Method used

The development of semi-synthetic organisms with an expanded genetic alphabet using unnatural base pairs (UBPs) to encode non-natural amino acids, enabling efficient replication, transcription, and translation of nucleic acids containing unnatural nucleotides.

Benefits of technology

This approach allows for the faithful replication and translation of DNA with unnatural nucleotides, facilitating the incorporation of multiple non-natural amino acids into proteins with high fidelity, thereby expanding the functional diversity of proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide compositions, methods, cells, engineered microorganisms, and kits for increasing the production of proteins or polypeptides comprising one or more unnatural amino acids.SOLUTION: The present invention provides a double-stranded oligonucleotide comprising a nucleobase of the following structure [where a wavy line indicates a point of bonding to a ribosyl, deoxyribosyl, or dideoxyribosyl moiety or an analog thereof].SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 62 / 861,901, filed on June 14, 2019, the disclosure of which is hereby incorporated by reference in its entirety.

[0002] Sequence Listing This application includes a sequence listing that was electronically submitted in ASCII format, which is hereby incorporated by reference in its entirety. The ASCII copy was created on June 10, 2020, has the name 36271 - 808_601_SL.txt, and is 18,162 bytes in size.

[0003] Statement Regarding Federally Sponsored Research The invention disclosed herein was made, at least in part, in support of grants Nos. 5R35 GM118178 and GM128376 and F31 GM128376 from the National Institutes of Health (NIH), and grant No. NSF / DGE - 1346837 from the National Science Foundation (NSF) of the United States government. Accordingly, the United States government has certain rights in this invention.

Background Art

[0004] Biological diversity enables life to adapt to different environments and, over time, evolve new forms and functions. The origin of this diversity lies in the variations within protein sequences provided by the 20 natural amino acids, which are encoded in an organism's genome by the four natural DNA nucleotides. Although the functional diversity provided by the natural amino acids is high, the vastness of sequence space severely limits the targets that can actually be explored, and some functions may not even be available in the first place. The fact that nature uses cofactors for hydride transfer, redox activity, and electrophilic bond formation, among others, attests to these limitations. Moreover, these limitations are problematic because the increasing attention to the development of proteins as therapeutic agents has shown that the physicochemical diversity of natural amino acids is significantly restricted compared to that of small molecule drugs designed by chemists. In principle, it should be possible to circumvent these limitations by expanding the genetic code to include additional non-standard amino acids (ncAAs or "non-natural amino acids") with desired physicochemical properties.

[0005] Nearly 20 years ago, a method was developed to increase the diversity available to living organisms by expanding the genetic code using the amber stop codon (UAG) that encodes an ncAA in Escherichia coli. This was achieved using a pair of tRNA - aminoacyl - tRNA synthetase (aaRS) from Methanococcus jannaschii, where the tRNA was recoded to suppress the stop codon and the aaRS evolved to charge the tRNA with the ncAA. This method of codon suppression has since been extended not only to other stop codons and even quadruplet codons, but also to the use of several other orthogonal tRNA - aaRS pairs (most notably, the Pyl tRNA synthetase pair from Methanosarcina barkeri / mazei), expanding the range of ncAAs incorporated into proteins. These methods have already begun to bring about a major revolution in both chemical biology and protein therapeutics.

[0006] These methods enable the incorporation of up to two different ncAAs in both prokaryotic and eukaryotic cells, but the heterologous recoded tRNAs compete with the endogenous release factor (RF) and, in the case of quadruplet codons, with normal decoding, limiting the efficiency and fidelity of ncAA incorporation. Efforts have been directed towards removing many or all of the amber stop codons that occur in the host genome, or modifying RF2 to allow deletion of RF1, in order to eliminate competition with RF1, which recognizes the amber stop codon and terminates translation. However, eukaryotes have only one release factor, which can be modified but not deleted, and in prokaryotes, deletion of RF1 results in higher mis-suppression of amber stop codons by other tRNAs, reducing the fidelity of ncAA incorporation. Efforts to further utilize codon redundancy to free up natural codons for reassignment to ncAAs are complicated by pleiotropic effects because codons are not truly redundant, for example due to their effects on translation speed and protein folding. In addition, codon reassignment is limited by the challenges of large-scale genome manipulation, particularly in eukaryotes.

[0007] An alternative approach to natural codon reassignment is the creation of entirely new codons without any natural function or constraint, and their recognition in the ribosome is inherently more orthogonal. This can be achieved by creating organisms with a fifth and sixth nucleotide that form unnatural base pairs (UBPs). Such semi-synthetic organisms (SSOs) would need to faithfully replicate DNA containing UBPs, efficiently transcribe it into mRNA and tRNA containing unnatural nucleotides, and then efficiently decode the unnatural codons with cognate unnatural anticodons. Such SSOs would have a virtually unlimited number of new codons encoding ncAAs. Summary of the Invention Means for Solving the Problems

[0008] In certain embodiments, methods, cells, engineered microorganisms, plasmids, and kits for the increased production of nucleic acid molecules comprising unnatural nucleotides are described herein.

[0009] The following embodiments are included.

[0010] Embodiment A1 is a nucleic base having the structure:

Chemical formula

[0011] Embodiment A2 is the nucleic base according to Embodiment A1, wherein X is carbon.

[0012] Embodiment A3 is the nucleic base according to Embodiment A1 or A2, wherein E is sulfur.

[0013] Embodiment A4 is the nucleic base according to any one of Embodiments A1 to A3, wherein Y is sulfur.

[0014] Embodiment A5 is the nucleobase described in Embodiment A1 having a structure [Chemical formula] and is the nucleobase described in Embodiment A1 having a structure

[0015] Embodiment A6 is the nucleobase according to any one of Embodiments A1 to A5, which binds to a complementary base pair-forming nucleobase to form a non-natural base pair (UBP).

[0016] Embodiment A7 is the nucleobase according to Embodiment A6, wherein the complementary base pair-forming nucleobase is selected from [Chemical formula] and is the nucleobase described in Embodiment A6.

[0017] Embodiment A8 is a double-stranded oligonucleotide, wherein the first oligonucleotide strand contains the nucleobase according to any one of Embodiments A1 to A5, and the second complementary oligonucleotide strand contains a complementary base pair-forming nucleobase at its complementary base pair-forming site.

[0018] Embodiment A9 is the double-stranded oligonucleotide according to Embodiment A8, wherein the first oligonucleotide strand contains [Chemical formula] and the second strand contains a complementary base pair-forming nucleobase selected from [Chemical formula] at its complementary base pair-forming site.

[0019] Embodiment A10 is the double-stranded oligonucleotide according to Embodiment A9, wherein the second strand contains a complementary base pair-forming nucleobase [Chemical formula] and is the double-stranded oligonucleotide described in Embodiment A9.

[0020] Embodiment A11 is a double strand of the double-stranded oligonucleotide according to Embodiment A9, wherein the second strand contains complementary base-pair-forming nucleobases

Chemical formula

[0021] Embodiment A12 is a plasmid containing a gene encoding transfer RNA (tRNA) and / or a gene encoding a protein of interest, wherein the gene is at least one nucleobase according to any one of Embodiments A1 to A5, or TPT3(

Chemical formula

Chemical formula

[0022] Embodiment A13 is an mRNA encoded by the plasmid according to Embodiment A10 that encodes tRNA

[0023] Embodiment A14 is an mRNA encoded by the plasmid according to Embodiment A10 that encodes a protein

[0024] Embodiment A15 is a transfer RNA (tRNA) containing a nucleobase according to any one of Embodiments A1 to A5, wherein the nucleobase is an anticodon containing a nucleobase, and optionally the nucleobase is at the first, second, or third position of the anticodon; and a recognition element that promotes selective charging of the tRNA with a non-natural amino acid by an aminoacyl-tRNA synthetase; which is a transfer RNA containing the above

[0025] Embodiment A16 is the tRNA according to Embodiment A15, wherein the aminoacyl tRNA synthetase is derived from Methanosarcina or a variant thereof, or Methanococcus (Methanocaldococcus) or a variant thereof.

[0026] Embodiment A17 is the tRNA according to Embodiment A15, wherein the unnatural amino acid contains an aromatic moiety.

[0027] Embodiment A18 is the tRNA according to Embodiment A15, wherein the unnatural amino acid is a derivative of lysine or phenylalanine.

[0028] Embodiment A19 has the formula: N1-Zx-N2 [wherein, each Z is independently a nucleobase according to any one of Embodiments A1 to A7, and is bonded to ribosyl or deoxyribosyl, or an analog thereof; N1 is one or more nucleotides or an analog thereof, or a terminal phosphate group, bonded to the 5'-end of the ribosyl or deoxyribosyl of Z, or an analog thereof; N2 is one or more nucleotides or an analog thereof, or a terminal hydroxyl group, bonded to the 3'-end of the ribosyl or deoxyribosyl of Z, or an analog thereof; x is an integer from 1 to 20] and has a structure comprising the same.

[0029] Embodiment A20 is the structure according to Embodiment A19, wherein the structure encodes a gene, and optionally Zx is located in the translation region of the gene, or Zx is located in the untranslated region of the gene.

[0030] Embodiment A21 is a polynucleotide library, the library contains at least 5000 unique polynucleotides, and each polynucleotide contains at least one nucleobase described in any one of Embodiments A1 - A5, which is a polynucleotide library.

[0031] Embodiment A22 is a nucleoside triphosphate containing a nucleobase, and the nucleobase is

Chem.

[0032] In Embodiment A23, the nucleobase is

Chem.

[0033] Embodiment A24 is the nucleoside triphosphate described in Embodiment A22 or A23, and the nucleoside contains ribose or deoxyribose.

[0034] Embodiment A25 is a DNA containing a nucleobase having the structure

Chem.

Chem.

[0035] Embodiment A26 is a DNA containing a nucleobase having the structure

Chem.

Chem.

[0036] Embodiment A27 is a method of transcribing DNA into tRNA or mRNA encoding a protein, comprising contacting DNA containing a gene encoding tRNA or a protein with ribonucleoside triphosphates and RNA polymerase, wherein the gene encoding tRNA or a protein forms a pair with a second unnatural base and comprises a first unnatural base pairing that forms a first unnatural base pair with the second unnatural base, the ribonucleoside triphosphates comprise a third unnatural base capable of forming a first unnatural base pair with the second unnatural base, and the first unnatural base pair and the second unnatural base pair are not the same.

[0037] Embodiment A28 is the method according to Embodiment A27, wherein the ribonucleoside triphosphates further comprise a fourth unnatural base, and the fourth unnatural base is capable of forming a second unnatural base pair with the third unnatural base.

[0038] Embodiment A29 is the method according to Embodiment A28, wherein the first unnatural base pair and the second unnatural base pair are not the same.

[0039] Embodiment A30 further comprises a step of replicating DNA by contacting DNA with deoxyribonucleoside triphosphates and DNA polymerase before the step of contacting DNA with ribonucleoside triphosphates and RNA polymerase, the ribonucleoside triphosphates comprise a fifth unnatural base capable of forming a first unnatural base pair with a fifth unnatural base, and the first unnatural base pair and the fifth unnatural base pair are not the same. The method is as described in any one of Embodiments A27 to A29.

[0040] Embodiment A30 is the method according to any one of Embodiments A27 to A30, wherein the first unnatural base contains TPT3, the second unnatural base contains CNMO or NaM, the third unnatural base contains TAT1, and the fourth unnatural base contains NaM or 5FM.

[0041] Embodiment A32 is the method according to any one of Embodiments A27 to A31, wherein the method includes the use of a semi-synthetic organism, optionally the organism is a bacterium, and optionally the bacterium is Escherichia coli.

[0042] Embodiment A33 is the method according to Embodiment A32, wherein the organism includes a microorganism.

[0043] Embodiment A34 is the method according to Embodiment A32, wherein the organism includes a bacterium.

[0044] Embodiment A35 is the method according to Embodiment A34, wherein the organism includes a gram-positive bacterium.

[0045] Embodiment A36 is the method according to Embodiment A34, wherein the organism includes a gram-negative bacterium.

[0046] Embodiment A37 is the method according to any one of Embodiments A27 to A34, wherein the organism includes Escherichia coli.

[0047] Embodiment A38 is such that at least one unnatural base is (i) 2-Thiouracil, 2-thio-thymine, 2'-deoxyuridine, 4-thio-uracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl (I), 5-halouracil; 5-propynyl-uracil, 6-azothymine, 6-azouracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, methyl uracil-5-oxoacetate, uracil-5-oxoacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarbonylmethyluracil, 5-methoxyuracil, uracil-5-oxyacetic acid, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil or dihydrouracil; (ii) 5-Hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytarabine, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azacytosine, azacytosine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one) or pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d]pyrimidin-2-one); (iii) Adenine substituted with 2-aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2'-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl, N6-isopentenyladenine, 2-methyladenine, 2,6-dia minopurine, 2-methylthio-N6-isopentenyladenine or 6-aza-adenine; (iv) 2-Methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thio-guanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, guanine substituted with 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-aza-guanine; and (v) The method according to any one of embodiments A27 to A37, selected from the group consisting of hypoxanthine, xanthine, 1-methylinosine, queuosine, beta-D-galactosyl queuosine, inosine, beta-D-mannosyl queuosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone.

[0048] Embodiment A39 is such that at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is

Chemical formula

[0049] Embodiment A40 is such that at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is

Chemical formula

[0050] In Embodiment A41, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0051] In Embodiment A42, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0052] In Embodiment A43, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0053] In Embodiment A44, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0054] In Embodiment A45, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0055] In Embodiment A46, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0056] In Embodiment A47, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0057] In Embodiment A48, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0058] In Embodiment A49, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0059] In Embodiment A50, the first or second unnatural base is

Chemical formula

[0060] In Embodiment A51, the first or second unnatural base is [Chemical formula] It is the method described in Embodiment A40.

[0061] In Embodiment A52, the first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] or the first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] It is the method described in any one of Embodiments A27 - A40.

[0062] In Embodiment A53, the third or fourth unnatural base is [Chemical formula] It is the method described in any one of Embodiments A27 - A40 and A52.

[0063] In Embodiment A54, the third unnatural base is [Chemical formula] It is the method described in Embodiment A53.

[0064] In Embodiment A55, the fourth unnatural base is [Chemical formula] It is the method described in Embodiment A54.

[0065] Embodiment A56 is such that the third or fourth unnatural base is [Chemical formula] the method according to any one of Embodiments A27 to A52.

[0066] Embodiment A57 is such that the third unnatural base is [Chemical formula] the method according to Embodiment A56.

[0067] Embodiment A58 is such that the fourth unnatural base is [Chemical formula] the method according to Embodiment A56.

[0068] Embodiment A59 is such that the first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] the method according to any one of Embodiments A27 to A40.

[0069] Embodiment A60 is such that the first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] It is the method according to any one of Embodiments A27 to A40.

[0070] In Embodiment A61, the first unnatural base is [Chemical formula] and the second unnatural base is

[0071] [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] It is the method according to any one of Embodiments A27 to A40.

[0072] In Embodiment A62, the third unnatural base is [Chemical formula] It is the method according to any one of Embodiments A27 to A40.

[0073] In Embodiment A63, the fourth unnatural base is [Chemical formula] It is the method according to any one of Embodiments A27 to A51.

[0074] Embodiment A64, the first unnatural base is

Chem.

Chem.

Chem.

Chem.

[0075] Embodiment A65, the third unnatural base and the fourth unnatural base contain ribose, and the method according to any one of Embodiments A27 to A64.

[0076] Embodiment A66, the third unnatural base and the fourth unnatural base contain deoxyribose, and the method according to any one of Embodiments A27 to A64.

[0077] Embodiment A67, the first and second unnatural bases contain deoxyribose, and the method according to any one of Embodiments A27 to A66.

[0078] Embodiment A68, the first and second unnatural bases contain deoxyribose, and the third and fourth unnatural bases contain ribose, and the method according to any one of Embodiments A27 to A64.

[0079] Embodiment A69, the DNA is

Chem.

[0080] Embodiment A70 is the method according to Embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-d5SICS.

[0081] Embodiment A71 is the method according to Embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTPT3.

[0082] Embodiment A72 is the method according to Embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3.

[0083] Embodiment A73 is the method according to Embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dPTMO-dTPT3.

[0084] Embodiment A74 is the method according to Embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1.

[0085] Embodiment A75 is the method according to Embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.

[0086] Embodiment A76, wherein the DNA

Chemical formula

Chemical formula

[0087] Embodiment A77 is the method according to Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-d5SICS.

[0088] Embodiment A78 is the method according to Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTPT3.

[0089] Embodiment A79 is the method according to Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3.

[0090] Embodiment A80 is the method according to Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dPTMO-dTPT3.

[0091] Embodiment A81 is the method according to Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1.

[0092] Embodiment A82 is the method according to Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.

[0093] Embodiment A83 is the method according to any one of Embodiments A76 to A82, wherein the mRNA and tRNA comprise

Chemical formula

[0094] Embodiment A84 is the method according to Embodiment A76, wherein the mRNA and tRNA comprise [Chemical formula] The method according to embodiment A83, comprising a non-natural base selected from

[0095] Embodiment A85, wherein the mRNA is [Chemical formula] The method according to embodiment A83, comprising a non-natural base represented by

[0096] Embodiment A86, wherein the mRNA is [Chemical formula] The method according to embodiment A83, comprising a non-natural base represented by

[0097] Embodiment A87, wherein the mRNA is [Chemical formula] The method according to embodiment A83, comprising a non-natural base represented by

[0098] Embodiment A88, wherein the tRNA is [Chemical formula] The method according to any one of embodiments A76 to A87, comprising a non-natural base selected from

[0099] Embodiment A89, wherein the tRNA is [Chemical formula] The method according to embodiment A88, comprising a non-natural base represented by

[0100] Embodiment A90, wherein the tRNA is [Chemical formula] The method according to embodiment A88, comprising a non-natural base represented by

[0101] Embodiment A91 is a method according to any one of Embodiments A76 - A87, wherein the tRNA contains unnatural bases represented by

Chemical formula

[0102] Embodiment A92 is a method according to any one of Embodiments A27 - A40, wherein the first unnatural base contains dCNMO and the second unnatural base contains dTPT3.

[0103] Embodiment A93 is a method according to any one of Embodiments A27 - A40 and A92, wherein the third unnatural base contains NaM and the second unnatural base contains TAT1.

[0104] Embodiment A94 is a method according to any one of Embodiments A27 - A93, wherein the first unnatural base or the second unnatural base is recognized by DNA polymerase.

[0105] Embodiment A95 is a method according to any one of Embodiments A27 - A94, wherein the third unnatural base or the fourth unnatural base is recognized by RNA polymerase.

[0106] Embodiment A96 is a method according to any one of Embodiments A27 - A95, wherein the mRNA is transcribed, and the method further includes a step of translating the mRNA into a protein, and the protein contains unnatural amino acids at positions corresponding to the codons of the mRNA containing the third unnatural base.

[0107] Embodiment A97 is a method according to any one of Embodiments A27 - A96, wherein the protein contains at least two unnatural amino acids.

[0108] Embodiment A98 is a method according to any one of Embodiments A27 - A96, wherein the protein contains at least three unnatural amino acids.

[0109] Embodiment A99 is the method according to any one of Embodiments A27 - A98, wherein the protein contains at least two different unnatural amino acids.

[0110] Embodiment A100 is the method according to any one of Embodiments A27 - A98, wherein the protein contains at least three different unnatural amino acids.

[0111] Embodiment A101 is that at least one unnatural amino acid is a lysine analog; contains an aromatic side chain; contains an azide group; contains an alkyne group; or contains an aldehyde or ketone group, which is the method according to any one of Embodiments A27 - A100.

[0112] Embodiment A102 is the method according to any one of Embodiments A27 - A101, wherein at least one unnatural amino acid does not contain an aromatic side chain.

[0113] Embodiment A103 is that at least one unnatural amino acid is N6 - azidoethoxy - carbonyl - L - lysine (AzK), N6 - propargylethoxy - carbonyl - L - lysine (PraK), BCN - L - lysine, norbornene lysine, TCO - lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2 - amino - 8 - oxononanoic acid, 2 - amino - 8 - oxooctanoic acid, p - acetyl - L - phenylalanine, p - azidomethyl - L - phenylalanine (pAMF), p - iodo - L - phenylalanine, m - acetylphenylalanine, 2 - amino - 8 - oxononanoic acid, p - propargyloxyphenylalanine, p - propargyl - phenylalanine, 3 - methyl - phenylalanine, L - dopa, fluorinated phenylalanine, isopropyl - L - phenylalanine, p - azido - L - phenylalanine, p - acyl - L - phenylalanine, p - benzoyl - L The method according to any one of Embodiments A27 to A102, comprising -phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid or selenocysteine.

[0114] Embodiment A104 is the method according to Embodiment A102 or A103, wherein at least one unnatural amino acid comprises N6-azidoethoxy-carbonyl-L-lysine (AzK) or N6-propynylethoxy-carbonyl-L-lysine (PraK).

[0115] Embodiment A105 is the method according to Embodiment A104, wherein at least one unnatural amino acid comprises N6-azidoethoxy-carbonyl-L-lysine (AzK).

[0116] Embodiment A106 is the method according to Embodiment A104, wherein at least one unnatural amino acid comprises N6-propynylethoxy-carbonyl-L-lysine (PraK).

[0117] Embodiment A107 is mRNA generated from the method according to any one of Embodiments A27 to A106.

[0118] Embodiment A108 is tRNA generated from the method according to any one of Embodiments A27 to A106.

[0119] Embodiment A109 is a protein encoded by the mRNA according to Embodiment A107, which contains an unnatural amino acid at a position corresponding to the codon of the mRNA containing a third unnatural base.

[0120] Embodiment A110 is a semi-synthetic organism comprising an expanded genetic alphabet, the genetic alphabet comprising at least three distinct unnatural bases.

[0121] Embodiment A111 is the semi-synthetic organism according to Embodiment A110, wherein the organism comprises a microorganism, optionally the microorganism is Escherichia coli.

[0122] Embodiment A112 is

Chemical formula

[0123] Embodiment A113 is that the DNA contains at least one unnatural base pair (UBP), The unnatural base pair (UBP) is dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1, d5FM-dTAT1 or dNaM-dTAT1, and it is the semi-synthetic organism according to any one of Embodiments A110 to A113.

[0124] Embodiment A114 is that the DNA

Chemical formula

[0125] Embodiment A115 is the semi-synthetic organism according to any one of Embodiments A110 to A115, which expresses a heterologous nucleoside triphosphate transporter.

[0126] Embodiment A116 is the semi-synthetic organism according to Embodiment A115, wherein the heterologous nucleoside triphosphate transporter is PtNTT2.

[0127] Embodiment A117 is a semi-synthetic organism according to any one of Embodiments A110 to A116, which further expresses a heterologous tRNA synthetase.

[0128] In Embodiment A118, the heterologous tRNA synthetase is the pyrrolysyl-tR NA synthetase (Mb PylRS) of M. barkeri, and it is the semi-synthetic organism according to Embodiment A117.

[0129] Embodiment A119 is a semi-synthetic organism according to any one of Embodiments A110 to A118, which further expresses a heterologous RNA polymerase.

[0130] In Embodiment A120, the heterologous RNA polymerase is T7 RNAP, and it is the semi-synthetic organism according to Embodiment A119.

[0131] Embodiment A121 is a semi-synthetic organism according to any one of Embodiments A110 to A120, which does not express a protein having a DNA recombination repair function.

[0132] Embodiment A122 is a semi-synthetic organism according to Embodiment A121, which does not express RecA.

[0133] Embodiment A123 is a semi-synthetic organism according to any one of Embodiments A110 to A122, which further contains a heterologous mRNA.

[0134] In Embodiment A124, the heterologous mRNA is

Chemical formula

[0135] Embodiment A125 is a semi-synthetic organism according to any one of Embodiments A110 to A125, which further contains a heterologous tRNA.

[0136] In Embodiment A126, the heterologous tRNA is [Chem.] The semi-synthetic organism according to Embodiment A125, comprising at least one unnatural base selected from the group consisting of

[0137] Embodiment A127 is a method of transcribing DNA, the method comprising: (1) providing one or more DNAs comprising a gene encoding a protein, wherein the template strand of the gene encoding the protein comprises a first unnatural base, and (2) a gene encoding a tRNA, wherein the template strand of the gene encoding the tRNA comprises a second unnatural base capable of forming a base pair with the first unnatural base; transcribing the gene encoding the protein to incorporate a third unnatural base into the mRNA, wherein the third unnatural base is capable of forming a first unnatural base pair with the first unnatural base; transcribing the gene encoding the tRNA to incorporate a fourth unnatural base into the tRNA, wherein the fourth unnatural base is capable of forming a second unnatural base pair with the second unnatural base and the first unnatural base pair and the second unnatural base pair are not the same.

[0138] Embodiment A128 further comprises the step of translating a protein from the mRNA using the tRNA, wherein the protein comprises an unnatural amino acid at a position corresponding to a codon comprising the third unnatural base in the mRNA, according to the method of Embodiment A127.

[0139] Embodiment A129 is a method of replicating DNA, the method comprising: (1) A gene encoding a protein, wherein the template strand of the gene encoding the protein contains a first unnatural base, and (2) a gene encoding a tRNA, wherein the template strand of the gene encoding the tRNA contains a second unnatural base capable of forming a base pair with the first unnatural base, and providing a DNA comprising the gene encoding the tRNA; and Replicating the DNA and incorporating a first alternative unnatural base instead of the first unnatural base and / or incorporating a second alternative unnatural base instead of the second unnatural base comprising; The method optionally further comprises Transcribing the gene encoding the protein and incorporating a third unnatural base into the mRNA, wherein the third unnatural base is capable of forming a first unnatural base pair with the first unnatural base and / or the first alternative unnatural base; and / or Transcribing the gene encoding the tRNA and incorporating a fourth unnatural base into the tRNA, wherein the fourth unnatural base is capable of forming a second unnatural base pair with the second unnatural base and / or the second alternative unnatural base and the first unnatural base pair and the second unnatural base pair are not the same.

[0140] Embodiment A130 is the method according to Embodiment A129, further comprising transcribing the gene encoding the protein and incorporating a third unnatural base into the mRNA, wherein the third unnatural base is capable of forming a first unnatural base pair with the first unnatural base and / or the first alternative unnatural base.

[0141] Embodiment A131 is a process of transcribing a gene encoding tRNA and incorporating a fourth unnatural base into the tRNA, the fourth unnatural base being capable of forming a second unnatural base pair with a second unnatural base and / or a second alternative unnatural base, further comprising the step, wherein the first unnatural base pair and the second unnatural base pair are not the same, which is the method according to Embodiment A129 or A130.

[0142] Embodiment A132 is the method according to any one of Embodiments A127 to A131, wherein the method comprises the use of a semi-synthetic organism.

[0143] Embodiment A133 is the method according to Embodiment A132, wherein the organism comprises a microorganism.

[0144] Embodiment A134 is the method according to Embodiment A132 or A133, which is an in vivo method comprising the use of a semi-synthetic organism that is a bacterium.

[0145] Embodiment A135 is the method according to Embodiment A134, wherein the organism comprises Gram-positive bacteria.

[0146] Embodiment A136 is the method according to Embodiment A134, wherein the organism comprises Gram-negative bacteria.

[0147] Embodiment A137 is the method according to any one of Embodiments A132 to A134, wherein the organism comprises Escherichia coli.

[0148] Embodiment A138 is such that at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is

Chemical formula

[0149] In Embodiment A139, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is [Chemical formula] The method according to Embodiment A138, which includes

[0150] In Embodiment A140, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is [Chemical formula] The method according to Embodiment A138, which includes

[0151] In Embodiment A141, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is [Chemical formula] The method according to Embodiment A138, which includes

[0152] In Embodiment A142, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is [Chemical formula] The method according to Embodiment A138, which is

[0153] In Embodiment A143, at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is [Chemical formula] The method according to Embodiment A138, which includes

[0154] In Embodiment A144, the first or second unnatural base is [Chemical formula] It is the method described in Embodiment A138.

[0155] In Embodiment A145, the first or second unnatural base is

Chemical formula

[0156] In Embodiment A146, the first unnatural base is

Chemical formula

Chemical formula

[0157] In Embodiment A147, the first unnatural base is

Chemical formula

Chemical formula

[0158] In Embodiment A148, the third or fourth unnatural base is

Chemical formula

[0159] In Embodiment A149, the third unnatural base is

Chemical formula

[0160] Embodiment A150 is such that the fourth unnatural base is [Chemical formula] the method according to Embodiment A148.

[0161] Embodiment A151 is such that the third or fourth unnatural base is [Chemical formula] the method according to any one of Embodiments A138, A146, and A147. .

[0162] Embodiment A152 is such that the third unnatural base is [Chemical formula] the method according to Embodiment A151.

[0163] Embodiment A153 is such that the fourth unnatural base is [Chemical formula] the method according to Embodiment A151.

[0164] Embodiment A154 is such that the first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] The method according to Embodiment A138, which is as follows.

[0165] In Embodiment A155, the first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] The method according to Embodiment A138, which is as follows.

[0166] In Embodiment A156, the first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] The method according to Embodiment A138, which is as follows.

[0167] In Embodiment A157, the third unnatural base is [Chemical formula] The method according to Embodiment A138, which is as follows.

[0168] Embodiment A158 is such that the fourth unnatural base is [Chemical formula] the method according to Embodiment A138.

[0169] Embodiment A159 is such that the first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] the method according to Embodiment A138.

[0170] Embodiment A160 is the method according to any one of Embodiments A127 to A159, wherein the third and fourth unnatural bases contain ribose.

[0171] Embodiment A161 is the method according to any one of Embodiments A127 to A159, wherein the third and fourth unnatural bases contain deoxyribose.

[0172] Embodiment A162 is the method according to any one of Embodiments A127 to A161, wherein the first and second unnatural bases contain deoxyribose.

[0173] Embodiment A163 is the method according to any one of Embodiments A127 to A159, wherein the first and second unnatural bases contain deoxyribose and the third and fourth unnatural bases contain ribose.

[0174] Embodiment A164 is the method according to any one of Embodiments A127 - A137, wherein the DNA template comprises at least one unnatural base pair (UBP) selected from the group consisting of

Chemical formula

[0175] Embodiment A165 is the method according to Embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM - d5SICS.

[0176] Embodiment A166 is the method according to Embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dCNMO - dTPT3.

[0177] Embodiment A167 is the method according to Embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM - dTPT3.

[0178] Embodiment A168 is the method according to Embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM - dTAT1.

[0179] Embodiment A169 is the method according to Embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dCNMO - dTAT1.

[0180] Embodiment A170 is such that the DNA template

Chemical formula

Chemical formula

[0181] Embodiment A171 is the method according to Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM-d5SICS.

[0182] Embodiment A172 is the method according to Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dCNMO-dTPT3.

[0183] Embodiment A173 is the method according to Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM-dTPT3.

[0184] Embodiment A174 is the method according to Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM-dTAT1.

[0185] Embodiment A175 is the method according to Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dCNMO-dTAT1.

[0186] Embodiment A176 is the method according to any one of Embodiments A127 to A175, wherein the mRNA and tRNA comprise an unnatural base selected from

Chemical formula

[0187] Embodiment A177 is the method according to Embodiment A176, wherein the mRNA and tRNA comprise an unnatural base selected from

Chemical formula

[0188] Embodiment A178 is the method according to Embodiment A176, wherein the mRNA contains unnatural bases represented by

Chem.

[0189] Embodiment A179 is the method according to Embodiment A176, wherein the mRNA contains unnatural bases represented by

Chem.

[0190] Embodiment A180 is the method according to Embodiment A176, wherein the mRNA contains unnatural bases represented by

Chem.

[0191] Embodiment A181 is the method according to Embodiment A176, wherein the tRNA contains unnatural bases selected from

Chem.

[0192] Embodiment A182 is the method according to Embodiment A176, wherein the tRNA contains unnatural bases represented by

Chem.

[0193] Embodiment A183 is the method according to Embodiment A176, wherein the tRNA contains unnatural bases represented by

Chem.

[0194] Embodiment A184 is the method according to Embodiment A176, wherein the tRNA contains unnatural bases represented by

Chem.

[0195] In embodiment A185, the first unnatural base includes dCNMO, and the second unnatural base includes dTPT3, and it is the method according to any one of embodiments A127 - A137.

[0196] In embodiment A186, the third unnatural base includes NaM, and the second unnatural base includes TAT1, and it is the method according to any one of embodiments A127 - A137.

[0197] In embodiment A187, the protein includes at least two unnatural amino acids, and it is the method according to any one of embodiments A127 - A186.

[0198] In embodiment A188, the protein includes at least three unnatural amino acids, and it is the method according to any one of embodiments A127 - A186.

[0199] In embodiment A189, the protein includes at least two different unnatural amino acids, and it is the method according to any one of embodiments A127 - A186.

[0200] In embodiment A190, the protein includes at least three different unnatural amino acids, and it is the method according to any one of embodiments A127 - A186.

[0201] In embodiment A191, at least one unnatural amino acid is a lysine analog; includes an aromatic side chain; includes an azide group; includes an alkyne group; or includes an aldehyde or ketone group and it is the method according to any one of embodiments A127 - A190.

[0202] Embodiment A192 is the method according to any one of Embodiments A127 - A191, wherein at least one unnatural amino acid does not contain an aromatic side chain.

[0203] Embodiment A193 is the method according to Embodiment A191 or A192, wherein at least one unnatural amino acid comprises N6 - azidoethoxy - carbonyl - L - lysine (AzK) or N6 - propargylethoxy - carbonyl - L - lysine (PraK).

[0204] Embodiment A194 is the method according to Embodiment A193, wherein at least one unnatural amino acid comprises N6 - azidoethoxy - carbonyl - L - lysine (AzK).

[0205] Embodiment A195 is the method according to Embodiment A193, wherein at least one unnatural amino acid comprises N6 - propargylethoxy - carbonyl - L - lysine (PraK).

[0206] Various aspects of the present invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description of exemplary embodiments in which the principles of the invention are utilized, and to the accompanying drawings.

Brief Description of the Drawings

[0207]

Figure 1A

Figure 1B

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 7A

Figure 7B

Figure 7C

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 8E

Figure 8F

Figure 8G

Mode for Carrying Out the Invention

[0208] Specific Terms Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claimed subject matter. In case of conflict between the explicit content of this disclosure and the materials incorporated herein by reference, the explicit content shall prevail. The In this application, unless otherwise specified, the use of the singular form includes the plural form. As used in this specification and the appended claims, it should be noted that the singular forms "a (indefinite article)", "an (indefinite article)", and "the (definite article)" include plural referents unless the context clearly indicates otherwise. In this application, the use of "or" means "and / or" unless otherwise specified. Further, the term "including", as well as the use of other forms such as "include", "includes", and "included", is not limiting.

[0209] As used herein, ranges and amounts may be expressed as "about" a particular value or range. "About" includes the exact amount. Thus, "about 5 μL" means "about 5 μL" and also serves as an explanation of "5 μL". Generally, the term "about" includes amounts that are expected to be within the range of experimental error.

[0210] As used herein, phrases such as "under conditions suitable for providing" or "under conditions sufficient to produce" in the context of a synthetic method refer to reaction conditions such as time, temperature, solvent, reactant concentration, etc., that are within the scope of ordinary techniques that an experimenter can vary to provide a useful amount or yield of the reaction product. The desired reaction product need not be the only reaction product, nor does the starting material need to be completely consumed, provided that the desired reaction product can be isolated or otherwise further used.

[0211] "Chemically feasible" means a bond arrangement or compound that does not violate the generally understood rules of organic structures. For example, in certain situations, it is understood that a structure within the scope of a claim that contains a pentavalent carbon atom, which does not exist in nature, is not within the scope of the claim. The structures disclosed herein are intended to include only "chemically feasible" structures in all of their embodiments. For example, in a structure represented by a variable atom or group, any recited structure that is not chemically feasible is not intended to be disclosed or claimed herein.

[0212] As used herein, a "chemical analog" of a chemical structure refers to a chemical structure that may not be readily derivable synthetically from the parent structure but retains a substantial similarity to the parent structure. In some embodiments, a nucleotide analog is a non-natural nucleotide. In some embodiments, a nucleoside analog is a non-natural nucleoside. A related chemical structure that is readily derivable synthetically from the parent chemical structure is called a "derivative."

[0213] As used herein, "base" or "nucleic acid base" refers to at least the nucleic acid base portion of a nucleoside or nucleotide (nucleosides and nucleotides include ribo or deoxyribo variants), and may in some cases include further modifications to the sugar portion of the nucleoside or nucleotide. In some cases, "base" is also used to represent the entire nucleoside or nucleotide (e.g., a "base" can be incorporated into DNA by a DNA polymerase or into RNA by an RNA polymerase). However, the term "base" should not necessarily be construed to represent the entire nucleoside or nucleotide unless required by the context. In the chemical structures of bases or nucleic acid bases provided herein, only the base of the nucleoside or nucleotide is shown, and the sugar portion, and optionally any phosphate residues for clarity, are omitted. When used in the chemical structures of bases or nucleic acid bases provided herein, a wavy line represents the connection to the nucleoside or nucleotide, and the sugar portion of the nucleoside or nucleotide can be further modified. In some embodiments, the wavy line represents the attachment of the base or nucleic acid base to a sugar portion such as the pentose of the nucleoside or nucleotide. In some embodiments, the pentose is ribose or deoxyribose.

[0214] In some embodiments, a nucleobase is generally the heterocyclic base portion of a nucleoside. The nucleobase may be naturally occurring, may be modified, may have no similarity to natural bases, and / or may be synthesized, for example, by organic synthesis. In certain embodiments, a nucleobase includes any atom or group of atoms in a nucleoside or nucleotide that can interact with the base of another nucleic acid, with or without the use of hydrogen bonding. In certain embodiments, an unnatural nucleobase is not derived from a natural nucleobase. Although an unnatural nucleobase does not necessarily possess basic properties, it should be noted that it is referred to as a nucleobase for simplicity. In some embodiments, when referring to a nucleobase, "(d)" indicates that the nucleobase can be attached to deoxyribose or ribose, and "d" without parentheses indicates that the nucleobase is attached to deoxyribose.

[0215] In some embodiments, a nucleoside is a compound that includes a nucleobase portion and a sugar portion. Nucleosides include, but are not limited to, naturally occurring nucleosides (found in DNA and RNA), abasic nucleosides, modified nucleosides, and nucleosides having mimetic bases and / or sugar groups. Nucleosides include nucleosides having any of a variety of substituents. A nucleoside can be a glycoside compound formed by a glycosidic linkage between a nucleobase and the reducing group of a sugar.

[0216] As used herein, "nucleotide" refers to a compound that includes a nucleoside moiety and a phosphate moiety. Exemplary natural nucleotides include, but are not limited to, adenosine triphosphate (ATP), uridine triphosphate (UTP), cytidine triphosphate (CTP), guanosine triphosphate (GTP), adenosine diphosphate (ADP), uridine diphosphate (UDP), cytidine diphosphate (CDP), guanosine diphosphate (GDP), adenosine monophosphate (AMP), uridine monophosphate (UMP), cytidine monophosphate (CMP), and guanosine monophosphate (GMP), deoxyadenosine triphosphate (dATP), deoxythymidine triphosphate (dTTP), deoxycytidine triphosphate (dCTP), deoxyguanosine triphosphate (dGTP), deoxyadenosine diphosphate (dADP), thymidine diphosphate (dTDP), deoxycytidine diphosphate (dCDP), deoxyguanosine diphosphate (dGDP), deoxyadenosine monophosphate (dAMP), deoxythymidine monophosphate (dTMP), deoxycytidine monophosphate (dCMP), and deoxyguanosine monophosphate (dGMP). Exemplary natural deoxyribonucleotides that include deoxyribose as the sugar moiety include dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Exemplary natural ribonucleotides that include ribose as the sugar moiety include ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, and GMP.

[0217] As used herein, polynucleotide refers to DNA, RNA, DNA or RNA-like polymers such as peptide nucleic acid (PNA), locked nucleic acid (LNA), phosphorothioate, examples of which are well known in the art and may include unnatural bases. Polynucleotides can be synthesized on an automated synthesizer using, for example, phosphoramidite chemistry or other chemical approaches compatible with the use of a synthesizer.

[0218] Examples of DNA include, but are not limited to, complementary DNA (cDNA) and genomic DNA (gDNA). The DNA can be attached to another molecule, including but not limited to RNA and peptides, by covalent or non-covalent means. Examples of RNA include coding RNA, such as messenger RNA (mRNA). Examples of RNA also include non-coding RNA, such as ribosomal RNA (rRNA). Examples of RNA also include transfer RNA (tRNA), RNA interference (RNAi), small nucleolar RNA (snoRNA), microRNA (miRNA), small interfering RNA (si RNA) (also referred to as short interfering RNA), small nuclear RNA (snRNA), extracellular RNA (exRNA), PIWI-interacting RNA (piRNA), and long non-coding RNA (long ncRNA). In some embodiments, the RNA is rRNA, tRNA, RNAi, snoRNA, microRNA, siRNA, snRNA, exRNA, piRNA, long ncRNA, or any combination or hybrid thereof. Optionally, the RNA is a component of a ribozyme. The DNA and RNA can be in any form including, but not limited to, linear, circular, supercoiled, single-stranded, and double-stranded.

[0219] Peptide nucleic acid (PNA) is a synthetic DNA / RNA analog in which a peptide-like backbone replaces the sugar-phosphate backbone of DNA or RNA. PNA oligomers exhibit higher binding strength and greater specificity in binding to complementary DNA, and PNA / DNA base mismatches are less stable than similar mismatches in DNA / DNA duplexes. This binding strength and specificity also apply to PNA / RNA duplexes. PNA is resistant to enzymatic degradation because it is not readily recognized by either nucleases or proteases. PNA is also stable over a wide pH range. See also Nielsen PE, Egholm M, Berg RH, Buchardt O (December 1991). "Sequence-selective recognition of DNA by strand displacement with a thymine-substituted polyamide", Science 254(5037):1497-500. doi:10.1126 / science.1962210. PMID 1962210; and Egholm M, Buchardt O, Christensen L, Behrens C, Freier SM, Driver DA, Berg RH, Kim SK, Norden B, and Nielsen PE (1993), "PNA Hybridizes to Complementary Oligonucleotides Obeying the Watson-Crick Hydrogen Bonding Rules". Nature 365(6446):566-8. doi:10.1038 / 365566a0. PMID 7692304; the disclosures of each are hereby incorporated by reference in their entirety.

[0220] ​Locked nucleic acids (LNAs) are modified RNA nucleotides, where the ribose moiety of the LNA nucleotide is modified with an additional bridge connecting the 2'-oxygen and 4'-carbon. This bridge "locks" the ribose in the 3'-endo (North) conformation, which is commonly seen in A-form duplexes. LNA nucleotides can be mixed with DNA or RNA residues of oligonucleotides as needed. Such oligomers may be chemically synthesized and are commercially available. The locked ribose conformation enhances base stacking and pre-organization of the backbone. For example, Kaur, H; Arora, A; Wengel, J; Maiti, S (2006), "Thermodynamic, Counterion, and Hydration Effects for the Incorporation of Locked Nucleic Acid Nucleotides into DNA Duplexes", Biochemistry 45(23):7347-55.doi:10.1021 / bi060307w.PMID 16752924; Owczarzy R.; You Y., Groth C.L., Tataurov A.V. (2011), "Stability and mismatch discrimination of locked nucleic acid-DNA duplexes", Biochem.50(43):9352-936 7.doi:10.1021 / bi200904e.PMC3201676.PMID21928795; Alexei A. Koshkin; Sanjay K. Singh, Poul Nielsen, Vivek K. Rajwanshi, Ravindra Kumar, Michael Meldgaard, Carl Erik Olsen, Jesper Wengel (1998), "LNA (Locked Nucleic Acids): Synthesis of the adenine, cytosine, guanine, 5-methylcytosine, thymine and uracil bicyclonucleoside monomers, oligomerisation, and unprecedented nucleic acid recognition (LNA (Locked Nucleic Acids): Synthesis of the adenine, cytosine, guanine, 5-methylcytosine, thymine and uracil bicyclonucleoside monomers, oligomerisation, and unprecedented nucleic acid recognition)", Tetrahedron 54(14):3607-30.doi:10.1016 / S0040-4020(98)00094-5; and, Satoshi Obika; Daishu Nanbu, Yoshiyuki Hari, Ken-ichiro Morio, Yasuko In, Toshimasa Ishida, Takeshi Imanishi (1997), "Synthesis of 2’-O,4’-C-methyleneuridine and -cytidine. Novel bicyclic nucleosides having a fixed C3’-endo sugar puckering)", Tetrahedron Lett. 38(50):8735-8.doi:10.1016 / S0040-4039(97)10322-7 (each disclosure is hereby incorporated by reference in its entirety).

[0221] As used herein, the term "gene" refers to a polynucleotide that encodes the synthesis of a gene product such as RNA or protein.

[0222] A molecular beacon or molecular beacon probe is an oligonucleotide hybridization probe that can detect the presence of a specific nucleic acid sequence in a homogeneous solution. A molecular beacon is a hairpin-shaped molecule with an internally quenched fluorophore that fluoresces upon binding to the target nucleic acid sequence. For example, see Tyagi S, Kramer FR (1996), "Molecular beacons: probes that fluoresce upon hybridization", Nat Biotechnol. 14(3):303-8. PMID 9630890; Tapp I, Malmberg L, Rennel E, Wik M, Syvanen AC (April 2000), "Homogeneous scoring of single-nucleotide polymorphisms: comparison of the 5'-nuclease TaqMan assay and Molecular Beacon probes", Biotechniques 28(4):732-8. PMID 10769752; and Akimitsu Okamoto (2011), "ECHO probes: a concept for fluorescence control for practical nucleic acid sensing", Chem. Soc. Rev. 40:5815-5828 (each disclosure is incorporated herein by reference in its entirety).

[0223] As used herein, the term "unnatural base" refers to a base other than A, C, G, T, U, and other naturally occurring bases (e.g., 5-methylcytosine, pseudouridine, and inosine).

[0224] ​As used herein, the term "unnatural base pair" refers to two bases that bind to each other and are on opposite strands of a double-stranded polynucleotide (which may be, for example, a molecule that is at least partially self-hybridized or a pair of molecules that are partially or fully hybridized), where at least one of the two bases is an unnatural base.

[0225] As used herein, a "semi-synthetic organism" is an organism that contains non-natural components, such as an expanded genetic alphabet that includes one or more unnatural bases.

[0226] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0227] Methods and compositions comprising unnatural base pairs In certain embodiments, in vitro and in vivo methods and compositions for generating nucleic acids using an expanded genetic alphabet are disclosed herein (Figure 1). In some examples, the nucleic acids encode non-natural proteins, where the non-natural proteins include non-natural amino acids. In some cases, the in vivo methods or compositions described herein utilize or comprise semi-synthetic organisms. In some examples, the methods include the step of incorporating at least one unnatural base pair (UBP) into one or more nucleic acids. Such base pairs are formed by base pairing between the nucleobases of two nucleosides. In an exemplary workflow, DNA 101 encoding protein 102 and tRNA 103, which is a template strand encoding a region for tRNA containing a protein and unnatural nucleobases (X, Y) that are complementary, can form base pairs and / or are configured to form base pairs, and are transcribed 104 to yield tRNA 106 and mRNA 107. After charging the tRNA with the non-natural amino acid 105, the mRNA 107 is translated 108 to yield a protein 110 that contains one or more non-natural amino acids 109. The methods and compositions described herein, in some examples, enable site-specific incorporation of non-natural amino acids with high fidelity and yield. Also described herein are methods for using semi-synthetic organisms that contain an expanded genetic alphabet and semi-synthetic organisms for generating protein products that contain at least one non-natural amino acid residue.

[0228] The selection of unnatural nucleobases enables the optimization of one or more steps in the methods described herein. For example, the nucleobases are selected for high-efficiency replication, transcription, and / or translation. In some examples, one or more unnatural nucleobase pairs are utilized for the methods described herein. For example, a first set of nucleobases containing deoxyribo moieties are used for DNA replication (e.g., a first nucleobase and a second nucleobase configured to form a first base pair), and a second set of nucleobases (e.g., a third and a fourth nucleobase that are attached to ribose and configured to form a second base pair) are used for transcription / translation. In some embodiments, a first set of nucleobases are used to construct a plasmid (e.g., a first nucleobase and a second nucleobase configured to form a first base pair), a second set of nucleobases are used for replication (e.g., a third nucleobase and a fourth nucleobase configured to form a second base pair), and a third set of bases are used for transcription / translation (e.g., a fifth and a sixth nucleobase configured to form a third base pair). Complementary base pairing between the first set of nucleobases and the second set of nucleobases enables, in some examples, the transcription of genes to yield tRNA or protein from a DNA template containing nucleobases from the first set. Complementary base pairing between the second set of nucleobases (the second base pair) enables, in some examples, the matching of tRNA and mRNA containing unnatural nucleic acids, thereby enabling translation. Enable translation. In some cases, the nucleobases in the first set bind to the deoxyribose moiety. In some cases, the nucleobases in the first set bind to the ribose moiety. In some examples, the nucleobases of both sets are unique. In some examples, at least one nucleobase is the same in both sets. In some examples, the first nucleobase and the third nucleobase are the same. In some embodiments, the first base pair and the second base pair are not the same. In some cases, the first base pair, the second base pair, and the third base pair are not the same.

[0229] In one aspect, an in vivo method for producing a protein comprising a non-natural amino acid, the method comprising: Transcribing a DNA template comprising a first unnatural base and a second unnatural base that is complementary to the first unnatural base, or capable of forming a base pair with the first unnatural base, and / or configured to form a base pair with the first unnatural base to incorporate a third unnatural base into the mRNA, wherein the third unnatural base is complementary to the first unnatural base, or capable of forming a base pair with the first unnatural base, and / or configured to form a base pair with the first unnatural base pair; Transcribing the DNA template to incorporate a fourth unnatural base into the tRNA, wherein the fourth unnatural base is complementary to the second unnatural base, or capable of forming a base pair with the second unnatural base, and / or configured to form a base pair with the second unnatural base pair, and the first unnatural base pair and the second unnatural base pair are not the same; and Translating a protein from the mRNA and the tRAN, wherein the protein comprises a non-natural amino acid. A method comprising the steps is provided herein.

[0230] Nucleic acid molecule In some embodiments, the nucleic acid (e.g., also referred to herein as the nucleic acid molecule of interest) is derived from any origin or composition, e.g., DNA, cDNA, gDNA (genomic DNA), RNA, siRNA (short interfering RNA), RNAi, tRNA, mRNA or rRNA (ribosomal RNA), and can be, for example, in any form (e.g., linear, circular, supercoiled, single-stranded, double-stranded, etc.). In some embodiments, the nucleic acid comprises nucleotides, nucleosides or polynucleotides. In some cases, the nucleic acid comprises natural nucleic acids and non-natural nucleic acids. In some cases, the nucleic acid also comprises non-natural nucleic acids such as DNA or RNA analogs (e.g., containing base analogs, sugar analogs and / or non-native backbones). It is understood that the term "nucleic acid" does not refer to, or imply, a polynucleotide chain of a specific length, and thus polynucleotides and oligonucleotides are also included within its definition. Exemplary natural nucleotides include, but are not limited to, ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, GMP, dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP and dGMP. Exemplary natural deoxyribonucleotides include dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP and dGMP. Exemplary natural ribonucleotides include ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP and GMP. For natural RNA, the uracil-containing nucleoside is uridine. The nucleic acid can be a vector, plasmid, phagemid, autonomously replicating sequence (ARS), centromere, artificial chromosome, yeast artificial chromosome (e.g., YAC), or other nucleic acid that is replicable in, or replicated in, a host cell. In some cases, the non-natural nucleic acid is a nucleic acid analog. In additional cases, the non-natural nucleic acid is of extracellular origin. In other cases, the non-natural nucleic acid is as described in this specification It is available in the intracellular space of organisms provided in the book, such as genetically modified organisms. In some embodiments, the unnatural nucleotide is not a natural nucleotide. In some embodiments, the nucleotide that does not contain a natural base contains an unnatural nucleobase.

[0231] Unnatural nucleic acid Nucleotide analogs or unnatural nucleotides include nucleotides containing some type of modification in any of the base, sugar, or phosphate moieties. The term "modification" (and related grammatical forms such as "modified") does not necessarily mean that the nucleotide analog or unnatural nucleotide is produced by direct modification of a natural nucleotide, rather, it means that the nucleotide analog or unnatural nucleotide is different from a natural nucleotide. In some embodiments, the modification includes chemical modification. In some cases, the modification occurs at the 3’OH or 5’OH group, the backbone, the sugar moiety, or the nucleotide base. In some examples, the modification optionally includes linker molecules that do not occur naturally and / or interstrand or intrastrand cross-links. In one aspect, the modified nucleic acid includes one or more modifications of the 3’OH or 5’OH group, the backbone, the sugar moiety, or the nucleotide base, and / or the addition of linker molecules that do not occur naturally. In one aspect, the modified backbone includes a backbone other than a phosphodiester backbone. In one aspect, the modified sugar includes a sugar other than deoxyribose (in modified DNA), or other than ribose (in modified RNA). In one aspect, the modified base includes a base other than adenine, guanine, cytosine, or thymine (in modified DNA), or other than adenine, guanine, cytosine, or uracil (in modified RNA). In some embodiments, the unnatural nucleotide includes an unnatural base. In some embodiments, the unnatural base is a base having a ring or ring system other than purine or pyrimidine (purine and pyrimidine including purines and pyrimidines having exocyclic substituents), or includes a ring or ring system containing one or more non-nitrogen heteroatoms and / or a ring or ring system that does not contain nitrogen.

[0232] In some embodiments, the nucleic acid comprises at least one modified base. In some examples, the nucleic acid comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more modified bases. In some cases, the modification to the base moiety includes not only natural and synthetic modifications of adenine (A), cytosine (C), guanine (G) and thymine (T) / uracil (U), but also different purine or pyrimidine bases. In some embodiments, the modification is a modified form of adenine, guanine, cytosine or thymine (in modified DNA), or a modified form of adenine, guanine, cytosine or uracil (modified RNA).

[0233] Modified bases of unnatural nucleic acids include, but are not limited to, uracil-5-yl, hypoxanthin-9-yl (I), 2-aminoadenin-9-yl, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azauracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, in particular, 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine. Certain unnatural nucleic acids, such as 5-substituted pyrimidines, 6-azapyrimidines and N-2 substituted purines, N-6 substituted purines, O-6 substituted pur n, 2-aminopropyladenine, 5-propynyluracil, 5-propynylcytosine, 5-methylcytosine, those that increase the stability of double-strand formation, universal nucleic acids, hydrophobic nucleic acids, random nucleic acids, nucleic acids with an enlarged size, fluorinated nucleic acids, 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines include 2-aminopropyladenine, 5-propynyluracil and 5-propynylcytosine, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl of adenine and guanine, other alkyl derivatives, 2-propyl of adenine and guanine and other alkyl derivatives, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyl (-C≡C-CH3) uracil, 5-propynylcytosine, other alkynyl derivatives of pyrimidine nucleic acids, 6-azouracil, 6-azocytosine, 6-azothymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl, other 5-substituted uracils and cytosines, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine, 8-azadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, tricyclic pyrimidines, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), G-clamps, phenoxazine cytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d] pyrimidin-2-one), purine or pyrimidine bases replaced by other heterocycles, 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, 2-pyridone, azacitidine, 5-bromocytosine, bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytarabine, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil and 5-iodouracil, 2-amino-adenine, 6-thio-guanine, 2-thio-thymine, 4-thio-thymine, 5-propynyl-uracil, 4-thio-uracil, N4-ethylcytosine, 7-deazaguanine, 7-deaza-8-azaguanine, 5-hydroxycytosine, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, and U.S. Patent Nos. 3,687,808; 4,845,205; 4,910,300; 4,948,882; 5,093,232; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121; 5,596,091; 5,614,617; 5,645,985; 5,681,941; 5,750,692; 5,763,588; 5,830,653 and 6,005,096; International Publication No. 99 / 62923; Kandimalla et al., (2001) Bioorg. Med. Chem. 9: 807-813; The Concise Encyclopedia of Polymer, Including those described in Science and Engineering, edited by Kroschwitz, J.I., John Wiley & Sons, 1990, pages 858 - 859; Englisch et al., Angewandte Chemie, International Edition, 1991, volume 30, page 613; and Sanghvi, chapter 15, Antisense Research and Applications, edited by Crooke and Lebleu, CRC Press, 1993, pages 273 - 288. Additional base modifications can be found, for example, in U.S. Patent No. 3,687,808; Englisch et al., Angewandte Chemie, International Edition, 1991, volume 30, page 613. In some examples, the unnatural nucleic acid contains the nucleobases of Figure 2. In some examples, the unnatural nucleic acid contains the nucleobases of Figure 4A. In some examples, the unnatural nucleic acid contains the nucleobases of Figure 4B.

[0234] Unnatural nucleic acids containing various heterocyclic bases and various sugar moieties (and sugar analogs) are available in the art, and in some cases, the nucleic acid contains one or several heterocyclic bases other than the five major base components of naturally occurring nucleic acids. For example, as heterocyclic bases, in some cases, uracil - 5 - yl, cytosine - 5 - yl, adenine - 7 - yl, adenine - 8 - yl, guanine - 7 - yl, guanine - 8 - yl, 4 - aminopyrrolo[2.3 - d]pyrimidin - 5 - yl, 2 - amino - 4 - oxopyrrolo[2,3 - d]pyrimidin - 5 - yl, 2 - amino - 4 - oxopyrrolo[2.3 - d]pyrimidin - 3 - yl groups are included, where purine is attached to the sugar moiety of the nucleic acid via the 9 - position, to pyrimidine via the 1 - position, to pyrrolopyrimidine via the 7 - position, and to pyrazolopyrimidine via the 1 - position.

[0235] In some embodiments, the modified bases of the unnatural nucleic acid are represented below, where the wavy line identifies the point of attachment to the sugar of the nucleoside or nucleotide (e.g., deoxyribose or ribose).

[0236]

Chem.

Chem.

Chem.

Chem.

[0237] In some embodiments, a non-natural base (e.g., at least one of the first non-natural base, the second non-natural base, the third non-natural base, or the fourth non-natural base for the method of producing a protein comprising the non-natural amino acids described herein) is

Chem.

Chem.

Chem.

Chem.

Chem.

Chem.

[0238] In some embodiments, the unnatural base (e.g., the first or second unnatural base of the method for producing a protein comprising an unnatural amino acid described herein) is

Chem.

Chem.

Chem.

Chem.

[0239] In some embodiments, the unnatural base (e.g., the third or fourth unnatural base of the method for producing a protein comprising an unnatural amino acid described herein) is

Chem.

Chem.

Chem.

Chem.

Chemical formula

Chemical formula

[0240] In some embodiments, the first unnatural base is

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

[0241] In some embodiments, the third and fourth unnatural bases include ribose. In some embodiments, the third and fourth unnatural bases include deoxyribose. In some embodiments, the first and second unnatural bases include deoxyribose. In some embodiments, the first and second unnatural bases include deoxyribose, and the third and fourth unnatural bases include ribose.

[0242] In some embodiments of methods of producing proteins comprising unnatural amino acids described herein, the DNA

Chemical formula

[0243] In some embodiments of methods of producing proteins comprising unnatural amino acids described herein, the DNA

Chemical formula

[0244] In some embodiments of the method for producing a protein comprising an unnatural amino acid described herein, the DNA is

Chemical formula

Chemical formula

[0245] In some embodiments of the method for producing a protein containing an unnatural amino acid described herein, the DNA is

Chemical formula

Chemical formula

[0246] In some embodiments, the mRNA and tRNA

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

[0247] In some embodiments of the method for producing a protein containing an unnatural amino acid described herein, the first unnatural base contains dCNMO, and the second unnatural base contains dTPT3. In some embodiments, the third unnatural base contains NaM, and the second unnatural base contains TAT1.

[0248] Also provided herein are proteins containing at least one unnatural amino acid, which are produced according to any of the methods disclosed herein. In some embodiments, the protein contains at least one unnatural amino acid. In some embodiments, the protein contains one unnatural amino acid. In some embodiments, the protein contains two or more unnatural amino acids. In some embodiments, the protein contains two unnatural amino acids. In some embodiments, the protein contains three or more unnatural amino acids.

[0249] In some embodiments, the nucleotide analogs are also modified in the phosphate moiety. Modified phosphate moieties include, but are not limited to, those having a modification in the linkage between two nucleotides, such as phosphorothioate, chiral phosphorothioate, phosphorodithioate, phosphotriester, aminoalkyl phosphotriester, 3'-alkylene phosphonate and methyl and other alkyl phosphonates including chiral phosphonate, phosphinate, phosphoramidate including 3'-aminophosphoramidate and aminoalkyl phosphoramidate, thionophosphoramidate, thionoalkyl phosphonate, thionoalkyl phosphotriester, and boranophosphate. The linkage of these phosphates or modified phosphates between two nucleotides is by a 3'-5' linkage or a 2'-5' linkage, and the linkage may be in the reverse direction such as from 3'-5' to 5'-3' or from 2'-5' to 5'-2'. It is understood to include various salts, mixed salts and free acid forms. A number of U.S. patents teach methods of making and using nucleotides containing modified phosphates, including, but not limited to, U.S. Patent Nos. 3,687,808; 4,469,863; 4,476,301; 5,023,243; 5,177,196; 5,188,897; 5,264,423; 5,276,019; 5,278,302; 5,286,717; 5,321,131; 5,399,676; 5,405,939; 5,453,496; 5,455,233; 5,466,677; 5,476,925; 5,519,126; 5,536,821; 5,541,306; 5,550,111; 5,563,253; 5,571,799; 5,587,361; and 5,625,050, the disclosures of each of which are hereby incorporated by reference in their entirety.

[0250] In some embodiments, the unnatural nucleic acids include 2′,3′-dideoxy-2′,3′-didehydro-nucleosides (PCT / US2002 / 006460), 5′-substituted DNA and RNA derivatives (PCT / US2011 / 033961; Saha et al., J. Org Chem., 1995, 60, pp. 788-789; Wang et al., Bioorganic & Medicinal Chemistry Letters, 1999, 9, pp. 885-890; and Mikhailov et al., Nucleosides & Nucleotides, 1991, 10(1-3), pp. 339-343; Leonid et al., 1995, 14(3-5), pp. 901-905; and Eppacher et al., Helvetica Chimica Acta, 2004, 87, pp. 3004-3020; PCT / JP2000 / 004720; PCT / JP2003 / 002342; PCT / JP2004 / 013216; PCT / JP2005 / 020435; PCT / JP2006 / 315479; PCT / JP2006 / 324484; PCT / JP2009 / 056718; PCT / JP2010 / 067560), or 5′-substituted monomers made as monophosphates using modified bases (Wang et al., Nucleosides Nucleotides & Nucleic Acids, 2004, 23(1 and 2), pp. 317-337), the disclosures of each of which are hereby incorporated by reference in their entirety.

[0251] In some embodiments, the unnatural nucleic acids include modifications at the 5' and 2' positions of the sugar ring (PCT / US94 / 02993), e.g., 5'-CH2 substituted 2'-O-protected nucleosides (Wu et al., Helvetica Chimica Acta, 2000, Vol. 83, pp. 1127-1143 and Wu et al., Bioconjugate Chem. 1999, Vol. 10, pp. 921-924). In some cases, the unnatural nucleic acids include amide-linked nucleoside dimers prepared for incorporation into oligonucleotides, where the 3'-linked nucleoside (5' to 3') in the dimer includes 2'-OCH3 and 5'-(S)-CH3 (Mesmaeker et al., Synlett, 1997, pp. 1287-1290). The unnatural nucleic acids can include 2'-substituted 5'-CH2(or O) modified nucleosides (PCT / US92 / 01020). The unnatural nucleic acids can include 5'-methylene phosphonate DNA and RNA monomers and dimers (Bohringer et al., Tet. Lett., 1993, Vol. 34, pp. 2723-2726; Collingwood et al., Synlett, 1995, Vol. 7, pp. 703-705; and Hutter et al., Helvetica Chimica Acta, 2002, Vol. 85, pp. 2777-2806). The unnatural nucleic acids can include 5'-phosphonate monomers with 2'-substitutions (U.S. Patent Application Publication No. 2006 / 0074035) and other modified 5'-phosphonate monomers (International Publication No. 1997 / 35869). The unnatural nucleic acids can include 5'-modified methylene phosphonate monomers (European Patent Application Publication No. 614907 and European Patent Application Publication No. 629633). The unnatural nucleic acids can include 5 Analogs of 5' or 6'-phosphonate ribonucleosides containing a hydroxyl group at the 'and / or 6' position (Chen et al., Phosphorus, Sulfur and Silicon, 2002, Vol. 777, pp. 1783 - 1786; Jung et al., Bioorg. Med. Chem., 2000, Vol. 8, pp. 2501 - 2509; Gallier et al., Eur. J. Org. Chem., 2007, pp. 925 - 933; and Hampton et al., J. Med. Chem., 1976, Vol. 19(8), pp. 1029 - 1033) can be included. The unnatural nucleic acids can include 5'-phosphonate deoxyribonucleoside monomers and dimers having a 5'-phosphate group (Nawrot et al., Oligonucleotides, 2006, Vol. 16(1), pp. 68 - 82). The unnatural nucleic acids can include nucleosides having a 6'-phosphonate group, where the 5' or / and 6' position is unsubstituted, or a thio-tert-butyl group (SC(CH3)3) (and its analogs); a methyleneamino group (CH2NH2) (and its analogs) or a cyano group (CN) (and its analogs) (Fairhurst et al., Synlett, 2001, Vol. 4, pp. 467 - 472; Kappler et al., J. Med. Chem., 1986, Vol. 29, pp. 1030 - 1038; Kappler et al., J. Med. Chem., 1982, Vol. 25, pp. 1179 - 1184; Vrudhula et al., J. Med. Chem., 1987, Vol. 30, pp. 888 - 894; Hampton et al., J. Med. Chem., 1976, Vol. 19, pp. 1371 - 1377; Geze et al., J. Am. Chem. Soc, 1983, Vol. 105(26), pp. 7638 - 7640; and Hampton et al., J. Am. Chem. Soc, 1973, Vol. 95(13), pp. 4404 - 4414). The disclosure of each of the references listed in this paragraph is hereby incorporated by reference in its entirety into this specification.

[0252] In some embodiments, the unnatural nucleic acid also includes modifications of the sugar moiety. In some cases, the nucleic acid contains one or more nucleosides with modified sugar groups. Such sugar-modified nucleosides may confer enhanced nuclease stability, increased binding affinity, or some other advantageous biological property. In certain embodiments, the nucleic acid includes a chemically modified ribofuranose ring moiety. Examples of chemically modified ribofuranose rings include, but are not limited to, addition of substituents (5' and / or 2' substituents; cross-linking of two ring atoms to form bicyclic nucleic acids (BNAs); replacement of the oxygen atom of the ribosyl ring with S, N(R) or C(R1)(R2) (R = H, C1-C 12 alkyl or protecting group); and combinations thereof). Examples of chemically modified sugars can be found in International Publication No. WO 2008 / 101157, U.S. Patent Application Publication No. 2005 / 0130923, and International Publication No. WO 2007 / 134181, the disclosures of each of which are incorporated herein by reference in their entirety.

[0253] In some examples, the modified nucleic acid includes a modified sugar or sugar analog. Thus, in addition to ribose and deoxyribose, the sugar moiety can be a pentose, deoxypentose, hexose, deoxyhexose, glucose, arabinose, xylose, lyxose, or a cyclopentyl group of a sugar "analog". The sugar can be in the form of pyranosyl or furanosyl. The sugar moiety can be a furanoside of ribose, deoxyribose, arabinose or 2'-O-alkyl ribose, and the sugar can be attached to each of the heterocyclic bases in either the [alpha] or [beta] anomeric configuration. Sugar modifications include, but are not limited to, 2'-alkoxy-RNA analogs, 2'-amino-RNA analogs, 2'-fluoro-DNA and 2'-alkoxy- or amino-RNA / DNA chimeras. For example, sugar modifications can include 2'-O-methyl-uridine or 2'-O-methyl-cytidine. Sugar modifications include 2'-O-alkyl substituted deoxyribonucleosides and 2'-O-ethylene glycol-like ribonucleosides. The production of these sugars or sugar analogs and their respective "nucleosides" when attached to heterocyclic bases (nucleobases) is known. Sugar modifications can also be made and combined with other modifications. The production of these and their respective "nucleosides" is known. Sugar modifications can also be made and combined with other modifications.

[0254] Modifications to the sugar moiety include natural and non-natural modifications of ribose and deoxyribose. Sugar modifications include, but are not limited to, the following modifications at the 2-position: OH; F; O-, S- or N-alkyl; O-, S- or N-alkenyl; O-, S- or N-alkynyl; or O-alkyl-O-alkyl, where alkyl, alkenyl and alkynyl are substituted or unsubstituted C1-C 10 alkyl or C2-C10 alkenyl and alkynyl may also be used. 2'-sugar modifications include, but are not limited to, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2)n NH2, -O(CH2) n CH3, -O(CH2) n ONH2 and -O(CH2) n ON[(CH2) n CH3)]2 (wherein n and m are from 1 to about 10) may also be mentioned.

[0255] Other modifications at the 2'-position include, but are not limited to, C1-C 10Lower alkyl, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl, O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, interfering substance, a group for improving the pharmacokinetic properties of an oligonucleotide, or a group for improving the pharmacodynamic properties of an oligonucleotide, and other substituents having similar properties. Similar modifications can also be made at other positions of the sugar, particularly at the 3'-position of the sugar in the 3'-terminal nucleotide or in a 2'-5' linked oligonucleotide, and at the 5'-position of the 5'-terminal nucleotide. Modified sugars also include those containing modifications at the bridging ring oxygen such as CH2 and S. Nucleotide sugar analogs can also have a sugar mimetic such as a cyclobutyl moiety instead of a pentofuranosyl sugar.U.S. Patent Nos. 4,981,957; 5,118,800; 5,319,080; 5,359,044; 5,393,878; 5,446,137; 5,466,786; 5,514,785; 5,519,134; 5,567,811; 5,576,427; 5,591,722; 5,597,909; 5,610,300; 5,627,053; 5,639,873; 5,646,265; 5,658,873; 5,670,633; 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121; 5,596,091; 5,614,617; 5,681,941; and 5,700,920, etc., teach the manufacture of such modified sugar structures and numerous U.S. patents detail and describe the scope of base modifications, the entire disclosure of each of which is incorporated herein by reference in its entirety.

[0256] Examples of nucleic acids having modified sugar moieties include, but are not limited to, nucleic acids containing 5'-vinyl, 5'-methyl (R or S), 4'-S, 2'-F, 2'-OCH3 and 2'-O(CH2)2OCH3 substituents. Substituents at the 2'-position are allyl, amino, azido, thio, O-allyl, O-(C1-C 1O alkyl), OCF3, O(CH2)2SCH3, O(CH2)2-O-N(R m )(R n ) and O-CH2-C(=O)-N(R m )(R n )(wherein each R m and R n is independently H, or substituted or unsubstituted C1-C 10 alkyl).

[0257] In certain embodiments, the nucleic acids described herein include one or more bicyclic nucleic acids. In certain such embodiments, the bicyclic nucleic acid includes a bridge between the 4' and 2' ribosyl ring atoms. In certain embodiments, the nucleic acids provided herein include one or more bicyclic nucleic acids, where the bridge includes a 4'-2' bicyclic nucleic acid. Examples of such 4'-2' bicyclic nucleic acids include, but are not limited to, the formulas: 4'-(CH2)-O-2' (LNA); 4'-(CH2)-S-2'; 4'-(CH2)2-O-2' (ENA); 4'-CH(CH3)-O-2' and 4'-CH(CH2OCH3)-O-2', and analogs thereof (see U.S. Patent No. 7,399,845); 4'-C(CH3)(CH3)-O-2' and analogs thereof (see International Publication No. 2009 / 006478, International Publication No. 2008 / 150729, U.S. Patent Application Publication No. 2004 / 0171570, U.S. Patent No. 7,427,672, Chattopadhyaya et al., J. Org. Chem., Vol. 209, No. 74, pp. 118-134, and International Publication No. 2008 / 154401).For example, Singh et al., Chem. Commun., 1998, Vol. 4, pp. 455-456; Koshkin et al., Tetrahedron, 1998, Vol. 54, pp. 3607-3630; Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 2000, Vol. 97, pp. 5633-5638; Kumar et al., Bioorg. Med. Chem. Lett., 1998, Vol. 8, pp. 2219-2222; Singh et al., J. Org. Chem., 1998, Vol. 63, pp. 10035-10039; Srivastava et al., J. Am. Chem. Soc., 2007, Vol. 129 (No. 26), pp. 8362-8379; Elayadi et al., Curr. Opinion Invens. Drugs, 2001, Vol. 2, pp. 558-561; Braasch et al., Chem. Biol, 2001, Vol. 8, pp. 1-7; Oram et al., Curr. Opinion Mol. Ther., 2001, Vol. 3, pp. 239-243; U.S. Patent Nos. 4,849,513; 5,015,733; 5,118,800; 5,118,802; 7,053,207; 6,268,490; 6,770,748; 6,794,499; 7,034,133; 6,525,191; 6,670,461; and 7,399,845; International Publication Nos. WO 2004 / 106356, WO 1994 / 14226, WO 2005 / 021570, WO 2007 / 090071 and WO 2007 / 134181; U.S. Patent Application Publication Nos. 2004 / 0171570, 2007 / 0287831 and 2008 / 0039618; U.S. Provisional Application Nos. 60 / 989,574, 61 / 026,995, 61 / 026,998, 61 / 056,564, 61 / 086,231, 61 / 097,787 and 61 / 099,844; and International Application Nos. PCT / US2008 / 064591, PCT US2008 / 066154, PCT US2008 / 068922 and PCT / DK98 / 00393 are also referred to, the disclosures of each of which are hereby incorporated herein by reference in their entirety.

[0258] In certain embodiments, the nucleic acid comprises linked nucleic acids. Nucleic acids can be linked together using any inter-nucleic acid linkage. Two main classes of internucleic acid linkers are defined by the presence or absence of a phosphorus atom. Representative phosphorus-containing internucleic acid linkages include, but are not limited to, phosphodiester, phosphorotriester, methylphosphonate, phosphoramidate, and phosphorothioate (P=S). Representative phosphorus-free internucleic acid linkers include, but are not limited to, methylene methylimino (-CH2-N(CH3)-O-CH2-), thiodiester (-O-C(O)-S-), thiocarbamate (-O-C(O)(NH)-S-); siloxane (-O-Si(H)2-O-); and N,N * -dimethylhydrazine (-CH2-N(CH3)-N(CH3)). In certain embodiments, internucleic acid linkages having chiral atoms can be produced as a racemic mixture or as separate enantiomers, such as alkylphosphonates and phosphorothioates. Unnatural nucleic acids can contain a single modification or can contain multiple modifications within one moiety or between different moieties.

[0259] Modifications of the phosphate of the backbone of the nucleic acid include, but are not limited to, methylphosphonate, phosphorothioate, phosphoramidate (bridged or unbridged), phosphorotriester, phosphorodithioate, phosphodithioate, and boranophosphate and can be used in any combination. Other non-phosphate linkages can also be used.

[0260] In some embodiments, backbone modifications (e.g., internucleotide linkages of methylphosphonate, phosphorothioate, phosphoramidate, and phosphorodithioate) can confer immunomodulatory activity in the modified nucleic acid and / or enhance their stability in vivo.

[0261] In some instances, the phosphorus derivative (or modified phosphate group) is attached to a sugar or sugar analog moiety and can be a monophosphate, diphosphate, triphosphate, alkyl phosphonate, phosphorothioate, phosphorodithioate, phosphoramidate, etc. Exemplary polynucleotides containing modified phosphate linkages or non-phosphate linkages can be found in Peyrottes et al., 1996, Nucleic Acids Res. 24:1841-1848; Chaturvedi et al., 1996, Nucleic Acids Res. 24:2318-2323; and Schultz et al., (1996) Nucleic Acids Res. 24:2966-2973; Matteucci, 1997, "Oligonucleotide Analogs: an Overview" in Oligonucleotides as Therapeutic Agents, (Chadwick and Cardew eds) John Wiley and Sons, New York, NY; Zon, 1993, "Oligonucleoside Phosphorothioates" in Protocols for Oligonucleotides and Analogs, Synthesis and Properties, Humana Press, 165-190; Miller et al., 1971, JACS 93:6657-6665; Jager et al., 1988, Biochem. 27:7247-7246; Nelson et al., 1997, JOC 62:7278-7287; U.S. Patent No. 5,453,496; and Micklefield, 2001, Curr. Med. Chem. 8:1157-1179, the disclosures of each of which are hereby incorporated by reference in their entirety.

[0262] In some cases, modification of the backbone involves replacing the phosphodiester linkage with an alternative moiety such as an anionic, neutral or cationic group. Examples of such modifications include anionic internucleoside linkages; N3’-P5’ phosphoramidate modifications; boranophosphate DNA; prooligonucleotides; neutral internucleoside linkages, such as methylphosphonates; amide-linked DNA; methylene(methylimino) linkages; formacetal and thioformacetal linkages; backbones containing sulfonyl groups; morpholino oligos; peptide nucleic acids (PNA); and positively charged deoxyribonucleic acid guanidine (DNG) oligos (Micklefield, 2001, Current Medicinal Chemistry 8:1157-1179), the disclosure of which is hereby incorporated by reference in its entirety. Modified nucleic acids can include chimeric or mixed backbones that include combinations of phosphate linkages, such as combinations of phosphodiester and phosphorothioate linkages, and one or more modifications.

[0263] Examples of alternatives to phosphate include, for example, short chain alkyl or cycloalkyl Examples include internucleoside linkages of nucleosides, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or internucleoside linkages of one or more short chain heteroatoms or heterocycles. These include those having a morpholino linkage (in part formed from the sugar moiety of the nucleoside); a siloxane backbone; sulfide, sulfoxide and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; alkene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S and CH2 components. Numerous U.S. patents disclose methods of making and using these types of phosphate replacements, including, but not limited to, U.S. Patent Nos. 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; 5,434,257; 5,466,677; 5,470,967; 5,489,677; 5,541,307; 5,561,225; 5,596,086; 5,602,240; 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439, the disclosures of each of which are incorporated herein by reference in their entirety. It is also understood that in nucleotide analogs, both the sugar and phosphate moieties of the nucleotide can be replaced, for example, by an amide-type linkage (aminoethylglycine) (PNA). U.S. Patent Nos. 5,539,082; 5,714,331; and 5,719,262 teach methods of making and using PNA molecules, each of which is incorporated herein by reference.See also Nielsen et al., Science, 1991, Vol. 254, pp. 1497-1500. For example, it is also possible to link (conjugate) other types of molecules to nucleotides or nucleotide analogs in order to enhance cellular uptake. The conjugate can be chemically linked to the nucleotide or nucleotide analog.Such conjugates include, but are not limited to, lipid moieties such as cholesterol moieties (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, 86, pp. 6553 - 6556), cholic acid (Manoharan et al., Bioorg. Med. Chem. Let., 1994, 4, pp. 1053 - 1060), thioethers such as hexyl - S - tritylthiol (Manoharan et al., Ann. KY. Acad. Sci., 1992, 660, pp. 306 - 309; Manoharan et al., Bioorg. Med. Chem. Let., 1993, 3, pp. 2765 - 2770), thiocolesterol (Oberhauser et al., Nucl. Acids Res., 1992, 20, pp. 533 - 538), aliphatic chains such as dodecanediol or undecyl residues (Saison - Behmoaras et al., EMBO J, 1991, 10, pp. 1111 - 1118; Kabanov et al., FEBS Lett., 1990, 259, pp. 327 - 330; Svinarchuk et al., Biochimie, 1993, 75, pp. 49 - 54), phospholipids such as di - hexadecyl - rac - glycerol or triethylammonium l - di - O - hexadecyl - rac - glycerol - S - H - phosphonate (Manoharan et al., Tetrahedron Lett., 1995, 36, pp. 3651 - 3654; Shea et al., Nucl. Acids Res., 1990, 18, pp. 3777 - 3783), polyamine or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides, 1995, 14, pp. 969 - 973), or adamantaneacetic acid (Manoharan et al., Tetrahedron Lett., 1995, 36, pp. 3651 - 3654), palmitoyl moieties (Mishra et al., Biochem. Biophys. Acta, 1995, 1264, pp. 229 - 237), or octadecylamine. Examples include cholesteryl-(6-oxohexanoyl)-oxycarbonyl-(6-oxohexanoyloxy) cholesteryl-(6-oxohexanoyl)-oxycarbonyl-cholesterol (Crooke et al., J. Pharmacol. Exp. Ther., 1996, 277, 923-937), the disclosures of each of which are incorporated herein by reference in their entirety. A number of U.S. patents teach the manufacture of such conjugates, including but not limited to, U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241; 5,391,723; 5,416,203; 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941, the disclosures of each of which are incorporated herein by reference in their entirety.

[0264] Nucleobases for use in compositions and methods for replication, transcription, translation, and incorporation of unnatural amino acids into proteins are described herein. In some embodiments, the nucleobases described herein have the structure: [Chemical formula] [wherein: each X is independently carbon or nitrogen; R2 is present when X is carbon and is independently hydrogen, alkyl, alkenyl, alkynyl, methoxy, methanethiol, methaneseleno, halogen, cyano or azide group; each Y is independently sulfur, oxygen, selenium or secondary amine; each E is independently oxygen, sulfur or selenium; The wavy line indicates the point of attachment to a ribosyl, deoxyribosyl or dideoxyribosyl moiety, or an analog thereof, and the ribosyl, deoxyribosyl or dideoxyribosyl moiety, or an analog thereof, is in free form or is optionally connected to a monophosphate, diphosphate or triphosphate group containing an α-thiotriphosphate, β-thiotriphosphate or γ-thiotriphosphate group, or is included in RNA or DNA, or in an RNA analog or DNA analog] comprises.

[0265] In some embodiments, R2 is lower alkyl (e.g., C1-C6), hydrogen or halogen. In some embodiments of the nucleobases described herein, R2 is fluoro. In some embodiments of the nucleobases described herein, X is carbon. In some embodiments of the nucleobases described herein, E is sulfur. In some embodiments of the nucleobases described herein, Y is sulfur.

[0266] In some embodiments of the nucleobases described herein, the nucleobase has the structure: [Chemical formula] has. In some embodiments, the nucleobases described herein have the structure: [Chemical formula] has. In some embodiments of the nucleobases described herein, E is sulfur and Y is sulfur. In some embodiments, the wavy line indicates a point of attachment to a ribosyl, deoxyribosyl or dideoxyribosyl moiety, or an analog thereof, and the ribosyl, deoxyribosyl or dideoxyribosyl moiety, or an analog thereof, is in free form or is connected to a monophosphate, diphosphate, triphosphate, α-thiotriphosphate, β-thiotriphosphate or γ-thiotriphosphate group, or is included in RNA or DNA, or in an RNA analog or DNA analog. In some embodiments of the nucleobases described herein, the wavy line indicates a point of attachment to a ribosyl or deoxyribosyl moiety. In some embodiments of the nucleobases described herein, the wavy line indicates a point of attachment to a ribosyl or deoxyribosyl moiety connected to a triphosphate group. In some embodiments of the nucleobases described herein, they are a component of a nucleic acid polymer. In some embodiments of the nucleobases described herein, the nucleobase is a component of tRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of the anticodon in tRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of mRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of the codon in mRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of RNA or DNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of the codon in DNA. In some embodiments of the nucleobases described herein, the nucleobase forms, or is capable of forming, or is configured to form, a nucleobase pair with another (e.g., complementary) nucleobase.

[0267] In some embodiments, the nucleobases described herein have the structure:

Chemical formula

[0268] In some embodiments, each X is carbon. In some embodiments, at least one X is carbon. In some embodiments, one X is carbon. In some embodiments, at least two Xs are carbon. In some embodiments, two Xs are carbon. In some embodiments, at least one X is nitrogen. In some embodiments, one X is nitrogen. In some embodiments, at least two Xs are nitrogen. In some embodiments, two Xs are nitrogen.

[0269] In some embodiments, Y is sulfur. In some embodiments, Y is oxygen. In some embodiments, Y is selenium. In some embodiments, Y is a secondary amine.

[0270] In some embodiments, E is sulfur. In some embodiments, E is oxygen. In some embodiments, E is selenium.

[0271] In some embodiments, R2 is present when X is carbon. In some embodiments, R2 is not present when X is nitrogen. In some embodiments, each R2, when present, is hydrogen. In some embodiments, R2 is an alkyl such as methyl, ethyl or propyl. In some embodiments, R2 is an alkenyl such as -CH=CH2. In some embodiments, R2 is an alkynyl such as ethynyl. In some embodiments, R2 is methoxy. In some embodiments, R2 is methanethiol. In some embodiments, R2 is methaneseleno. In some embodiments, R2 is a halogen such as chloro, bromo or fluoro. In some embodiments, R2 is cyano. In some embodiments, R2 is azide.

[0272] In some embodiments, E is sulfur, Y is sulfur, and each X is independently carbon or nitrogen. In some embodiments, E is sulfur, Y is sulfur, and each X is carbon.

[0273] In some embodiments, the nucleobase has the structure

Chemical formula

Chemical formula

[0274] In some embodiments, the nucleobase has a structure

Chemical formula

[0275] In some embodiments, the nucleobases disclosed herein are capable of binding (e.g., non-covalently) to complementary base-pairing nucleobases to form a nucleobase and a non-natural base pair (UBP), or are capable of base-pairing with or are configured to base-pair with a nucleobase. In some embodiments, the complementary base-pairing nucleobases are

Chemical formula

[0276] In one aspect, provided herein is a double-stranded double-stranded oligonucleotide, wherein the first oligonucleotide strand comprises a nucleobase disclosed herein, and the second complementary oligonucleotide strand comprises a complementary base-pairing nucleobase at its complementary base-pairing site. In some embodiments, the first oligonucleotide strand

Chemical formula

Chemical formula

[0277] In a further aspect, provided herein is a transfer RNA (tRNA) comprising a nucleobase as described herein, the tRNA comprising an anticodon comprising the nucleobase; and a recognition element that facilitates selective charging of the tRNA with a non-natural amino acid by an aminoacyl-tRNA synthetase. In some embodiments, the nucleobase is located in the anticodon region of the tRNA. In some embodiments, the nucleobase is located at the first position of the anticodon. In some embodiments, the nucleobase is located at the second position of the anticodon. In some embodiments, the nucleobase is located at the third position of the anticodon. In some embodiments, the aminoacyl-tRNA synthetase is derived from Methanosarcina or a variant thereof. In some embodiments, the aminoacyl-tRNA synthetase is derived from Methanococcus (Methanocaldococcus) or a variant thereof. In some embodiments, the non-natural amino acid comprises an aromatic moiety. In some embodiments, the non-natural amino acid is a lysine derivative. In some embodiments, the non-natural amino acid is a phenylalanine derivative.

[0278] Formula: N1-Zx-N2 [wherein, Z is a nucleobase as described herein, and is attached to ribosyl or deoxyribosyl, or an analog thereof; N1 is one or more nucleotides or an analog thereof, or a terminal phosphate group, attached at the 5'-end of the ribosyl or deoxyribosyl of Z, or an analog thereof; N2 is one or more nucleotides or an analog thereof, or a terminal hydroxyl group, attached at the 3'-end of the ribosyl or deoxyribosyl of Z, or an analog thereof; x is an integer from 1 to 20] Also provided herein is a structure comprising the same.

[0279] In some embodiments, N1 is one or more nucleotides or analogs thereof attached at the 5'-end of the ribosyl or deoxyribosyl moiety of Z, or an analog thereof. The attachment to the 5'-end of the ribosyl or deoxyribosyl moiety is by a phosphodiester. In some embodiments, N1 is a terminal phosphate group attached at the 5'-end of the ribosyl or deoxyribosyl moiety of Z, or an analog thereof. In some embodiments, N2 is one or more nucleotides or analogs thereof attached at the 3'-end of the ribosyl or deoxyribosyl moiety of Z, or an analog thereof. The attachment to the 3'-end of the ribosyl or deoxyribosyl moiety is by a phosphodiester. In some embodiments, N2 is a terminal hydroxyl group attached at the 3'-end of the ribosyl or deoxyribosyl moiety of Z, or an analog thereof.

[0280] In some embodiments, x is an integer from 1 to 20. In some embodiments, x is an integer from 1 to 15. In some embodiments, x is an integer from 1 to 10. In some embodiments, x is an integer from 1 to 5. In some embodiments, x is 1. In some embodiments, x is 2. In some embodiments, x is 3. In some embodiments, x is 4. In some embodiments, x is 5. In some embodiments, x is 6. In some embodiments, x is 7. In some embodiments, x is 8. In some embodiments, x is 9. In some embodiments, x is 10. In some embodiments, x is 11. In some embodiments, x is 12. In some embodiments, x is 13. In some embodiments, x is 14. In some embodiments, x is 15. In some embodiments, x is 16. In some embodiments, x is 17. In some embodiments, x is 18. In some embodiments, x is 19. In some embodiments, x is 20.

[0281] In some embodiments, Z has the structure detailed herein

Chemical formula

Chemical formula

[0282] In some embodiments, the structure of formula N1-Zx-N2 encodes a gene. In some embodiments, Zx is located in the translation region of the gene. In some embodiments, Zx is located in the untranslated region of the gene. In some embodiments, the structure further comprises a 5' or 3' untranslated region (UTR). In some embodiments, the structure further comprises a terminator region. In some embodiments, the structure further comprises a promoter region.

[0283] In a further aspect, provided herein is a polynucleotide library, the library comprising at least 5000 unique polynucleotides, each polynucleotide comprising at least one nucleobase disclosed herein. In some embodiments, the polynucleotide library encodes at least one gene.

[0284] In yet another aspect, provided herein is a nucleoside triphosphate, wherein the nucleobase is

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

[0285] Base pairing properties of nucleic acids; exemplary base pairs In some embodiments, non-natural nucleotides form base pairs (unnatural base pairs; UBPs) with another non-natural nucleotide during or after incorporation into DNA or RNA. In some embodiments, a stably incorporated non-natural nucleic acid is a non-natural nucleic acid that can form base pairs with another nucleic acid, e.g., a natural or non-natural nucleic acid. In some embodiments, a stably incorporated non-natural nucleic acid is a non-natural nucleic acid that can form base pairs (unnatural nucleic acid base pairs (UBPs)) with another non-natural nucleic acid. For example, a first non-natural nucleic acid can form base pairs with a second non-natural nucleic acid. For example, one pair of non-natural nucleoside triphosphates that can form base pairs during or after incorporation into a nucleic acid includes the triphosphate of (d)5SICS ((d)5SICSTP) and the triphosphate of (d)NaM ((d)NaMTP). Other examples include, but are not limited to, the triphosphate of (d)CNMO ((d)CNMOTP) and the triphosphate of (d)TPT3 ((d)TPT3TP). Such non-natural nucleotides can have a ribose or deoxyribose sugar moiety (indicated by "(d)"). For example, one pair of non-natural nucleoside triphosphates that can form base pairs when incorporated into a nucleic acid includes the triphosphate of TAT1 ((d)TAT1TP) and the triphosphate of NaM ((d)NaMTP). In some embodiments, one pair of non-natural nucleoside triphosphates that can form base pairs when incorporated into a nucleic acid includes the triphosphate of dCNMO (dCNMOTP) and the triphosphate of TAT1 (TAT1TP). In some embodiments, one pair of non-natural nucleoside triphosphates that can form base pairs when incorporated into a nucleic acid includes the triphosphate of dTPT3 (dTPT3TP) and the triphosphate of NaM (NaMTP). In some embodiments, non-natural nucleic acids do not substantially form base pairs with natural nucleic acids (A, T, G, C). In some embodiments, a stably incorporated non-natural nucleic acid can form base pairs with a natural nucleic acid.

[0286] In some embodiments, the stably incorporated unnatural (deoxy)ribonucleotides are unnatural (deoxy)ribonucleotides that can form UBP, but do not substantially form base pairs with any of the respective natural (deoxy)ribonucleotides. In some embodiments, the stably incorporated unnatural (deoxy)ribonucleotides are unnatural (deoxy)ribonucleotides that can form UBP, but are actually Substantially, it does not form base pairs with one or more natural nucleic acids. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with A, T, and C, but can form base pairs with G. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with A, T, and G, but can form base pairs with C. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with C, G, and A, but can form base pairs with T. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with C, G, and T, but can form base pairs with A. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with A and T, but can form base pairs with C and G. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with A and C, but can form base pairs with T and G. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with A and G, but can form base pairs with C and T. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with C and T, but can form base pairs with A and G. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with C and G, but can form base pairs with T and G. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with T and G, but can form base pairs with A and G. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with G, but can form base pairs with A, T, and C. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with A, but can form base pairs with G, T, and C. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with T, but can form base pairs with G, A, and C. For example, a stably incorporated unnatural nucleic acid cannot substantially form base pairs with C, but can form base pairs with G, T, and A.

[0287] Exemplarily, non-natural nucleotides that can form non-natural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, 5SICS, d5SICS, NaM, dNaM, dTPT3, dMTMO, dCNMO, TAT1, and combinations thereof. In some embodiments, non-natural nucleotides that can form non-natural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, 5SICS, NaM, TPT3, MTMO, CNMO, TAT1, and combinations thereof, wherein the sugar moiety of the nucleotide is deoxyribose sugar. In some embodiments, non-natural nucleotides that can form non-natural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, 5SICS, NaM, TPT3, MTMO, CNMO, TAT1, and combinations thereof, wherein the sugar moiety of the nucleotide is ribose sugar. In some embodiments, non-natural nucleotides that can form non-natural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, (d)5SICS, (d)NaM, (d)TPT3, (d)MTMO, (d)CNMO, (d)TAT1, and combinations thereof. In some embodiments, non-natural nucleotide base pairs include, but are not limited to, [Chemical formula] are included, wherein the sugar moiety is any embodiment or variation described herein. In some embodiments, non-natural nucleotide base pairs include, but are not limited to, [Chemical formula] are included. In any such embodiment, one or both of the deoxyriboses attached to the non-natural base can be replaced with ribose.

[0288] engineered organism In some embodiments, the methods and plasmids disclosed herein are further used to generate engineered organisms, e.g., organisms that incorporate and replicate unnatural nucleotides or unnatural base pairs (UBPs), and to transcribe mRNAs and tRNAs that are used to translate proteins containing unnatural amino acid residues, which can also use nucleic acids containing unnatural nucleotides. In some examples, the organism is a non-human semi-synthetic organism (SSO). In some examples, the organism is a semi-synthetic organism (SSO). In some examples, the SSO is a cell. In some examples, the in vivo methods involve semi-synthetic organisms (SSOs). In some examples, the semi-synthetic organisms include microorganisms. In some examples, the organism includes bacteria. In some examples, the organism includes gram-negative bacteria. In some examples, the organism includes gram-positive bacteria. In some examples, the organism includes Escherichia coli. Such modified organisms variously include additional components such as DNA repair mechanisms, modified polymerases, nucleotide transporters, or other components. In some examples, the SSO includes E. coli strain YZ3. In some examples, the SSO includes E. coli strains ML1 or ML2, e.g., strains described in FIGS. 1(B-D) of Ledbetter et al., J. Am. Chem. Soc. 2018, 140(2), 758, the disclosure of which is incorporated herein by reference in its entirety. In some examples, the cells used are genetically transformed with expression cassettes encoding heterologous proteins, e.g., nucleoside triphosphate transporters that can transport unnatural nucleoside triphosphates into the cell, and optionally the CRISPR / Cas9 system for eliminating DNA lacking unnatural nucleotides (e.g., E. coli strains YZ3, ML1 or ML2). In some examples, the cells further include enhanced activity for the uptake of unnatural nucleic acids. In some cases, the cells further include enhanced activity for the import of unnatural nucleic acids.

[0289]

[0290] ​ In some embodiments, Cas9 and the appropriate guide RNA(sgRNA) are encoded on separate plasmids. In some examples, CAS9 and sgRNA are encoded on the same plasmid. In some cases, the nucleic acid molecule containing Cas9, the sgRNA-encoding nucleic acid molecule, or non-natural nucleotides is located on one or more plasmids. In some examples, Cas9 is encoded on a first plasmid, and the sgRNA-encoding nucleic acid molecule and the nucleic acid molecule containing non-natural nucleotides are encoded on a second plasmid. In some examples, the nucleic acid molecule containing Cas9, sgRNA, and non-natural nucleotides is encoded on the same plasmid. In some examples, the nucleic acid molecule contains two or more non-natural nucleotides. In some examples, Cas9 is integrated into the genome of the host organism, and the sgRNA is encoded on a plasmid or in the genome of the organism.

[0291] In some examples, the first plasmid encoding Cas9 and sgRNA, and the second plasmid encoding the nucleic acid molecule containing non-natural nucleotides are introduced into the engineered microorganism. In some examples, the first plasmid encoding Cas9, and the second plasmid encoding the nucleic acid molecule containing sgRNA and non-natural nucleotides are introduced into the engineered microorganism. In some examples, the plasmid encoding the nucleic acid molecule containing Cas9, sgRNA, and non-natural nucleotides is introduced into the engineered microorganism. In some examples, the nucleic acid molecule contains two or more non-natural nucleotides.

[0292] In some embodiments, viable cells are created that have at least one unnatural nucleotide and / or at least one unnatural base pair (UBP) incorporated within their DNA (plasmid or genome). In some examples, the unnatural base pairs include pairs of unnatural base-pairing nucleotides that can form unnatural base pairs under in vivo conditions when their respective triphosphates, the unnatural base-pairing nucleotides, are taken up into the cell by the action of a nucleotide triphosphate transporter. In some examples, the unnatural base pairs include pairs of unnatural base-pairing nucleotides configured to form unnatural base pairs under in vivo conditions when their respective triphosphates, the unnatural base-pairing nucleotides, are taken up into the cell by the action of a nucleotide triphosphate transporter. The cells can be genetically transformed by an expression cassette encoding a nucleotide triphosphate transporter such that the nucleotide triphosphate transporter is expressed and transport of the unnatural nucleotides into the cells is available. The cells are prokaryotic or eukaryotic cells, and the pairs of unnatural base-pairing nucleotides can be the triphosphates of dTPT3 (dTP3TP) and dNaM (dNaMTP) or the triphosphates of dCNMO (dCNMOTP) as their respective triphosphates.

[0293] In some embodiments, the cells are cells that are genetically transformed with a nucleic acid, such as an expression cassette encoding a nucleotide triphosphate transporter that can transport such unnatural nucleotides into the cell. The cells can include a heterologous nucleoside triphosphate transporter, where the heterologous nucleoside triphosphate transporter can transport natural and unnatural nucleoside triphosphates into the cell.

[0294] In some cases, the methods described herein also include contacting the genetically transformed cells with the respective triphosphates in the presence of potassium phosphate and / or an inhibitor of phosphatase or nuclease. During or after such contact, the cells can be placed in a maintenance medium suitable for cell growth and replication. The cells can be maintained in the maintenance medium such that the respective triphosphate forms of the unnatural nucleotides are incorporated into the nucleic acids within the cells through at least one replication cycle of the cells. The pairs of unnatural base-pairing nucleotides can include, as the respective triphosphates, the triphosphate of dTPT3 (dTPT3TP) and the triphosphate of dCNMO or dNaM (dCNOM or dNaMTP), the cells are Escherichia coli, dTPT3TP and dNaMTP can be imported into Escherichia coli by the transporter PtNTT2, and an Escherichia coli polymerase such as Pol III or Pol II can use the unnatural triphosphates to replicate DNA containing UBP, thereby incorporating unnatural nucleotides and / or unnatural base pairs into the nucleic acids of the cells within the cell environment. Additionally, ribonucleotides such as NaMTP and TAT1TP, 5FMTP, and TPT3TP are imported into Escherichia coli by the transporter PtNTT2 in some examples.

[0295] Compositions and methods are described herein that include the use of three or more non-natural base-pairing nucleotides. Such base-pairing nucleotides enter cells in some cases by the use of nucleotide transporters or by standard nucleic acid transformation methods known in the art (e.g., electroporation, chemical transformation or other methods). In some cases, the base-pairing non-natural nucleotides enter cells as part of a polynucleotide such as a plasmid. One or more base-pairing non-natural nucleotides that enter cells as part of a polynucleotide (RNA or DNA) need not themselves be replicated in vivo. For example, a double-stranded DNA plasmid or other nucleic acid containing a first unnatural deoxyribonucleotide and a second unnatural deoxyribonucleotide having bases configured to form a first non-natural base pair is electroporated into cells. The cell medium is treated with a third unnatural deoxyribonucleotide and a fourth unnatural deoxyribonucleotide having bases configured to form a second non-natural base pair with each other, where the base of the first unnatural deoxyribonucleotide and the base of the third unnatural deoxyribonucleotide form the second non-natural base pair, and the base of the second unnatural deoxyribonucleotide and the base of the fourth unnatural deoxyribonucleotide form the third non-natural base pair. In some examples, in vivo replication of the initially transformed double-stranded DNA plasmid subsequently results in replicated plasmids containing the third unnatural deoxyribonucleotide and the fourth unnatural deoxyribonucleotide. Alternatively, or in combination, variants of the ribonucleotides of the third unnatural deoxyribonucleotide and the fourth unnatural deoxyribonucleotide are added to the cell medium. These ribonucleotides are incorporated into RNA such as mRNA or tRNA in some examples. In some examples, the first, second, third and fourth deoxynucleotides contain different bases. In some examples, the first, third and fourth deoxynucleotides contain different bases. In some examples, the first and third deoxynucleotides contain the same base.

[0296] By practicing the method of the present invention, one of ordinary skill in the art can maintain within at least some of the individual cells a population of living and proliferating cells having at least one unnatural nucleotide and / or at least one unnatural base pair (UBP) within at least one nucleic acid that is maintained, wherein the at least one nucleic acid is stably propagated within the cell, and the cell expresses a nucleotide triphosphate transporter suitable for providing cellular uptake of the triphosphate form of one or more unnatural nucleotides when contacted (e.g., grown in the presence of) with the unnatural nucleotide in a maintenance medium suitable for the growth and replication of the organism.

[0297] After transport into the cell by the nucleotide triphosphate transporter, the unnatural base pair-forming nucleotides are incorporated into the nucleic acids within the cell by cellular machinery, such as the cell's own DNA and / or RNA polymerase, a heterologous polymerase, or a polymerase evolved using directed evolution (Chen T, Romesberg FE, FEBS Lett. Jan 21, 2014; 588(2):219-29; Betz K et al., J Am Chem Soc. Dec 11, 2013; 135(49):18637-43; the disclosures of each of which are hereby incorporated by reference in their entirety). The unnatural nucleotides can be incorporated into cellular nucleic acids such as genomic DNA, genomic RNA, mRNA, tRNA, structural RNA, microRNA, and self-replicating nucleic acids (e.g., plasmids, viruses or vectors).

[0298] In some cases, genetically engineered cells are produced by introduction of nucleic acids, such as heterologous nucleic acids, into cells. Any cell described herein can be a host cell and can contain an expression vector. In one embodiment, the host cell is a prokaryotic cell. In another embodiment, the host cell is Escherichia coli. In some embodiments, the cell contains one or more heterologous polynucleotides. Nucleic acid reagents can be introduced into microorganisms using a variety of techniques. Non-limiting examples of methods used to introduce heterologous nucleic acids into various organisms include transformation, transfection, transduction, electroporation, sonication-mediated transformation, conjugation, particle guns, and the like. In some examples, the addition of carrier molecules (see, e.g., bis-benzoimidazolyl compounds, e.g., U.S. Patent No. 5,595,899) can increase the uptake of DNA in cells that are typically considered difficult to transform by conventional methods. Conventional methods of transformation are readily available to those of skill in the art and can be found in Maniatis, T., E.F. Fritsch and J. Sambrook (1982) Molecular Cloning: a Laboratory Manual; Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y., the disclosure of which is incorporated herein by reference in its entirety.

[0299] In some instances, genetic transformation can be obtained, but is not limited to, by direct introduction of an expression cassette in a plasmid, viral vector, viral nucleic acid, phage nucleic acid, phage, cosmid, and artificial chromosome, or via introduction of genetic material in a cell or a carrier such as a cationic liposome. Such methods are available in the art and can be readily adapted for use in the methods described herein. An introduction vector is any nucleotide construct (e.g., plasmid) used to deliver a gene to a cell, or as part of a general strategy for delivering a gene, e.g., as part of a recombinant retrovirus or adenovirus (Ram et al., Cancer Res. 53:83-88, (1993)). Suitable means for transfection, including viral vectors, chemical transfectants, or physical and mechanical methods such as electroporation and direct diffusion of DNA, are described, for example, in Wolff, J.A. et al., Science, 247, 1465-1468, (1990); and, Wolff, J.A., Nature, 352, 815-818, (1991 ), which are each incorporated herein by reference in their entirety.

[0300] For example, DNA encoding a nucleoside triphosphate transporter or polymerase expression cassette, and / or a vector can be introduced into cells by any method including, but not limited to, calcium-mediated transformation, electroporation, microinjection, lipofection, particle gun, etc.

[0301] In some cases, cells contain unnatural nucleoside triphosphates incorporated into one or more nucleic acids within the cell. For example, the cells are living cells that can incorporate at least one unnatural nucleotide into DNA or RNA maintained within the cell. The cells can also incorporate at least one unnatural base pair (UBP) containing a pair of unnatural base-pairing nucleotides into nucleic acids within the cell under in vivo conditions, where the unnatural base-pairing nucleotides, for example, their respective triphosphates, are taken up by the cell by the action of a nucleoside triphosphate transporter and the genes are present (e.g., introduced) into the cell by genetic transformation. For example, upon incorporation into nucleic acids maintained within the cell, dTPT3 and dCNMO can form stable unnatural base pairs that are stably propagated by the organism's DNA replication machinery when grown in a growth medium containing, for example, dTPT3TP and dCNMOTP.

[0302] In some cases, cells can replicate nucleic acids containing unnatural nucleotides. Such methods can include genetically transforming cells with an expression cassette encoding a nucleoside triphosphate transporter that can transport one or more unnatural nucleotides as respective triphosphates under in vivo conditions. Alternatively, previously genetically transformed cells can be utilized with an expression cassette capable of expressing the encoded nucleoside triphosphate transporter. The method can also include contacting or exposing the genetically transformed cells in a maintenance medium suitable for cell growth and replication to the respective triphosphate forms of potassium phosphate and at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides capable of forming an unnatural base pair (UBP)), and maintaining the transformed cells in the maintenance medium in the presence of the respective triphosphate forms of at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides capable of forming an unnatural base pair (UBP)) by at least one replication cycle of the cells under in vivo conditions. The method can also include contacting or exposing the genetically transformed cells in a maintenance medium suitable for cell growth and replication to the respective triphosphate forms of potassium phosphate and at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides configured to form an unnatural base pair (UBP)), and maintaining the transformed cells in the maintenance medium in the presence of the respective triphosphate forms of at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides configured to form an unnatural base pair (UBP)) by at least one replication cycle of the cells under in vivo conditions.

[0303] In some embodiments, the cell contains a stably integrated unnatural nucleic acid. Some embodiments include cells (e.g., as E. coli) that stably incorporate nucleotides other than A, G, T, and C within nucleic acids maintained intracellularly. For example, the nucleotides other than A, G, T, and C are d5SICS, dCNMO, dNaM, and / or dTPT3, which can form stable unnatural base pairs within the nucleic acid upon incorporation into the nucleic acid of the cell. In one aspect, the unnatural nucleotides and unnatural base pairs are such that an organism transformed with a gene for a triphosphate transporter propagates stably by the replication machinery of the organism in a growth medium containing potassium phosphate, as well as the triphosphate forms of d5SICS, dNaM, dCNMO, and / or dTPT3 when growing.

[0304] In some cases, the cell contains an expanded genetic alphabet. The cell can contain a stably integrated unnatural nucleic acid. In some embodiments, a cell having an expanded genetic alphabet includes an unnatural nucleic acid containing an unnatural nucleotide that can base pair with another unnatural nucleotide. In some embodiments, a cell having an expanded genetic alphabet includes an unnatural nucleic acid that hydrogen bonds to another nucleic acid. In some embodiments, a cell having an expanded genetic alphabet includes an unnatural nucleic acid that does not hydrogen bond to another base-paired nucleic acid. In some embodiments, a cell having an expanded genetic alphabet includes an unnatural nucleic acid containing an unnatural nucleotide having a nucleobase that base pairs with a nucleobase or another unnatural nucleotide via hydrophobic and / or packing interactions. In some embodiments, a cell having an expanded genetic alphabet includes an unnatural nucleic acid that base pairs with another nucleic acid via non-hydrogen bonding interactions. A cell having an expanded genetic alphabet is a cell that can copy a homologous nucleic acid to form a nucleic acid containing an unnatural nucleic acid. A cell having an expanded genetic alphabet is a cell that includes an unnatural nucleic acid base paired with another unnatural nucleic acid (unnatural base pair (UBP)).

[0305] In some embodiments, the cells form unnatural DNA base pairs (UBPs) from the introduced unnatural nucleotides under in vivo conditions. In some embodiments, the activity of potassium phosphate and / or inhibitors of phosphatase and / or nucleotidase can promote the transport of unnatural nucleotides. The method includes the use of cells that express a heterologous nucleoside triphosphate transporter. When such cells are contacted with one or more nucleoside triphosphates, the nucleoside triphosphates are transported into the cells. The cells are in the presence of potassium phosphate and / or inhibitors of phosphatase and nucleotidase. Unnatural nucleoside triphosphates can be incorporated into the nucleic acids within the cells by the natural machinery (i.e., polymerase) of the cells, for example, by forming mutual base pairs to form unnatural base pairs within the nucleic acids of the cells. In some embodiments, the UBPs are formed between DNA and RNA nucleotides having unnatural bases.

[0306] In some embodiments, the UBPs are incorporated into a cell or a population of cells when exposed to unnatural triphosphates. In some embodiments, the UBPs are incorporated into a cell or a population of cells when substantially always exposed to unnatural triphosphates.

[0307] In some embodiments, induction of the expression of a heterologous gene, such as a nucleoside triphosphate transporter (NTT), in the cells can result in slower cell growth and increased uptake of unnatural triphosphates compared to the growth and uptake of one or more unnatural triphosphates in cells without induction of the expression of the heterologous gene. Uptake can variously include the transport of nucleotides into the cells, such as by diffusion, osmosis, or via the action of transporters. In some embodiments, induction of the expression of a heterologous gene, such as NTT, in the cells can result in increased cell growth and increased uptake of unnatural nucleic acids compared to the growth and uptake of cells without induction of the expression of the heterologous gene.

[0308] In some embodiments, the UBP is incorporated during the logarithmic growth phase. In some embodiments, the UBP is incorporated during the non-logarithmic growth phase. In some embodiments, the UBP is incorporated during the substantially linear growth phase. In some embodiments, the UBP is stably incorporated into a cell or population of cells after growing over a period. For example, the UBP can be stably incorporated into a cell or population of cells after growing over at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45 or 50, or more replications. For example, the UBP can be stably incorporated into a cell or population of cells after growing over at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24 hours of growth. For example, the UBP can be stably incorporated into a cell or population of cells after growing over at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or 31 days of growth. For example, the UBP can be stably incorporated into a cell or population of cells after growing over at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months of growth. For example, the UBP can be stably incorporated into a cell or population of cells after growing over at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45 or 50 years of growth.

[0309] In some embodiments, the semi-synthetic organisms disclosed herein are [Chemical formula] It contains DNA that includes at least one unnatural nucleobase selected from the group consisting of. In some embodiments, the DNA of a semi-synthetic organism that includes at least one of the unnatural bases forms an unnatural base pair (UBP). In some embodiments, the unnatural base pair (UBP) is dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1 or dNaM-dTAT1. In some embodiments, the DNA is [Chemical formula] It includes at least one unnatural nucleobase selected from the group consisting of. In some embodiments, the DNA is [Chemical formula] It includes at least one unnatural nucleobase selected from the group consisting of. In some embodiments, the DNA is [Chemical formula] It includes at least one unnatural nucleobase selected from the group consisting of. In some embodiments, the DNA is [Chemical formula] It includes at least one unnatural nucleobase selected from the group consisting of. In some embodiments, the DNA is [Chemical formula] It includes at least one unnatural nucleobase selected from. In some embodiments, the DNA is [Chemical formula] It includes at least two unnatural nucleobases selected from. In some embodiments, the DNA includes two strands, and the first strand is [Chemical formula] comprises at least one nucleobase which is, [Chemical formula] comprises at least one nucleobase which is. In some embodiments, the DNA [Chemical formula] comprises at least one unnatural nucleobase which is.

[0310] In some embodiments, the cell further utilizes RNA polymerase to generate an mRNA containing one or more unnatural nucleotides. In some embodiments, the RNA polymerase is a heterologous RNA polymerase. In some examples, the cell further utilizes a polymerase to generate a tRNA containing an anticodon that includes one or more unnatural nucleotides. In some embodiments, the tRNA is a heterologous tRNA. In some examples, the tRNA is charged with an unnatural amino acid. In some examples, the unnatural anticodon of the tRNA base pairs with an unnatural codon of the mRNA during translation to synthesize a protein containing the unnatural amino acid.

[0311] In some embodiments, the semi-synthetic organisms disclosed herein express a heterologous nucleoside triphosphate transporter. In some embodiments, the heterologous nucleoside triphosphate transporter is PtNTT2. In some embodiments, the semi-synthetic organisms further express a heterologous tRNA synthetase. In some embodiments, the heterologous tRNA synthetase is the pyrrolysyl-tRNA synthetase of M. barkeri (Mb PylRS). In some embodiments, the semi-synthetic organisms express the nucleoside triphosphate transporter PtNTT2 and further express the pyrrolysyl-tRNA synthetase of M. barkeri (Mb PylRS) as the tRNA synthetase. In some embodiments, the semi-synthetic organisms further express a heterologous RNA polymerase. In some embodiments, the heterologous RNA polymerase is T7 RNAP. In some embodiments, the semi-synthetic organisms do not express a protein having a DNA recombination repair function. In some embodiments, the semi-synthetic organisms are Escherichia coli and the organisms do not express RecA.

[0312] In some embodiments, the semi-synthetic organisms further contain a heterologous mRNA. In some embodiments, the heterologous mRNA

Chemical formula

Chemical formula

Chemical formula

Chemical formula

[0313] In some embodiments, the semi-synthetic organism further comprises a heterologous tRNA. In some embodiments, the heterologous tRNA comprises [Chemical formula] at least one unnatural base selected from [Chemical formula] at least one unnatural base that is [Chemical formula] at least one unnatural base that is [Chemical formula] at least one unnatural base that is

[0314] In some embodiments, the semi-synthetic organisms disclosed herein further comprise a heterologous mRNA and a heterologous tRNA. In some embodiments, the semi-synthetic organism further comprises (a) a heterologous nucleoside triphosphate transporter, (b) a heterologous mRNA, (c) a heterologous tRNA, (d) a heterologous tRNA synthetase, and (e) a heterologous RNA polymerase, and the organism does not express a protein having a DNA recombination repair function. In some embodiments, the nucleoside triphosphate transporter is PtNTT2, the tRNA synthetase is the pyrrolysyl-tRNA synthetase of M. barkeri (Mb PylRS), and the RNA polymerase is T7 RNAP. In some embodiments, the semi-synthetic organism is Escherichia coli, and the organism does not express RecA. In some embodiments, the semi-synthetic organism overexpresses one or more DNA polymerases. In some embodiments, the organism overexpresses DNA Pol II.

[0315] Natural and Unnatural Amino Acids As used herein, an amino acid residue may refer to a molecule containing both an amino group and a carboxyl group. Suitable amino acids include, but are not limited to, both the D and L isomers of naturally occurring amino acids, as well as amino acids not found in nature that are prepared by organic synthesis or any other method. The term amino acid as used herein includes, but is not limited to, α-amino acids, β-amino acids, naturally occurring amino acids, non-standard amino acids, unnatural amino acids, and amino acid analogs.

[0316] The term "α-amino acid" may refer to a molecule containing both an amino group and a carboxyl group attached to a carbon, called the α-carbon. For example:

Chemical formula

[0317] The term "β-amino acid" can refer to a molecule containing both an amino group and a carboxyl group in the β configuration.

[0318] The term "naturally occurring amino acid" may refer to any one of the 20 amino acids commonly found in peptides synthesized in nature, known by the single-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V.

[0319] The following table shows a summary of the properties of naturally occurring amino acids.

[0320]

Table 1

[0321] Examples of "hydrophobic amino acids" include small hydrophobic amino acids and large hydrophobic amino acids. "Small hydrophobic amino acids" may be glycine, alanine, proline, and their analogs. "Large hydrophobic amino acids" may be valine, leucine, isoleucine, phenylalanine, methionine, tryptophan, and their analogs. "Polar amino acids" may be serine, threonine, asparagine, glutamine, cysteine, tyrosine, and their analogs. "Charged amino acids" may be lysine, arginine, histidine, aspartic acid, glutamate, and their analogs.

[0322] An "amino acid analog" is a molecule that is structurally similar to an amino acid and can substitute for an amino acid in the formation of peptide-mimicking macrocyclic molecules. Examples of amino acid analogs include, but are not limited to, β-amino acids and amino acids in which the amino or carboxy group is replaced by a similarly reactive group (e.g., substitution of a primary amine with a secondary or tertiary amine, or substitution of a carboxy group with an ester).

[0323] A "non-standard amino acid (ncAA)" or "non-natural amino acid" or "unnatural amino acid" is an amino acid that is not one of the 20 amino acids commonly found in peptides synthesized in nature and is known by the one-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. In some cases, unnatural amino acids are a subset of non-standard amino acids.

[0324] Amino acid analogs can include β - amino acid analogs. Examples of β - amino acid analogs include, but are not limited to, the following: cyclic β - amino acid analogs; β - alanine; (R)-β - phenylalanine; (R)-1,2,3,4 - tetrahydro - isoquinoline - 3 - acetic acid; (R)-3 - amino - 4-(1 - naphthyl)-butyric acid; (R)-3 - amino - 4-(2,4 - dichlorophenyl)butyric acid; (R)-3 - amino - 4-(2 - chlorophenyl)-butyric acid; (R)-3 - amino - 4-(2 - cyanophenyl)-butyric acid; (R)-3 - amino - 4-(2 - fluorophenyl)-butyric acid; (R)-3 - amino - 4-(2 - furyl)-butyric acid; (R)-3 - amino - 4-(2 - methylphenyl)-butyric acid; (R)-3 - amino - 4-(2 - naphthyl)-butyric acid; (R)-3 - amino - 4-(2 - thienyl)-butyric acid; (R)-3 - amino - 4-(2 - trifluoromethylphenyl)-butyric acid; (R)-3 - amino - 4-(3,4 - dichlorophenyl)butyric acid; (R)-3 - amino - 4-(3,4 - difluorophenyl)butyric acid; (R)-3 - amino - 4-(3 - benzothienyl)-butyric acid; (R)-3 - amino - 4-(3 - chlorophenyl)-butyric acid; (R)-3 - amino - 4-(3 - cyanophenyl)-butyric acid; (R)-3 - amino - 4-(3 - fluorophenyl)-butyric acid; (R)-3 - amino - 4-(3 - methylphenyl)-butyric acid; (R)-3 - amino - 4-(3 - pyridyl)-butyric acid; (R)-3 - amino - 4-(3 - thienyl)-butyric acid; (R)-3 - amino - 4-(3 - trifluoromethylphenyl)-butyric acid; (R)-3 - amino - 4-(4 - bromophenyl)-butyric acid; (R)-3 - amino - 4-(4 - chlorophenyl)-butyric acid; (R)-3 - amino - 4-(4 - cyanophenyl)-butyric acid; (R)-3 - amino - 4-(4 - fluorophenyl)-butyric acid; (R)-3 - amino - 4-(4 - iodophenyl)-butyric acid; (R)-3 - amino - 4-(4 - methylphenyl)-butyric acid; (R)-3 - amino - 4-(4 - nitrophenyl)-butyric acid; (R)-3 - amino - 4-(4 - pyridyl)-butyric acid; (R)-3 - amino - 4-(4 - trifluoromethylphenyl)-butyric acid; (R)-3 - amino - 4 - pentafluoro - phenylbutyric acid; (R)-3 - amino - 5 - hexenoic acid; (R)-3 - amino - 5 - hexyne acid;(R)-3-Amino-5-phenylpentanoic acid; (R)-3-Amino-6-phenyl-5-hexenoic acid; (S)-1,2,3,4-Tetrahydroisoquinoline-3-acetic acid; (S)-3-Amino-4-(1-naphthyl)butyric acid; (S)-3-Amino-4-(2,4-dichlorophenyl)butyric acid; (S)-3-Amino-4-(2-chlorophenyl)butyric acid; (S)-3-Amino-4-(2-cyanophenyl)butyric acid; (S)-3-Amino-4-(2-fluorophenyl)butyric acid; (S)-3-Amino-4-(2-furyl)butyric acid; (S)-3-Amino-4-(2-methylphenyl)butyric acid; (S)-3-Amino-4-(2-naphthyl)butyric acid; (S)-3-Amino-4-(2-thienyl)butyric acid; (S)-3-Amino-4-(2-trifluoromethylphenyl)butyric acid; (S)-3-Amino-4-(3,4-dichlorophenyl)butyric acid; (S)-3-Amino-4-(3,4-difluorophenyl)butyric acid; (S)-3-Amino-4-(3-benzothienyl)butyric acid; (S)-3-Amino-4-(3-chlorophenyl)butyric acid; (S)-3-Amino-4-(3-cyanophenyl)butyric acid; (S)-3-Amino-4-(3-fluorophenyl)butyric acid; (S)-3-Amino-4-(3-methylphenyl)butyric acid; (S)-3-Amino-4-(3-pyridyl)butyric acid; (S)-3-Amino-4-(3-thienyl)butyric acid; (S)-3-Amino-4-(3-trifluoromethylphenyl)butyric acid; (S)-3-Amino-4-(4-bromophenyl)butyric acid; (S)-3-Amino-4-(4-chlorophenyl)butyric acid; (S)-3-Amino-4-(4-cyanophenyl)butyric acid; (S)-3-Amino-4-(4-fluorophenyl)butyric acid; (S)-3-Amino-4-(4-iodophenyl)butyric acid; (S)-3-Amino-4-(4-methylphenyl)butyric acid; (S)-3-Amino-4-(4-nitrophenyl)butyric acid; (S)-3-Amino-4-(4-pyridyl)butyric acid; (S)-3-Amino-4-(4-trifluoromethylphenyl)butyric acid; (S)-3-Amino-4-pentafluoro-phenylbutyric acid; (S)-3-Amino-5-hexenoic acid; (S)-3-Amino-5-hexynoic acid; (S)-3-Amino-5-phenylpentanoic acid; (S)-3-Amino-6-phenyl-5-hexenoic acid;1,2,5,6-Tetrahydropyridine-3-carboxylic acid; 1,2,5,6-tetrahydropyridine; 4-carboxylic acid; 3-amino-3-(2-chlorophenyl)-propionic acid; 3-amino-3-(2-thienyl)-propionic acid; 3-amino-3-(3-bromophenyl)-propionic acid; 3-amino-3-(4-chlorophenyl)-propionic acid; 3-amino-3-(4-methoxyphenyl)-propionic acid; 3-amino-4,4,4-trifluoro-butyric acid; 3-aminoadipic acid; D-β-phenylalanine; β-leucine; L-β-homoalanine; L-β-homol-aspartic acid γ-benzyl ester; L-β-homoglutamic acid δ-benzyl ester; L-β-homoisoleucine; L-β-homoleucine; L-β-homomethionine; L-β-homophenylalanine; L-β-homoproline; L-β-homotryptophan; L-β-homovaline; L-Nω-benzyloxycarbonyl-β-homolysine; Nω-L-β-homoarginine; O-benzyl-L-β-homohydroxyproline; O-benzyl-L-β-homoserine; O-benzyl-L-β-homothreonine; O-benzyl-L-β-homotyrosine; γ-trityl-L-β-homol-aspartic acid; (R)-β-phenylalanine; L-β-homol-aspartic acid γ-t-butyl ester; L-β-homoglutamic acid δ-t-butyl ester; L-Nω-β-homolysine; Nδ-trityl-L-β-homoglutamic acid; Nω-2,2,4,6,7-pentamethyl-dihydrobenzofuran-5-sulfonyl-L-β-homoarginine; O-t-butyl-L-β-homohydroxy-proline; O-t-butyl-L-β-homoserine; O-t-butyl-L-β-homothreonine; O-t-butyl-L-β-homotyrosine; 2-aminocyclopentanecarboxylic acid; and 2-aminocyclohexanecarboxylic acid.

[0325] Examples of amino acid analogs may include analogs of alanine, valine, glycine, or leucine. Examples of amino acid analogs of alanine, valine, glycine, and leucine include, but are not limited to, the following: α-methoxy glycine; α-allyl-L-alanine; α-aminoisobutyric acid; α-methyl-leucine; β-(1-naphthyl)-D-alanine; β-(1-naphthyl)-L-alanine; β-(2-naphthyl)-D-alanine; β-(2-naphthyl)-L-alanine; β-(2-pyridyl)-D-alanine; β-(2-pyridyl)-L-alanine; β-(2-thienyl)-D-alanine; β-(2-thienyl)-L-alanine; β-(3-benzothienyl)-D-alanine; β-(3-benzothienyl)-L-alanine; β-(3-pyridyl)-D-alanine; β-(3-pyridyl)-L-alanine; β-(4-pyridyl)-D-alanine; β-(4-pyridyl)-L-alanine; β-chloro-L-alanine; β-cyano-L-alanine; β-cyclohexyl-D-alanine; β-cyclohexyl-L-alanine; β-cyclopenten-1-yl-alanine; β-cyclopentyl-alanine; β-cyclopropyl-L-Ala-OH.dicyclohexylammonium salt; β-t-butyl-D-alanine; β-t-butyl-L-alanine; γ-aminobutyric acid; L-α,β-diaminopropionic acid; 2,4-dinitrophenylglycine; 2,5-dihydro-D-phenylglycine; 2-amino-4,4,4-trifluorobutyric acid; 2-fluoro-phenylglycine; 3-amino-4,4,4-trifluoro-butyric acid; 3-fluorovaline; 4,4,4-trifluorovaline; 4,5-dehydro-L-leu-OH.dicyclohexylammonium salt; 4-fluoro-D-phenylglycine; 4-fluoro-L-phenylglycine; 4-hydroxy-D-phenylglycine; 5,5,5-trifluoroleucine; 6-aminohexanoic acid; cyclopentyl-D-Gly-OH.dicyclohexylammonium salt; cyclopentyl-Gly-OH.Dicyclohexylammonium salts; D-α,β-diaminopropionic acid; D-α-aminobutyric acid; D-α-t-butylglycine; D-(2-thienyl)glycine; D-(3-thienyl)glycine; D-2-aminocaproic acid; D-2-indanylglycine; D-allylglycine-dicyclohexylammonium salt; D-cyclohexylglycine; D-norvaline; D-phenylglycine; β-aminobutyric acid; β-aminoisobutyric acid; (2-bromophenyl)glycine; (2-methoxyphenyl)glycine; (2-methylphenyl)glycine; (2-thiazolyl)glycine; (2-thienyl)glycine; 2-amino-3-(dimethylamino)-propionic acid; L-α,β-diaminopropionic acid; L-α-aminobutyric acid; L. -α-t-butylglycine; L-(3-thienyl)glycine; L-2-amino-3-(dimethylamino)-propionic acid; L-2-aminocaproic acid dicyclohexyl-ammonium salt; L-2-indanylglycine; L-allylglycine dicyclohexylammonium salt; L-cyclohexylglycine; L-phenylglycine; L-propargylglycine; L-norvaline; N-α-aminomethyl-L-alanine; D-α,γ-diaminobutyric acid; L-α,γ-diaminobutyric acid; β-cyclopropyl-L-alanine; (N-β-(2,4-dinitrophenyl))-L-α,β-diaminopropionic acid; (N-β-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-D-α,β-diaminopropionic acid; (N-β-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-L-α,β-diaminopropionic acid; (N-β-4-methyltrityl)-L-α,β-diaminopropionic acid; (N-β-allyloxycarbonyl)-L-α,β-diaminopropionic acid; (N-γ-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-D-α,γ-diaminobutyric acid; (N-γ-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-L-α,γ-diaminobutyric acid; (N-γ-4-methyltrityl)-D-α,γ-diaminobutyric acid; (N-γ-4-methyltrityl)-L-α,γ-diaminobutyric acid; (N-γ-allyloxycarbonyl)-L-α,γ-diaminobutyric acid; D-α,γ-diaminobutyric acid; 4,5-dehydro-L-leucine; cyclopentyl-D-Gly-OH; cyclopentyl-Gly-OH; D-allylglycine; D-homocyclohexylalanine; L-1-pyrenylalanine; L-2-aminocaproic acid; L-allylglycine; L-homocyclohexylalanine; and N-(2-hydroxy-4-methoxy-Bzl)-Gly-OH.

[0326] Examples of amino acid analogs may include analogs of arginine or lysine. Examples of amino acid analogs of arginine and lysine include, but are not limited to, the following: citrulline; L-2-amino-3-guanidinopropionic acid; L-2-amino-3-ureidopropionic acid; L-citrulline; Lys(Me)2-OH; Lys(N3)-OH; Nδ-benzyloxycarbonyl-L-ornithine; Nω-nitro-D-arginine; Nω-nitro-L-arginine; α-methyloornithine; 2,6-diaminoheptanedioic acid; L-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxo-cyclohex-1-ylidene)ethyl)-D-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxo-cyclohexen-1-ylidene)ethyl)-L-ornithine; (Nδ-4-methyltrityl)-D-ornithine; (Nδ-4-methyltrityl)-L-ornithine; D-ornithine; L-ornithine; Arg(Me)(Pbf)-OH; Arg(Me)2-OH (asymmetric); Arg(Me)2-OH (symmetric); Lys(ivDde)-OH; Lys(Me)2-OH.HCl; Lys(Me3)-OH chloride; Nω-nitro-D-arginine; and Nω-nitro-L-arginine.

[0327] Examples of amino acid analogs may include analogs of aspartic acid or glutamic acid. Examples of amino acid analogs of aspartic acid and glutamic acid include, but are not limited to, the following: α-methyl-D-aspartic acid; α-methyl-glutamic acid; α-methyl-L-aspartic acid; γ-methylene-glutamic acid; (N-γ-ethyl)-L-glutamic acid; [N-α-(4-aminobenzoyl)]-L-glutamic acid; 2,6-diaminopimelic acid; L-α-amino-suberic acid; D-2-aminoadipic acid; D-α-amino-suberic acid; α-aminopimelic acid; iminodiacetic acid; L-2-aminoadipic acid; threo-β-methyl-aspartic acid; γ-carboxy-D-glutamic acid γ,γ-di-t-butyl ester; γ-carboxy-L-glutamic acid γ,γ-di-t-butyl ester; Glu(OAll)-OH; L-Asu(OtBu)-OH; and pyroglutamic acid.

[0328] Amino acid analogs can include analogs of cysteine and methionine. Cysteine and examples of amino acid analogs of methionine include, but are not limited to, the following: Cys(farnesyl)-OH, Cys(farnesyl)-OMe, α-methyl-methionine, Cys(2-hydroxyethyl)-OH, Cys(3-aminopropyl)-OH, 2-amino-4-(ethylthio)butyric acid, butionine, butionine sulfoximine, ethionine, methionine methylsulfonium chloride, selenomethionine, cysteic acid, [2-(4-pyridyl)ethyl]-DL-penicillamine, [2-(4-pyridyl)ethyl]-L-cysteine, 4-methoxybenzyl-D-penicillamine, 4-methoxybenzyl-L-penicillamine, 4-methylbenzyl-D-penicillamine, 4-methylbenzyl-L-penicillamine, benzyl-D-cysteine, benzyl-L-cysteine, benzyl-DL-homocysteine, carbamoyl-L-cysteine, carboxyethyl-L-cysteine, carboxymethyl-L-cysteine, diphenylmethyl-L-cysteine, ethyl-L-cysteine, methyl-L-cysteine, t-butyl-D-cysteine, trityl-L-homocysteine, trityl-D-penicillamine, cystathionine, homocystine, L-homocystine, (2-aminoethyl)-L-cysteine, seleno-L-cystine, cystathionine, Cys(StBu)-OH, and acetamidomethyl-D-penicillamine.

[0329] Examples of amino acid analogs may include analogs of phenylalanine and tyrosine. Examples of amino acid analogs of phenylalanine and tyrosine include, but are not limited to, the following: β-methyl-phenylalanine, β-hydroxyphenylalanine, α-methyl-3-methoxy-DL-phenylalanine, α-methyl-D-phenylalanine, α-methyl-L-phenylalanine, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, 2,4-dichloro-phenylalanine, 2-(trifluoromethyl)-D-phenylalanine, 2-(trifluoromethyl)-L-phenylalanine, 2-bromo-D-phenylalanine, 2-bromo-L-phenylalanine, 2-chloro-D-phenylalanine, 2-chloro-L-phenylalanine, 2-cyano-D-phenylalanine, 2-cyano-L-phenylalanine, 2-fluoro-D-phenylalanine, 2-fluoro-L-phenylalanine, 2-methyl-D-phenylalanine, 2-methyl-L-phenylalanine, 2-nitro-D-phenylalanine, 2-nitro-L-phenylalanine, 2,4,5-trihydroxy-phenylalanine, 3,4,5-trifluoro-D-phenylalanine, 3,4,5-trifluoro-L-phenylalanine, 3,4-dichloro-D-phenylalanine, 3,4-dichloro-L-phenylalanine, 3,4-difluoro-D-phenylalanine, 3,4-difluoro-L-phenylalanine, 3,4-dihydroxy-L-phenylalanine, 3,4-dimethoxy-L-phenylalanine, 3,5,3’-triiodo-L-thyronine, 3,5-diiodo-D-tyrosine, 3,5-diiodo-L-tyrosine, 3,5-Iodo-L-thyronine, 3-(trifluoromethyl)-D-phenylalanine, 3-(trifluoromethyl)-L-phenylalanine, 3-amino-L-tyrosine, 3-bromo-D-phenylalanine, 3-bromo-L-phenylalanine, 3-chloro-D-phenylalanine, 3-chloro-L-phenylalanine, 3-chloro-L-tyrosine, 3-cyano-D-phenylalanine, 3-cyano-L-phenylalanine, 3-fluoro-D-phenylalanine, 3-fluoro-L-phenylalanine, 3-fluoro-thyrosine, 3-iodo-D-phenylalanine, 3-iodo-L-phenylalanine, 3-iodo-L-thyronine, 3-methoxy-L-thyronine, 3-methyl-D-phenylalanine, 3-methyl-L-phenylalanine, 3-nitro-D-phenylalanine, 3-nitro-L-phenylalanine, 3-nitro-L-thyronine, 4-(trifluoromethyl)-D-phenylalanine, 4-(trifluoromethyl)-L-phenylalanine, 4-amino-D-phenylalanine, 4-amino-L-phenylalanine, 4-benzoyl-D-phenylalanine, 4-benzoyl-L-phenylalanine, 4-bis(2-chloroethyl)amino-L-phenylalanine, 4-bromo-D-phenylalanine, 4-bromo-L-phenylalanine, 4-chloro-D-phenylalanine, , 4-chloro-L-phenylalanine, 4-cyano-D-phenylalanine, 4-cyano-L-phenylalanine, 4-fluoro-D-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-D-phenylalanine, 4-iodo-L-phenylalanine, homophenylalanine, thyroxine, 3,3-diphenylalanine, thyronine, ethyl-tyrosine, and methylthyronine.

[0330] Examples of amino acid analogs can include analogs of proline. Examples of amino acid analogs of proline include, but are not limited to: 3,4-dehydro-proline, 4-fluoro-proline, cis-4-hydroxy-proline, thiazolidine-2-carboxylic acid, and trans-4-fluoro-proline.

[0331] Examples of amino acid analogs can include analogs of serine and threonine. Examples of amino acid analogs of serine and threonine include, but are not limited to, the following: 3-amino-2-hydroxy-5-methylhexanoic acid, 2-amino-3-hydroxy-4-methylpentanoic acid, 2-amino-3-ethoxybutanoic acid, 2-amino-3-methoxybutanoic acid, 4-amino-3-hydroxy-6-methylheptanoic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-ethoxypropionic acid, 4-amino-3-hydroxybutanoic acid, and α-methylserine.

[0332] Examples of amino acid analogs can include analogs of tryptophan. Examples of amino acid analogs of tryptophan include, but are not limited to, the following: α-methyl-tryptophan; β-(3-benzothienyl)-D-alanine; β-(3-benzothienyl)-L-alanine; 1-methyl-tryptophan; 4-methyl-tryptophan; 5-benzyloxy-tryptophan; 5-bromo-tryptophan; 5-chloro-tryptophan; 5-fluoro-tryptophan; 5-hydroxytryptophan; 5-hydroxy-L-tryptophan; 5-methoxy-tryptophan; 5-methoxy-L-tryptophan; 5-methyl-tryptophan; 6-bromo-tryptophan; 6-chloro-D-tryptophan; 6-chloro-tryptophan; 6-fluoro-tryptophan; 6-methyl-tryptophan; 7-benzyloxy-tryptophan; 7-bromo-tryptophan; 7-methyl-tryptophan; D-1,2,3,4-tetrahydro-norharman-3-carboxylic acid; 6-methoxy-1,2,3,4-tetrahydronorharman-1-carboxylic acid; 7-azatryptophan; L-1,2,3,4-tetrahydro-norharman-3-carboxylic acid; 5-methoxy-2-methyl-tryptophan; and 6-chloro-L-tryptophan.

[0333] The amino acid analog may be a racemate. In some cases, the D-isomer of the amino acid analog is used. In some cases, the L-isomer of the amino acid analog is used. In some cases, the amino acid analog contains a chiral center in the R or S configuration. Sometimes, the amino group of the β-amino acid analog is substituted with a protecting group such as tert-butyloxycarbonyl (BOC group), 9-fluorenylmethyloxycarbonyl (FMOC), tosyl, etc. Sometimes, the carboxylic acid functional group of the β-amino acid analog is protected, for example, as its ester derivative. In some cases, salts of the amino acid analog are used.

[0334] In some embodiments, the non-natural amino acid is a non-natural amino acid as described in Liu C.C., Schultz, P.G. Annu. Rev. Biochem. 2010, 79, 413 (the disclosure of which is incorporated herein by reference in its entirety). In some embodiments, the non-natural amino acid includes N6(2-azidoethoxy)-carbonyl-L-lysine.

[0335] In some embodiments, the amino acid residues described herein (e.g., within a protein) are mutated to unnatural amino acids prior to attachment to the conjugate moiety. In some cases, the mutation to an unnatural amino acid prevents or minimizes an autoimmune response of the immune system. As used herein, the term “unnatural amino acid” refers to an amino acid other than the 20 amino acids that are naturally present in proteins. Non-limiting examples of unnatural amino acids include: p-acetyl-L-phenylalanine, p-iodo-L-phenylalanine, p-methoxyphenylalanine, O-methyl-L-tyrosine, p-propynyloxyphenylalanine, p-propynyl-phenylalanine, L-3-(2-naphthyl)alanine, 3-methyl-phenylalanine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, tri-O-acetyl-GlcNAcp-serine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-boronophenylalanine, O-propynyl tyrosine, L-phosphoserine, phosphonoserine, phosphonotyrosine, p-bromophenylalanine, selenocysteine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, azido-lysine (N6-azidoethoxy-carbonyl-L-lysine, AzK), unnatural analogs of tyrosine amino acids; unnatural analogs of glutamine amino acids; unnatural analogs of phenylalanine amino acids; unnatural analogs of serine amino acids; unnatural analogs of threonine amino acids; alkyl, aryl, acyl, azide, cyano, halo, hydrazine, hydrazide, hydroxyl, alkenyl, alkynyl, ether, thiol, sulfonyl, seleno, ester, thioacid, borate, boronate, phospho, phosphono, phosphine, heterocyclic, enone, imine, aldehyde, hydroxylamine, keto, or amino-substituted amino acids, or combinations thereof; amino acids having photoactivatable crosslinkers; spin-labeled amino acids; fluorescent amino acids; metal-binding amino acids; metal-containing amino acids; radioactive amino acids; photocaged and / or photo-isomerizable amino acids; biotin or biotin analogs containing amino acids;Ketones containing amino acids; Amino acids containing polyethylene glycol or polyethers; Heavy atom-substituted amino acids; Chemically cleavable or photocleavable amino acids; Amino acids having an elongated side chain; Amino acids containing a toxic group; Sugar-substituted amino acids; Carbon-bonded sugar-containing amino acids; Redox-active amino acids; α-hydroxy-containing acids; Aminothio acids; α,α-disubstituted amino acids; β-amino acids; Cyclic amino acids other than proline or histidine, and aromatic amino acids other than phenylalanine, tyrosine or tryptophan.;

[0336] In some embodiments, the unnatural amino acid comprises a selectively reactive group, or a reactive group for site-selective labeling of a target protein or polypeptide. Optionally, the chemical reaction is a bioorthogonal reaction (e.g., biocompatible and selective reaction). Optionally, the chemical reaction is a Cu(I)-catalyzed or "copper-free" alkyne-azide triazole formation reaction, Staudinger ligation, inverse electron demand Diels-Alder (IEDDA) reaction, "photoclick" chemistry, or a metal-mediated process such as olefin metathesis and Suzuki-Miyaura or Sonogashira cross-coupling. In some embodiments, the unnatural amino acid comprises a photoreactive group that crosslinks upon irradiation with, for example, UV. In some embodiments, the unnatural amino acid comprises a photocaged amino acid. Optionally, the unnatural amino acid is a para-substituted, meta-substituted, or ortho-substituted amino acid derivative.

[0337] Optionally, the unnatural amino acid is p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, O-methyl-L-tyrosine, p-methoxyphenylalanine, p-propynyloxyphenylalanine, p-propynyl-phenylalanine, L-3-(2-naphthyl)alanine, 3-methyl-phenylalanine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, tri-O-acetyl-GlcNAcp-serine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalan It contains p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, L-phosphoserine, phosphonoserine, phosphonotyrosine, p-bromophenylalanine, p-amino-L-phenylalanine, or isopropyl-L-phenylalanine.

[0338] In some cases, the non-natural amino acid is 3-aminotyrosine, 3-nitrotyrosine, 3,4-dihydroxy-phenylalanine, or 3-iodotyrosine. In some cases, the non-natural amino acid is phenylselenocysteine. In some cases, the non-natural amino acid is benzophenone, ketone, iodide, methoxy, acetyl, benzoyl, or azide (including phenylalanine derivatives). In some cases, the non-natural amino acid is a benzophenone, ketone, iodide, methoxy, acetyl, benzoyl, or azide-containing lysine derivative. In some cases, the non-natural amino acid contains an aromatic side chain. In some cases, the non-natural amino acid does not contain an aromatic side chain. In some cases, the non-natural amino acid contains an azide group. In some cases, the non-natural amino acid contains a Michael acceptor group. In some cases, the Michael acceptor group contains an unsaturated moiety capable of forming a covalent bond via a 1,2-addition reaction. In some cases, the Michael acceptor group contains an electron-deficient alkene or alkyne. In some cases, the Michael acceptor group includes, but is not limited to, alpha, beta-unsaturated: ketone, aldehyde, sulfoxide, sulfone, nitrile, imine, or aromatic. In some cases, the non-natural amino acid is dehydroalanine. In some cases, the non-natural amino acid contains an aldehyde or ketone group. In some cases, the non-natural amino acid is a lysine derivative containing an aldehyde or ketone group. In some cases, the non-natural amino acid is a lysine derivative containing one or more O, N, Se, or S atoms at the beta, gamma, or delta position. In some cases, the non-natural amino acid is a lysine derivative containing an O, N, Se, or S atom at the gamma position. In some cases, the non-natural amino acid is a lysine derivative in which the epsilon N atom is replaced by an oxygen atom. In some cases, the non-natural amino acid is a lysine derivative that is a post-translationally modified lysine not naturally occurring.

[0339] In some cases, the unnatural amino acid is an amino acid containing a side chain, and the sixth atom from the alpha position contains a carbonyl group. In some cases, the unnatural amino acid is an amino acid containing a side chain, the sixth atom from the alpha position contains a carbonyl group, and the fifth atom from the alpha position is nitrogen. In some cases, the unnatural amino acid is an amino acid containing a side chain, and the seventh atom from the alpha position is an oxygen atom.

[0340] In some cases, the unnatural amino acid is a serine derivative containing selenium. In some cases, the unnatural amino acid is selenoserine (2-amino-3-hydroxyselenopropanoic acid). In some cases, the unnatural amino acid is 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid. In some cases, the unnatural amino acid is 2-amino-3-(phenylselanyl)propanoic acid. In some cases, the unnatural amino acid contains selenium, and the oxidation of selenium results in the formation of an unnatural amino acid containing an alkene.

[0341] In some cases, the unnatural amino acid contains a cyclooctynyl group. In some cases, the unnatural amino acid contains a trans-cyclooctenyl group. In some cases, the unnatural amino acid contains a norbornenyl group. In some cases, the unnatural amino acid contains a cyclopropenyl group. In some cases, the unnatural amino acid contains a diazirinyl group. In some cases, the unnatural amino acid contains a tetrazinyl group.

[0342] In some cases, the unnatural amino acid is a lysine derivative in which the side-chain nitrogen is carbamylated. In some cases, the unnatural amino acid is a lysine derivative in which the side-chain nitrogen is acylated. In some cases, the unnatural amino acid is 2-amino-6-{[(tert-but It is (xy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(tert-butoxy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is N6-Boc-N6-methyllysine. In some cases, the unnatural amino acid is N6-acetyllysine. In some cases, the unnatural amino acid is pyrrolysine. In some cases, the unnatural amino acid is N6-trifluoroacetyllysine. In some cases, the unnatural amino acid is 2-amino-6-{[(benzyloxy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(p-iodobenzyloxy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(p-nitrobenzyloxy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is N6-prolyllysine. In some cases, the unnatural amino acid is 2-amino-6-{[(cyclopentyloxy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is N6-(cyclopentanecarbonyl) lysine. In some cases, the unnatural amino acid is N6-(tetrahydrofuran-2-carbonyl) lysine. In some cases, the unnatural amino acid is N6-(3-ethynyltetrahydrofuran-2-carbonyl) lysine. In some cases, the unnatural amino acid is N6-((prop-2-yn-1-yloxy) carbonyl) lysine. In some cases, the unnatural amino acid is 2-amino-6-{[(2-azidocyclopentyloxy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is N6-((2-azidoethoxy) carbonyl) lysine. In some cases, the unnatural amino acid is 2-amino-6-{[(2-nitrobenzyloxy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(2-cyclooctynyl oxy) carbonyl] amino} hexanoic acid. In some cases, the unnatural amino acid is N6-(2-aminobut-3-ynoyl) lysine. In some cases, the unnatural amino acid is 2-amino-6-((2-aminobut-3-ynoyl) oxy) hexanoic acid.In some cases, the unnatural amino acid is N6-(allyloxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-(butenyl-4-oxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-(pentenyl-5-oxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-((but-3-yn-1-yloxy)carbonyl)-lysine. In some cases, the unnatural amino acid is N6-((penta-4-yn-1-yloxy)carbonyl)-lysine. In some cases, the unnatural amino acid is N6-(thiazolidine-4-carbonyl)lysine. In some cases, the unnatural amino acid is 2-amino-8-oxononanoic acid. In some cases, the unnatural amino acid is 2-amino-8-oxooctanoic acid. In some cases, the unnatural amino acid is N6-(2-oxoacetyl)lysine.

[0343] In some cases, the unnatural amino acid is N6-propionyllysine. In some cases, the unnatural amino acid is N6-butyryllysine. In some cases, the unnatural amino acid is N6-(but-2-enoyl)lysine. In some cases, the unnatural amino acid is N6-((bicyclo[2.2.1]hept-5-en-2-yloxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((spiro[2.3]hex-1-en-5-ylmethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-(((4-(1-(trifluoromethyl)cycloprop-2-en-1-yl)benzyl)oxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((bicyclo[2.2.1]hept-5-en-2-ylmethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is cysteinyllysine. In some cases, the unnatural amino acid is N6-((1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((3-(3- It is methyl-3H-diazirin-3-yl)propoxy)carbonyl)lysine. Optionally, the unnatural amino acid is N6-((methanitrobenzyloxy)N6-methylcarbonyl)lysine. Optionally, the unnatural amino acid is N6-((bicyclo[6.1.0]non-4-in-9-ylmethoxy)carbonyl)-lysine. Optionally, the unnatural amino acid is N6-((cyclohepta-3-en-1-yloxy)carbonyl)-L-lysine.

[0344] Optionally, the unnatural amino acid is 2-amino-3-(((((benzyloxy)carbonyl)amino)methyl)selanyl)propanoic acid. In some embodiments, the unnatural amino acid is incorporated into the protein by a reassigned amber, opal, or ochre stop codon. In some embodiments, the unnatural amino acid is incorporated into the protein by a four-base codon. In some embodiments, the unnatural amino acid is incorporated into the protein by a reassigned rare sense codon.

[0345] In some embodiments, the unnatural amino acid is incorporated into the protein by an unnatural codon containing an unnatural nucleotide.

[0346] In some embodiments, the protein comprises at least two non-natural amino acids. In some embodiments, the protein comprises at least three non-natural amino acids. In some embodiments, the protein comprises at least two different non-natural amino acids. In some embodiments, the protein comprises at least three different non-natural amino acids. At least one non-natural amino acid is: an analog of lysine; contains an aromatic side chain; contains an azide group; contains an alkyne group; or contains an aldehyde or ketone group. In some embodiments, at least one non-natural amino acid does not contain an aromatic side chain. In some embodiments, at least one non-natural amino acid comprises N6-azidoethoxy-carbonyl-L-lysine (AzK) or N6-propargylethoxy-carbonyl-L-lysine (PraK). In some embodiments, at least one non-natural amino acid comprises N6-azidoethoxy-carbonyl-L-lysine (AzK). In some embodiments, at least one non-natural amino acid comprises N6-propargylethoxy-carbonyl-L-lysine (PraK).

[0347] In some cases, the incorporation of non-natural amino acids into proteins is mediated by orthogonal pairs of modified synthetases / tRNAs. Such orthogonal pairs include natural or mutant synthetases that can charge a specific non-natural amino acid to a non-natural tRNA, while often minimizing a) the charging of other endogenous amino acids or alternative non-natural amino acids to the non-natural tRNA and b) the charging of any other (including endogenous) tRNAs. Such orthogonal pairs include tRNAs that can be charged by a synthetase while avoiding the charging of other endogenous amino acids by endogenous synthetases. In some embodiments, such pairs are identified from various organisms such as bacteria, yeast, archaea, or human sources. In some embodiments, the orthogonal synthetase / tRNA pair comprises components from a single organism. In some embodiments, the orthogonal synthetase / tRNA pair comprises components from two different organisms. In some embodiments, the orthogonal synthetase / tRNA pair comprises components that promote the translation of different amino acids prior to modification. In some embodiments, the orthogonal synthetase is a modified alanine synthetase. In some embodiments, the orthogonal synthetase is a modified arginine synthetase. In some embodiments, the orthogonal synthetase is a modified asparagine synthetase. In some embodiments, the orthogonal synthetase is a modified aspartic acid synthetase. In some embodiments, the orthogonal synthetase is a modified cysteine synthetase. In some embodiments, the orthogonal synthetase is a modified glutamine synthetase. In some embodiments, the orthogonal synthetase is It is a modified glutamate synthetase. In some embodiments, the orthogonal synthetase is a modified alanine glycine. In some embodiments, the orthogonal synthetase is a modified histidine synthetase. In some embodiments, the orthogonal synthetase is a modified leucine synthetase. In some embodiments, the orthogonal synthetase is a modified isoleucine synthetase. In some embodiments, the orthogonal synthetase is a modified lysine synthetase. In some embodiments, the orthogonal synthetase is a modified methionine synthetase. In some embodiments, the orthogonal synthetase is a modified phenylalanine synthetase. In some embodiments, the orthogonal synthetase is a modified proline synthetase. In some embodiments, the orthogonal synthetase is a modified serine synthetase. In some embodiments, the orthogonal synthetase is a modified threonine synthetase. In some embodiments, the orthogonal synthetase is a modified tryptophan synthetase. In some embodiments, the orthogonal synthetase is a modified tyrosine synthetase. In some embodiments, the orthogonal synthetase is a modified valine synthetase. In some embodiments, the orthogonal synthetase is a modified phosphoserine synthetase. In some embodiments, the orthogonal tRNA is a modified alanine tRNA. In some embodiments, the orthogonal tRNA is a modified arginine tRNA. In some embodiments, the orthogonal tRNA is a modified asparagine tRNA. In some embodiments, the orthogonal tRNA is a modified aspartic acid tRNA. In some embodiments, the orthogonal tRNA is a modified cysteine tRNA. In some embodiments, the orthogonal tRNA is a modified glutamine tRNA. In some embodiments, the orthogonal tRNA is a modified glutamate tRNA. In some embodiments, the orthogonal tRNA is a modified alanine glycine. In some embodiments, the orthogonal tRNA is a modified histidine tRNA.In some embodiments, the orthogonal tRNA is a modified leucine tRNA. In some embodiments, the orthogonal tRNA is a modified isoleucine tRNA. In some embodiments, the orthogonal tRNA is a modified lysine tRNA. In some embodiments, the orthogonal tRNA is a modified methionine tRNA. In some embodiments, the orthogonal tRNA is a modified phenylalanine tRNA. In some embodiments, the orthogonal tRNA is a modified proline tRNA. In some embodiments, the orthogonal tRNA is a modified serine tRNA. In some embodiments, the orthogonal tRNA is a modified threonine tRNA. In some embodiments, the orthogonal tRNA is a modified tryptophan tRNA. In some embodiments, the orthogonal tRNA is a modified tyrosine tRNA. In some embodiments, the orthogonal tRNA is a modified valine tRNA. In some embodiments, the orthogonal tRNA is a modified phosphoserine tRNA. In any of these embodiments, the tRNA can be a heterologous tRNA.

[0348] In some embodiments, the unnatural amino acid is incorporated into the protein by an aminoacyl (aaRS or RS)-tRNA synthetase-tRNA pair. Exemplary aaRS-tRNA pairs include, but are not limited to, the Methanococcus jannaschii (Mj-Tyr) aaRS / tRNA pair, the E. coli TyrRS (Ec-Tyr) / B. stearothermophilus tRNACUA pair, the E. coli LeuRS (Ec-Leu) / B. stearothermophilus tRNACUA pair, and the pyrrolysyl-tRNA pair. In some cases, the unnatural amino acid is incorporated into the protein by the Mj-TyrRS / tRNA pair. Exemplary unnatural amino acids (UAAs) that can be incorporated by the Mj-TyrRS / tRNA pair include, but are not limited to, para-substituted phenylalanine derivatives such as p-aminophenylalanine and p-methoxyphenylalanine; meta-substituted tyrosine derivatives such as 3-aminotyrosine, 3-nitrotyrosine, 3,4-dihydroxyphenylalanine , and 3-iodotyrosine; phenylselenocysteine; p-boronophenylalanine; and o-nitrobenzyltyrosine.

[0349] In some cases, the unnatural amino acid is incorporated into the protein by the Ec-Tyr / tRNACUA or Ec-Leu / tRNACUA pair. Exemplary UAAs that can be incorporated by the Ec-Tyr / tRNACUA or Ec-Leu / tRNACUA pair include, but are not limited to, phenylalanine derivatives containing a benzophenone, ketone, iodide, or azide substitution; O-propargyltyrosine; α-aminocaprylic acid, O-methyltyrosine, O-nitrobenzylcysteine; and 3-(naphthalen-2-ylamino)-2-amino-propanoic acid.

[0350] In some cases, non-natural amino acids are incorporated into proteins by the pyrrolyl-tRNA pair. In some cases, PylRS is obtained from archaeal species such as methanogenic bacteria. In some cases, PylRS is obtained from Methanosarcina barkeri, Methanosarcina mazei, or Methanosarcina acetivorans. Exemplary UAAs that can be incorporated by the pyrrolyl-tRNA pair include, but are not limited to, amide and carbamate substituted lysines such as 2-amino-6-((R)-tetrahydrofuran-2-carboxamide)hexanoic acid, N-ε-D-prolyl-L-lysine, and N-ε-cyclopentyloxycarbonyl-L-lysine; N-ε-acryloyl-L-lysine; N-ε-[(1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl]-L-lysine; and N-ε-(1-methylcycloprop-2-ene-carboxamide)lysine.

[0351] In some cases, non-natural amino acids are incorporated into the proteins described herein by the synthetases disclosed in U.S. Patent No. 9,988,619 and U.S. Patent No. 9,938,516, the disclosures of each of which are incorporated herein by reference in their entireties. Exemplary UAAs that can be incorporated by such synthetases include para-methylazido-L-phenylalanine, aralkyl, heterocyclyl, heteroaralkyl non-natural amino acids, and the like. In some embodiments, such UAAs include pyridyl, pyrazinyl, pyrazolyl, triazolyl, oxazolyl, thiazolyl, thiophenyl, or other heterocycles. Such amino acids in some embodiments include other chemical groups that can conjugate to coupling partners such as azide, tetrazine, or water-soluble moieties. In some embodiments, such synthetases are expressed and used to incorporate UAAs into proteins in vivo. In some embodiments, such synthetases are used to incorporate UAAs into proteins using a cell-free translation system.

[0352] In some cases, the unnatural amino acid is incorporated into the proteins described herein by a naturally occurring synthetase. In some embodiments, the unnatural amino acid is incorporated into the protein by an organism that is auxotrophic for one or more amino acids. In some embodiments, the synthetase corresponding to the auxotrophic amino acid can charge the corresponding tRNA with the unnatural amino acid. In some embodiments, the unnatural amino acid is selenocysteine or a derivative thereof. In some embodiments, the unnatural amino acid is selenomethionine or a derivative thereof. In some embodiments, the unnatural amino acid is an aromatic amino acid, and the aromatic amino acid contains an aryl halide such as iodide. In an embodiment, the unnatural amino acid is structurally similar to the auxotrophic amino acid.

[0353] In some cases, the unnatural amino acid includes the unnatural amino acid shown in FIG. 8A.

[0354] In some cases, the unnatural amino acid includes a lysine or phenylalanine derivative or analog. In some cases, the unnatural amino acid includes a lysine derivative or lysine analog. In some cases, the unnatural amino acid includes pyrrolysine (Pyl). In some cases, the unnatural amino acid includes a phenylalanine derivative or phenylalanine analog. In some cases, the unnatural amino acid is the unnatural amino acid described in Wan et al., "Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool", Biocheim Biophys Aceta 1844(6):1059 - 4070(2014). In some cases, the unnatural amino acid includes the unnatural amino acids shown in FIGS. 8B and 8C.

[0355] In some embodiments, the unnatural amino acids include the unnatural amino acids shown in FIGS. 8D-8G (adopted from Table 1 of Dumas et al., Chemical Science 2015, 6, 50-69).

[0356] In some embodiments, the unnatural amino acids incorporated into the proteins described herein are disclosed in U.S. Patent No. 9,840,493; U.S. Patent No. 9,682,934; U.S. Patent Application Publication No. 2017 / 0260137; U.S. Patent No. 9,938,516; or U.S. Patent Application Publication No. 2018 / 0086734 (the disclosures of each are incorporated herein by reference in their entirety). Exemplary UAAs that can be incorporated by such synthetases include para-methylazido-L-phenylalanine, aralkyl, heterocyclyl, and heteroaralkyl, as well as unnatural amino acids of lysine derivatives. In some embodiments, such UAAs include pyridyl, pyrazinyl, pyrazolyl, triazolyl, oxazolyl, thiazolyl, thiophenyl, or other heterocycles. Such amino acids in some embodiments include other chemical groups that can bind to coupling partners such as azide, tetrazine, or a water-soluble moiety. In some embodiments, the UAA includes an azide bonded to an aromatic moiety via an alkyl linker. In some embodiments, the alkyl linker is a C1-C10 linker. In some embodiments, the UAA includes a tetrazine bonded to an aromatic moiety via an alkyl linker. In some embodiments, the UAA includes a tetrazine bonded to an aromatic moiety via an amino group. In some embodiments, the UAA includes a tetrazine bonded to an aromatic moiety via an alkylamino group. In some embodiments, the UAA includes an azide bonded to the terminal nitrogen of an amino acid side chain (e.g., N6 of a lysine derivative, or N5, N4, or N3 of a derivative containing a shorter alkyl side chain) via an alkyl chain. In some embodiments, the UAA includes a tetrazine bonded to the terminal nitrogen of an amino acid side chain via an alkyl chain. In some embodiments, the UAA includes an azide or tetrazine bonded to an amide via an alkyl linker. In some embodiments, the UAA is an azide- or tetrazine-containing carbamate or an amide of 3-aminoalanine, serine, lysine, or derivatives thereof. In some embodiments, such UAAs are incorporated into proteins in vivo.In some embodiments, such UAAs are incorporated into cell-free proteins.

[0357] Cell type In some embodiments, many types of cells / microorganisms are used, for example, for transformation or genetic engineering. In some embodiments, the cells are prokaryotic or eukaryotic cells. In some embodiments, the prokaryotic cells are bacterial cells. In some embodiments, the eukaryotic cells are fungal cells or unicellular protozoa. In some embodiments, the fungal cells are yeast cells. In other cases, the eukaryotic cells are cultured animal, plant, or human cells. Additionally, eukaryotic cells are present in organisms such as plants, multicellular fungi, or animals.

[0358] In some embodiments, the engineered microorganisms are single-celled organisms and can often divide and proliferate. As used herein, an "engineered microorganism" is a microorganism whose genetic material has been altered using genetic engineering techniques (i.e., recombinant DNA technology). The microorganism can include one or more of the following characteristics: aerobic, anaerobic, filamentous, non-filamentous, haploid, diploid, auxotrophic and / or non-auxotrophic. In certain embodiments, the engineered microorganism is a prokaryotic microorganism (e.g., a bacterium), and in certain embodiments, the engineered microorganism is a non-prokaryotic microorganism such as a eukaryotic microorganism. In some embodiments, the engineered microorganism is a eukaryotic microorganism (e.g., yeast, other fungi, amoeba). In some embodiments, the engineered microorganism is a fungus. In some embodiments, the engineered organism is yeast.

[0359] Any suitable yeast can be selected as a source of host microorganisms, engineered microorganisms, genetically modified organisms, or heterologous or modified polynucleotides. Yeasts include, but are not limited to, Yarrowia yeasts (e.g., Y. lipolytica (formerly classified as Candida lipolytica)), Candida yeasts (e.g., C. revkaufi, C. viswanathii, C. pulcherrima, C. tropicalis, C. utilis), Rhodotorula yeasts (e.g., R. glutinus, R. graminis), Rhodosporidium yeasts (e.g., R. toruloides), Saccharomyces yeasts (e.g., S. cerevisiae, S. bayanus, S. pastorianus, S. carlsbergensis), Cryptococcus yeasts, Trichosporon yeasts (e.g., T. pullans, T. cutaneum), Pichia yeasts (e.g., P. pastoris, K. phaffii), and Lipomyces yeasts (e.g., L. starkeyii, L. lipoferus). In some embodiments, suitable yeasts are yeasts of the genus Arachniotus, Aspergillus, Aureobasidium, Auxarthron, Blastomyces, Candida, Chrysosporuim, Chrysosporium, Debaryomyces, Coccidiodes, Cryptococcus, Gymnoascus, Hansenula, Histoplasma, Issatchenkia, Kluyveromyces, Lipomyces, Lssatchenkia, Microsporum, Myxotrichum, Myxozyma, Oidiodendron, Pachysolen, Penicillium, Pichia, Rhodosporidium, Rhodotorula, Rhodotorula, Saccharomyces, Schizosaccharomyces, Scopulariopsis, Sepedonium, Trichosporon, or Yarrowia.In some embodiments, suitable yeasts are Arachniotus flavoluteus, Aspergillus flavus, Aspergillus fumigatus, Aspergillus niger, Aureobasidium pullulans, Auxarthron thaxteri, Blastomyces dermatitidis, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xes. tobii, Chrysosporuim keratinophilum, Coccidiodes immitis, Cryptococcus albidus var. diffluens, Cryptococcus laurentii, Cryptococcus neofomans, Debaryomyces hansenii, Gymnoascus dugwayensis, Hansenula anomala, Histoplasma capsulatum, Issatchenkia occidentalis, Isstachenkia orientalis, Kluyveromyces lactis, Kluyveromyces marxianus, Kluyveromyces thermotolerans, Kluyveromyces waltii, Lipomyces lipoferus, Lipomyces starkeyii, Microsporum gypseum, Myxotrichum deflexum, Oidiodendron echinulatum, Pachysolen tannophilis, Penicillium notatum, Pichia anomala, Pichia pastoris, Pichia stipitis, Rhodosporidium toruloides, Rhodotorula glutinus, Rhodotorula graminis, Saccharomyces cerevisiae, Saccharomyces kluyveri, Schizosaccharomyces pombe, Scopulariopsis a yeast of the species of Acrimonium, Sepedonium chrysospermum, Trichosporon cutaneum, Trichosporon pullans, Yarrowia lipolytica, or Yarrowia lipolytica (formerly classified as Candida lipolytica). In some embodiments, the yeast is, but is not limited to, ATCC20362, ATCC8862, ATCC18944, ATCC20228, ATCC76982, and LGAM S(7)1 strain (Papanikolaou S. and Aggelis It is a Y. lipolytica strain including G., Bioresor. Technol. 82(1):43-9(2002)). In certain embodiments, the yeast is a Candida species (i.e., Candida genus) yeast. Any suitable Candida species may be used and / or may be genetically modified for the production of fatty dicarboxylic acids (e.g., octanedioic acid, decanedioic acid, dodecanedioic acid, tetradecanedioic acid, hexadecanedioic acid, octadecanedioic acid, eicosanedioic acid). In some embodiments, suitable Candida species include, but are not limited to, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xestobii, and any other Candida species yeast described herein. Non-limiting examples of Candida species strains include, but are not limited to, sAA001 (ATCC20336), sAA002 (ATCC20913), sAA003 (ATCC20962), sAA496 (U.S. Patent Application Publication No. 2012 / 0077252), sAA106 (U.S. Patent Application Publication No. 2012 / 0077252), SU-2 (ura3- / ura3-), H5343 (beta-oxidation block; U.S. Patent No. 5648247). Any suitable strain yeast derived from Candida species can be used as a parent strain for genetic recombination.

[0360] Since the genera, species, and strains of yeast are often very closely related in genetic content, it can be difficult to distinguish, classify, and / or name them. In some cases, C It can be difficult to distinguish, classify, and / or name strains of .lipolytica and Y. lipolytica, and in some cases, they may be considered the same organism. In some cases, it can be difficult to distinguish, classify, and / or name various strains of C. tropicalis and C. viswanathii (see, for example, Arie et al., J. Gen. Appl. Microbiol., 46, 257-262 (2000)). Some C. tropicalis and C. viswanathii strains obtained from ATCC and other commercial or academic sources may be considered equivalent and equally suitable for the embodiments described herein. In some embodiments, some parental strains of C. tropicalis and C. viswanathii are considered to differ only in name.

[0361] Any suitable fungus can be selected as a source of host microorganisms, engineered microorganisms, or heterologous polynucleotides. Non-limiting examples of fungi include, but are not limited to, Aspergillus fungi (e.g., A. parasiticus, A. nidulans), Thraustochytrium fungi, Schizochytrium fungi, and Rhizopus fungi (e.g., R. arrhizus, R. oryzae, R. nigricans). In some embodiments, the fungus is an A. parasiticus strain, including but not limited to the ATCC 24690 strain, and in certain embodiments, the fungus is an A. nidulans strain, including but not limited to the ATCC 38163 strain.

[0362] Any suitable prokaryote can be selected as a source of host microorganism, engineered microorganism, or heterologous polynucleotide. Gram-negative or Gram-positive bacteria can be selected. Examples of bacteria include, but are not limited to, Bacillus spp. (e.g., B. subtilis, B. megaterium), Acinetobacter spp., Norcardia spp., Xanthobacter spp., Escherichia spp. (e.g., E. coli (e.g., strains DH10B, Stbl2, DH5-alpha, DB3, DB3.1), DB4, DB5, JDP682 and ccdA-over (e.g., U.S. application No. 09 / 518,188)), Streptomyces spp., Erwinia spp., Klebsiella spp., Serratia spp. (e.g., S. marcessans), Pseudomonas spp. (e.g., P. aeruginosa), Salmonella spp. (e.g., S. typhimurium, S. typhi), Megasphaera spp. (e.g., Megasphaera elsdenii). Bacteria also include, but are not limited to, photosynthetic bacteria (e.g., green non-sulfur bacteria (e.g., Choroflexus spp. (e.g., C. aurantiacus), Chloronema spp. (e.g., C. gigateum)), green sulfur bacteria (e.g., Chlorobium spp. (e.g., C. limicola), Pelodictyon spp. (e.g., P. luteolum)), purple sulfur bacteria (e.g., Chromatium spp. (e.g., C. okenii)), and purple non-sulfur bacteria (e.g., Rhodospirillum spp. (e.g., R. rubrum), Rhodobacter spp. (e.g., R. sphaeroides, R. capsulatus), and Rhodomicrobium spp. (e.g., R. vanellii)).

[0363] Non-microbe-derived cells can be used as a source of host microbes, engineered microbes, or heterologous polynucleotides. Examples of such cells include, but are not limited to, insect cells (e.g., Drosophila (e.g., D. melanogaster), Spodoptera (e.g., S. frugiperda Sf9 or Sf21 cells ), and Trichoplusa (e.g., High-Five cells); nematode cells (e.g., C. elegans cells); avian cells; amphibian cells (e.g., Xenopus laevis cells); reptilian cells; mammalian cells (e.g., NIH3T3, 293, CHO, COS, VERO, C127, BHK, Per-C6, Bowes melanoma, and HeLa cells); and plant cells (e.g., Arabidopsis thaliana, Nicotania tabacum, Cuphea acinifolia, Cuphea aequipetala, Cuphea angustifolia, Cuphea appendiculata, Cuphea avigera, Cuphea avigera var. pulcherrima, Cuphea axilliflora, Cuphea bahiensis, Cuphea baillonis, Cuphea brachypoda, Cuphea bustamanta, Cuphea calcarata, Cuphea calophylla, Cuphea calophylla subsp. mesostemon, Cuphea carthagenensis, Cuphea circaeoides, Cuphea confertiflora, Cuphea cordata, Cuphea crassiflora, Cuphea cyanea, Cuphea decandra, Cuphea denticulata, Cuphea disperma, Cuphea epilobiifolia, Cuphea ericoides, Cuphea flava, Cuphea flavisetula, Cuphea fuchsiifolia, Cuphea gaumeri, Cuphea glutinosa, Cuphea heterophylla, Cuphea hookeriana, Cuphea hyssopifolia (Mexican - heather), Cuphea hyssopoides, Cuphea ignea, Cuphea ingrata, Cuphea jorullensis, Cuphea lanceolata, Cuphea linarioides, Cuphea llavea, Cuphea lophostoma, Cuphea lutea, Cuphea lutescens, Cuphea melanium, Cuphea melvilla, Cuphea micrantha, Cuphea micropetala, Cuphea mimuloides, Cuphea nitidula, Cuphea palustris, Cuphea parsonsia, Cuphea pascuorum, Cuphea paucipetala, Cuphea procumbens, Cuphea pseudosilene, Cuphea pseudovaccinium, Cuphea pulchra, Cuphea racemosa, Cuphea repens, Cuphea salicifolia, Cuphea salvadorensis, Cuphea schumannii, Cuphea sessiliflora, Cuphea sessilifolia, Cuphea setosa, Cuphea spectabilis, Cuphea spermacoce, Cuphea splendida, Cuphea splendida var. viridiflava, Cuphea strigulosa, Cuphea subuligera, Cuphea teleandra, Cuphea thymoides, Cuphea tolucana, Cuphea urens, Cuphea utriculosa, Cuphea such as viscosissima, Cuphea watsoniana, Cuphea wrightii, and Cuphea lanceolata).

[0364] Microorganisms or cells used as a source of a host organism or a heterologous polynucleotide are commercially available. The microorganisms and cells described herein, as well as other suitable microorganisms and cells, are available, for example, from Invitrogen Corporation (Carlsbad, CA), the American Type Culture Collection (Manassas, Virginia), and the Agricultural Research Culture re Collection (NRRL; Peoria, Illinois). The host microorganism and the engineered microorganism can be provided in any suitable form. For example, such microorganisms can be provided in liquid culture or solid culture (e.g., an agar-based medium), which can be a primary culture or can have been passaged one or more times (e.g., diluted and cultured). Microorganisms can also be provided in frozen or dried form (e.g., lyophilized). Microorganisms can be provided at any suitable concentration.

[0365] Polymerase A particularly useful function of a polymerase is to catalyze the polymerization of a nucleic acid strand using an existing nucleic acid as a template. Other useful functions are described elsewhere herein. Examples of useful polymerases include DNA polymerases and RNA polymerases.

[0366] The ability to improve the specificity, processability, or other characteristics of a polymerase with respect to an unnatural nucleic acid is highly desirable in a variety of situations where the incorporation of an unnatural nucleic acid is desired, including, for example, amplification, sequencing, labeling, detection, cloning, and the like.

[0367] In some examples, disclosed herein is, for example, a polymerase that incorporates a non-natural nucleic acid into an elongating template copy during DNA amplification. In some embodiments, the polymerase can be modified such that the active site of the polymerase is modified to reduce steric hindrance of the non-natural nucleic acid to the active site. In some embodiments, the polymerase can be modified to provide complementarity to one or more non-natural features of the non-natural nucleic acid. Such polymerases may be expressed or engineered intracellularly to stably incorporate UBP into cells. Accordingly, the present invention includes compositions comprising heterologous or recombinant polymerases, and methods of using the same.

[0368] The polymerase can be modified using methods related to protein engineering. For example, molecular modeling can be performed based on the crystal structure to identify positions in the polymerase where mutations can be made to alter the target activity. Residues identified as targets for substitution may be substituted with residues selected using energy minimization modeling, homology modeling, and / or conservative amino acid substitutions as described in Bordo et al., J Mol Biol 217:721-729 (1991) and Hayes et al., Proc Natl Acad Sci, USA 99:15926-15931 (2002), the disclosures of each of which are incorporated herein by reference in their entirety).

[0369] Any of a variety of polymerases can be used in the methods or compositions described herein, including, for example, protein-based enzymes isolated from biological systems and functional variants thereof. References to specific polymerases as exemplified below will be understood to include their functional variants unless otherwise indicated. In some embodiments, the polymerase is a wild-type polymerase. In some embodiments, the polymerase is a modified or mutant polymerase. In some embodiments, the polymerase can be a heterologous polymerase.

[0370] Polymerases having features for improving the entry of unnatural nucleic acids into the active site region and for cooperating with unnatural nucleotides in the active site region can also be used. In some embodiments, the modified polymerase has a modified nucleotide binding site.

[0371] In some embodiments, the modified polymerase has at least about 10%, 20%, 30%, 40%, 50%, 60%, 7 of the specificity of the wild-type polymerase for the unnatural nucleic acid It has specificity for unnatural nucleic acids that are 0%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99%. In some embodiments, the modified or wild-type polymerase has at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for unnatural nucleic acids that do not contain natural nucleic acids and / or modified sugars. In some embodiments, the modified or wild-type polymerase has at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for unnatural nucleic acids that do not contain natural nucleic acids and / or modified bases. In some embodiments, the modified or wild-type polymerase has at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for unnatural nucleic acids that contain triphosphates. For example, the modified or wild-type polymerase can have at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for unnatural nucleic acids that contain triphosphates and contain diphosphates or monophosphates or do not contain phosphates or combinations thereof.

[0372] In some embodiments, the modified or wild-type polymerase has relaxed specificity for non-natural nucleic acids. In some embodiments, the modified or wild-type polymerase has specificity for non-natural nucleic acids and at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity for natural nucleic acids of the wild-type polymerase for natural nucleic acids. In some embodiments, the modified or wild-type polymerase has specificity for non-natural nucleic acids containing modified sugars and at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity for natural nucleic acids of the wild-type polymerase for natural nucleic acids. In some embodiments, the modified or wild-type polymerase has specificity for non-natural nucleic acids containing modified bases and at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity for natural nucleic acids of the wild-type polymerase for natural nucleic acids.

[0373] Absence of exonuclease activity can be a wild-type characteristic or a characteristic conferred by a variant or engineered polymerase. For example, exo-minus Klenow fragment is a mutant version of the Klenow fragment that lacks 3’ to 5’ proofreading exonuclease activity.

[0374] The method of the present invention can be used to expand the substrate range of any DNA polymerase that lacks intrinsic 3→5’ exonuclease proofreading activity or has its 3→5’ exonuclease proofreading activity inactivated (e.g., through mutation). Examples of DNA polymerases include polA, polB (see, e.g., Parrel&Loeb, Nature Struc Biol 2001), polC, polD, polY, polX, and reverse transcriptase (RT), with polymerases that are evolutionarily and highly faithful (PCT / GB2004 / 004643) being preferred. In some embodiments, the modified or wild-type polymerase substantially lacks 3’→5’ proofreading exonuclease activity. In some embodiments, the modified or wild-type polymerase substantially lacks 3’→5’ proofreading exonuclease activity against non-natural nucleic acids. In some embodiments, the modified or wild-type polymerase has 3’→5’ proofreading exonuclease activity. In some embodiments, the modified or wild-type polymerase has 3’→5’ proofreading exonuclease activity against natural nucleic acids and substantially lacks 3’→5’ proofreading exonuclease activity against non-natural nucleic acids.

[0375] In some embodiments, the modified polymerase has 3'→5' proofreading exonuclease activity that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5% or 99.99% of the proofreading exonuclease activity of the wild-type polymerase. In some embodiments, the modified polymerase has 3'→5' proofreading exonuclease activity for non-natural nucleic acids that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the proofreading exonuclease activity of the wild-type polymerase for natural nucleic acids. In some embodiments, the modified polymerase has 3'→5' proofreading exonuclease activity for non-natural nucleic acids and 3'→5' proofreading exonuclease activity for natural nucleic acids that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the proofreading exonuclease activity of the wild-type polymerase for natural nucleic acids. In some embodiments, the modified polymerase has 3'→5' proofreading exonuclease activity for natural nucleic acids that is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the proofreading exonuclease activity of the wild-type polymerase for natural nucleic acids.

[0376] In some embodiments, the polymerases are characterized by their dissociation rates from nucleic acids. In some embodiments, the polymerases have a relatively low dissociation rate with respect to one or more natural and non-natural nucleic acids. In some embodiments, the polymerases have a relatively high dissociation rate with respect to one or more natural and non-natural nucleic acids. This dissociation rate is an activity of the polymerase and can be adjusted to adjust the reaction rate in the methods described herein.

[0377] In some embodiments, polymerases are characterized according to their fidelity when used with certain natural and / or non-natural nucleic acids or collections of natural and / or non-natural nucleic acids. Fidelity generally refers to the accuracy with which a polymerase incorporates the correct nucleic acid into the growing nucleic acid strand when making a copy of a nucleic acid template. The fidelity of a DNA polymerase can be measured as the ratio of correct to incorrect incorporation of natural and non-natural nucleic acids when the natural and non-natural nucleic acids are present, for example, at equal concentrations and compete for strand synthesis at the same site of the polymerase-strand-template nucleic acid binary complex. The fidelity of a DNA polymerase can be calculated as the ratio (kcat / Km) for natural and non-natural nucleic acids and (kcat / Km) for incorrect natural and non-natural nucleic acids; where kcat and Km are Michaelis-Menten parameters in steady-state enzyme kinetics (Fersht, A.R. (1985) Enzyme Structure and Mechanism, 2nd ed., p350, W.H. Freeman & Co., New York. (incorporated herein by reference)). In some embodiments, the polymerase has a fidelity value of at least about 100, 1000, 10,000, 100,000, or 1 × 106, with or without proofreading activity.

[0378] In some embodiments, polymerases from natural sources or variants thereof Screening is performed using an assay that detects the incorporation of unnatural nucleic acids having a specific structure. In one example, the polymerase may be screened for its ability to incorporate an unnatural nucleic acid or UBP (e.g., d5SICSTP, dCNMOTP, dTPT3TP, dNaMTP, dCNMOTP-dTPT3TP, or d5SICSTP-dNaMTP UBP). A polymerase that exhibits modified properties with respect to the unnatural nucleic acid, e.g., a heterologous polymerase, may be used as compared to the wild-type polymerase. For example, the modified properties may be, e.g., Km, kcat, Vmax, polymerase processivity in the presence of the unnatural nucleic acid (or naturally occurring nucleotides), average template read length by the polymerase in the presence of the unnatural nucleic acid, polymerase specificity for the unnatural nucleic acid, binding rate of the unnatural nucleic acid, release rate of the product (pyrophosphate, triphosphate, etc.), branching rate, or any combination thereof. In one embodiment, the modified properties are a decreased Km for the unnatural nucleic acid and / or an increased kcat / Km or Vmax / Km for the unnatural nucleic acid. Similarly, the polymerase optionally has an increased binding rate of the unnatural nucleic acid, an increased rate of product release, and / or a decreased branching rate as compared to the wild-type polymerase.

[0379] Simultaneously, the polymerase may incorporate natural nucleic acids, e.g., A, C, G, and T, into the elongating nucleic acid copy. For example, the polymerase optionally has a specific activity for natural nucleic acids that is at least about 5% higher (e.g., 5%, 10%, 25%, 50%, 75%, 100% or more) than the corresponding wild-type polymerase, and processivity by the natural nucleic acid in the presence of the template that is at least 5% higher (e.g., 5%, 10%, 25%, 50%, 75%, 100% or more) than the wild-type polymerase in the presence of the natural nucleic acid. Optionally, the polymerase exhibits a kcat / Km or Vmax / Km for naturally occurring nucleotides that is at least about 5% higher (e.g., about 5%, 10%, 25%, 50%, 75% or 100% or more) than the wild-type polymerase.

[0380] The polymerases used herein that may have the ability to incorporate non-natural nucleic acids of a specific structure may also be generated using a directed evolution approach. Nucleic acid synthesis assays may be used to screen for polymerase variants that are specific for any of a variety of non-natural nucleic acids. For example, polymerase variants may be screened for their ability to incorporate unnatural nucleoside triphosphates opposite unnatural nucleotides on a DNA template (e.g., dTPT3TP opposite dCNMO, dCNMOTP opposite dTPT3, NaMTP opposite dTPT3, or TAT1TP opposite dCNMO or dNaM). In some embodiments, such assays are, for example, in vitro assays using recombinant polymerase variants. In some embodiments, such assays are, for example, in vivo assays that express polymerase variants in cells. Using such directed evolution techniques, variants of any suitable polymerase may be screened for activity against any of the non-natural nucleic acids described herein. Optionally, the polymerases used herein have the ability to incorporate unnatural ribonucleotides into nucleic acids such as RNA. For example, NaM or TAT1 ribonucleotides are incorporated into nucleic acids using the polymerases described herein.

[0381] The modified polymerase of the described composition can optionally be a modified and / or recombinant Φ29-type DNA polymerase. Optionally, the polymerase can be a modified and / or recombinant Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, or L17 polymerase.

[0382] The modified polymerase of the described composition can optionally be a modified and / or recombinant prokaryotic DNA polymerase, for example, DNA polymerase II (Pol II), DN It may be polymerase III (Pol III), DNA polymerase IV (Pol IV), or DNA polymerase V (Pol V). In some embodiments, the modified polymerase includes a polymerase that mediates DNA synthesis across non-instructive damaged nucleotides. In some embodiments, the genes encoding Pol I, Pol II (polB), Poll IV (dinB), and / or Pol V (umuCD) are constitutively expressed or overexpressed in engineered cells or SSOs. In some embodiments, the increased expression or overexpression of Pol II contributes to the increased retention of unnatural base pairs (UBPs) in engineered cells or SSOs.

[0383] Nucleic acid polymerases that are generally useful in the present invention include DNA polymerases, RNA polymerases, reverse transcriptases, and their mutant or modified forms. DNA polymerases and their properties are described in detail, inter alia, in DNA Replication, 2nd Edition, Kornberg and Baker, W.H. Freeman, New York, N.Y. (1991). Known conventional DNA polymerases useful in the present invention include, but are not limited to, Pyrocococcus furiosus (Pfu) DNA polymerase (Lundberg et al., 1991, Gene, 108:1, Stratagene), Pyrocococcus woesei (Pwo) DNA polymerase (Hinnisdaels et al., 1996, Biotechniques, 20:186-8, Boehringer Mannheim), Thermus Thermophilus (Tth) DNA polymerase (Myers and Gelfand 1991, Biochemistry 30:7661), Bacillus stearothermophilus DNA polymerase (Stenesh and McGowan, 1977, Biochim Biophys Acta 475:32), Thermococcus litoralis (TIi) DNA polymerase (also called Vent™ DNA polymerase, Cariello et al., 1991, Polynucleotides Res, 19:4193, New England Biolabs), 9°Nm™ DNA polymerase (New England Biolabs), Stoffel fragment, Thermo Sequenase® (Amersham Pharmacia Biotech UK), Therminator™ (New England Biolabs), Thermotoga maritima (Tma) DNA polymerase (Diaz and Sabino, 1998 Braz J Med.Res, 31:1239), Thermus aquaticus (Taq) DNA polymerase (Chien et al., 1976, J.Bacteoriol, 127:1550), DNA polymerase, Pyrococcus kodakaraensis KOD DNA polymerase (Takagi et al., 1997, Appl.Environ.Microbiol. 63:4504), JDF-3 DNA polymerase (thermococcus sp. JDF-3, from patent application WO0132887), Pyrococcus GB-D (PGB-D) DNA polymerase (also called Deep Vent™ DNA polymerase, Juncosa-Ginesta et al., 1994, Biotechniques, 16:820, New England Biolabs), Ultma DNA polymerase (from the thermophilic bacterium Thermotoga maritima; Diaz and Sabino, 1998 Braz J. Med. Res, 31:1239; PE Applied Biosystems), Tgo DNA polymerase (thermococcus gorgonarius, Roche from Molecular Biochemicals), DNA polymerase I of E. coli (Lecomte and Doubleday, 1983, Polynucleotides Res. 11:7505), T7 DNA polymerase (Nordstrom et al., 1981, J Biol. Chem. 256:3112), and archaeal DP1I / DP2 DNA polymerase II (Cann et al., 1998, Proc. Natl. Acad. Sci. USA 95:14250). Mesophilic polymerases and Both thermophilic polymerases are contemplated. Examples of thermophilic DNA polymerases include, but are not limited to, ThermoSequenase®, 9°Nm™, Therminator™, Taq, Tne, Tma, Pfu, TfI, Tth, Tli, Stoffel fragment, Vent® and DeepVent® DNA polymerases, KOD DNA polymerase, Tgo, JDF-3, as well as variants, mutants and derivatives thereof. Polymerases that are 3'exonuclease-deficient mutants are also contemplated. Examples of reverse transcriptases useful in the present invention include, but are not limited to, reverse transcriptases derived from HIV, HTLV-I, HTLV-II, FeLV, FIV, SIV, AMV, MMTV, MoMuLV and other retroviruses (see Levin, Cell 88:5-8 (1997); Verma, Biochim Biophys Acta. 473:1-38 (1977); Wu et al., CRC Crit Rev Biochem. 3:289-347 (1975)). Further examples of polymerases include, but are not limited to, 9°N DNA polymerase, Taq DNA polymerase, Phusion® DNA polymerase, Pfu DNA polymerase, RB69 DNA polymerase, KOD DNA polymerase, and VentR® DNA polymerase. Gardner et al. (2004) "Comparative Kinetics of Nucleotide Analog Incorporation by Vent DNA Polymerase" (J. Biol. Chem., 279(12), 11834-11842; Gardner and Jack "Determinants of nucleotide sugar recognition in an archaeon DNA polymerase" Nucleic Acids Research, 27(12)2545-2553). Polymerases isolated from non-thermophilic organisms can be heat-inactivated.Examples include DNA polymerases from phages. It is understood that polymerases from any of a variety of sources can be modified to increase or decrease their resistance to high temperature conditions. In some embodiments, the polymerase can be thermophilic. In some embodiments, the thermophilic polymerase can be heat-inactivatable. Thermophilic polymerases are typically useful in thermal cycling conditions such as high temperature conditions or the conditions used in polymerase chain reaction (PCR) techniques.

[0384] In some embodiments, the polymerase includes Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, L17, ThermoSequenase®, 9°Nm™, Therminator™ DNA polymerase, Tne, Tma, TfI, Tth, TIi, Stoffel fragment, Vent® and DeepVent® DNA polymerases, KOD DNA polymerase, Tgo, JDF-3, Pfu, Taq, T7 DNA polymerase, T7 RNA polymerase, PGB-D, UlTma DNA polymerase, E. coli DNA polymerase I, E. coli DNA polymerase III, archaeal DP1I / DP2 DNA polymerase II, 9°N DNA polymerase, Taq DNA polymerase, Phusion® DNA polymerase, Pfu DNA polymerase, SP6 RNA polymerase, RB69 DNA polymerase, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, or SuperScript® III reverse transcriptase.

[0385] In some embodiments, the polymerase is DNA polymerase I (or the Klenow fragment), Vent polymerase, Phusion® DNA polymerase, KOD DNA polymerase, Taq polymerase, T7 DNA polymerase, T7 RNA polymerase, Therminator™ DNA polymerase, POLB polymerase, SP6 RNA polymerase, E. coli DNA polymerase I, E. coli DNA polymerase III, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, or SuperScript® III reverse transcriptase.

[0386] Nucleotide transporter Nucleotide transporters (NTs) are a group of membrane transport proteins that facilitate the movement of nucleotide substrates across cell membranes and vesicles. In some embodiments, there are two types of NTs, concentrative nucleoside transporters and equilibrative nucleoside transporters. Optionally, NTs also include organic anion transporters (OATs) and organic cation transporters (OCTs). Optionally, the nucleotide transporter is a nucleoside triphosphate transporter (NTT).

[0387] In some embodiments, the nucleoside triphosphate transporter (NTT) is derived from bacteria, plants, or algae. In some embodiments, the nucleotide nucleoside triphosphate transporters are TpNTT1, TpNTT2, TpNTT3, TpNTT4, TpNTT5, TpNTT6, TpNTT7, TpNTT8 (T. pseudonana), PtNTT1, PtNTT2, PtNTT3, PtNTT4, PtNTT5, PtNTT6 (P. tricornutum), GsNTT (Galdieria sulphuraria), AtNTT1, AtNTT2 (Arabidopsis thaliana), CtNTT1, CtNTT2 (Chlamydia trachomatis), PamNTT1, PamNTT2 (Protochlamydia amoebophila), CcNTT (Caedibacter caryophilus), or RpNTT1 (Rickettsia prowazekii).

[0388] In some embodiments, the NTT is CNT1, CNT2, CNT3, ENT1, ENT2, OAT1, OAT3, or OCT1.

[0389] In some embodiments, the NTT introduces a non-natural nucleic acid into an organism, such as a cell. In some embodiments, the NTT may be modified such that the nucleotide binding site of the NTT is modified to reduce steric hindrance of the non-natural nucleic acid to the nucleotide binding site. In some embodiments, the NTT may be modified to provide an increase in the interaction of one or more natural or non-natural features of the non-natural nucleic acid. Such NTTs may be expressed or engineered intracellularly to stably introduce the UBP into the cell. Accordingly, the present invention includes compositions comprising heterologous or recombinant NTTs and methods of using the same.

[0390] The NTT can be modified using methods related to protein engineering. For example, molecular modeling can be performed based on the crystal structure to identify the position of the NTT where mutations can be made to alter the target activity or binding site. Residues identified as targets for substitution can be substituted with the selected residues using energy minimization modeling, homology modeling, and / or conservative amino acid substitutions as described in Bordo et al., J Mol Biol 217:721-729 (1991) and Hayes et al., Proc Natl Acad Sci, USA 99:15926-15931 (2002), the disclosures of each of which are incorporated herein by reference in their entirety.

[0391] Any of the various NTTs may be used in the methods or compositions described herein, including, for example, protein-based enzymes isolated from biological systems and functional variants thereof. Reference to a particular NTT as exemplified below is understood to include its functional variants unless otherwise indicated. In some embodiments, the NTT is a wild-type NTT. In some embodiments, the NTT is a modified or mutant NTT.

[0392] NTT with characteristics for improving the entry of unnatural nucleic acids into cells and cooperating with unnatural nucleotides in the nucleotide-binding region may also be used. In some embodiments, the modified NTT has a modified nucleotide-binding site. In some embodiments, the modified or wild-type NTT has relaxed specificity for unnatural nucleic acids. For example, the NTT optionally exhibits a ratio activity of incorporation into unnatural nucleotides that is at least about 0.1% higher (e.g., about 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.1%, 1.2%, 1.5%, 1.8%, 2%, 3%, 4%, 5%, 10%, 25%, 50%, 75%, 100% or more) than the corresponding wild-type NTT. Optionally, the NTT exhibits a kcat / Km or Vmax / Km for unnatural nucleotides that is at least about 0.1% higher (e.g., about 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.1%, 1.2%, 1.5%, 1.8%, 2%, 3%, 4%, 5%, 10%, 25%, 50%, 75%, or 100% or more) than the wild-type NTT.

[0393] NTTs can be characterized according to their affinity for triphosphates (i.e., Km) and / or incorporation rate (i.e., Vmax). In some embodiments, the NTT has a relative Km or Vmax for one or more natural and unnatural triphosphates. In some embodiments, the NTT has a relatively high Km or Vmax for one or more natural and unnatural triphosphates.

[0394] NTTs derived from natural sources or variants thereof can be screened using assays that detect the amount of triphosphate (either mass spectrometry or radioactivity if the triphosphate is appropriately labeled). In one example, NTTs may be screened for their ability to incorporate unnatural triphosphates (e.g., dTPT3TP, dCNMOTP, d5SICSTP, dNaMTP, NaMTP, and / or TPT1TP). Heterologous NTTs that exhibit modified properties with respect to unnatural nucleic acids compared to, for example, wild-type NTTs may be used. For example, the modified properties can be, for example, Km, kcat, Vmax in the case of incorporation of triphosphates. In one embodiment, the modified property is a decreased Km in the case of unnatural triphosphates and / or an increased kcat / Km or Vmax / Km in the case of unnatural triphosphates. Similarly, NTTs optionally have an increased rate of binding of unnatural triphosphates, an increased rate of intracellular release, and / or an increased rate of cell entry compared to wild-type NTTs.

[0395] At the same time, NTTs can introduce natural triphosphates, such as dATP, dCTP, dGTP, dTTP, ATP, CTP, GTP, and / or TTP, into cells. Optionally, in some cases, NTTs exhibit an optional specific activity for the introduction of natural nucleic acids that can support replication and transcription. In some embodiments, NTTs optionally exhibit a kcat / Km or Vmax / Km for natural nucleic acids that can support replication and transcription.

[0396] The NTTs used herein that can have the ability to incorporate unnatural triphosphates of a particular structure can also be generated using a directed evolution approach. Nucleic acid synthesis assays can be used to screen for NTT variants that are specific for any of a variety of unnatural triphosphates. For example, NTT variants can be screened for their ability to incorporate unnatural triphosphates (e.g., d5SICSTP, dNaMTP, dCNMOTP, dTPT3TP, NaMTP, and / or TPT1TP). In some embodiments, such assays are in vitro assays using, for example, recombinant NTT variants. In one embodiment, such an assay is, for example, an in vivo assay that expresses the NTT variant in cells. Using such techniques, any suitable variant of NTT can be screened for activity against any of the unnatural triphosphates described herein.

[0397] Nucleic Acid Reagents and Tools The nucleotides and / or nucleic acid reagents (or polynucleotides) for use with the methods, cells, or engineered microorganisms described herein include one or more open reading frames (ORFs) with or without unnatural nucleotides. The ORF can be derived from any suitable source, sometimes genomic DNA, mRNA, reverse transcribed RNA or complementary DNA (cDNA), or a nucleic acid library containing one or more of the foregoing, and is from any biological species containing the nucleic acid sequence of interest, the protein of interest, or the activity of interest. Non-limiting examples of organisms from which the ORF can be obtained include, for example, bacteria, yeast, fungi, humans, insects, nematodes, cows, horses, dogs, cats, rats or mice. In some embodiments, the nucleotides and / or nucleic acid reagents or other reagents described herein are isolated or purified. ORFs containing unnatural nucleotides can be generated by published in vitro methods. In some cases, the nucleotides or nucleic acid reagents contain unnatural nucleobases.

[0398] The nucleic acid reagent may include a nucleotide sequence adjacent to the ORF that is translated together with the ORF and encodes an amino acid tag. The nucleotide sequence encoding the tag is located 3' and / or 5' of the ORF in the nucleic acid reagent, thereby encoding the tag at the C-terminus or N-terminus of the protein or peptide encoded by the ORF. Any tag that does not inactivate in vitro transcription and / or translation may be utilized and may be appropriately selected by the skilled person. The tag may facilitate the isolation and / or purification of the desired ORF product from the culture or fermentation medium. In some cases, a library of nucleic acid reagents is used with the methods and compositions described herein. For example, a library of at least 100, 1000, 2000, 5000, 10,000, or more than 50,000 unique polynucleotides is present in the library, and each polynucleotide contains at least one unnatural nucleobase.

[0399] Nucleic acids or nucleic acid reagents, whether containing or not containing unnatural nucleotides, may contain specific elements, such as regulatory elements that are often selected according to the purpose of use of the nucleic acid. Any of the following elements may be included in or excluded from the nucleic acid reagent. For example, the nucleic acid reagent may contain one or more or all of the following nucleotide elements: one or more promoter elements, one or more 5' untranslated regions (5' UTRs), one or more regions ( "insertion elements") into which the target nucleotide sequence can be inserted, one or more target nucleotide sequences, one or more 3' untranslated regions (3' UTRs), and one or more selection elements. The nucleic acid reagent may be provided with one or more of such elements, and other elements may be inserted into the nucleic acid before the nucleic acid is introduced into the desired organism. In some embodiments, the provided nucleic acid reagent contains a promoter, 5' UTR, any 3' UTR, and an insertion element into which the target nucleotide sequence is inserted (i.e., cloned) into the nucleic acid reagent. In certain embodiments, the provided nucleic acid reagent contains a promoter, an insertion element, and any 3' UTR, and the 5' UTR / target nucleotide sequence is inserted together with any 3' UTR. The elements may be arranged in any order suitable for expression in the selected expression system (e.g., expression in the selected organism, or expression in a cell-free system, for example), and in some embodiments, the nucleic acid reagent contains the following elements in the 5'→3' direction: (1) a promoter element, 5' UTR, and insertion element; (2) a promoter element, 5' UTR, and target nucleotide sequence; (3) a promoter element, 5' UTR, insertion element and 3' UTR; and (4) a promoter -ter element, 5' UTR, target nucleotide sequence and 3' UTR. In some embodiments, the UTR can be fully natural or optimized to alter or increase the transcription or translation of the ORF, including unnatural nucleotides.

[0400] Nucleic acid reagents, such as expression cassettes and / or expression vectors, may contain various regulatory elements including promoters, enhancers, translation initiation sequences, transcription termination sequences, and other elements. A "promoter" is generally a DNA sequence that functions when located at a relatively fixed position with respect to the transcription start site. For example, a promoter can be upstream of a nucleotide triphosphate transporter nucleic acid segment. A "promoter" contains the core elements necessary for the basic interaction of RNA polymerase and transcription factors and may also include upstream elements and response elements. An "enhancer" generally refers to a DNA sequence that functions at a non-fixed distance from the transcription start site and can be either 5' or 3' to the transcription unit. Furthermore, enhancers may be present within introns and within the coding sequence itself. They are usually between 10 and 300 in length and function in cis. Enhancers function to increase transcription from nearby promoters. Like promoters, enhancers often contain response elements that mediate the regulation of transcription. Enhancers often determine the regulation of expression and can be used to alter or optimize ORF expression, such as those that are completely natural or contain non-natural nucleotides.

[0401] As described above, the nucleic acid reagent may also include one or more 5’UTRs and one or more 3’UTRs. For example, expression vectors used in eukaryotic host cells (e.g., yeast, fungi, insects, plants, animals, humans or nucleated cells) and prokaryotic host cells (e.g., viruses, bacteria) may include sequences that signal for termination of transcription, which can affect the expression of mRNA. These regions can be transcribed as polyadenylation segments of the untranslated portion of the mRNA encoding the tissue factor protein. The 3’’ untranslated region also includes the transcription termination site. In some preferred embodiments, the transcription unit includes a polyadenylation region. One advantage of this region is that it increases the likelihood that the transcribed unit will be processed and transported like mRNA. The identification and use of polyadenylation signals in expression constructs are well established. In some preferred embodiments, homologous polyadenylation signals can be used in transgene constructs.

[0402] The 5’UTR may contain one or more elements endogenous to the nucleotide sequence from which it is derived and sometimes contains one or more exogenous elements. The 5’UTR may be derived from any suitable nucleic acid such as genomic DNA, plasmid DNA, RNA or mRNA, for example, from any suitable organism (e.g., virus, bacterium, yeast, fungus, plant, insect or mammal). The skilled person may select suitable elements for the 5’UTR based on the selected expression system (e.g., expression in a selected organism or, for example, expression in a cell-free system). The 5’UTR may sometimes contain one or more of the following elements known to the skilled person: enhancer sequences (e.g., for transcription or translation), transcription start sites, transcription factor binding sites, translation regulatory sites, translation start sites, translation factor binding sites, accessory protein binding sites, feedback regulator binding sites, Pribnow box, TATA box, -35 element, E-box (helix-loop-helix binding element), ribosome binding sites, replicon, internal ribosome entry site (IRES), silencer elements, etc. In some embodiments, the promoter element may be separated such that all 5’UTR elements necessary for appropriate conditional regulation are contained within the promoter element fragment or within a functional partial sequence of the promoter element fragment.

[0403] The 5’UTR in a nucleic acid reagent may contain a translation enhancer nucleotide sequence. Translation Enhancer nucleotide sequences are often located between the promoter and the target nucleotide sequence of a nucleic acid reagent. Translational enhancer sequences often bind to ribosomes and may be 18S rRNA-binding ribonucleotide sequences (i.e., 40S ribosome-binding sequences) or may be internal ribosome entry sequences (IRESs). IRESs generally form an RNA scaffold with an accurately arranged RNA tertiary structure that contacts the 40S ribosome subunit through a number of specific intermolecular interactions. Examples of ribosome enhancer sequences are known and can be identified by those skilled in the art (e.g., Mignone et al., Nucleic Acids Research 33:D141-D146 (2005); Paulous et al., Nucleic Acids Research 31:722-733 (2003); Akbergenov et al., Nucleic Acids Research 32:239-247 (2004); Mignone et al., Genome Biology 3(3):reviews0004.1-0001.10 (2002); Gallie, Nucleic Acids Research 30:3401-3411 (2002); Shaloiko et al., DOI:10.1002 / bit.20267; and Gallie et al., Nucleic Acids Research 15:3257-3273 (1987); the disclosures of each are incorporated herein by reference in their entirety).

[0404] The translation enhancer sequence may be a eukaryotic sequence such as a Kozak consensus sequence or other sequences (e.g., Hydra polyp sequence, GenBank accession number U07128). The translation enhancer sequence may be a prokaryotic sequence such as a Shine-Dalgarno consensus sequence. In certain embodiments, the translation enhancer sequence is a viral nucleotide sequence. The translation enhancer sequence may be derived from, for example, the 5’ UTR of plant viruses such as tobacco mosaic virus (TMV), alfalfa mosaic virus (AMV); tobacco etch virus (ETV); potato virus Y (PVY); turnip mosaic (poty) virus and pea seed-borne mosaic virus. In certain embodiments, an omega sequence of approximately 67 bases in length from TMV is included in the nucleic acid reagent as a translation enhancer sequence (e.g., lacking guanosine nucleotides and including a 25 nucleotide long poly(CAA) central region).

[0405] The 3’UTR may contain one or more elements endogenous to the nucleotide sequence from which it is derived and may also contain one or more exogenous elements. The 3’UTR can be derived from any suitable nucleic acid such as genomic DNA, plasmid DNA, RNA or mRNA, for example, from any suitable organism (e.g., virus, bacterium, yeast, fungus, plant, insect or mammal). A skilled person can select suitable elements for the 3’UTR based on the selected expression system (e.g., expression in the selected organism, etc.). The 3’UTR may contain one or more of the following elements known to the skilled person: transcriptional regulatory site, transcription start site, transcription termination site, transcription factor binding site, translational regulatory site, translation termination site, translation start site, translation factor binding site, ribosome binding site, replicon, enhancer element, silencer element and polyadenosine tail. The 3’UTR often contains a polyadenosine tail and may not, and when a polyadenosine tail is present, one or more adenosine moieties may be added or removed (e.g., about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45 or about 50 adenosine moieties may be added or removed).

[0406] In some embodiments, modifications to the 5’UTR and / or 3’UTR are used to alter (e.g., increase, add, decrease, or substantially eliminate) the activity of a promoter. The alteration of promoter activity can then change the activity (e.g., enzymatic activity) of a peptide, polypeptide, or protein by a change in the transcription of the nucleotide sequence of interest from an operably linked promoter element containing the modified 5’ or 3’UTR. For example, a microorganism may be genetically engineered to express a nucleic acid reagent comprising a modified 5' or 3' UTR that can confer a novel activity (e.g., an activity not normally found in the host organism), or in certain embodiments, increase the expression of an existing activity by increasing transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest (e.g., a homologous or heterologous nucleotide sequence of interest). In some embodiments, a microorganism may be genetically engineered to express a nucleic acid reagent comprising a modified 5' or 3' UTR that can decrease the expression of an activity, in certain embodiments, by decreasing or substantially eliminating transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest.

[0407] The expression of the nucleotide triphosphate transporter from the expression cassette or expression vector can be controlled by any promoter that can be expressed in prokaryotic or eukaryotic cells. Promoter elements are usually required for DNA synthesis and / or RNA synthesis. A promoter element often includes a region of DNA that can promote the transcription of a specific gene by providing a starting site for the synthesis of RNA corresponding to a gene. Promoters are generally located near the genes they regulate, upstream of the gene (e.g., 5' of the gene), and in some embodiments, on the same DNA strand as the sense strand of the gene. In some embodiments, the promoter element may be isolated from a gene or organism and inserted in functional association with a polynucleotide sequence to enable altered and / or regulated expression. Non-native promoters used for the expression of nucleic acids (e.g., promoters that are not normally associated with a given nucleic acid sequence) are often referred to as heterologous promoters. In certain embodiments, a heterologous promoter and / or 5' UTR can be inserted in functional association with a polynucleotide encoding a polypeptide having the desired activity as described herein. As used herein with respect to a promoter, the terms "operably linked" and "functionally associated" refer to the relationship between a coding sequence and a promoter element. A promoter is operably linked to, or functionally associated with, a coding sequence if the expression from the coding sequence via transcription is regulated or controlled by the promoter element. The terms "operably linked" and "functionally associated" are used interchangeably herein with respect to a promoter element.

[0408] Promoters often interact with RNA polymerase. Polymerase is an enzyme that catalyzes the synthesis of nucleic acids using existing nucleic acid reagents. When the template is a DNA template, an RNA molecule is transcribed before the protein is synthesized. Enzymes having polymerase activity suitable for use in the methods of the present invention include any polymerase that is active in a selected system using a template selected for synthesizing a protein. In some embodiments, a promoter (e.g., a heterologous promoter), also referred to herein as a promoter element, can be operably linked to a nucleotide sequence or open reading frame (ORF). Transcription from the promoter element can catalyze the synthesis of RNA corresponding to the nucleotide sequence or ORF sequence operably linked to the promoter, which in turn results in the synthesis of the desired peptide, polypeptide, or protein.

[0409] Promoter elements may exhibit responsiveness to regulatory control. Promoter elements may also be regulated by a selective agent. That is, transcription from the promoter element can be turned on, off, upregulated, or downregulated in response to changes in environmental, nutritional, or internal conditions or signals (e.g., heat-inducible promoters, light-regulated promoters, feedback-regulated promoters, hormone-responsive promoters, tissue-specific promoters, oxygen- and pH-responsive promoters, promoters responsive to a selective agent (e.g., kanamycin), etc.). Promoters that are affected by environmental, nutritional, or internal signals are often affected by signals (direct or indirect) that bind to the promoter or near it and increase or decrease the expression of the target sequence under specific conditions. As with all methods disclosed herein, the inclusion of natural or modified promoters can be used to alter or optimize the expression of a completely natural ORF (e.g., NTT or aaRS) or an ORF containing non-natural nucleotides (e.g., mRNA or tRNA). As with all methods disclosed herein, the inclusion of natural or modified promoters can be used to alter or optimize the expression of a completely natural ORF (e.g., NTT or aaRS) or an ORF containing non-natural nucleotides (e.g., mRNA or tRNA).

[0410] Non-limiting examples of selective agents or regulatory agents that affect transcription from promoter elements used in the embodiments described herein include, but are not limited to, the following: (1) nucleic acid segments encoding products that confer resistance to otherwise toxic compounds (e.g., antibiotics); (2) nucleic acid segments encoding products that are otherwise lacking in the recipient cell (e.g., essential products, tRNA genes, auxotrophic markers); (3) nucleic acid segments encoding products that suppress the activity of gene products; (4) nucleic acid segments encoding products that can be readily identified (e.g., phenotypic markers such as antibiotics (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and cell surface proteins); (5) nucleic acid segments that bind to products that are otherwise harmful to cell survival and / or function; (6) nucleic acid segments that inhibit the activity of any of the nucleic acid segments described in items 1 to 5 above (e.g., antisense oligonucleotides); (7) nucleic acid segments that bind to products that modify substrates (e.g., restriction endonucleases); (8) nucleic acid segments that can be used to isolate or identify desired molecules (e.g., specific protein binding sites); (9) nucleic acid segments encoding specific nucleotide sequences that cannot otherwise function (e.g., for PCR amplification of subpopulations of molecules); (10) nucleic acid segments that directly or indirectly confer resistance or sensitivity to specific compounds when absent; (11) nucleic acid segments encoding products that convert a compound that is toxic or relatively non-toxic in the recipient cell into a toxic compound (e.g., herpes simplex thymidine kinase, cytosine deaminase); (12) nucleic acid segments that inhibit the replication, segregation, or heritability of the nucleic acid molecules containing them; (13) nucleic acid segments encoding conditional replication functions, e.g., replication in a specific host or host cell line or under specific environmental conditions (e.g., temperature, nutrient conditions, etc.); and / or (14) nucleic acids encoding one or more mRNAs or tRNAs containing non-natural nucleotides.In some embodiments, a modulator or a selective agent may be added to alter the existing growth conditions to which the organism is subjected (e.g., growth in liquid culture, growth in a fermenter, growth on a solid nutrient plate, etc.).

[0411] In some embodiments, the regulation of promoter elements can be used to alter (e.g., increase, add, decrease, or substantially eliminate) the activity of a peptide, polypeptide, or protein (e.g., enzymatic activity, etc.). For example, a microorganism can be engineered by genetic modification to add a new activity (e.g., an activity not normally found in the host organism), or in certain embodiments, to increase the expression of an existing activity by increasing transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest (e.g., a homologous or heterologous nucleotide sequence of interest). In some embodiments, a microorganism can be engineered by genetic modification to express a nucleic acid reagent that can decrease or substantially eliminate transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest, thereby decreasing the expression of the activity, in certain embodiments.

[0412] A heterologous protein, e.g., a nucleic acid encoding a nucleotide triphosphate transporter, may be inserted into or otherwise used in any suitable expression system. In some embodiments, the nucleic acid reagent may sometimes be stably integrated into the chromosome of the host organism, or in certain embodiments, the nucleic acid reagent may be a deletion of a portion of the host chromosome (e.g., in a genetically modified organism In addition, modification of the host genome confers the ability to selectively or preferentially maintain a desired organism that has the genetic modification. Such nucleic acid reagents (e.g., nucleic acids that confer a selectable trait on an organism with the modified genome or genetically modified organisms) can be selected for their ability to induce the production of a desired protein or nucleic acid molecule. Optionally, the nucleic acid reagent may be modified such that the codon encodes the same amino acid using a tRNA that is (i) different from that specified by the native sequence, or (ii) encodes a different amino acid, including an unnatural amino acid (including a detectable labeled amino acid), that is different from the usual one.

[0413] Recombinant expression is usefully achieved using an expression cassette that can be part of a vector such as a plasmid. The vector can include a promoter operably linked to a nucleic acid encoding a nucleotide triphosphate transporter. The vector may also include other elements necessary for transcription and translation, as described herein. The expression cassette, expression vector, and sequences in the cassette or vector can be heterologous to the cell in which the unnatural nucleotide is contacted. For example, the nucleotide triphosphate transporter sequence can be heterologous to the cell.

[0414] A variety of prokaryotic and eukaryotic expression vectors can be generated that carry, encode, and / or express nucleotide triphosphate transporters. Such expression vectors include, for example, pET, pET3d, pCR2.1, pBAD, pUC, and yeast vectors. The vectors can be used, for example, in a variety of in vivo and in vitro situations. Non-limiting examples of prokaryotic promoters that can be used include SP6, T7, T5, tac, bla, trp, gal, lac, or the maltose promoter. Non-limiting examples of eukaryotic promoters that can be used include constitutive promoters, such as viral promoters like CMV, SV40, and RSV promoters, and regulatable promoters, such as inducible or repressible promoters like the tet promoter, hsp70 promoter, and synthetic promoters regulated by CRE. Vectors for bacterial expression include pGEX-5X-3, and vectors for eukaryotic expression include pCIneo-CMV. Viral vectors that can be used include those related to lentivirus, adenovirus, adeno-associated virus, herpesvirus, vaccinia virus, poliovirus, AIDS virus, neurotrophic virus, Sindbis, and other viruses. Any viral family that shares the characteristics of these viruses and is suitable for use as a vector is also useful. Retroviral vectors that can be used include those described in Verma, American Society for Microbiology, pp. 229-232, Washington, (1985). For example, such retroviral vectors include murine Moloney leukemia virus, MMLV, and other retroviruses that express desirable characteristics. Usually, viral vectors include non-structural early genes, structural late genes, RNA polymerase III transcripts, inverted terminal repeats necessary for replication and capsid formation, and promoters that control transcription and replication of the viral genome.When designed as a vector, the virus typically has one or more initial genes removed, and a gene or gene / promoter cassette is inserted into the viral genome in place of the removed viral nucleic acid.

[0415] Cloning Any convenient cloning strategy known in the art may be utilized to incorporate elements such as ORFs into nucleic acid reagents. Elements may be inserted into a template independently of the inserted element using known methods such as, for example, (1) cutting the template at one or more existing restriction enzyme sites and ligating the element of interest, and (2) adding a restriction enzyme site to the template by hybridizing an oligonucleotide primer containing one or more appropriate restriction enzyme sites and amplifying by polymerase chain reaction (described in more detail herein). Other cloning strategies utilize one or more insertion sites present in or inserted into nucleic acid reagents, such as, for example, oligonucleotide primer hybridization sites for PCR and others described herein. In some embodiments, the cloning strategy may be combined with genetic manipulations such as recombination (e.g., recombination of a nucleic acid reagent having a nucleic acid sequence of interest into the genome of the organism to be modified, as further described herein). In some embodiments, the cloned ORF may (directly or indirectly) generate a modified or wild-type nucleotide triphosphate transporter and / or polymerase by manipulating a microorganism having one or more ORFs of interest, and the microorganism may have an altered activity of nucleotide triphosphate transporter activity or polymerase activity. Hybridizing an oligonucleotide primer and amplifying by polymerase chain reaction (described in more detail herein) to add a restriction enzyme site to the template. Other cloning strategies utilize one or more insertion sites present in or inserted into nucleic acid reagents, such as, for example, oligonucleotide primer hybridization sites for PCR and others described herein. In some embodiments, the cloning strategy may be combined with genetic manipulations such as recombination (e.g., recombination of a nucleic acid reagent having a nucleic acid sequence of interest into the genome of the organism to be modified, as further described herein). In some embodiments, the cloned ORF may (directly or indirectly) generate a modified or wild-type nucleotide triphosphate transporter and / or polymerase by manipulating a microorganism having one or more ORFs of interest, and the microorganism may have an altered activity of nucleotide triphosphate transporter activity or polymerase activity.

[0416] A nucleic acid can be specifically cleaved by contacting it with one or more specific cleavage agents. Certain cleavage agents often specifically cleave according to a specific nucleotide sequence at a specific site. Examples of enzyme-specific cleavage agents include, but are not limited to, endonucleases (e.g., DNase (e.g., DNase I, II); RNase (e.g., RNase E, F, H, P); Cleavase™ enzyme; Taq DNA polymerase; E. coli DNA polymerase I and eukaryotic structure-specific endonucleases; mouse FEN-1 endonuclease; type I, II or III restriction endonucleases, e.g., Acc I, Afl III, Alu I, Alw44 I, Apa I, Asn I, Ava I, Ava II, BamH I, Ban II, Bcl I, Bgl I.Bgl II, Bln I, BsaI, Bsm I, BsmBI, BssH II, BstE II, Cfo I, CIa I, Dde I, Dpn I, Dra I, EcIX I, EcoR I, EcoR I, EcoR II, EcoR V, Hae II, Hae II, Hind II, Hind III, Hpa I, Hpa II, Kpn I, Ksp I, Mlu I, MIuN I, Msp I, Nci I, Nco I, Nde I, Nde II, Nhe I, Not I, Nru I, Nsi I, Pst I, Pvu I, Pvu II, Rsa I, Sac I, Sal I, Sau3A I, Sca I, ScrF I, Sfi I, Sma I, Spe I, Sph I, Ssp I, Stu I, Sty I, Swa I, Taq I, Xba I, Xho I); glycosylase (e.g., uracil-DNA glycosylase (UDG), 3-methyladenine DNA glycosylase, 3-methyladenine DNA glycosylase II, pyrimidine hydrate-DNA glycosylase, FaPy-DNA glycosylase, thymine mismatch-DNA glycosylase, hypoxanthine-DNA glycosylase, 5-hydroxymethyluracil DNA glycosylase (HmUDG), 5-hydroxymethylcytosine DNA glycosylase, or 1,N6-etheno-adenine DNA glycosylase); exonuclease (e.g., exonuclease III); ribozyme; and DNAzymes. The sample nucleic acid may be treated with a chemical or may be synthesized using modified nucleotides and may cleave the modified nucleic acid. In non-limiting examples, the sample nucleic acid may be (i) an alkylating agent such as methylnitrosourea that generates several alkylated bases including N3-methyladenine and N3-methylguanine that are recognized and cleaved by alkylpurine DNA glycosylase; (ii) sodium bisulfite, which causes the cytosine residue of DNA to be deaminated to form a uracil residue that can be cleaved by uracil N-glycosylase; and (iii) a chemical that converts guanine to its oxidized form, 8-hydroxyguanine, which can be cleaved by formamidopyrimidine DNA N-glycosylase. Examples of chemical cleavage processes include, but are not limited to, alkylation (e.g., alkylation of phosphorothioate-modified nucleic acids); acid-labile cleavage of P3’-N5’-phosphoroamidate-containing nucleic acids; and osmium tetroxide of nucleic acids Examples include treatment with um and piperidine.

[0417] In some embodiments, the nucleic acid reagent comprises one or more recombinase insertion sites. A recombinase insertion site is a recognition sequence on a nucleic acid molecule that participates in an integration / recombination reaction by a recombinant protein. For example, the recombination site for Cre recombinase is loxP, which is a 34-base pair sequence composed of two 13-base pair inverted repeat sequences (functioning as recombinase binding sites) flanking an 8-base pair core sequence (e.g., Sauer, Curr. Opin. Biotech. 5:521-527 (1994)). Other examples of recombination sites include attB, attP, attL, and attR sequences, and their variants, fragments, variants, and derivatives recognized by the recombinant protein λInt and by accessory proteins integration host factor (IHF), FIS, and excisionase (Xis) (e.g., U.S. Patent Nos. 5,888,732; 6,143,557; 6,171,861; 6,270,969; 6,277,608; and 6,720,140; U.S. Patent Application Publication Nos. 09 / 517,466, and 09 / 732,914; U.S. Patent Application Publication No. 2002 / 0007051; and Landy, Curr. Opin. Biotech. 3:699-707 (1993) (the disclosures of each are incorporated herein by reference in their entireties)).

[0418] Examples of recombinase cloning nucleic acids are found in the Gateway™ system (Invitrogen, California), which includes at least one recombination site for cloning a desired nucleic acid molecule in vivo or in vitro. In some embodiments, the system often includes at least two different site-specific recombination sites, based on the bacteriophage lambda system (e.g., att1 and att2), and utilizes vectors that are mutated from the wild-type (att0) site. Each mutant site has specificity unique to its cognate partner att site of the same type (i.e., its binding partner recombination site) (e.g., attB1 and attP1, or attL1 and attR1), and does not cross-react with other mutant forms of the recombination site or with the wild-type att0 site. The different site specificities allow for directional cloning or ligation of the desired molecule and thus provide the desired orientation of the cloned molecule. Nucleic acid fragments adjacent to the recombination site are cloned and subcloned by replacing a selectable marker (e.g., ccdB) adjacent to the att site on a recipient plasmid molecule, sometimes called the Destination Vector, using the Gateway™ system. The desired clone is then selected by transformation of a ccdB-sensitive host strain and positive selection of the marker on the recipient molecule. Similar strategies for negative selection (e.g., use of a toxic gene) can be used in other organisms, such as mammalian and insect thymidine kinase (TK).

[0419] The nucleic acid reagent may contain one or more origin of replication (ORI) elements. In some embodiments, the template contains two or more ORIs, one of which functions efficiently in one organism (e.g., bacteria), and another functions efficiently in another organism (e.g., a eukaryote such as yeast). In some embodiments, an ORI may function efficiently in one species (e.g., S. cerevisiae, etc.), and another ORI may function efficiently in a different species (e.g., S. pombe, etc.). The nucleic acid reagent may also contain one or more transcriptional regulatory sites.

[0420] The nucleic acid reagent, e.g., an expression cassette or vector, may contain a nucleic acid sequence encoding a marker product. The marker product is used to determine whether a certain gene has been delivered to a cell and is being expressed after delivery. Examples of marker genes include the E. coli lacZ gene encoding β-galactosidase and green fluorescent protein. In some embodiments, the marker can be a selectable marker. When such a selectable marker is successfully transferred into a host cell, the transformed host cell can survive when placed under selective pressure. There are two distinct categories of selection regimes that are widely used. The first category is based on the use of mutant cell lines that lack the ability to metabolize and grow independently of supplemented media. The second category is dominant selection, which refers to a selection scheme used in any cell type and does not require the use of mutant cell lines. These schemes typically use drugs to block the growth of host cells. Those cells with the new gene express a protein that confers drug resistance and survive the selection. Examples of such dominant selection are the use of the drug neomycin (Southern et al., J. Molec. Appl. Genet. 1:327 (1982)), mycophenolic acid (Mulligan et al., Science 209:1422 (1980)) or hygromycin (Sugden et al., Mol. Cell. Biol. 5:410 - 413 (1985); the disclosures of each of them are incorporated herein by reference in their entirety).

[0421] A nucleic acid reagent can include one or more selection elements (e.g., elements for selecting the presence of the nucleic acid reagent, elements not for activating the activity of a promoter element that can be selectively regulated). The selection elements are often utilized using known processes to determine whether the nucleic acid reagent is contained in a cell. In some embodiments, the nucleic acid reagent includes two or more selection elements, where one element functions efficiently in one organism and another element functions efficiently in another organism.Examples of selectable elements include, but are not limited to, the following: (1) a nucleic acid segment encoding a product that confers resistance to a compound that is toxic by other means (e.g., an antibiotic); (2) a nucleic acid segment encoding a product that is lacking in the recipient cell by other means (e.g., an essential product, a tRNA gene, an auxotrophic marker); (3) a nucleic acid segment encoding a product that suppresses the activity of a gene product; (4) a nucleic acid segment encoding a product that can be easily identified (e.g., an antibiotic (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and phenotypic markers such as cell surface proteins); (5) a nucleic acid segment that binds to a product that is harmful to cell survival and / or function by other means; (6) a nucleic acid segment that inhibits the activity of any of the nucleic acid segments described in (1)-(5) above (e.g., an antisense oligonucleotide); (7) a nucleic acid segment that binds to a product that modifies a substrate (e.g., a restriction endonuclease); (8) a nucleic acid segment that can be used to isolate or identify a desired molecule (e.g., a specific protein binding site); (9) a nucleic acid segment encoding a specific nucleotide sequence that may not function by other means (e.g., for PCR amplification of a subpopulation of molecules); (10) a nucleic acid segment that directly or indirectly confers resistance or sensitivity to a specific compound when absent; (11) a nucleic acid segment encoding a product that converts a toxic or relatively non-toxic compound into a toxic compound in the recipient cell (e.g., herpes simplex thymidine kinase, cytosine deaminase); (12) a nucleic acid segment that inhibits the replication, partitioning, or heritability of the nucleic acid molecule containing them; and / or (13) a nucleic acid segment encoding a conditional replication function, e.g., replication in a specific host or host cell line, or under specific environmental conditions (e.g., temperature, nutrient status, etc.).

[0422] The nucleic acid reagent can be in any form useful for in vivo transcription and / or translation. The nucleic acid can be a plasmid, such as a supercoiled plasmid, or a yeast artificial chromosome (e.g., YAC), or a linear nucleic acid (e.g., a linear nucleic acid produced by PCR or restriction digestion), or single-stranded, and sometimes double-stranded. The nucleic acid reagent can be prepared by an amplification process such as the polymerase chain reaction (PCR) process or the transcription-mediated amplification process ( TMA). In TMA, two enzymes are used in an isothermal reaction to generate amplification products that are detected by luminescence (e.g., Biochemistry 1996 Jun 25;35(25):8429-38). The standard PCR process is known (e.g., U.S. Patent Nos. 4,683,202; 4,683,195; 4,965,188; and 5,565,493) and is generally carried out in cycles. Each cycle includes heat denaturation (where the hybrid nucleic acid dissociates), cooling (where the primer oligonucleotide hybridizes); and elongation of the oligonucleotide by a polymerase (i.e., Taq polymerase). An example of a PCR cycling process is to treat the sample at 95°C for 5 minutes; repeat 45 cycles of 1 minute at 95°C, 1 minute at 59°C, 10 seconds, and 1 minute 30 seconds at 72°C; then treat the sample at 72°C for 5 minutes. Multiple cycles are frequently performed using a commercially available thermal cycler. The PCR amplification product may be temporarily stored at a low temperature (e.g., 4°C) or frozen before analysis (e.g., -20°C).

[0423] Using a cloning strategy similar to the above, DNA containing unnatural nucleotides can be generated. For example, an oligonucleotide containing an unnatural nucleotide at a desired position is synthesized using standard solid-phase synthesis and purified by HPLC. The oligonucleotide is then inserted into a plasmid containing the necessary sequence context (i.e., UTR and coding sequence) using a cloning method (such as Golden Gate assembly) that uses a cloning site such as a BsaI site (although other ones described above may be used).

[0424] Kit / Manufactured Product In certain embodiments, disclosed herein are kits and manufactured products for use in one or more of the methods described herein. Such kits include a carrier, package, or container compartmentalized to receive one or more containers such as vials, tubes, etc., each of the containers containing one of the distinct elements used in the methods described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. In one embodiment, the containers are formed from a variety of materials such as glass or plastic.

[0425] In some embodiments, the kit comprises a suitable packaging material for containing the contents of the kit. Optionally, the packaging material is constructed by well-known methods, preferably to provide a sterile and contaminant-free environment. Packaging materials used herein include, for example, those customarily utilized in commercially available kits sold for use in nucleic acid sequencing systems. Exemplary packaging materials include, but are not limited to, glass, plastic, paper, foil, etc. that can hold the components described herein within a defined range.

[0426] The packaging material may comprise a label indicating the specific use of the component. The use of the kit indicated by this label may be one or more of the methods described herein that are appropriate for a specific combination of components present in the kit. For example, the label may indicate that the kit is useful for a method of synthesizing polynucleotides or for a method of determining the sequence of nucleic acids.

[0427] Instructions for use of the packaged reagent or component may also be included in the kit. Such instructions typically include specific statements explaining the relative amounts of kit components and samples to be mixed, the maintenance period of the reagent / sample mixture, reaction parameters such as temperature, buffer conditions, and the like.

[0428] It is understood that not all components necessary for a particular reaction need to be present in a particular kit. Instead, one or more additional components may be provided from other sources. The instructions attached to the kit may identify the additional components provided and where they can be obtained.

[0429] In some embodiments, for example, kits useful for stably integrating non-natural nucleic acids into cellular nucleic acids are provided using the methods provided by the present invention for preparing genetically engineered cells. In one embodiment, the kits described herein comprise genetically engineered cells and one or more non-natural nucleic acids. In another embodiment, the kits described herein comprise an isolated and purified plasmid comprising a sequence selected from SEQ ID NOs: 1-2. In a further embodiment, the kits described herein comprise a primer comprising a sequence selected from SEQ ID NOs: 3-20.

[0430] In additional embodiments, the kits described herein provide a cell and a nucleic acid molecule comprising a heterologous gene for introduction into the cell, thereby providing a genetically engineered cell such as an expression vector comprising the nucleic acid of any of the above embodiments described in this paragraph.

[0431] Exemplary Embodiments The present disclosure is further described by the following embodiments. Each configuration of the embodiments is combined with any of the other embodiments if appropriate and practical.

[0432] Embodiment 1. An in vivo method for producing a protein containing a non-natural amino acid, the method comprising: transcribing a DAN template containing a first non-natural base and a complementary second non-natural base to incorporate a third non-natural base into mRNA, the third non-natural base being configured to form a first non-natural base pair with the first non-natural base; transcribing the DNA template to incorporate a fourth non-natural base into tRNA, the fourth non-natural base being configured to form a second non-natural base pair with the second non-natural base, and the first non-natural base pair and the second non-natural base pair being different; and translating a protein from the mRNA and the tRAN, the protein containing a non-natural amino acid. A method comprising the steps of.

[0433] Embodiment 2. The in vivo method according to Embodiment 1, comprising the use of a semi-synthetic organism.

[0434] Embodiment 3. The method according to Embodiment 2, wherein the organism comprises a microorganism.

[0435] Embodiment 4. The method according to Embodiment 3, wherein the organism comprises bacteria.

[0436] Embodiment 5. The method according to Embodiment 4, wherein the organism comprises Gram-positive bacteria.

[0437] Embodiment 6. The method according to Embodiment 4, wherein the organism comprises Gram-negative bacteria.

[0438] Embodiment 7. The method according to any one of Embodiments 2 to 4, wherein the organism comprises Escherichia coli.

[0439] Embodiment 8. At least one unnatural base is (i) 2-thiouracil, 2-thio-thymine, 2'-deoxyuridine, 4-thio-uracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl (I), 5-halouracil; 5-propynyl-uracil, 6-azo-thymine, 6-azo-uracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, shu dourasil, methyl uracil-5-oxoacetate, uracil-5-oxoacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, uracil-5-oxoacetic acid, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil or dihydrouracil; (ii) 5-hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytarabine, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azacytosine, azacytosine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one) or pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d]pyrimidin-2-one); (iii) Adenine substituted with 2-aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2'-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-aza-adenine; (iv) 2-methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thio-guanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, guanine substituted with 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-aza-guanine; and (v) The method according to any one of Embodiments 1 to 7, selected from the group consisting of hypoxanthine, xanthine, 1-methylinosine, queuosine, beta-D-galactosylqueuosine, inosine, beta-D-mannosylqueuosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone.

[0440] Embodiment 9. At least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is [Chemical formula] The method according to any one of Embodiments 1 to 7, selected from the group consisting of.

[0441] Embodiment 10. At least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is [Chemical formula] The method according to any one of Embodiments 1 to 7, selected from the group consisting of.

[0442] Embodiment 11. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0443] Embodiment 12. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0444] Embodiment 13. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0445] Embodiment 14. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0446] Embodiment 15. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0447] Embodiment 16. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0448] Embodiment 17. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0449] Embodiment 18. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0450] Embodiment 19. At least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is

Chemical formula

[0451] Embodiment 20. The first or second unnatural base is

Chemical formula

[0452] Embodiment 21. The first or second unnatural base is

Chemical formula

[0453] Embodiment 22. The first unnatural base is

Chemical formula

Chemical formula

[0454] Embodiment 23. The first unnatural base is

Chemical formula

Chemical formula

[0455] Embodiment 24. The third or fourth unnatural base is

Chemical formula

[0456] Embodiment 25. The third unnatural base is

Chemical formula

[0457] Embodiment 26. The fourth unnatural base is

Chemical formula

[0458] Embodiment 27. The third or fourth unnatural base is

Chemical formula

[0459] Embodiment 28. The third unnatural base is

Chemical formula

[0460] Embodiment 29. The fourth unnatural base is

Chemical formula

[0461] Embodiment 30. The first unnatural base is

Chemical formula

Chemical formula

Chemical formula

Chemical formula

[0462] Embodiment 31. The first unnatural base is

Chemical formula

Chemical formula

Chemical formula

[0463] Embodiment 32. The first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] The method according to Embodiment 9, which is as follows.

[0464] Embodiment 33. The third unnatural base is [Chemical formula] The method according to Embodiment 9, which is as follows.

[0465] Embodiment 34. The fourth unnatural base is [Chemical formula] The method according to Embodiment 9, which is as follows.

[0466] Embodiment 35. The first unnatural base is [Chemical formula] and the second unnatural base is [Chemical formula] and the third unnatural base is [Chemical formula] and the fourth unnatural base is [Chemical formula] The method according to Embodiment 9, which is

[0467] Embodiment 36. The method according to any one of Embodiments 9 to 35, wherein the third unnatural base and the fourth unnatural base contain ribose.

[0468] Embodiment 37. The method according to any one of Embodiments 9 to 35, wherein the third unnatural base and the fourth unnatural base contain deoxyribose.

[0469] Embodiment 38. The method according to any one of Embodiments 9 to 35, wherein the first and second unnatural bases contain deoxyribose.

[0470] Embodiment 39. The method according to any one of Embodiments 9 to 35, wherein the first and second unnatural bases contain deoxyribose, and the third and fourth unnatural bases contain ribose.

[0471] Embodiment 40. The DNA template is [Chemical formula] The method according to Embodiment 9, which contains at least one unnatural base pair (UBP) selected from the group consisting of

[0472] Embodiment 41. The method according to Embodiment 40, wherein the DNA template contains at least one unnatural base pair (UBP) which is dNaM-d5SICS.

[0473] Embodiment 42. The method according to Embodiment 40, wherein the DNA template contains at least one unnatural base pair (UBP) which is dCNMO-dTPT3.

[0474] Embodiment 43. The method according to Embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3.

[0475] Embodiment 44. The method according to Embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dPTMO-dTPT3.

[0476] Embodiment 45. The method according to Embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1.

[0477] Embodiment 46. The method according to Embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.

[0478] Embodiment 47. The DNA template

Chemical formula

Chemical formula

[0479] Embodiment 48. The method according to Embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-d5SICS.

[0480] Embodiment 49. The method according to Embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTPT3.

[0481] Embodiment 50. The method according to Embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3.

[0482] Embodiment 51. The method according to Embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dPTMO-dTPT3.

[0483] Embodiment 52. The method according to Embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1.

[0484] Embodiment 53. The method according to Embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.

[0485] Embodiment 54. The mRNA and tRNA are 【Chemical ...

Claims

1. 1. A method for producing a tRNA or an mRNA encoding a protein, the method comprising the step of contacting or exposing DNA comprising a gene encoding the tRNA or protein to a first ribonucleoside triphosphate and an RNA polymerase, wherein the gene encoding the tRNA or protein comprises a first unnatural base on a first strand of the DNA and a second unnatural base on a second strand of the DNA, the first unnatural base and the second unnatural base form a first unnatural base pair, and the first ribonucleoside triphosphate comprises a third unnatural base, which can form a second unnatural base pair with the first unnatural base, and the first unnatural base and the second unnatural base are not identical, thereby producing a tRNA or an mRNA encoding a protein comprising the third unnatural base.

2. 2. The method of claim 1, further comprising contacting or exposing the DNA to a second ribonucleoside triphosphate comprising a fourth unnatural base, wherein the fourth unnatural base can form a third unnatural base pair with the third unnatural base.

3. 3. The method of claim 2, wherein the first unnatural base pair and the third unnatural base pair are not identical.

4. 4. The method of any one of claims 1 to 3, further comprising replicating the DNA by contacting the DNA with a first deoxyribonucleoside triphosphate and a DNA polymerase prior to contacting or exposing the DNA with the first ribonucleoside triphosphate and the RNA polymerase, wherein the first deoxyribonucleoside triphosphate comprises a fifth unnatural base, the fifth unnatural base being capable of forming a fourth unnatural base pair with the first unnatural base, and the first unnatural base pair and the fourth unnatural base pair are not identical.

5. At least one of the first unnatural base, the second unnatural base, the third unnatural base, and, if present, the fourth unnatural base 【Chemistry 1】 is a nucleobase having the structure During the ceremony, each X is independently carbon or nitrogen; R 2 is present when X is carbon and independently C 1 ~C 6 alkyl, hydrogen, methoxy, methanethiol, methaneseleno, halogen, cyano, or azido group; Y is sulfur, oxygen, or selenium; E is oxygen, sulfur or selenium; 4. The method of any one of claims 1 to 3, wherein a wavy line indicates the point of attachment of the nucleobase to a ribosyl, deoxyribosyl or dideoxyribosyl moiety, wherein the ribosyl, deoxyribosyl or dideoxyribosyl moiety is in free form, attached to a monophosphate, diphosphate, triphosphate, α-thiotriphosphate, β-thiotriphosphate or γ-thiotriphosphate group, or comprised in a polynucleotide, optionally wherein the polynucleotide is RNA, DNA, a bicyclic nucleic acid, a linked nucleic acid, peptide nucleic acid (PNA), locked nucleic acid (LNA) or a phosphorothioate-containing nucleic acid.

6. (i) X is carbon; and / or (ii) E is sulfur; and / or (iii) Y is sulfur; The method of claim 5.

7. 4. The method of any one of claims 1 to 3, wherein the first unnatural base is TPT3, the second unnatural base is CNMO or NaM, the third unnatural base is TAT1, and the fourth unnatural base, if present, is NaM or 5FM.

8. The method of any one of claims 1 to 3, wherein the method comprises the use of a cell, optionally the cell is a bacterium, and optionally the bacterium is Escherichia coli.

9. At least one of the first unnatural base, the second unnatural base, the third unnatural base, and, if present, the fourth unnatural base 【Chemistry 2】 The method according to any one of claims 1 to 3, wherein

10. At least one of the first unnatural base, the second unnatural base, the third unnatural base, and, if present, the fourth unnatural base 【Transformation 3】 The method of claim 9, wherein

11. (i) at least one of the first unnatural base, the second unnatural base, the third unnatural base, and, if present, the fourth unnatural base is 【Chemistry 4】 is; or (ii) at least one of the first unnatural base, the second unnatural base, the third unnatural base, and, if present, the fourth unnatural base is 【Transformation 5】 is; or (iii) at least one of the first unnatural base, the second unnatural base, the third unnatural base, and, if present, the fourth unnatural base is 【Transformation 6】 is; or (iv) at least one of the first unnatural base, the second unnatural base, the third unnatural base, and, if present, the fourth unnatural base is 【Transformation 7】 The method of claim 9, wherein

12. (i) the first or second unnatural base is 【Transformation 8】 is; or (ii) the first unnatural base is 【Chemistry 9】 and the second unnatural base is 【Chemistry 10】 or the first unnatural base is 【Chemistry 11】 and the second unnatural base is 【Chemistry 12】 is; or (iii) the first unnatural base is 【Chemistry 13】 and the second unnatural base is 【Chemistry 14】 and the third unnatural base is 【Chemistry 15】 and said fourth unnatural base, if present, is 【Chemistry 16】 is; or (iv) the first unnatural base is 【Chemistry 17】 and the second unnatural base is [Chemistry 18] and the third unnatural base is 【Chemistry 19】 and said fourth unnatural base, if present, is 【Chemistry 20】 is; or (v) the first unnatural base is 【Chemistry 21】 and the second unnatural base is 【Chemistry 22】 and the third unnatural base is 【Chemistry 23】 and said fourth unnatural base, if present, is 【Chemistry 24】 is; or (vi) the first unnatural base is 【Chemistry 25】 and the second unnatural base is 【Chemistry 26】 and the third unnatural base is 【Chemistry 27】 and said fourth unnatural base, if present, is 【Chemistry 28】 The method of claim 11, wherein

13. (i) the first or second unnatural base is 【Chemistry 29】 and / or (ii) the third or, if present, fourth unnatural base is 【Transformation 30】 and / or (iii) the third or, if present, fourth unnatural base is 【Chemistry 31】 and / or (iv) the third unnatural base is 【Chemistry 32】 and / or (v) the method includes contacting or exposing the DNA to a second ribonucleoside triphosphate that includes a fourth unnatural base, wherein the fourth unnatural base is 【Transformation 33】 The method of claim 11, wherein

14. the first unnatural base pair 【Transformation 34】 The method of claim 9, wherein the compound is selected from the group consisting of:

15. the first unnatural base pair 【Chemistry 35】 and the third unnatural base is selected from 【Transformation 36】 15. The method of claim 14, wherein the compound is selected from the group consisting of:

16. (i) the first unnatural base is dCNMO and the second unnatural base is dTPT3; and / or (ii) the third unnatural base is NaM and the second unnatural base is TAT1; The method according to any one of claims 1 to 3.

17. 4. The method of any one of claims 1 to 3, wherein when an mRNA encoding the protein is produced, the mRNA is translated into a protein comprising at least one unnatural amino acid at a position corresponding to the codon of the mRNA that contains the third unnatural base.

18. The protein (i) comprises at least two unnatural amino acids or at least three unnatural amino acids; and / or (ii) at least two different unnatural amino acids or at least three different unnatural amino acids Contains amino acids, 18. The method of claim 17.

19. the at least one unnatural amino acid (i) is a lysine analog; (ii) containing an aromatic side chain; (iii) contains an azide group; (iv) contains an alkyne group, or (v) containing an aldehyde or ketone group; The method according to any one of claims 1 to 3.

20. the at least one unnatural amino acid (i) does not contain aromatic side chains; (ii) N6-azidoethoxy-carbonyl-L-lysine (AzK) or N6-propargylethoxy-carbonyl-L-lysine (PraK), (iii) N6-azidoethoxy-carbonyl-L-lysine (AzK), or (iv) N6-propargylethoxy-carbonyl-L-lysine (PraK), 20. The method of claim 19.

21. mRNA produced by the method according to any one of claims 1 to 3.

22. 22. A protein encoded by the mRNA of claim 21, comprising an unnatural amino acid at a position corresponding to the codon of the mRNA that contains the third unnatural base.

23. A cell comprising an expanded genetic alphabet, wherein the genetic alphabet comprises at least three distinct unnatural bases.

24. 1. A method for transcribing DNA, comprising: providing one or more DNAs including: (1) a gene encoding a protein, wherein the template strand of the gene encoding the protein comprises a first unnatural base; and (2) a gene encoding a tRNA, wherein the template strand of the gene encoding the tRNA comprises a second unnatural base, and the second unnatural base can form a first unnatural base pair with the first unnatural base; transcribing a gene encoding the protein to produce mRNA that includes a third unnatural base, wherein the third unnatural base is capable of forming a second unnatural base pair with the first unnatural base; transcribing a gene encoding the tRNA to produce a tRNA that includes a fourth unnatural base, wherein the fourth unnatural base can form a third unnatural base pair with the second unnatural base, and the second unnatural base pair and the third unnatural base pair are not identical; The method comprising:

25. 1. A method for replicating DNA, comprising: Providing DNA comprising: (1) a gene encoding a protein, wherein the template strand of the gene encoding the protein comprises a first unnatural base; and (2) a gene encoding a tRNA, wherein the template strand of the gene encoding the tRNA comprises a second unnatural base that can form a first unnatural base pair with the first unnatural base; replicating the DNA to generate new DNA strands that include a first alternative unnatural base in place of the first unnatural base and / or a second alternative unnatural base in place of the second unnatural base. a process of and the method optionally further comprises: transcribing a gene encoding the protein to produce mRNA that includes a third unnatural base, wherein the third unnatural base is capable of forming a second unnatural base pair with the first unnatural base and / or the first alternative unnatural base; and / or transcribing a gene encoding the tRNA to produce a tRNA that includes a fourth unnatural base, wherein the fourth unnatural base can form a third unnatural base pair with the second unnatural base and / or the second alternative unnatural base, and wherein the second unnatural base pair and the third unnatural base pair are not identical.