Reagents and methods for replication, transcription, and translation in semi-synthetic organisms
By introducing non-natural base pairs into organisms to expand the genetic code, using specific tRNA-amino acid synthetase and heterologous RNA/DNA polymerase, the efficient and fidelity incorporation of non-natural amino acids is achieved, solving the problems of low incorporation efficiency and fidelity in the prior art, and expanding protein diversity.
Patent Information
- Application Number
- CN202080056659.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-14
- Filing Date
- 2020-06-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-06-12
AI Technical Summary
The prior art is difficult to effectively extend the genetic code to incorporate non-natural amino acids, resulting in limited protein diversity, especially in eukaryotes with low incorporation efficiency and fidelity, and complex genomic engineering.
By introducing non-natural base pairs (UBPs), forming new base pairs, expanding the genetic code, re-encoding non-natural amino acids in organisms using specific tRNA-amino acid synthetases, combining heterologous RNA and DNA polymerases, achieving high-fidelity incorporation of non-natural amino acids.
It realizes efficient and fidelity incorporating non-natural amino acids in organisms, expands protein diversity, solves the problems of low incorporation efficiency and fidelity in the prior art, and is suitable for protein therapy for semi-synthetic organisms.
Smart Images

Figure CN114207129B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the priority of U.S. Provisional Application No. 62 / 861,901, filed on June 14, 2019, the disclosure of which is hereby incorporated by reference in its entirety.
[0003] Sequence Listing
[0004] This application contains a sequence listing that has been electronically submitted in ASCII format and is hereby incorporated by reference in its entirety. The ASCII copy, created on June 10, 2020, is named 36271 - 808_601_SL.txt and is 18,162 bytes in size.
[0005] Statement Regarding Federally Sponsored Research
[0006] The invention disclosed herein was made, at least in part, in under the support of the U.S. government by grants 5R35 GM118178 and GM128376 and F31 GM128376 from the National Institutes of Health (NIH) and grant NSF / DGE - 1346837 from the National Science Foundation (NSF). Accordingly, the U.S. government has certain rights in this invention. Background of the Invention
[0007] Biodiversity enables life to adapt to different environments and to evolve new forms and functions over time. The source of this diversity is the variation within protein sequences provided by the twenty natural amino acids, which are encoded by the four natural DNA nucleotides in an organism's genome. While the functional diversity provided by natural amino acids may be high, the vastness of sequence space severely limits what can be practically explored. Moreover, certain functions are simply not available. The use of cofactors for hydride transfer, redox activity, and electrophilic bond formation, etc. in nature demonstrates these limitations. Additionally, as there is increasing interest in developing proteins as therapeutic agents, these limitations are problematic because the physicochemical diversity of natural amino acids is severely restricted compared to small - molecule drugs designed by chemists. In principle, it should be possible to circumvent these limitations by expanding the genetic code to include additional non - canonical amino acids (ncAAs, or "unnatural" amino acids) with desired physicochemical properties.
[0008] Approximately 20 years ago, an approach to increase the diversity available to organisms was created by expanding the genetic code through encoding ncAAs in Escherichia coli using the amber stop codon (UAG). This was achieved using a tRNA - aminoacyl tRNA synthetase (aaRS) pair from Methanococcus jannaschii, where the tRNA was recoded to suppress the stop codon and the aaRS was evolved to load the tRNA with the ncAA. This codon suppression approach has been extended to other stop codons and even quadruplet codons, as well as using several other orthogonal tRNA - aaRS pairs (most notably the Pyl tRNA - synthetase pair from Methanosarcina barkeri / M. mazei), expanding the range of ncAAs that can be incorporated into proteins. These methods have begun to revolutionize both chemical biology and protein therapeutics.
[0009] Although these methods enable incorporation of up to two different ncAAs in both prokaryotic and eukaryotic cells, the heterologously recoded tRNAs necessarily compete with endogenous release factors (RFs), or in the case of quadruplet codons, with normal decoding, which limits the efficiency and fidelity of ncAA incorporation. To eliminate competition with RF1, which recognizes the amber stop codon and terminates translation, efforts have been made to remove many or all amber stop codons from the host genome or to modify RF2 to allow for the absence of RF1. However, eukaryotes have only one release factor, which although it can be modified, cannot be absent, and for prokaryotes, the absence of RF1 results in more mis - suppression of amber stop codons by other tRNAs, thus reducing the fidelity of ncAA incorporation. Further efforts to utilize codon redundancy to free up natural codons for reassignment to ncAAs can be complicated by pleiotropic effects, as codons are not truly redundant as they have an impact on translation rate and protein folding, for example. Additionally, codon reassignment is limited by the challenges of large - scale genome engineering, especially for eukaryotes.
[0010] An alternative route for natural codon reassignment is to create entirely new codons that have no native function or constraints and are inherently more orthogonal in recognition at the ribosome. This can be achieved by creating organisms with fifth and sixth nucleotides that form unnatural base pairs (UBPs). Such semi - synthetic organisms (SSOs) would need to faithfully replicate DNA containing the UBP, efficiently transcribe it into mRNA and tRNA containing the unnatural nucleotides, and then effectively decode the unnatural codons with the associated unnatural anticodons. Such SSOs would have an almost infinite number of new codons to encode ncAAs. Summary of the Invention
[0011] In certain embodiments, methods, cells, engineered microorganisms, plasmids, and kits for increasing the production of nucleic acid molecules comprising unnatural nucleotides are described herein.
[0012] The following embodiments are included.
[0013] Embodiment A1 is a nucleobase having the following structure:
[0014]
[0015] Wherein:
[0016] Each X is independently carbon or nitrogen;
[0017] When X is carbon, R2 is present and is independently hydrogen, alkyl, alkenyl, alkynyl, methoxy, methanethiol, methylseleno, halogen, cyano, or azido;
[0018] Y is sulfur, oxygen, selenium, or secondary amine; and
[0019] E is oxygen, sulfur, or selenium;
[0020] Where the wavy line indicates the point of attachment to a ribosyl, deoxyribosyl, or dideoxyribosyl moiety or an analogue thereof, wherein the ribosyl, deoxyribosyl, or dideoxyribosyl moiety or an analogue thereof is in free form, linked to a monophosphate, diphosphate, triphosphate, α-thiotriphosphate, β-thiotriphosphate, or γ-thiotriphosphate group, or incorporated in RNA or DNA or an RNA analogue or DNA analogue.
[0021] Embodiment A2 is the nucleobase of Embodiment A1, wherein X is carbon.
[0022] Embodiment A3 is the nucleobase of Embodiment A1 or A2, wherein E is sulfur.
[0023] Embodiment A4 is the nucleobase of any one of Embodiments A1 to A3, wherein Y is sulfur.
[0024] Embodiment A5 is the nucleobase of Embodiment A1 having the structure
[0025] Embodiment A6 is the nucleobase of any one of Embodiments A1 to A5, which binds to a complementary base-pairing nucleobase to form an unnatural base pair (UBP).
[0026] Embodiment A7 is the nucleobase of Embodiment A6, wherein the complementary base-pairing nucleobase is selected from:
[0027]
[0028] Embodiment A8 is a double-stranded oligonucleotide duplex, wherein the first oligonucleotide strand contains the nucleobases of any one of Embodiments A1 to A5, and the second complementary oligonucleotide strand contains complementary base-pairing nucleobases at its complementary base-pairing sites.
[0029] Embodiment A9 is the double-stranded oligonucleotide duplex of Embodiment A8, wherein the first oligonucleotide strand contains and the second strand contains complementary base-pairing nucleobases selected from the following at its complementary base-pairing sites:
[0030]
[0031] Embodiment A10 is the double-stranded oligonucleotide duplex of Embodiment A9, wherein the second strand contains complementary base-pairing nucleobases
[0032] Embodiment A11 is the double-stranded oligonucleotide duplex of Embodiment A9, wherein the second strand contains complementary base-pairing nucleobases
[0033] Embodiment A12 is a plasmid containing a gene encoding a transfer RNA (tRNA) and / or a gene encoding a protein of interest, wherein the gene contains at least one nucleobase of any one of Embodiments A1 to A5 or TPT3 and at least one complementary base-pairing nucleobase of Embodiment A7 or wherein the complementary base-pairing nucleobases are at complementary base-pairing sites.
[0034] Embodiment A13 is an mRNA encoded by the plasmid of Embodiment A10, and the mRNA encodes the tRNA.
[0035] Embodiment A14 is an mRNA encoded by the plasmid of Embodiment A10, and the mRNA encodes the protein.
[0036] Embodiment A15 is a transfer RNA (tRNA) containing the nucleobases of any one of Embodiments A1 to A5, and the transfer RNA contains:
[0037] an anticodon, wherein the anticodon contains the nucleobase, optionally wherein the nucleobase is located at the first position, the second position, or the third position of the anticodon; and
[0038] a recognition element, wherein the recognition element promotes selective loading of a non-natural amino acid by the tRNA by an aminoacyl-tRNA synthetase.
[0039] Embodiment A16 is the tRNA of Embodiment A15, wherein the aminoacyl-tRNA synthetase is derived from Methanosarcina or a variant thereof, or Methanococcus / Methanocaldococcus or a variant thereof.
[0040] Embodiment A17 is the tRNA of Embodiment A15, wherein the unnatural amino acid contains an aromatic moiety.
[0041] Embodiment A18 is the tRNA of Embodiment A15, wherein the unnatural amino acid is a lysine or phenylalanine derivative.
[0042] Embodiment A19 is a structure comprising the following formula:
[0043] N1-Zx-N2
[0044] wherein:
[0045] each Z is independently a nucleobase of any one of Embodiments A1 to A7, which is bonded to a ribosyl or deoxyribosyl or an analogue thereof;
[0046] N1 is one or more nucleotides or an analogue thereof or a terminal phosphate group attached at the 5'-end of the ribosyl or deoxyribosyl or an analogue thereof of Z;
[0047] N2 is one or more nucleotides or an analogue thereof or a terminal hydroxyl group attached at the 3'-end of the ribosyl or deoxyribosyl or an analogue thereof of Z; and
[0048] x is an integer from 1 to 20.
[0049] Embodiment A20 is the structure of Embodiment A19, wherein the structure encodes a gene, optionally wherein Zx is located in the translation region of the gene, or wherein Zx is located in the untranslated region of the gene.
[0050] Embodiment A21 is a polynucleotide library, wherein the library comprises at least 5000 unique polynucleotides, and wherein each polynucleotide comprises at least one nucleobase of any one of Embodiments A1 to A5.
[0051] Embodiment A22 is a nucleoside triphosphate comprising a nucleobase, wherein the nucleobase is selected from:
[0052]
[0053] Embodiment A23 is the nucleoside triphosphate of Embodiment A22, wherein the nucleobase is
[0054] Embodiment A24 is a nucleoside triphosphate of Embodiment A22 or A23, wherein the nucleoside comprises ribose or deoxyribose.
[0055] Embodiment A25 is a DNA that comprises a nucleobase having the structure and a complementary base-pairing nucleobase having the structure
[0056] Embodiment A26 is a DNA that comprises a nucleobase having the structure and a complementary base-pairing nucleobase having the structure
[0057] Embodiment A27 is a method of transcribing DNA into tRNA or mRNA encoding a protein, the method comprising:
[0058] contacting DNA comprising a gene encoding the tRNA or protein with ribonucleoside triphosphates and an RNA polymerase, wherein the gene encoding the tRNA or protein comprises a first unnatural base that pairs with a second unnatural base and forms a first unnatural base pair with the second unnatural base, and wherein the ribonucleoside triphosphates comprise a third unnatural base that is capable of forming a second unnatural base pair with the first unnatural base, wherein the first unnatural base pair and the second unnatural base pair are different.
[0059] Embodiment A28 is the method of Embodiment A27, wherein the ribonucleoside triphosphates further comprise a fourth unnatural base, wherein the fourth unnatural base is capable of forming a second unnatural base pair with the third unnatural base.
[0060] Embodiment A29 is the method of Embodiment A28, wherein the first unnatural base pair and the second unnatural base pair are different.
[0061] Embodiment A30 is the method of any one of Embodiments A27 - A29, the method further comprising, prior to contacting the DNA with the ribonucleoside triphosphates and the RNA polymerase, replicating the DNA by contacting the DNA with deoxyribonucleoside triphosphates and a DNA polymerase, wherein the ribonucleoside triphosphates comprise a fifth unnatural base that is capable of forming a fifth unnatural base pair with the first unnatural base, wherein the first unnatural base pair and the fifth unnatural base pair are different.
[0062] Embodiment A30 is the method of any one of Embodiments A27 - A30, wherein the first unnatural base comprises TPT3, the second unnatural base comprises CNMO or NaM, the third unnatural base comprises TAT1, and the fourth unnatural base comprises NaM or 5FM.
[0063] Embodiment A32 is the method of any one of embodiments A27 - A31, wherein the method comprises using a semi-synthetic organism, optionally wherein the organism is a bacterium, and optionally wherein the bacterium is Escherichia coli.
[0064] Embodiment A33 is the method of embodiment A32, wherein the organism comprises a microorganism.
[0065] Embodiment A34 is the method of embodiment A32, wherein the organism comprises a bacterium.
[0066] Embodiment A35 is the method of embodiment A34, wherein the organism comprises a Gram-positive bacterium.
[0067] Embodiment A36 is the method of embodiment A34, wherein the organism comprises a Gram-negative bacterium.
[0068] Embodiment A37 is the method of any one of embodiments A27 - A34, wherein the organism comprises Escherichia coli.
[0069] Embodiment A38 is the method of any one of embodiments A27 - A37, wherein at least one unnatural base is selected from
[0070] (i) 2-thiouracil, 2-thiothymine, 2'-deoxyuridine, 4-thio-uracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl (I), 5-halouracil; 5-propynyl-uracil, 6-azathymine, 6-azauracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, uracil-5-oxoacetic acid methyl ester, uracil-5-oxoacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, uracil-5-oxyacetic acid, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiothymidine, 5-carboxymethylaminomethyluracil or dihydrouracil;
[0071] (ii) 5-Hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytidine, cytarabine, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azidocytosine, azacytidine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one) or pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d]pyrimidin-2-one);
[0072] (iii) 2-Aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2’-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxy substituted adenines, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-aza-adenine;
[0073] (iv) 2-Methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thioguanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxy substituted guanines, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-aza-guanine; and
[0074] (v) Hypoxanthine, xanthine, 1-methylinosine, queosine, β-D-galactosylqueosine, inosine, β-D-mannosylqueosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone.
[0075] Embodiment A39 is the method of any one of Embodiments A27 - A37, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base comprises
[0076]
[0077] Embodiment A40 is the method of any one of Embodiments A27 - A39, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0078]
[0079] Embodiment A41 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0080]
[0081] Embodiment A42 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0082] Embodiment A43 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0083]
[0084] Embodiment A44 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises:
[0085]
[0086] Embodiment A45 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises:
[0087]
[0088] Embodiment A46 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises:
[0089]
[0090] Embodiment A47 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0091]
[0092] Embodiment A48 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is
[0093] Embodiment A49 is the method of Embodiment A40, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises:
[0094]
[0095] Embodiment A50 is the method of Embodiment A40, wherein the first unnatural base or the second unnatural base is
[0096] Embodiment A51 is the method of Embodiment A40, wherein the first unnatural base or the second unnatural base is
[0097] Embodiment A52 is the method of any one of Embodiments A27 - A40, wherein the first unnatural base is and the second unnatural base is or the first unnatural base is and the second unnatural base is
[0098] Embodiment A53 is the method of any one of Embodiments A27 - A40 and A52, wherein the third unnatural base or the fourth unnatural base is
[0099] Embodiment A54 is the method of Embodiment A53, wherein the third unnatural base is
[0100] Embodiment A55 is the method of Embodiment A54, wherein the fourth unnatural base is
[0101] Embodiment A56 is the method of any one of Embodiments A27 - A52, wherein the third unnatural base or the fourth unnatural base is
[0102] Embodiment A57 is the method of Embodiment A56, wherein the third unnatural base is
[0103] Embodiment A58 is the method of Embodiment A56, wherein the fourth unnatural base is
[0104] Embodiment A59 is the method of any one of Embodiments A27 - A40, wherein the first unnatural base is the second unnatural base is the third unnatural base is and the fourth unnatural base is
[0105] Embodiment A60 is the method of any one of Embodiments A27 - A40, wherein the first unnatural base is the second unnatural base is the third unnatural base is and the fourth unnatural base is
[0106] Embodiment A61 is the method of any one of Embodiments A27 - A40, wherein the first unnatural base is the second unnatural base is the third unnatural base is and the fourth unnatural base is
[0107] Embodiment A62 is the method of any one of Embodiments A27 - A40, wherein the third unnatural base is
[0108] Embodiment A63 is the method of any one of Embodiments A27 - A51, wherein the fourth unnatural base is
[0109] Embodiment A64 is the method of any one of Embodiments A27 - A40, wherein the first unnatural base is the second unnatural base is the third unnatural base is and the fourth unnatural base is
[0110] Embodiment A65 is a method of any one of embodiments A27 - A64, wherein the third unnatural base and the fourth unnatural base comprise ribose.
[0111] Embodiment A66 is a method of any one of embodiments A27 - A64, wherein the third unnatural base and the fourth unnatural base comprise deoxyribose.
[0112] Embodiment A65 is a method of any one of embodiments A27 - A66, wherein the first unnatural base and the second unnatural base comprise deoxyribose.
[0113] Embodiment A68 is a method of any one of embodiments A27 - A64, wherein the first unnatural base and the second unnatural base comprise deoxyribose, and the third unnatural base and the fourth unnatural base comprise ribose.
[0114] Embodiment A69 is a method of any one of embodiments A27 - A40, wherein the DNA comprises at least one unnatural base pair (UBP) selected from the following:
[0115]
[0116] in
[0117]
[0118] Embodiment A70 is a method of embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM - d5SICS.
[0119] Embodiment A71 is a method of embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO - dTPT3.
[0120] Embodiment A72 is a method of embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM - dTPT3.
[0121] Embodiment A73 is a method of embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dPTMO - dTPT3.
[0122] Embodiment A74 is a method of embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM - dTAT1.
[0123] Embodiment A75 is the method of Embodiment A69, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.
[0124] Embodiment A76 is the method of any one of Embodiments A27 - A40, wherein the DNA comprises at least one unnatural base pair (UBP) selected from:
[0125]
[0126] and
[0127] wherein the mRNA and / or the tRNA comprises at least one unnatural base selected from:
[0128]
[0129] Embodiment A77 is the method according to Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-d5SICS.
[0130] Embodiment A78 is the method of Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTPT3.
[0131] Embodiment A79 is the method of Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3.
[0132] Embodiment A80 is the method of Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dPTMO-dTPT3.
[0133] Embodiment A81 is the method of Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1.
[0134] Embodiment A82 is the method of Embodiment A76, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.
[0135] Embodiment A83 is the method of any one of Embodiments A76 - A82, wherein the mRNA and the tRNA comprise unnatural bases selected from the unnatural bases.
[0136] Embodiment A84 is the method of Embodiment A83, wherein the mRNA and the tRNA comprise unnatural bases selected from unnatural base.
[0137] Embodiment A85 is the method of Embodiment A83, wherein the mRNA comprises being unnatural base.
[0138] Embodiment A86 is the method of Embodiment A83, wherein the mRNA comprises being unnatural base.
[0139] Embodiment A87 is the method of Embodiment A83, wherein the mRNA comprises being unnatural base.
[0140] Embodiment A88 is the method of any one of Embodiments A76 - A87, wherein the tRNA comprises being selected from unnatural base.
[0141] Embodiment A89 is the method of Embodiment A88, wherein the tRNA comprises being unnatural base.
[0142] Embodiment A90 is the method of A88, wherein the tRNA comprises being unnatural base.
[0143] Embodiment A91 is the method of any one of Embodiments A76 - A87, wherein the tRNA comprises being unnatural base.
[0144] Embodiment A92 is the method of any one of Embodiments A27 - A40, wherein the first unnatural base comprises dCNMO, and the second unnatural base comprises dTPT3.
[0145] Embodiment A93 is the method of any one of Embodiments A27 - 40 and A92, wherein the third unnatural base comprises NaM, and the second unnatural base comprises TAT1.
[0146] Embodiment A94 is the method of any one of Embodiments A27 - A93, wherein the first unnatural base or the second unnatural base is recognized by DNA polymerase.
[0147] Embodiment A95 is the method of any one of Embodiments A27 - A94, wherein the third unnatural base or the fourth unnatural base is recognized by RNA polymerase.
[0148] Embodiment A96 is the method of any one of embodiments A27 - A95, wherein the mRNA is transcribed, and the method further comprises translating the mRNA into a protein, wherein the protein comprises an unnatural amino acid at a position corresponding to the mRNA codon comprising the third unnatural base.
[0149] Embodiment A97 is the method of any one of embodiments A27 - A96, wherein the protein comprises at least two unnatural amino acids.
[0150] Embodiment A98 is the method of any one of embodiments A27 - A96, wherein the protein comprises at least three unnatural amino acids.
[0151] Embodiment A99 is the method of any one of embodiments A27 - A98, wherein the protein comprises at least two different unnatural amino acids.
[0152] Embodiment A100 is the method of any one of embodiments A27 - A98, wherein the protein comprises at least three different unnatural amino acids.
[0153] Embodiment A101 is the method of any one of embodiments A27 - A100, wherein the at least one unnatural amino acid:
[0154] is a lysine analogue;
[0155] comprises an aromatic side chain;
[0156] comprises an azide group;
[0157] comprises an alkynyl group; or
[0158] comprises an aldehyde or ketone group.
[0159] Embodiment A102 is the method of any one of embodiments A27 - A101, wherein the at least one unnatural amino acid does not comprise an aromatic side chain.
[0160] Embodiment A103 is the method of any one of embodiments A27 - A102, wherein the at least one unnatural amino acid includes N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propynylethoxy-carbonyl-L-lysine (PraK), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propynyloxy phenylalanine, p-propynyl-phenylalanine, 3-methyl-phenylalanine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)seleno)propanoic acid, 2-amino-3-(phenylseleno)propanoic acid, or selenocysteine.
[0161] Embodiment A104 is the method of embodiment A102 or A103, wherein the at least one unnatural amino acid includes N6-azidoethoxy-carbonyl-L-lysine (AzK) or N6-propynylethoxy-carbonyl-L-lysine (PraK).
[0162] Embodiment A105 is the method of embodiment A104, wherein the at least one unnatural amino acid includes N6-azidoethoxy-carbonyl-L-lysine (AzK).
[0163] Embodiment A106 is the method of embodiment A104, wherein the at least one unnatural amino acid includes N6-propynylethoxy-carbonyl-L-lysine (PraK).
[0164] Embodiment A107 is an mRNA produced by the method of any one of embodiments A27 - A106.
[0165] Embodiment A108 is a tRNA produced by the method of any one of embodiments A27 - A106.
[0166] Embodiment A109 is a protein encoded by the mRNA of Embodiment A107, and the protein contains unnatural amino acids at positions corresponding to the mRNA codons containing the third unnatural base.
[0167] Embodiment A110 is a semi-synthetic organism that contains an expanded genetic alphabet, wherein the genetic alphabet contains at least three different unnatural bases.
[0168] Embodiment A111 is the semi-synthetic organism of Embodiment A110, wherein the organism includes a microorganism, optionally wherein the microorganism is Escherichia coli.
[0169] Embodiment A112 is the semi-synthetic organism of Embodiment A110 or A111, wherein the organism contains DNA containing at least one unnatural nucleobase selected from the following:
[0170]
[0171] Embodiment A113 is the semi-synthetic organism of any one of Embodiments A110 - A113, wherein the DNA contains at least one unnatural base pair (UBP),
[0172] wherein the unnatural base pair (UBP) is dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1, d5FM-dTAT1 or dNaM-dTAT1.
[0173] Embodiment A114 is the semi-synthetic organism of Embodiment A112, wherein the DNA contains at least one unnatural nucleobase that is
[0174] Embodiment A115 is the semi-synthetic organism of any one of Embodiments A110 - A115, wherein the organism expresses a heterologous nucleoside triphosphate transporter.
[0175] Embodiment A116 is the semi-synthetic organism of Embodiment A115, wherein the heterologous nucleoside triphosphate transporter is PtNTT2.
[0176] Embodiment A117 is the semi-synthetic organism of any one of Embodiments A110 - A116, wherein the organism also expresses a heterologous tRNA synthetase.
[0177] Embodiment A118 is the semi-synthetic organism of Embodiment A117, wherein the heterologous tRNA synthetase is Methanosarcina barkeri pyrrolysyl-tRNA synthetase (Mb PylRS).
[0178] Embodiment A119 is a semi-synthetic organism of any one of embodiments A110 - A118, wherein the organism further expresses a heterologous RNA polymerase.
[0179] Embodiment A120 is a semi-synthetic organism of embodiment A119, wherein the heterologous RNA polymerase is T7 RNAP.
[0180] Embodiment A121 is a semi-synthetic organism of any one of embodiments A110 - A120, wherein the organism does not express a protein having DNA recombination repair function.
[0181] Embodiment A122 is a semi-synthetic organism of embodiment A121, wherein the organism does not express RecA.
[0182] Embodiment A123 is a semi-synthetic organism of any one of embodiments A110 - A122, the semi-synthetic organism further comprising heterologous mRNA.
[0183] Embodiment A124 is a semi-synthetic organism of embodiment A123, wherein the heterologous mRNA comprises at least one unnatural base selected from the group consisting of.
[0184] Embodiment A125 is a semi-synthetic organism of any one of embodiments A110 - A125, the semi-synthetic organism further comprising heterologous tRNA.
[0185] Embodiment A126 is a semi-synthetic organism of embodiment A125, wherein the heterologous tRNA comprises at least one unnatural base selected from the group consisting of.
[0186] Embodiment A127 is a method for transcribing DNA, the method comprising:
[0187] providing one or more DNAs, the one or more DNAs comprising (1) a gene encoding a protein, wherein the template strand of the gene encoding the protein comprises a first unnatural base and (2) a gene encoding a tRNA, wherein the template strand of the gene encoding the tRNA comprises a second unnatural base capable of forming a base pair with the first unnatural base;
[0188] transcribing the gene encoding the protein to incorporate a third unnatural base into the mRNA, the third unnatural base capable of forming a first unnatural base pair with the first unnatural base;
[0189] Transcribing the gene encoding the tRNA to incorporate a fourth unnatural base into the tRNA, wherein the fourth unnatural base is capable of forming a second unnatural base pair with the second unnatural base, and wherein the first unnatural base pair and the second unnatural base pair are different.
[0190] Embodiment A128 is the method of embodiment A127, the method further comprising translating a protein from the mRNA using the tRNA, wherein the protein comprises an unnatural amino acid at a position corresponding to a codon in the mRNA that contains the third unnatural base.
[0191] Embodiment A129 is a method of replicating DNA, the method comprising:
[0192] Providing DNA that comprises (1) a gene encoding a protein, wherein the template strand of the gene encoding the protein comprises a first unnatural base and (2) a gene encoding a tRNA, wherein the template strand of the gene encoding the tRNA comprises a second unnatural base that is capable of forming a base pair with the first unnatural base; and
[0193] Replicating the DNA to incorporate a first alternative unnatural base in place of the first unnatural base, and / or to incorporate a second alternative unnatural base in place of the second unnatural base;
[0194] wherein the method optionally further comprises:
[0195] Transcribing the gene encoding the protein to incorporate a third unnatural base into the mRNA, the third unnatural base being capable of forming a first unnatural base pair with the first unnatural base and / or the first alternative unnatural base; and / or
[0196] Transcribing the gene encoding the tRNA to incorporate a fourth unnatural base into the tRNA, wherein the fourth unnatural base is capable of forming a second unnatural base pair with the second unnatural base and / or the second alternative unnatural base, and wherein the first unnatural base pair and the second unnatural base pair are different.
[0197] Embodiment A130 is the method of embodiment A129, the method further comprising transcribing the gene encoding the protein to incorporate a third unnatural base into the mRNA, the third unnatural base being capable of forming a first unnatural base pair with the first unnatural base and / or the first alternative unnatural base.
[0198] Embodiment A131 is the method of Embodiment A129 or A130, the method further comprising transcribing the gene encoding the tRNA to incorporate a fourth unnatural base into the tRNA, wherein the fourth unnatural base is capable of forming a second unnatural base pair with the second unnatural base and / or the second alternative unnatural base, wherein the first unnatural base pair and the second unnatural base pair are different.
[0199] Embodiment A132 is the method of any one of Embodiments A127 - A131, wherein the method comprises using a semi-synthetic organism.
[0200] Embodiment A133 is the method of Embodiment A132, wherein the organism comprises a microorganism.
[0201] Embodiment A134 is the method of Embodiment A132 or A133, wherein the method is an in vivo method, comprising using a semi-synthetic organism that is a bacterium.
[0202] Embodiment A135 is the method of Embodiment A134, wherein the organism comprises a Gram-positive bacterium.
[0203] Embodiment A136 is the method of Embodiment A134, wherein the organism comprises a Gram-negative bacterium.
[0204] Embodiment A137 is the method of Embodiments A132 - A134, wherein the organism comprises Escherichia coli.
[0205] Embodiment A138 is the method of any one of Embodiments A127 - A137, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises
[0206]
[0207] Embodiment A139 is the method of Embodiment A138, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises:
[0208]
[0209] Embodiment A140 is the method of Embodiment A138, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises:
[0210]
[0211] Embodiment A141 is the method of Embodiment A138, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises
[0212] Embodiment A142 is the method of Embodiment A138, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is
[0213] Embodiment A143 is the method of Embodiment A138, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base comprises:
[0214]
[0215] Embodiment A144 is the method of Embodiment A138, wherein the first unnatural base or the second unnatural base is
[0216] Embodiment A145 is the method of Embodiment A138, wherein the first unnatural base or the second unnatural base is
[0217] Embodiment A146 is the method of Embodiment A138, wherein the first unnatural base is and the second unnatural base is
[0218] Embodiment A147 is the method of Embodiment A138, wherein the first unnatural base is and the second unnatural base is
[0219] Embodiment A148 is the method of any one of Embodiments A138, A146, and A147, wherein the third unnatural base or the fourth unnatural base is
[0220] Embodiment A149 is the method of Embodiment A148, wherein the third unnatural base is
[0221] Embodiment A150 is the method of Embodiment A148, wherein the fourth unnatural base is
[0222] Embodiment A151 is the method of any one of embodiments A138, A146, and A147, wherein the third unnatural base or the fourth unnatural base is
[0223] Embodiment A152 is the method of embodiment A151, wherein the third unnatural base is
[0224] Embodiment A153 is the method of embodiment A151, wherein the fourth unnatural base is
[0225] Embodiment A154 is the method of any one of embodiments A138, wherein the first unnatural base is The second unnatural base is The third unnatural base is And the fourth unnatural base is
[0226] Embodiment A155 is the method of embodiment A138, wherein the first unnatural base is The second unnatural base is The third unnatural base is And the fourth unnatural base is
[0227] Embodiment A is the method of embodiment A138, wherein the first unnatural base is The second unnatural base is The third unnatural base is And the fourth unnatural base is
[0228] Embodiment A157 is the method of embodiment A138, wherein the third unnatural base is
[0229] Embodiment A158 is the method of embodiment A138, wherein the fourth unnatural base is
[0230] Embodiment A159 is the method of embodiment A138, wherein the first unnatural base is The second unnatural base is The third unnatural base is And the fourth unnatural base is
[0231] Embodiment A160 is a method of any one of embodiments A127 - A159, wherein the third unnatural base and the fourth unnatural base comprise ribose.
[0232] Embodiment A161 is a method of any one of embodiments A127 - A159, wherein the third unnatural base and the fourth unnatural base comprise deoxyribose.
[0233] Embodiment A162 is a method of any one of embodiments A127 - A161, wherein the first unnatural base and the second unnatural base comprise deoxyribose.
[0234] Embodiment A163 is a method of any one of embodiments A127 - A159, wherein the first unnatural base and the second unnatural base comprise deoxyribose, and the third unnatural base and the fourth unnatural base comprise ribose.
[0235] Embodiment A164 is a method of any one of embodiments A127 - A137, wherein the DNA template comprises at least one unnatural base pair (UBP) selected from the following:
[0236]
[0237] Embodiment A165 is a method of embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) of dNaM - d5SICS.
[0238] Embodiment A166 is a method of embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) of dCNMO - dTPT3.
[0239] Embodiment A167 is a method of embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) of dNaM - dTPT3.
[0240] Embodiment A168 is a method of embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) of dNaM - dTAT1.
[0241] Embodiment A169 is a method of embodiment A164, wherein the DNA template comprises at least one unnatural base pair (UBP) of dCNMO - dTAT1.
[0242] Embodiment A170 is a method of any one of embodiments A127 - A137, wherein the DNA template comprises at least one unnatural base pair (UBP) selected from the following:
[0243]
[0244] and
[0245] wherein the mRNA and the tRNA comprise at least one unnatural base selected from the following:
[0246]
[0247] Embodiment A171 is the method according to Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM-d5SICS.
[0248] Embodiment A172 is the method of Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dCNMO-dTPT3.
[0249] Embodiment A173 is the method of Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM-dTPT3.
[0250] Embodiment A174 is the method of Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dNaM-dTAT1.
[0251] Embodiment A175 is the method of Embodiment A170, wherein the DNA template comprises at least one unnatural base pair (UBP) which is dCNMO-dTAT1.
[0252] Embodiment A176 is the method of any one of Embodiments A127 - A175, wherein the mRNA and the tRNA comprise unnatural bases selected from unnatural bases.
[0253] Embodiment A177 is the method of Embodiment A176, wherein the mRNA and the tRNA comprise unnatural bases selected from unnatural bases.
[0254] Embodiment A178 is the method of Embodiment A176, wherein the mRNA comprises an unnatural base which is unnatural base.
[0255] Embodiment A179 is the method of Embodiment A176, wherein the mRNA comprises an unnatural base which is unnatural base.
[0256] Embodiment A180 is the method of Embodiment A176, wherein the mRNA comprises an unnatural base which is non-natural bases.
[0257] Embodiment A181 is the method of Embodiment A176, wherein the tRNA comprises a non-natural base selected from non-natural bases.
[0258] Embodiment A182 is the method of Embodiment A176, wherein the tRNA comprises a non-natural base that is non-natural bases.
[0259] Embodiment A183 is the method of Embodiment A176, wherein the tRNA comprises a non-natural base that is non-natural bases.
[0260] Embodiment A184 is the method of Embodiment A176, wherein the tRNA comprises a non-natural base that is non-natural bases.
[0261] Embodiment A185 is the method of any one of Embodiments A127 - A137, wherein the first non-natural base comprises dCNMO, and the second non-natural base comprises dTPT3.
[0262] Embodiment A186 is the method of any one of Embodiments A127 - A137, wherein the third non-natural base comprises NaM, and the second non-natural base comprises TAT1.
[0263] Embodiment A187 is the method of any one of Embodiments A127 - A186, wherein the protein comprises at least two non-natural amino acids.
[0264] Embodiment A188 is the method of any one of Embodiments A127 - A186, wherein the protein comprises at least three non-natural amino acids.
[0265] Embodiment A189 is the method of any one of Embodiments A127 - A186, wherein the protein comprises at least two different non-natural amino acids.
[0266] Embodiment A190 is the method of any one of Embodiments A127 - A186, wherein the protein comprises at least three different non-natural amino acids.
[0267] Embodiment A191 is the method of any one of Embodiments A127 - A190, wherein the at least one non-natural amino acid:
[0268] is a lysine analogue;
[0269] comprises an aromatic side chain;
[0270] Containing an azide group;
[0271] Containing an alkynyl group; or
[0272] Containing an aldehyde group or a ketone group.
[0273] Embodiment A192 is the method of any one of embodiments A127 - A191, wherein the at least one unnatural amino acid does not contain an aromatic side chain.
[0274] Embodiment A193 is the method of embodiment A191 or A192, wherein the at least one unnatural amino acid comprises N6 - azidoethoxy - carbonyl - L - lysine (AzK) or N6 - propargylethoxy - carbonyl - L - lysine (PraK).
[0275] Embodiment A194 is the method of embodiment A193, wherein the at least one unnatural amino acid comprises N6 - azidoethoxy - carbonyl - L - lysine (AzK).
[0276] Embodiment A195 is the method of embodiment A193, wherein the at least one unnatural amino acid comprises N6 - propargylethoxy - carbonyl - L - lysine (PraK). Brief Description of the Drawings
[0277] Aspects of the present invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description of illustrative embodiments that set forth the principles of the invention, in the accompanying drawings, in which:
[0278] Figure 1A Shows a workflow for using unnatural base pairs (UBPs) to site - specifically incorporate non - canonical amino acids (ncAAs) into proteins using unnatural X - Y base pairs. The incorporation of three ncAAs into the protein is shown only as an example; any number of ncAAs can be incorporated.
[0279] Figure 1B Depicts an unnatural base pair (UBP).
[0280] Figure 2 Depicts a dXTP analogue. For clarity, ribose and phosphate are omitted.
[0281] Figure 3A Shows a UBP retention (%) plot in the sfGFP gene to optimize the incorporation of AzK into sfGFP using various dXTPs. Each bar represents the mean, where the error bars indicate the standard error (n = 3). The open circles represent the data from each independent experiment. Asterisks indicate that the cells could not grow under the indicated conditions.
[0282] Figure 3B Shows the UBP retention (%) profile in the tRNAPyl gene to optimize the incorporation of AzK into sfGFP using various dXTPs. Each bar represents the mean, where the error bars indicate the standard error (n = 3). Open circles represent the data from each independent experiment. Asterisks indicate that the cells were unable to grow under the indicated conditions.
[0283] Figure 3C Shows the relative sfGFP fluorescence (normalized to cell growth (relative fluorescence units (RFU) / OD600)) observed in the presence or absence of AzK to optimize the incorporation of AzK into sfGFP using various dXTPs. Each bar represents the mean, where the error bars indicate the standard error (n = 3). Open circles represent the data from each independent experiment. Asterisks indicate that the cells were unable to grow under the indicated conditions.
[0284] Figure 3D Shows the relative protein shift (%) measured by Western blot to optimize the incorporation of AzK into sfGFP using various dXTPs. Each bar represents the mean, where the error bars indicate the standard error (n = 3). Open circles represent the data from each independent experiment. Asterisks indicate that the cells were unable to grow under the indicated conditions.
[0285] Figure 4A Depicts ribonucleotide XTP analogs. For clarity, ribose and phosphate esters are omitted.
[0286] Figure 4B Depicts ribonucleotide YTP analogs. For clarity, ribose and phosphate esters are omitted.
[0287] Figure 5A Shows the SAR analysis of translation using various unnatural ribonucleotides (for the incorporation of AzK into sfGFP), where the total sfGFP fluorescence (RFU) observed for the XTP analogs in the presence of AzK is on the y-axis. Each bar represents the mean, where the error bars indicate the standard error (n = 4). Open circles represent the data from each independent experiment.
[0288] Figure 5B Shows the SAR analysis of translation using various unnatural ribonucleotides (for the incorporation of AzK into sfGFP), where the protein shift (%) measured by Western blot for the XTP analogs is on the y-axis. Each bar represents the mean, where the error bars indicate the standard error (n = 4). Open circles represent the data from each independent experiment.
[0289] Figure 5CShows a SAR analysis plot of translation using various unnatural ribonucleotides (for incorporation of AzK into sfGFP), where the total sfGFP fluorescence (RFU) observed for YTP analogs in the presence of AzK is on the y - axis. Each bar represents the mean, where the error bars indicate the standard error (n = 4). Open circles represent the data for each independent experiment.
[0290] Figure 5D Shows a SAR analysis plot of translation using various unnatural ribonucleotides (for incorporation of AzK into sfGFP), where the protein shift (%) measured by western blot for YTP analogs is on the y - axis. Each bar represents the mean, where the error bars indicate the standard error (n = 4). Open circles represent the data for each independent experiment.
[0291] Figure 6A Shows an optimization plot of the unnatural ribonucleotide triphosphate concentration against the total sfGFP fluorescence (RFU) varying with the NaMTP and TAT1TP concentrations (μM). Error bars indicate the standard error for each value (n = 3).
[0292] Figure 6B Shows an optimization plot of the unnatural ribonucleotide triphosphate concentration against the total sfGFP fluorescence (RFU) varying with the 5FMTP and TAT1TP (μM) concentrations. Error bars indicate the standard error for each value (n = 3).
[0293] Figure 6C Shows an optimization plot of the unnatural ribonucleotide triphosphate concentration against the protein shift (%) varying with the NaMTP and TAT1TP concentrations (μM). Error bars indicate the standard error for each value (n = 3).
[0294] Figure 6D Shows an optimization plot of the unnatural ribonucleotide triphosphate concentration against the protein shift (%) varying with the 5FMTP and TAT1TP concentrations (μM). Error bars indicate the standard error for each value (n = 3).
[0295] Figure 7A Is a plot of storage and retrieval of higher - density unnatural information for various unnatural base pairs, codon positions, and total sfGFP fluorescence (RFU) observed in the presence of AzK. For the bar graph, each bar represents the mean, where the error bars indicate the standard error (n = 4), and open circles represent the data for each independent experiment.
[0296] Figure 7BGraph showing the storage and retrieval of higher density unnatural information for various unnatural base pairs, codon positions, and protein shifts (%) (measured by Western blotting) observed in the presence of AzK. For bar graphs, each bar represents the mean, where error bars indicate standard error (n = 4), and open circles represent the data for each independent experiment.
[0297] Figure 7C Depicts a representative spectrum of the quantitative HRMS analysis of a triple-labeled protein generated using dCNMOdTPT3 / NaMTP,TAT1TP. Peak labels show the deconvoluted molecular weight of the intact protein, where amino acid residues at positions 149, 151, and 153 are shown and the quantification (%) of each peak is shown below (n = 3).
[0298] Figure 8A Shows exemplary unnatural amino acids. This figure is adapted from Young et al., “Beyond the canonical 20 amino acids: expanding the genetic lexicon,” J. of Biological Chemistry 285(15):11039 - 11044 (2010) Figure 2 。
[0299] Figure 8B Shows an exemplary unnatural amino acid lysine derivative.
[0300] Figure 8C Shows an exemplary unnatural amino acid phenylalanine derivative.
[0301] Figure 8D - Figure 8G Shows exemplary unnatural amino acids. These unnatural amino acids (UAAs) have been genetically encoded into proteins ( Figure 8D -UAA#1 - 42; Figure 8E -UAA#43 - 89; Figure 8F -UAA#90 - 128; Figure 8G -UAA#129 - 167). Figure 8D - Figure 8G Taken from Table 1 of Dumas et al., Chemical Science 2015, 6, 50 - 69. Detailed Description
[0302] Specific Terms
[0303] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and do not limit any claimed subject matter. In the event of any inconsistency between any material incorporated herein by reference and the explicit content of the present disclosure text, the explicit content shall prevail. In this application, unless otherwise explicitly stated, the use of the singular includes the plural meaning. It must be noted that as used in the specification and the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural referents. In this application, unless otherwise stated, the use of "or" means "and / or". In addition, the use of the term "including" and other forms such as "include", "includes", and "included" is non-restrictive.
[0304] As used herein, ranges and amounts can be expressed as "about" a particular value or range. About also includes the exact amount. Thus, "about 5 μL" refers to "about 5 μL" and also constitutes a description of "5 μL". Generally, the term "about" includes amounts that can be expected within experimental error.
[0305] As used herein, in the context of synthetic methods, phrases such as "under conditions suitable to provide..." or "under conditions sufficient to produce..." refer to reaction conditions that can be varied within the ordinary skill of the experimenter, such as time, temperature, solvent, reactant concentration, etc., to provide a useful amount or yield of the reaction product. The desired reaction product does not necessarily have to be the only reaction product or the starting materials do not necessarily have to be completely consumed, as long as the desired reaction product can be isolated or otherwise further used.
[0306] "Chemically feasible" refers to a bonding arrangement or compound that does not violate the generally understood rules of organic structure; for example, in some cases, a structure within the claim definition that contains a pentavalent carbon atom not found in nature should be understood to be outside the scope of the claim. The structures disclosed herein, in all of their embodiments, are intended to include only "chemically feasible" structures, and any enumerated structures that are not chemically feasible, such as structures shown to have variable atoms or groups, are not intended to be disclosed or claimed herein.
[0307] As used herein, the term "analogue" of a chemical structure refers to a chemical structure that maintains substantial similarity to a parent structure but may not be readily synthesized from the parent structure. In some embodiments, nucleotide analogues are unnatural nucleotides. In some embodiments, nucleoside analogues are unnatural nucleosides. Related chemical structures that are readily synthesized from a parent chemical structure are referred to as "derivatives".
[0308] As used herein, "base" or "nucleobase" refers to at least the nucleobase portion of a nucleoside or nucleotide (nucleosides and nucleotides encompassing ribose or deoxyribose variants), where the nucleoside or nucleotide may in some instances contain further modifications to the sugar portion of the nucleoside or nucleotide. In some instances, "base" is also used to represent the entire nucleoside or nucleotide (e.g., a "base" may be incorporated into DNA by a DNA polymerase or into RNA by an RNA polymerase). However, unless the context requires, the term "base" should not be construed to necessarily represent the entire nucleoside or nucleotide. In the chemical structures of bases or nucleobases provided herein, only the base of the nucleoside or nucleotide is shown, and the sugar portion and any optional phosphate residues are omitted for clarity. As used in the chemical structures of bases or nucleobases provided herein, a wavy line represents a linkage to a nucleoside or nucleotide, where the sugar portion of the nucleoside or nucleotide may be further modified. In some embodiments, the wavy line represents the attachment of a base or nucleobase to the sugar portion (such as a pentose) of a nucleoside or nucleotide. In some embodiments, the pentose is ribose or deoxyribose.
[0309] In some embodiments, a nucleobase is generally the heterocyclic base portion of a nucleoside. Nucleobases can be naturally occurring, can be modified, can have no similarity to natural bases, and / or can be synthetic, such as by organic synthesis. In certain embodiments, a nucleobase comprises any atom or group of atoms in a nucleoside or nucleotide, where the atom or group of atoms is capable of interacting with the base of another nucleic acid with or without hydrogen bonding. In certain embodiments, an unnatural nucleobase is not derived from a natural nucleobase. It should be noted that unnatural nucleobases do not necessarily have base characteristics, but for simplicity they are referred to as nucleobases. In some embodiments, when referring to a nucleobase, "(d)" indicates that the nucleobase can be attached to either deoxyribose or ribose, while "d" without parentheses indicates that the nucleobase is attached to deoxyribose.
[0310] In some embodiments, a nucleoside is a compound comprising a nucleobase portion and a sugar portion. Nucleosides include, but are not limited to, naturally occurring nucleosides (such as those found in DNA and RNA), abasic nucleosides, modified nucleosides, and nucleosides having mimetic bases and / or sugar groups. Nucleosides include nucleosides containing any kind of substituents. A nucleoside can be a glycoside compound formed by a glycosidic linkage between a nucleic acid base and a reducing group of a sugar.
[0311] As used herein, "nucleotide" refers to a compound comprising a nucleoside moiety and a phosphate moiety. Exemplary natural nucleotides include, but are not limited to, adenosine triphosphate (ATP), uridine triphosphate (UTP), cytidine triphosphate (CTP), guanosine triphosphate (GTP), adenosine diphosphate (ADP), uridine diphosphate (UDP), cytidine diphosphate (CDP), guanosine diphosphate (GDP), adenosine monophosphate (AMP), uridine monophosphate (UMP), cytidine monophosphate (CMP), and guanosine monophosphate (GMP), deoxyadenosine triphosphate (dATP), deoxythymidine triphosphate (dTTP), deoxycytidine triphosphate (dCTP), deoxyguanosine triphosphate (dGTP), deoxyadenosine diphosphate (dADP), thymidine diphosphate (dTDP), deoxycytidine diphosphate (dCDP), deoxyguanosine diphosphate (dGDP), deoxyadenosine monophosphate (dAMP), deoxythymidine monophosphate (dTMP), deoxycytidine monophosphate (dCMP), and deoxyguanosine monophosphate (dGMP). Exemplary natural deoxyribonucleotides that contain deoxyribose as the sugar moiety include dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Exemplary natural ribonucleotides that contain ribose as the sugar moiety include ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, and GMP.
[0312] As used herein, polynucleotide refers to DNA, RNA, DNA-like or RNA-like polymers (such as peptide nucleic acid (PNA), locked nucleic acid (LNA), phosphorothioate, etc.), examples of which are well known in the art and may contain unnatural bases. Polynucleotides can be synthesized in an automated synthesizer, for example, using phosphoramidite chemistry or other chemical routes suitable for use in the synthesizer.
[0313] DNA includes, but is not limited to, complementary DNA (cDNA) and genomic DNA (gDNA). DNA can be attached to another molecule (including, but not limited to, RNA and peptides) by covalent or non-covalent means. RNA includes coding RNA, such as messenger RNA (mRNA). RNA also includes non-coding RNA, such as ribosomal RNA (rRNA). RNA also includes transfer RNA (tRNA), RNA interference (RNAi), small nucleolar RNA (snoRNA), microRNA (miRNA), small interfering RNA (siRNA) (also known as short interfering RNA), small nuclear RNA (snRNA), extracellular RNA (exRNA), PIWI-interacting RNA (piRNA), and long non-coding RNA (long ncRNA). In some embodiments, the RNA is rRNA, tRNA, RNAi, snoRNA, microRNA, siRNA, snRNA, exRNA, piRNA, long ncRNA, or any combination or hybrid thereof. In some cases, the RNA is a component of a ribozyme. DNA and RNA can be in any form, including, but not limited to, linear, circular, supercoiled, single-stranded, and double-stranded.
[0314] Peptide nucleic acid (PNA) is a synthetic DNA / RNA analogue in which a peptide-like backbone replaces the sugar-phosphate backbone of DNA or RNA. PNA oligomers show higher binding strength and higher specificity when binding to complementary DNA, where PNA / DNA base mismatches result in more destabilization compared to similar mismatches in DNA / DNA duplexes. This binding strength and specificity also apply to PNA / RNA duplexes. PNA is not readily recognized by nucleases or proteases, making them resistant to enzymatic degradation. PNA is also stable over a wide pH range. See also Nielsen PE, Egholm M, Berg RH, Buchardt O (December 1991). “Sequence-selective recognition of DNA by strand displacement with a thymine-substituted polyamide”, Science 254(5037):1497-500. doi:10.1126 / science.1962210. PMID 1962210; and Egholm M, Buchardt O, Christensen L, Behrens C, Freier SM, Driver DA, Berg RH, Kim SK, Norden B, and Nielsen PE (1993), “PNA Hybridizes to Complementary Oligonucleotides Obeying the Watson-Crick Hydrogen Bonding Rules”. Nature 365(6446):566-8. doi:10.1038 / 365566a0. PMID 7692304; the disclosures of each of these documents are hereby incorporated by reference in their entirety.
[0315] Locked nucleic acid (LNA) is a modified RNA nucleotide in which the ribose part of the LNA nucleotide is modified with an extra bridge connecting the 2'-oxygen and 4'-carbon. The bridge "locks" the ribose in the C3'-endo (North) conformation, which is typically found in A-form duplexes. LNA nucleotides can be mixed with DNA or RNA residues in an oligonucleotide whenever desired. Such oligomers can be chemically synthesized and are commercially available. The locked ribose conformation enhances base stacking and backbone pre-organization. See, for example, Kaur, H; Arora, A; Wengel, J; Maiti, S (2006), “Thermodynamic, Counterion, and Hydration Effects for the Incorporation of Locked Nucleic Acid Nucleotides into DNA Duplexes”, Biochemistry 45(23):7347-55.doi:10.1021 / bi060307w.PMID 16752924; Owczarzy R.; You Y., Groth C.L., Tataurov A.V. (2011), “Stability and mismatch discrimination of locked nucleic acid-DNA duplexes.”, Biochem. 50(43):9352-9367.doi:10.1021 / bi200904e.PMC 3201676.PMID 21928795; Alexei A. Koshkin; Sanjay K. Singh, Poul Nielsen, Vivek K. Rajwanshi, Ravindra Kumar, Michael Meldgaard, Carl Erik Olsen, Jesper Wengel (1998), “LNA (Locked Nucleic Acids): Synthesis of the adenine, cytosine, guanine, 5-methylcytosine, thymine and uracil bicyclonucleoside monomers, oligomerisation, and unprecedented nucleic acid recognition”, Tetrahedron 54(14):3607-30.doi:10.1016 / S0040-4020(98)00094-5; and Satoshi Obika; Daishu Nanbu, Yoshiyuki Hari, Ken-ichiro Morio, Yasuko In, Toshimasa Ishida, Takeshi Imanishi (1997), “Synthesis of 2′-O,4′-C-methyleneuridine and -cytidine. Novel bicyclic nucleosides having a fixed C3'-endo sugar puckering”, Tetrahedron Lett. 38(50):8735-8. doi:10.1016 / S0040-4039(97)10322-7; the disclosures of each of these documents are hereby incorporated by reference in their entirety.
[0316] As used herein, the term “gene” refers to a polynucleotide encoding the synthesis of a gene product such as RNA or protein.
[0317] A molecular beacon or molecular beacon probe is an oligonucleotide hybridization probe that can detect the presence of a specific nucleic acid sequence in a homogeneous solution. A molecular beacon is a hairpin-shaped molecule having an internally quenched fluorophore, the fluorescence of which is restored when they bind to a target nucleic acid sequence. See, e.g., Tyagi S, Kramer FR (1996), “Molecular beacons: probes that fluoresce upon hybridization”, Nat Biotechnol. 14(3):303-8. PMID 9630890; I, Malmberg L, Rennel E, Wik M, AC (April 2000), “Homogeneous scoring of single-nucleotide polymorphisms: comparison of the 5'-nuclease TaqMan assay and Molecular Beacon probes”, Biotechniques 28(4):732-8. PMID 10769752; and Akimitsu Okamoto (2011), “ECHO probes: a concept of fluorescence control for practical nucleic acid sensing”, Chem. Soc. Rev. 40:5815-5828; the disclosures of each of these documents are hereby incorporated by reference in their entirety.
[0318] As used herein, the term “unnatural base” refers to a base other than A, C, G, T, U, and other naturally occurring bases (e.g., 5-methylcytosine, pseudouridine, and inosine).
[0319] As used herein, the term “unnatural base pair” refers to two bases that are bonded to each other and are located on opposite strands of a double-stranded polynucleotide, which can be, for example, at least a partially self-hybridizing molecule or a pair of partially or fully hybridized molecules, where at least one of the two bases is an unnatural base.
[0320] As used herein, a “semicynthetic organism” is an organism that contains unnatural components, such as an expanded genetic alphabet that includes one or more unnatural bases.
[0321] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.
[0322] Methods and compositions containing unnatural base pairs
[0323] Disclosed herein, in certain embodiments, are in vitro and in vivo methods and compositions for generating nucleic acids having an expanded genetic alphabet (Figure 1). In some cases, the nucleic acids encode unnatural proteins, where the unnatural proteins include unnatural amino acids. In some instances, the in vivo methods or compositions described herein use or comprise semi-synthetic organisms. In some cases, the methods include incorporating at least one unnatural base pair (UBP) into one or more nucleic acids. Such base pairs are formed by the pairing between the nucleobases of two nucleosides. In an exemplary workflow, DNA 101 encoding protein 102 and tRNA 103 is transcribed 104 to produce tRNA 106 and mRNA 107, where the template strand coding regions of the protein and tRNA include complementary unnatural nucleobases (X, Y) capable of forming base pairs and / or configured to form base pairs. After the tRNA is loaded with an unnatural amino acid 105, mRNA 107 is translated 108 to produce protein 110 comprising one or more unnatural amino acids 109. In some cases, the methods and compositions described herein allow for site-specific incorporation of unnatural amino acids with high fidelity and yield. Also described herein are semi-synthetic organisms comprising an expanded genetic alphabet, and methods of using semi-synthetic organisms to produce protein products, including those comprising at least one unnatural amino acid residue.
[0324] The selection of non-natural nucleobases allows for the optimization of one or more steps in the methods described herein. For example, nucleobases are selected for efficient replication, transcription, and / or translation. In some cases, more than one non-natural base pair is used in the methods described herein. For example, a first set of nucleobases containing a deoxyribose moiety is used for DNA replication (e.g., a first nucleobase and a second nucleobase configured to form a first base pair), while a second set of nucleobases (e.g., a third nucleobase and a fourth nucleobase, where the third nucleobase and the fourth nucleobase are attached to ribose and are configured to form a second base pair) is used for transcription / translation. In some embodiments, a first set of nucleobases is used to construct a plasmid (e.g., a first nucleobase and a second nucleobase configured to form a first base pair), a second set of nucleobases is used for replication (e.g., a third nucleobase and a fourth nucleobase configured to form a second base pair), and a third set of bases is used for transcription / translation (e.g., a fifth nucleobase and a sixth nucleobase configured to form a third base pair). In some cases, complementary pairing between the nucleobases in the first set and the nucleobases in the second set allows for gene transcription to produce tRNA or protein from a DNA template containing nucleobases from the first set. In some cases, complementary pairing between the nucleobases in the second set (the second base pair) allows for translation by matching tRNA containing non-natural nucleic acids with mRNA. In some instances, the nucleobases in the first set are attached to a deoxyribose moiety. In some instances, the nucleobases in the first set are attached to a ribose moiety. In some cases, the nucleobases in both sets are unique. In some cases, at least one nucleobase is the same in both sets. In some cases, the first nucleobase and the third nucleobase are the same. In some embodiments, the first base pair and the second base pair are not the same. In some instances, the first base pair, the second base pair, and the third base pair are not the same.
[0325] In one aspect, provided herein is a method for in vivo production of a protein comprising a non-natural amino acid, the method comprising:
[0326] transcribing a DNA template comprising a first non-natural base and a second non-natural base to incorporate a third non-natural base into an mRNA, the second non-natural base being complementary to the first non-natural base, capable of forming a base pair with the first non-natural base, and / or being configured to form a base pair with the first non-natural base, the third non-natural base being complementary to the first non-natural base, capable of forming a base pair with the first non-natural base, and / or being configured to form a first non-natural base pair with the first non-natural base;
[0327] Transcribing the DNA template to incorporate a fourth unnatural base into the tRNA, wherein the fourth unnatural base is complementary to the second unnatural base, capable of forming a base pair with the second unnatural base and / or configured to form a second unnatural base pair with the second unnatural base, wherein the first unnatural base pair and the second unnatural base pair are different; and
[0328] Translating a protein from the mRNA and the tRNA, wherein the protein comprises an unnatural amino acid.
[0329] Nucleic acid molecule
[0330] In some embodiments, the nucleic acid (e.g., also referred to herein as the nucleic acid molecule of interest) is from any source or composition, such as DNA, cDNA, gDNA (genomic DNA), RNA, siRNA (short interfering RNA), RNAi, tRNA, mRNA, or rRNA (ribosomal RNA), and is in any form (e.g., linear, circular, supercoiled, single-stranded, double-stranded, etc.). In some embodiments, the nucleic acid comprises nucleotides, nucleosides, or polynucleotides. In some cases, the nucleic acid comprises natural nucleic acids and non-natural nucleic acids. In some cases, the nucleic acid further comprises non-natural nucleic acids, such as DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-natural backbones, etc.). It should be understood that the term "nucleic acid" does not refer to or imply a polynucleotide chain of a specific length, and thus polynucleotides and oligonucleotides are also included in the definition. Exemplary natural nucleotides include, but are not limited to, ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, GMP, dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Exemplary natural deoxyribonucleotides include dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Exemplary natural ribonucleotides include ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, and GMP. For natural RNA, the nucleoside containing uracil is uridine. The nucleic acid is sometimes a vector, plasmid, phagemid, autonomously replicating sequence (ARS), centromere, artificial chromosome, yeast artificial chromosome (e.g., YAC), or other nucleic acid capable of replicating or being replicated in a host cell. In some cases, the non-natural nucleic acid is a nucleic acid analog. In additional cases, the non-natural nucleic acid is from an extracellular source. In other cases, the non-natural nucleic acid can be used in the intracellular space of the organisms provided herein (e.g., genetically modified organisms). In some embodiments, the non-natural nucleotide is not a natural nucleotide. In some embodiments, the nucleotide that does not contain a natural base contains a non-natural nucleobase.
[0331] Non-natural nucleic acid
[0332] Nucleotide analogs or unnatural nucleotides include nucleotides containing a certain type of modification to the base, sugar, or phosphate moiety. The term "modification" (and related grammatical forms such as "modified") does not necessarily imply that the nucleotide analog or unnatural nucleotide is prepared by directly altering a natural nucleotide, but rather that the nucleotide analog or unnatural nucleotide is different from a natural nucleotide. In some embodiments, the modification includes chemical modification. In some cases, the modification occurs at the 3'OH or 5'OH group, at the backbone, at the sugar moiety, or at the nucleobase. In some cases, the modification optionally includes non-naturally occurring linker molecules and / or interstrand or intrastrand crosslinks. In one aspect, the modified nucleic acid includes a modification of one or more of the following: the 3'OH or 5'OH group, the backbone, the sugar moiety, or the nucleobase, and / or the addition of a non-naturally occurring linker molecule. In one aspect, the modified backbone includes a backbone other than the phosphodiester backbone. In one aspect, the modified sugar includes a sugar other than deoxyribose (in modified DNA) or ribose (modified RNA). In one aspect, the modified base includes a base other than adenine, guanine, cytosine, or thymine (in modified DNA) or adenine, guanine, cytosine, or uracil (in modified RNA). In some embodiments, the unnatural nucleotide contains an unnatural base. In some embodiments, the unnatural base is a base having a ring or ring system other than purine or pyrimidine (where purine and pyrimidine encompass purines and pyrimidines having exocyclic substituents), or contains a ring or ring system containing one or more non-nitrogen heteroatoms and / or no nitrogen.
[0333] In some embodiments, the nucleic acid contains at least one modified base. In some cases, the nucleic acid contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more modified bases. In some cases, the modification to the base moiety includes natural and synthetic modifications to adenine (A), cytosine (C), guanine (G), and thymine (T) / uracil (U) as well as different purine or pyrimidine bases. In some embodiments, the modification is a modified form of adenine, guanine, cytosine, or thymine (in modified DNA) or adenine, guanine, cytosine, or uracil (modified RNA).
[0334] Modified bases of unnatural nucleic acids include, but are not limited to, uracil-5-yl, hypoxanthin-9-yl (I), 2-aminoadenin-9-yl, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azauracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other 8-substituted adenines and guanines, 5-halo (especially 5-bromo), 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azadenine, 7-deazaguanine and 7-deazapurine, and 3-deazaguanine and 3-deazapurine. Certain unnatural nucleic acids, such as 5-substituted pyrimidines, 6-azapyrimidines and N-2-substituted purines, N-6-substituted purines, O-6-substituted purines, 2-aminopropyladenine, 5-propynyluracil, 5-propynylcytosine, 5-methylcytosine, those that increase the stability of duplex formation, general nucleic acids, hydrophobic nucleic acids, chimeric nucleic acids, size-expanded nucleic acids, fluorinated nucleic acids, 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6-substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl derivatives and other alkyl derivatives of adenine and guanine, 2-propyl derivatives and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyl (-C≡C-CH3) uracil, 5-propynylcytosine, other alkynyl derivatives of pyrimidine nucleic acids, 6-azauracil, 6-azacytosine, 6-azathymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-mercapto, 8-thioalkyl, 8-hydroxy and other 8-substituted adenines and guanines, 5-halo (especially 5-bromo), 5-trifluoromethyl, other 5-substituted uracils and cytosines, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine, 8-azadenine, 7-deazaguanine, 7-deazapurine, 3-deazaguanine, 3-deazapurine, tricyclic pyrimidines, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4] Benzothiazin-2(3H)-one), G-clamp, phenoxazine cytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one), pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d]pyrimidin-2-one), those in which the purine or pyrimidine base is replaced by other heterocycles, 7-deaza-adenine, 7-deaza-guanine, 2-aminopyridine, 2-pyridone, azacytidine, 5-bromocytidine, bromouracil, 5-chlorocytidine, chlorocytosine, cyclocytidine, cytarabine, 5-fluorocytidine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytidine, 5-iodocytidine, hydroxyurea, iodouracil, 5-nitro-cytidine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil and 5-iodouracil, 2-amino-adenine, 6-thio-guanine, 2-thio-thymine, 4-thio-thymine, 5-propynyl-uracil, 4-thio-uracil, N4-ethylcytidine, 7-deaza-guanine, 7-deaza-8-aza-guanine, 5-hydroxycytidine, 2’-deoxyuridine, 2-amino-2’-deoxyadenosine, and those described in U.S. Patent Nos. 3,687,808; 4,845,205; 4,910,300; 4,948,882; 5,093,232; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121; 5,596,091; 5,614,617; 5,645,985; 5,681,941; 5,750,692; 5,763,588; 5,830,653 and 6,005,096; WO 99 / 62923; Kandimalla et al., (2001) Bioorg. Med. Chem. 9:807-813; The Concise Encyclopedia of Polymer Science and Engineering, Kroschwitz, J.I., ed., John Wiley & Sons, 1990, 858-859; Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613; and Sanghvi, Chapter 15, Antisense Research and Applications, Crooke and Lebleu eds., CRC Press,Those in 1993, 273 - 288. Additional base modifications can be found, for example, in U.S. Patent No. 3,687,808; Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613. In some cases, the unnatural nucleic acids contain, Figure 2 nucleobases. In some cases, the unnatural nucleic acids contain Figure 4A nucleobases. In some cases, the unnatural nucleic acids contain Figure 4B nucleobases.
[0335] Unnatural nucleic acids containing various heterocyclic bases and various sugar moieties (and sugar analogs) are available in the art, and in some instances, the nucleic acids include one or more heterocyclic bases in addition to the five major base components of naturally occurring nucleic acids. For example, in some instances, the heterocyclic bases include uracil - 5 - yl, cytosine - 5 - yl, adenine - 7 - yl, adenine - 8 - yl, guanine - 7 - yl, guanine - 8 - yl, 4 - aminopyrrolo[2.3 - d]pyrimidin - 5 - yl, 2 - amino - 4 - oxopyrrolo[2,3 - d]pyrimidin - 5 - yl, 2 - amino - 4 - oxopyrrolo[2.3 - d]pyrimidin - 3 - yl, wherein the purine is attached to the sugar moiety of the nucleic acid via the 9 - position, the pyrimidine via the 1 - position, the pyrrolopyrimidine via the 7 - position, and the pyrazolopyrimidine via the 1 - position.
[0336] In some embodiments, the modified bases of the unnatural nucleic acids are depicted below, where the wavy line identifies the point of attachment to the sugar (e.g., deoxyribose or ribose) of the nucleoside or nucleotide.
[0337]
[0338]
[0339]
[0340] In some embodiments, the unnatural base (such as at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base in the method for producing a protein containing an unnatural amino acid as described herein) is selected from:
[0341] In some embodiments, the unnatural base is selected from:
[0342] In some embodiments, the unnatural base is selected from: In some embodiments, the unnatural base is selected from
[0343]
[0344] In some embodiments, the unnatural base is In some embodiments, the unnatural base is selected from:
[0345]
[0346] In some embodiments, the unnatural base (such as the first unnatural base or the second unnatural base in the method for producing a protein comprising an unnatural amino acid as described herein) is
[0347] In some embodiments, the unnatural base (such as the first unnatural base or the second unnatural base in the method for producing a protein comprising an unnatural amino acid as described herein) is
[0348] In some embodiments, the first unnatural base is and the second unnatural base is
[0349] In some embodiments, the unnatural base (such as the third unnatural base or the fourth unnatural base in the method for producing a protein comprising an unnatural amino acid as described herein) is In some embodiments, the third unnatural base is In some embodiments, the fourth unnatural base is In some embodiments, the unnatural base (such as the third unnatural base or the fourth unnatural base in the method for producing a protein comprising an unnatural amino acid as described herein) is In some embodiments, the third unnatural base is In some embodiments, the fourth unnatural base is
[0350] In some embodiments, the first unnatural base is The second unnatural base is The third unnatural base is and the fourth unnatural base is In some embodiments, the first unnatural base is The second unnatural base is The third unnatural base is and the fourth unnatural base is In some embodiments, the first unnatural base is The second unnatural base is The third unnatural base is and the fourth unnatural base is In some embodiments, the third unnatural base is In some embodiments, the fourth unnatural base is In some embodiments, the first unnatural base is The second unnatural base is The third unnatural base is and the fourth unnatural base is
[0351] In some embodiments, the third unnatural base and the fourth unnatural base contain ribose. In some embodiments, the third unnatural base and the fourth unnatural base contain deoxyribose. In some embodiments, the first unnatural base and the second unnatural base contain deoxyribose. In some embodiments, the first unnatural base and the second unnatural base contain deoxyribose, and the third unnatural base and the fourth unnatural base contain ribose.
[0352] In some embodiments of the methods for producing proteins containing unnatural amino acids described herein, the DNA comprises at least one unnatural base pair (UBP) selected from the following:
[0353] wherein each sugar moiety is independently any embodiment or variation described herein. In some embodiments, the sugar moieties of both bases in the base pair both contain ribose. In some embodiments, the sugar moieties of both bases in the base pair both contain deoxyribose. In some embodiments, the sugar moiety of one base in the base pair contains ribose, and the sugar moiety of the other base in the base pair contains deoxyribose. In some embodiments, the DNA template comprises at least one unnatural base pair (UBP) that is NaM-5SICS. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is CNMO-TPT3. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is NaM-TPT3. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is NaM-TAT1. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is CNMO-TAT1.
[0354] In some embodiments of the methods for producing proteins containing unnatural amino acids described herein, the DNA comprises at least one unnatural base pair (UBP) selected from the following:
[0355] In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-d5SICS. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dCNMO-dTPT3. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.
[0356] In some embodiments of the methods for producing proteins comprising unnatural amino acids described herein, the DNA comprises at least one unnatural base pair (UBP) selected from:
[0357] wherein each sugar moiety is independently any embodiment or variation described herein; and wherein the mRNA and the tRNA comprise at least one unnatural base selected from:
[0358] In some embodiments, the sugar moieties of both bases in the base pair both comprise ribose. In some embodiments, the sugar moieties of both bases in the base pair both comprise deoxyribose. In some embodiments, the sugar moiety of one base in the base pair comprises ribose, and the sugar moiety of the other base in the base pair comprises deoxyribose. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-d5SICS. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dCNMO-dTPT3. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.
[0359] In some embodiments of the methods for producing proteins comprising unnatural amino acids described herein, the DNA comprises at least one unnatural base pair (UBP) selected from:
[0360] and wherein the mRNA and the tRNA comprise at least one unnatural base selected from:
[0361] In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-d5SICS. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dCNMO-dTPT3. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1. In some embodiments, the DNA comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.
[0362] In some embodiments, the mRNA and the tRNA comprise unnatural bases selected from In some embodiments, the mRNA and the tRNA comprise unnatural bases selected from In some embodiments, the mRNA comprises an unnatural base that is In some embodiments, the mRNA comprises an unnatural base that is In some embodiments, the mRNA comprises an unnatural base that is In some embodiments, the tRNA comprises unnatural bases selected from In some embodiments, the tRNA comprises an unnatural base that is In some embodiments, the tRNA comprises an unnatural base that is In some embodiments, the tRNA comprises an unnatural base that is In some embodiments, the tRNA comprises an unnatural base that is
[0363] In some embodiments of the methods for producing proteins comprising unnatural amino acids described herein, the first unnatural base comprises dCNMO and the second unnatural base comprises dTPT3. In some embodiments, the third unnatural base comprises NaM and the second unnatural base comprises TAT1.
[0364] The present disclosure also provides proteins comprising at least one unnatural amino acid, wherein the proteins are produced according to any of the methods disclosed herein. In some embodiments, the protein comprises at least one unnatural amino acid. In some embodiments, the protein comprises one unnatural amino acid. In some embodiments, the protein comprises two or more unnatural amino acids. In some embodiments, the protein comprises two unnatural amino acids. In some embodiments, the protein comprises three or more unnatural amino acids.
[0365] In some embodiments, the nucleotide analogs are also modified at the phosphate moiety. Modified phosphate moieties include, but are not limited to, those modified at the linkage between two nucleotides and contain, for example, phosphorothioates, chiral phosphorothioates, dithiophosphates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonates (including 3'-alkylene phosphonates) and chiral phosphonates, phosphinates, amino phosphates (including 3'-aminoamino phosphates and aminoalkylamino phosphates, thiocarbamyl phosphates), thiocarbamyl alkyl phosphonates, thiocarbamyl alkyl phosphotriesters, and borane phosphates. It is understood that these phosphate or modified phosphate linkages between two nucleotides are by 3'-5' linkages or 2'-5' linkages and that the linkages have opposite polarities, such as 3'-5' to 5'-3' or 2'-5' to 5'-2'. Also included are the various salt, mixed salt, and free acid forms. Many U.S. patents teach how to prepare and use nucleotides containing modified phosphates and the U.S. patents include, but are not limited to, 3,687,808; 4,469,863; 4,476,301; 5,023,243; 5,177,196; 5,188,897; 5,264,423; 5,276,019; 5,278,302; 5,286,717; 5,321,131; 5,399,676; 5,405,939; 5,453,496; 5,455,233; 5,466,677; 5,476,925; 5,519,126; 5,536,821; 5,541,306; 5,550,111; 5,563,253; 5,571,799; 5,587,361; and 5,625,050; the disclosures of each of these documents are hereby incorporated by reference in their entirety.
[0366] In some embodiments, the unnatural nucleic acids include 2′,3′-dideoxy-2′,3′-didehydro-nucleosides (PCT / US2002 / 006460), 5′-substituted DNA and RNA derivatives (PCT / US2011 / 033961; Saha et al., J. Org. Chem., 1995, 60, 788-789; Wang et al., Bioorganic & Medicinal Chemistry Letters, 1999, 9, 885-890; and Mikhailov et al., Nucleosides & Nucleotides, 1991, 10(1-3), 339-343; Leonid et al., 1995, 14(3-5), 901-905; and Eppacher et al., Helvetica Chimica Acta, 2004, 87, 3004-3020; PCT / JP2000 / 004720; PCT / JP2003 / 002342; PCT / JP2004 / 013216; PCT / JP2005 / 020435; PCT / JP2006 / 315479; PCT / JP2006 / 324484; PCT / JP2009 / 056718; PCT / JP2010 / 067560), or 5′-substituted monomers prepared as monophosphates with modified bases (Wang et al., Nucleosides Nucleotides & Nucleic Acids, 2004, 23(1&2), 317-337); the disclosures of each of these documents are hereby incorporated by reference in their entirety.
[0367] In some embodiments, the unnatural nucleic acids include modifications at the 5'- and 2'-positions of the sugar ring (PCT / US94 / 02993), such as 5'-CH2-substituted 2'-O-protected nucleosides (Wu et al., Helvetica Chimica Acta, 2000, 83, 1127-1143 and Wu et al., Bioconjugate Chem. 1999, 10, 921-924). In some cases, the unnatural nucleic acids include amide-linked nucleoside dimers that have been prepared for incorporation into oligonucleotides, wherein the 3'-linked nucleoside (5' to 3') in the dimer contains 2'-OCH3 and 5'-(S)-CH3 (Mesmaeker et al., Synlett, 1997, 1287-1290). The unnatural nucleic acids can include 2'-substituted 5'-CH2(or O)-modified nucleosides (PCT / US92 / 01020). The unnatural nucleic acids can include 5'-methylenephosphonate DNA and RNA monomers, and dimers (Bohringer et al., Tet. Lett., 1993, 34, 2723-2726; Collingwood et al., Synlett, 1995, 7, 703-705; and Hutter et al., Helvetica Chimica Acta, 2002, 85, 2777-2806). The unnatural nucleic acids can include 5'-phosphonate monomers with 2'-substituents (US 2006 / 0074035) and other modified 5'-phosphonate monomers (WO 1997 / 35869). The unnatural nucleic acids can include 5'-modified methylenephosphonate monomers (EP614907 and EP629633). The unnatural nucleic acids can include analogs of 5'- or 6'-phosphonoribonucleosides that contain a hydroxyl group at the 5' and / or 6'-position (Chen et al., Phosphorus, Sulfur and Silicon, 2002, 777, 1783-1786; Jung et al., Bioorg. Med. Chem., 2000, 8, 2501-2509; Gallier et al., Eur. J. Org. Chem., 2007, 925-933; and Hampton et al., J. Med. Chem., 1976, 19(8), 1029-1033). The unnatural nucleic acids can include 5'-phosphonodeoxyribonucleoside monomers and dimers with a 5'-phosphate group (Nawrot et al., Oligonucleotides, 2006, 16(1), 68-82).Unnatural nucleic acids can include nucleosides having a 6'-phosphonate group (wherein the 5' or / and 6' position is unsubstituted or substituted with tert-butylthio (SC(CH3)3) (and its analogs); methyleneamino (CH2NH2) (and its analogs) or cyano (CN) (and its analogs) (Fairhurst et al., Synlett, 2001, 4, 467-472; Kappler et al., J. Med. Chem., 1986, 29, 1030-1038; Kappler et al., J. Med. Chem., 1982, 25, 1179-1184; Vrudhula et al., J. Med. Chem., 1987, 30, 888-894; Hampton et al., J. Med. Chem., 1976, 19, 1371-1377; Geze et al., J. Am. Chem. Soc, 1983, 105(26), 7638-7640; and Hampton et al., J. Am. Chem. Soc, 1973, 95(13), 4404-4414). The disclosure of each reference listed in this paragraph is hereby incorporated by reference in its entirety.
[0368] In some embodiments, the unnatural nucleic acid further includes modifications of the sugar moiety. In some cases, the nucleic acid contains one or more nucleosides in which the sugar group has been modified. Such sugar-modified nucleosides can confer enhanced nuclease stability, increased binding affinity, or some other beneficial biological property. In certain embodiments, the nucleic acid comprises a chemically modified furanose ring moiety. Examples of chemically modified furanose rings include, but are not limited to, adding substituents (including 5' and / or 2' substituents); bridging two ring atoms to form a bicyclic nucleic acid (BNA); replacing the ribose ring oxygen atom with S, N(R) or C(R1)(R2) (R = H, C1-C 12 alkyl or a protecting group); and combinations thereof. Examples of chemically modified sugars can be found in WO2008 / 101157, US 2005 / 0130923 and WO 2007 / 134181; the disclosure of each of these documents is hereby incorporated by reference in its entirety.
[0369] In some cases, the modified nucleic acid comprises a modified sugar or sugar analogue. Thus, in addition to ribose and deoxyribose, the sugar moiety may be a pentose, deoxypentose, hexose, deoxyhexose, glucose, arabinose, xylose, lyxose or the sugar "analogue" cyclopentyl. The sugar may be in the pyranosyl or furanosyl form. The sugar moiety may be a furanoside of ribose, deoxyribose, arabinose or 2'-O-alkylribose, and the sugar may be attached to the corresponding heterocyclic base in the [α] or [β] anomeric configuration. Sugar modifications include, but are not limited to, 2'-alkoxy-RNA analogues, 2'-amino-RNA analogues, 2'-fluoro-DNA and 2'-alkoxy- or amino-RNA / DNA chimeras. For example, sugar modifications may include 2'-O-methyl-uridine or 2'-O-methyl-cytidine. Sugar modifications include 2'-O-alkyl-substituted deoxyribonucleosides and 2'-O-ethylene glycol-like ribonucleosides. The preparation of these sugars or sugar analogues and the corresponding "nucleosides" in which such sugars or analogues are attached to heterocyclic bases (nucleic acid bases) is known. Sugar modifications can also be made and combined with other modifications.
[0370] Modifications of the sugar moiety include both natural and non-natural modifications of ribose and deoxyribose. Sugar modifications include, but are not limited to, the following modifications at the 2'-position: OH; F; O-, S- or N-alkyl; O-, S- or N-alkenyl; O-, S- or N-alkynyl; or O-alkyl-O-alkyl, where the alkyl, alkenyl and alkynyl may be substituted or unsubstituted C1 to C 10 alkyl or C2 to C 10 alkenyl and alkynyl. 2'-sugar modifications also include, but are not limited to, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n ONH2 and -O(CH2) n ON[(CH2)n CH3)]2, where n and m are from 1 to about 10.
[0371] Other modifications at the 2'-position include, but are not limited to: C1 to C 10Lower alkyl, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl, O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, a group for improving the pharmacokinetic properties of an oligonucleotide or a group for improving the pharmacodynamic properties of an oligonucleotide, and other substituents having similar properties. Similar modifications can also be made at other positions of the sugar (particularly at the 3'-position of the sugar in the 3'-terminal nucleotide or in the 2'-5'-linked oligonucleotide and at the 5'-position of the 5'-terminal nucleotide). Modified sugars also include those sugars containing modifications (such as CH2 and S) at the bridging epoxy. Nucleotide sugar analogs can also have sugar mimetics, such as a cyclobutyl moiety replacing the pentafuranosyl sugar. Many U.S. patents teach the preparation of such modified sugar structures and detail and describe a range of base modifications, such U.S. patents being, for example, U.S. Patent Nos. 4,981,957; 5,118,800; 5,319,080; 5,359,044; 5,393,878; 5,446,137; 5,466,786; 5,514,785; 5,519,134; 5,567,811; 5,576,427; 5,591,722; 5,597,909; 5,610,300; 5,627,053; 5,639,873; 5,646,265; 5,658,873; 5,670,633; 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121, 5,596,091; 5,614,617; 5,681,941; and 5,700,920, the disclosures of each of these documents are incorporated herein by reference in their entirety.
[0372] Examples of nucleic acids having a modified sugar moiety include, but are not limited to, nucleic acids containing 5'-vinyl, 5'-methyl (R or S), 4'-S, 2'-F, 2'-OCH3, and 2'-O(CH2)2OCH3 substituents. Substituents at the 2'-position can also be selected from allyl, amino, azido, thio, O-allyl, O-(C1-C 1O alkyl), OCF3, O(CH2)2SCH3, O(CH2)2-O-N(R m )(R nand O-CH2-C(=O)-N(R m )(R n ), wherein R m and R n are each independently H or a substituted or unsubstituted C1-C 10 alkyl group.
[0373] In certain embodiments, the nucleic acids described herein include one or more bicyclic nucleic acids. In certain such embodiments, the bicyclic nucleic acid comprises a bridge between the 4' and 2' ribosyl ring atoms. In certain embodiments, the nucleic acids provided herein include one or more bicyclic nucleic acids, wherein the bridge comprises a 4'-to-2' bicyclic nucleic acid. Examples of such 4'-to-2' bicyclic nucleic acids include, but are not limited to, one of the following formulas: 4'-(CH2)-O-2'(LNA); 4'-(CH2)-S-2'; 4'-(CH2)2-O-2'(ENA); 4'-CH(CH3)-O-2' and 4'-CH(CH2OCH3)-O-2' and their analogs (see U.S. Patent No. 7,399,845); 4'-C(CH3)(CH3)-O-2' and its analogs (see WO 2009 / 006478, WO 2008 / 150729, US2004 / 0171570, U.S. Patent No. 7,427,672, Chattopadhyaya et al., J. Org. Chem., 209, 74, 118-134, and WO 2008 / 154401).See also, for example: Singh et al., Chem. Commun., 1998, 4, 455-456; Koshkin et al., Tetrahedron, 1998, 54, 3607-3630; Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 2000, 97, 5633-5638; Kumar et al., Bioorg. Med. Chem. Lett., 1998, 8, 2219-2222; Singh et al., J. Org. Chem., 1998, 63, 10035-10039; Srivastava et al., J. Am. Chem. Soc., 2007, 129(26)8362-8379; Elayadi et al., Curr. Opinion Invens. Drugs, 2001, 2, 558-561; Braasch et al., Chem. Biol, 2001, 8, 1-7; Oram et al., Curr. Opinion Mol. Ther., 2001, 3, 239-243; U.S. Patent Nos. 4,849,513; 5,015,733; 5,118,800; 5,118,802; 7,053,207; 6,268,490; 6,770,748; 6,794,499; 7,034,133; 6,525,191; 6,670,461; and 7,399,845; International Publication Nos. WO 2004 / 106356, WO1994 / 14226, WO 2005 / 021570, WO 2007 / 090071 and WO 2007 / 134181; U.S. Patent Publication Nos. US2004 / 0171570, US 2007 / 0287831 and US 2008 / 0039618; U.S. Provisional Application Nos. 60 / 989,574, 61 / 026,995, 61 / 026,998, 61 / 056,564, 61 / 086,231, 61 / 097,787 and 61 / 099,844; and International Application Nos. PCT / US2008 / 064591, PCT US2008 / 066154, PCT US2008 / 068922 and PCT / DK98 / 00393; the disclosures of each of these documents are hereby incorporated by reference in their entirety.
[0374] In certain embodiments, the nucleic acids comprise linked nucleic acids. The nucleic acids can be linked together using any internucleic acid linkage. Two major classes of internucleic acid linking groups are defined by the presence or absence of a phosphorus atom. Representative phosphorus-containing internucleic acid linkages include, but are not limited to, phosphodiester, phosphotriester, methylphosphonate, phosphoramidate, and phosphorothioate (P=S). Representative phosphorus-free internucleic acid linking groups include, but are not limited to, methylene methylimino (-CH2-N(CH3)-O-CH2-), thiodiester (-O-C(O)-S-), thiocarbamate (-O-C(O)(NH)-S-); siloxane (-O-Si(H)2-O-); and N,N*-dimethylhydrazine (-CH2-N(CH3)-N(CH3)). In certain embodiments, internucleic acid linkages having chiral atoms can be prepared as racemic mixtures, as separate enantiomers, such as alkylphosphonates and phosphorothioates. The unnatural nucleic acids can contain a single modification. The unnatural nucleic acids can contain multiple modifications within one of the moieties or between different moieties.
[0375] Backbone phosphate modifications to the nucleic acids include, but are not limited to, methylphosphonate, phosphorothioate, phosphoramidate (bridged or unbridged), phosphotriester, phosphorodithioate, phosphodithioate, and boranophosphate, and can be used in any combination. Other non-phosphate linkages can also be used.
[0376] In some embodiments, backbone modifications (e.g., methylphosphonate, phosphorothioate, phosphoramidate, and phosphorodithioate internucleotide linkages) can confer immunomodulatory activity and / or enhance the in vivo stability of the modified nucleic acids.
[0377] In some cases, a phosphorus derivative (or a modified phosphate group) is attached to a sugar or sugar analog moiety and can be a monophosphate, diphosphate, triphosphate, alkyl phosphonate, thiophosphate, dithiophosphate, aminophosphate, etc. Exemplary polynucleotides containing modified phosphate linkages or non-phosphate linkages can be found in: Peyrottes et al., 1996, Nucleic Acids Res. 24:1841-1848; Chaturvedi et al., 1996, Nucleic Acids Res. 24:2318-2323; and Schultz et al., (1996) Nucleic Acids Res. 24:2966-2973; Matteucci, 1997, “Oligonucleotide Analogs: an Overview” in Oligonucleotides as Therapeutic Agents, (Chadwick and Cardew, eds.) John Wiley and Sons, New York, NY; Zon, 1993, “Oligonucleoside Phosphorothioates” in Protocols for Oligonucleotides and Analogs, Synthesis and Properties, Humana Press, pp. 165-190; Miller et al., 1971, JACS 93:6657-6665; Jager et al., 1988, Biochem. 27:7247-7246; Nelson et al., 1997, JOC 62:7278-7287; U.S. Patent No. 5,453,496; and Micklefield, 2001, Curr. Med. Chem. 8:1157-1179; the disclosures of each of these references are hereby incorporated by reference in their entirety.
[0378] In some instances, backbone modifications include replacing the phosphodiester linkage with an alternative moiety such as an anionic group, a neutral group, or a cationic group. Examples of such modifications include: anionic internucleoside linkages; N3’ to P5’ phosphoramidate modifications; boranophosphate DNA; morpholino oligomers; neutral internucleoside linkages such as methylphosphonates; amide-linked DNA; methylene (methylimino) linkages; formacetal and thioformacetal linkages; sulfonyl-containing backbones; peptide nucleic acids (PNA); and positively charged deoxyribonucleic acid guanidine (DNG) oligomers (Micklefield, 2001, Current Medicinal Chemistry 8:1157-1179), the disclosure of which is hereby incorporated by reference in its entirety. The modified nucleic acids can contain chimeric or mixed backbones that contain one or more modifications (e.g., combinations of phosphodiester linkages such as combinations of phosphodiester and phosphorothioate linkages).
[0379] The substituents of the phosphate esters include, for example, short-chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short-chain heteroatom or heterocyclic internucleoside linkages. These include those having: a morpholino linkage (partially formed from the sugar moiety of the nucleoside); a siloxane backbone; sulfide, sulfoxide, and sulfone backbones; formylacetyl and thioformylacetyl backbones; methyleneformylacetyl and thioformylacetyl backbones; olefin-containing backbones; sulfamate backbones; methyleneimine and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and other backbones having mixed N, O, S, and CH2 components. Many U.S. patents disclose how to prepare and use these types of phosphate ester replacements, and the U.S. patents include, but are not limited to, U.S. Patent Nos. 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; 5,434,257; 5,466,677; 5,470,967; 5,489,677; 5,541,307; 5,561,225; 5,596,086; 5,602,240; 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439; the disclosures of each of these documents are hereby incorporated by reference in their entirety. It should also be understood that in nucleotide substituents, both the sugar and phosphate ester moieties of the nucleotide can be replaced, for example, by an amide-type linkage (aminoethylglycine) (PNA). U.S. Patent Nos. 5,539,082; 5,714,331; and 5,719,262 teach how to prepare and use PNA molecules, and each of these documents is incorporated herein by reference. See also Nielsen et al., Science, 1991, 254, 1497-1500. Other types of molecules (conjugates) can also be linked to the nucleotides or nucleotide analogs to enhance, for example, cellular uptake. The conjugate can be chemically linked to the nucleotide or nucleotide analog.Such conjugates include, but are not limited to, lipid moieties such as cholesterol moieties (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, 86, 6553-6556), bile acids (Manoharan et al., Bioorg. Med. Chem. Let., 1994, 4, 1053-1060), thioethers such as hexyl-S-tritylthiol (Manoharan et al., Ann. KY. Acad. Sci., 1992, 660, 306-309; Manoharan et al., Bioorg. Med. Chem. Let., 1993, 3, 2765-2770), cholesteryl thioether (Oberhauser et al., Nucl. Acids Res., 1992, 20, 533-538), aliphatic chains such as dodecanediol or undecyl residues (Saison-Behmoaras et al., EMBO J, 1991, 10, 1111-1118; Kabanov et al., FEBS Lett., 1990, 259, 327-330; Svinarchuk et al., Biochimie, 1993, 75, 49-54), phospholipids such as di-hexadecyl-rac-glycerol or l-di-O-hexadecyl-rac-glycerol-S-H-phosphonate triethylammonium (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651-3654; Shea et al., Nucl. Acids Res., 1990, 18, 3777-3783), polyamines or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides, 1995, 14, 969-973), or adamantane acetic acid (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651-3654), palmitoyl moieties (Mishra et al., Biochem. Biophys. Acta, 1995, 1264, 229-237), or octadecylamine or hexylamino-carbonyloxy cholesterol moieties (Crooke et al., J. Pharmacol. Exp. Ther., 1996, 277, 923-937); the disclosures of each of these references are hereby incorporated by reference in their entirety.Numerous U.S. patents teach the preparation of such conjugates, and such U.S. patents include, but are not limited to, U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717, 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241, 5,391,723; 5,416,203, 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941; the disclosures of each of these documents are hereby incorporated by reference in their entirety.
[0380] Nucleobases used in compositions and methods for replication, transcription, translation, and incorporation of unnatural amino acids into proteins are described herein. In some embodiments, the nucleobases described herein comprise the structure:
[0381] wherein
[0382] each X is independently carbon or nitrogen;
[0383] when X is carbon, R2 is present and is independently hydrogen, alkyl, alkenyl, alkynyl, methoxy, methanethiol, methaneselenyl, halogen, cyano, or azido;
[0384] wherein each Y is independently sulfur, oxygen, selenium, or secondary amine;
[0385] wherein each E is independently oxygen, sulfur, or selenium; and
[0386] wherein the wavy line indicates the point of attachment to a ribosyl, deoxyribosyl, or dideoxyribosyl moiety or an analogue thereof, wherein the ribosyl, deoxyribosyl, or dideoxyribosyl moiety or an analogue thereof is in free form, linked to a monophosphate, diphosphate, or triphosphate group (optionally including an α-thiotriphosphate, β-thiotriphosphate, or γ-thiotriphosphate group), or is incorporated in RNA or DNA or an RNA analogue or a DNA analogue.
[0387] In some embodiments, R2 is a lower alkyl (e.g., C1-C6), hydrogen, or halogen. In some embodiments of the nucleobases described herein, R2 is fluorine. In some embodiments of the nucleobases described herein, X is carbon. In some embodiments of the nucleobases described herein, E is sulfur. In some embodiments of the nucleobases described herein, Y is sulfur.
[0388] In some embodiments of the nucleobases described herein, the nucleobase has the structure: In some embodiments, the nucleobases described herein have the structure: In some embodiments of the nucleobases described herein, E is sulfur and Y is sulfur. In some embodiments, the wavy line indicates the point of attachment to a ribosyl, deoxyribosyl, or dideoxyribosyl moiety or an analogue thereof, wherein the ribosyl, deoxyribosyl, or dideoxyribosyl moiety or an analogue thereof is in free form, or linked to a monophosphate, diphosphate, triphosphate, α-thiotriphosphate, β-thiotriphosphate, or γ-thiotriphosphate group, or is incorporated in RNA or DNA or an RNA analogue or a DNA analogue. In some embodiments of the nucleobases described herein, the wavy line indicates the point of attachment to a ribosyl or deoxyribosyl moiety. In some embodiments of the nucleobases described herein, the wavy line indicates the point of attachment to a ribosyl or deoxyribosyl moiety that is linked to a triphosphate group. In some embodiments of the nucleobases described herein, it is a component of a nucleic acid polymer. In some embodiments of the nucleobases described herein, the nucleobase is a component of tRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of the anticodon in tRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of mRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of the codon in mRNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of RNA or DNA. In some embodiments of the nucleobases described herein, the nucleobase is a component of the codon in DNA. In some embodiments of the nucleobases described herein, the nucleobase forms, is capable of forming, or is configured to form a nucleobase pair with another (e.g., complementary) nucleobase.
[0389] In some embodiments, the nucleobases described herein have the structure:
[0390]
[0391] Wherein:
[0392] Each X is independently carbon or nitrogen;
[0393] When X is nitrogen, R2 is absent, and when X is carbon, R2 is present and is independently hydrogen, alkyl, alkenyl, alkynyl, methoxy, methylthiol, methylseleno, halogen, cyano or azido;
[0394] Y is sulfur, oxygen, selenium or secondary amine;
[0395] E is oxygen, sulfur or selenium; and
[0396] The wavy line indicates the point of attachment to a ribosyl, deoxyribosyl or dideoxyribosyl moiety or an analogue thereof, wherein the ribosyl, deoxyribosyl or dideoxyribosyl moiety or an analogue thereof is in free form, linked to a monophosphate, diphosphate, triphosphate, α-thiotriphosphate, β-thiotriphosphate or γ-thiotriphosphate group, or is incorporated in RNA or DNA or an RNA analogue or DNA analogue.
[0397] In some embodiments, each X is carbon. In some embodiments, at least one X is carbon. In some embodiments, one X is carbon. In some embodiments, at least two Xs are carbon. In some embodiments, two Xs are carbon. In some embodiments, at least one X is nitrogen. In some embodiments, one X is nitrogen. In some embodiments, at least two Xs are nitrogen. In some embodiments, two Xs are nitrogen.
[0398] In some embodiments, Y is sulfur. In some embodiments, Y is oxygen. In some embodiments, Y is selenium. In some embodiments, Y is secondary amine.
[0399] In some embodiments, E is sulfur. In some embodiments, E is oxygen. In some embodiments, E is selenium.
[0400] In some embodiments, R2 is present when X is carbon. In some embodiments, when X is nitrogen, R 2Absent. In some embodiments, each R2, when present, is hydrogen. In some embodiments, R2 is an alkyl group such as methyl, ethyl, or propyl. In some embodiments, R2 is an alkenyl group such as -CH2=CH2. In some embodiments, R2 is an alkynyl group such as ethynyl. In some embodiments, R2 is a methoxy group. In some embodiments, R2 is a methylthiol group. In some embodiments, R2 is a methylseleno group. In some embodiments, R2 is a halogen such as chlorine, bromine, or fluorine. In some embodiments, R2 is a cyano group. In some embodiments, R2 is an azide group.
[0401] In some embodiments, E is sulfur, Y is sulfur, and each X is independently carbon or nitrogen. In some embodiments, E is sulfur, Y is sulfur, and each X is carbon.
[0402] In some embodiments, the nucleobase has the structure In some embodiments, the nucleobase has the structure
[0403]
[0404] In some embodiments, the nucleobase has the structure
[0405] In some embodiments, the nucleobases disclosed herein bind (e.g., non-covalently) to complementary base-pairing nucleobases to form unnatural base pairs (UBPs), or are capable of base pairing or are configured to base pair with nucleobases. In some embodiments, the complementary base-pairing nucleobases are selected from:
[0406]
[0407] In one aspect, provided herein are double-stranded oligonucleotide duplexes, wherein the first oligonucleotide strand comprises a nucleobase disclosed herein, and the second complementary oligonucleotide strand comprises a complementary base-pairing nucleobase at its complementary base-pairing site. In some embodiments, the first oligonucleotide strand comprises and the second strand comprises at its complementary base-pairing site a complementary base-pairing nucleobase selected from:
[0408] On the other hand, the present disclosure provides transfer RNAs (tRNAs) comprising the nucleobases described herein, the transfer RNAs comprising: an anticodon, wherein the anticodon comprises the nucleobase; and an identity element, wherein the identity element promotes selective loading of a non-natural amino acid by an aminoacyl-tRNA synthetase. In some embodiments, the nucleobase is located in the anticodon region of the tRNA. In some embodiments, the nucleobase is located at the first position of the anticodon. In some embodiments, the nucleobase is located at the second position of the anticodon. In some embodiments, the nucleobase is located at the third position of the anticodon. In some embodiments, the aminoacyl-tRNA synthetase is derived from Methanosarcina or a variant thereof. In some embodiments, the aminoacyl-tRNA synthetase is derived from Methanococcus / Methanocaldococcus or a variant thereof. In some embodiments, the non-natural amino acid comprises an aromatic moiety. In some embodiments, the non-natural amino acid is a lysine derivative. In some embodiments, the non-natural amino acid is a phenylalanine derivative.
[0409] The present disclosure also provides a structure comprising the formula:
[0410] N1-Zx-N2
[0411] Wherein:
[0412] Z is a nucleobase as described herein, bonded to a ribosyl or deoxyribosyl or an analogue thereof;
[0413] N1 is one or more nucleotides or analogues thereof or a terminal phosphate group attached at the 5' end of the ribosyl or deoxyribosyl or an analogue thereof of Z;
[0414] N2 is one or more nucleotides or analogues thereof or a terminal hydroxyl group attached at the 3' end of the ribosyl or deoxyribosyl or an analogue thereof of Z; and
[0415] x is an integer from 1 to 20.
[0416] In some embodiments, N1 is one or more nucleotides or analogues thereof attached at the 5’ end of the ribosyl or deoxyribosyl moiety or an analogue thereof of Z. Attachment to the 5’ end of the ribosyl or deoxyribosyl moiety can be by a phosphodiester. In some embodiments, N1 is a terminal phosphate group attached at the 5’ end of the ribosyl or deoxyribosyl moiety or an analogue thereof of Z. In some embodiments, N2 is one or more nucleotides or analogues thereof attached at the 3’ end of the ribosyl or deoxyribosyl moiety or an analogue thereof of Z. Attachment to the 3’ end of the ribosyl or deoxyribosyl moiety can be by a phosphodiester. In some embodiments, N2 is a terminal hydroxyl attached at the 3’ end of the ribosyl or deoxyribosyl moiety or an analogue thereof of Z.
[0417] In some embodiments, x is an integer from 1 to 20. In some embodiments, x is an integer from 1 to 15. In some embodiments, x is an integer from 1 to 10. In some embodiments, x is an integer from 1 to 5. In some embodiments, x is 1. In some embodiments, x is 2. In some embodiments, x is 3. In some embodiments, x is 4. In some embodiments, x is 5. In some embodiments, x is 6. In some embodiments, x is 7. In some embodiments, x is 8. In some embodiments, x is 9. In some embodiments, x is 10. In some embodiments, x is 11. In some embodiments, x is 12. In some embodiments, x is 13. In some embodiments, x is 14. In some embodiments, x is 15. In some embodiments, x is 16. In some embodiments, x is 17. In some embodiments, x is 18. In some embodiments, x is 19. In some embodiments, x is 20.
[0418] In some embodiments, Z has the structure as detailed herein In some embodiments, Z has the structure
[0419] In some embodiments, the structure of formula N1-Zx-N2 encodes a gene. In some embodiments, Zx is located in the translation region of the gene. In some embodiments, Zx is located in the untranslated region of the gene. In some embodiments, the structure further comprises a 5' or 3' untranslated region (UTR). In some embodiments, the structure further comprises a terminator region. In some embodiments, the structure further comprises a promoter region.
[0420] In a further aspect, the present disclosure provides a polynucleotide library, wherein the library comprises at least 5000 unique polynucleotides, and wherein each polynucleotide comprises at least one nucleobase disclosed herein. In some embodiments, the polynucleotide library encodes at least one gene.
[0421] In yet another aspect, the present disclosure provides a nucleoside triphosphate, wherein the nucleobase is selected from In some embodiments, the nucleobase is In some embodiments, the nucleobase is In some embodiments, the nucleobase is In some embodiments, the nucleobase is In some embodiments, the nucleoside comprises ribose. In some embodiments, the nucleoside comprises deoxyribose.
[0422] Nucleic acid base pairing properties; exemplary base pairs
[0423] In some embodiments, non-natural nucleotides form base pairs (non-natural base pairs; UBP) with another non-natural nucleotide during or after incorporation into DNA or RNA. In some embodiments, a stably incorporated non-natural nucleic acid is a non-natural nucleic acid that can form base pairs with another nucleic acid (e.g., a natural or non-natural nucleic acid). In some embodiments, a stably incorporated non-natural nucleic acid is a non-natural nucleic acid that can form base pairs (non-natural nucleic acid base pairs (UBP)) with another non-natural nucleic acid. For example, a first non-natural nucleic acid can form a base pair with a second non-natural nucleic acid. For example, a pair of non-natural nucleoside triphosphates that can base pair during and after incorporation of nucleic acids includes the triphosphate of (d)5SICS ((d)5SICSTP) and the triphosphate of (d)NaM ((d)NaMTP). Other examples include, but are not limited to: the triphosphate of (d)CNMO ((d)CNMOTP) and the triphosphate of (d)TPT3 ((d)TPT3TP). Such non-natural nucleotides can have a ribose or deoxyribose sugar moiety (indicated by "(d)"). For example, a pair of non-natural nucleoside triphosphates that can base pair upon incorporation of nucleic acids includes the triphosphate of TAT1 ((d)TAT1TP) and the triphosphate of NaM ((d)NaMTP). For example, a pair of non-natural nucleoside triphosphates that can base pair upon incorporation of nucleic acids includes the triphosphate of dCNMO (dCNMOTP) and the triphosphate of TAT1 (TAT1TP). For example, a pair of non-natural nucleoside triphosphates that can base pair upon incorporation of nucleic acids includes the triphosphate of dTPT3 (dTPT3TP) and the triphosphate of NaM (NaMTP). In some embodiments, non-natural nucleic acids do not substantially form base pairs with natural nucleic acids (A, T, G, C). In some embodiments, a stably incorporated non-natural nucleic acid can form base pairs with natural nucleic acids.
[0424] In some embodiments, a stably incorporated unnatural (deoxy)ribonucleotide is an unnatural (deoxy)ribonucleotide that can form a UBP but that substantially does not form base pairs with any of the natural (deoxy)ribonucleotides. In some embodiments, a stably incorporated unnatural (deoxy)ribonucleotide is an unnatural (deoxy)ribonucleotide that can form a UBP but that substantially does not form base pairs with one or more natural nucleic acids. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with A, T, and C, but may form base pairs with G. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with A, T, and G, but may form base pairs with C. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with C, G, and A, but may form base pairs with T. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with C, G, and T, but may form base pairs with A. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with A and T, but may form base pairs with C and G. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with A and C, but may form base pairs with T and G. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with A and G, but may form base pairs with C and T. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with C and T, but may form base pairs with A and G. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with C and G, but may form base pairs with T and G. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with T and G, but may form base pairs with A and G. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with G, but may form base pairs with A, T, and C. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with A, but may form base pairs with G, T, and C. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with T, but may form base pairs with G, A, and C. For example, a stably incorporated unnatural nucleic acid may substantially not form base pairs with C, but may form base pairs with G, T, and A.
[0425] Exemplary unnatural nucleotides capable of forming unnatural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, 5SICS, d5SICS, NaM, dNaM, dTPT3, dMTMO, dCNMO, TAT1, and combinations thereof. In some embodiments, unnatural nucleotides capable of forming unnatural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, 5SICS, NaM, TPT3, MTMO, CNMO, TAT1, and combinations thereof, wherein the sugar moiety of the nucleotide is deoxyribose sugar. In some embodiments, unnatural nucleotides capable of forming unnatural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, 5SICS, NaM, TPT3, MTMO, CNMO, TAT1, and combinations thereof, wherein the sugar moiety of the nucleotide is ribose sugar. In some embodiments, unnatural nucleotides capable of forming unnatural DNA or RNA base pairs (UBPs) under in vivo conditions include, but are not limited to, (d)5SICS, (d)NaM, (d)TPT3, (d)MTMO, (d)CNMO, (d)TAT1, and combinations thereof. In some embodiments, unnatural nucleobase pairs include, but are not limited to: wherein the sugar moiety is any embodiment or variation described herein. In some embodiments, unnatural nucleobase pairs include, but are not limited to:
[0426] In any such embodiment, one or both of the deoxyriboses attached to the unnatural base may be replaced by ribose.
[0427] Engineered organisms
[0428] In some embodiments, the methods and plasmids disclosed herein are further used to generate engineered organisms, such as an organism that incorporates and replicates unnatural nucleotides or unnatural nucleic acid base pairs (UBPs), and can also use nucleic acids containing unnatural nucleotides to transcribe mRNA and tRNA for the translation of proteins containing unnatural amino acid residues. In some cases, the organism is a non-human semi-synthetic organism (SSO). In some cases, the organism is a semi-synthetic organism (SSO). In some cases, the SSO is a cell. In some cases, in vivo methods include semi-synthetic organisms (SSOs). In some cases, the semi-synthetic organism includes a microorganism. In some cases, the organism includes a bacterium. In some cases, the organism includes a Gram-negative bacterium. In some cases, the organism includes a Gram-positive bacterium. In some cases, the organism includes Escherichia coli (E. coli). Such modified organisms differently contain additional components, such as DNA repair machinery, modified polymerases, nucleotide transporters, or other components. In some cases, the SSO includes Escherichia coli strain YZ3. In some cases, the SSO includes Escherichia coli strains ML1 or ML2, such as those shown in FIGS. 1 (B-D) of Ledbetter, et al. J. Am. Chem. Soc. 2018, 140(2), 758, the disclosure of which is hereby incorporated by reference in its entirety.
[0429] In some cases, the cells to be used are genetically transformed with an expression cassette encoding a heterologous protein and optionally a CRISPR / Cas9 system (to eliminate DNA that has lost unnatural nucleotides), the heterologous protein such as a nucleoside triphosphate transporter (e.g., Escherichia coli strains YZ3, ML1, or ML2) capable of transporting unnatural nucleoside triphosphates into the cell. In some cases, the cells further contain enhanced activity for the uptake of unnatural nucleic acids. In some instances, the cells further contain enhanced activity for the import of unnatural nucleic acids.
[0430] In some embodiments, Cas9 and an appropriate guide RNA (sgRNA) are encoded on separate plasmids. In some cases, Cas9 and sgRNA are encoded on the same plasmid. In some instances, the nucleic acid molecule encoding Cas9, sgRNA, or a nucleic acid molecule comprising unnatural nucleotides is located on one or more plasmids. In some cases, Cas9 is encoded on a first plasmid, and sgRNA and the nucleic acid molecule comprising unnatural nucleotides are encoded on a second plasmid. In some cases, Cas9, sgRNA, and the nucleic acid molecule comprising unnatural nucleotides are encoded on the same plasmid. In some cases, the nucleic acid molecule comprises two or more unnatural nucleotides. In some cases, Cas9 is integrated into the genome of a host organism, and sgRNA is encoded on a plasmid or in the genome of the organism.
[0431] In some cases, a first plasmid encoding Cas9 and sgRNA and a second plasmid encoding a nucleic acid molecule comprising unnatural nucleotides are introduced into an engineered microorganism. In some cases, a first plasmid encoding Cas9 and a second plasmid encoding sgRNA and a nucleic acid molecule comprising unnatural nucleotides are introduced into an engineered microorganism. In some cases, a plasmid encoding Cas9, sgRNA, and a nucleic acid molecule comprising unnatural nucleotides is introduced into an engineered microorganism. In some cases, the nucleic acid molecule comprises two or more unnatural nucleotides.
[0432] In some embodiments, live cells are generated that incorporate at least one unnatural nucleotide and / or at least one unnatural base pair (UBP) into their DNA (plasmid or genome). In some cases, the unnatural base pair comprises a pair of unnaturally base-pairing nucleotides that are capable of forming an unnatural base pair under in vivo conditions when the unnaturally base-pairing nucleotides are taken up into the cell as their corresponding triphosphates by the action of a nucleotide triphosphate transporter. In some cases, the unnatural base pair comprises a pair of unnaturally base-pairing nucleotides that are configured to form an unnatural base pair under in vivo conditions when the unnaturally base-pairing nucleotides are taken up into the cell as their corresponding triphosphates by the action of a nucleotide triphosphate transporter. The cells can be genetically transformed with an expression cassette encoding a nucleotide triphosphate transporter such that the nucleotide triphosphate transporter is expressed and available for transporting unnatural nucleotides into the cells. The cells can be prokaryotic or eukaryotic cells, and the pair of unnaturally base-pairing nucleotides as their corresponding triphosphates can be the triphosphates of dTPT3 (dTP3TP) and dNaM (dNaMTP) or the triphosphates of dCNMO (dCNMOTP).
[0433] In some embodiments, the cell is a cell genetically transformed with a nucleic acid, such as an expression cassette encoding a nucleotide triphosphate transporter capable of transporting such unnatural nucleotides into the cell. The cell can comprise a heterologous nucleoside triphosphate transporter, wherein the heterologous nucleoside triphosphate transporter can transport natural and unnatural nucleoside triphosphates into the cell.
[0434] In some instances, the methods described herein further comprise contacting the genetically transformed cell with the corresponding triphosphate in the presence of potassium phosphate and / or an inhibitor of a phosphatase or a nucleotidase. During or after this contact, the cell can be placed within a life support medium suitable for growth and replication of the cell. The cell can be maintained in the life support medium such that the corresponding triphosphate form of the unnatural nucleotide is incorporated into the nucleic acid within the cell and through at least one replication cycle of the cell. The unnatural mutually base-pairing nucleotide pair as the corresponding triphosphate can comprise the triphosphate of dTPT3 or (dTPT3TP) and the triphosphate of dCNMO or dNaM (dCNOM or dNaMTP), the cell can be Escherichia coli, and dTPT3TP and dNaMTP can be imported into Escherichia coli by the transporter PtNTT2, wherein an Escherichia coli polymerase such as Pol III or Pol II can use the unnatural triphosphates to replicate DNA containing UBP, thereby incorporating unnatural nucleotides and / or unnatural base pairs into the cellular nucleic acid within the cellular environment. In addition, ribonucleotides such as NaMTP and TAT1TP, 5FMTP and TPT3TP are imported into Escherichia coli by the transporter PtNTT2 in some cases.
[0435] This disclosure describes compositions and methods that include the use of three or more unnatural base-pairing nucleotides. In some instances, such base-pairing nucleotides enter cells by using nucleotide transporters or by standard nucleic acid transformation methods known in the art (e.g., electroporation, chemical transformation, or other methods). In some instances, the base-pairing unnatural nucleotides enter cells as part of a polynucleotide, such as a plasmid. One or more base-pairing unnatural nucleotides that enter cells as part of a polynucleotide (RNA or DNA) do not themselves need to replicate in vivo. For example, a double-stranded DNA plasmid or other nucleic acid that includes a first unnatural deoxyribonucleotide and a second unnatural deoxyribonucleotide, the bases of which are configured to form a first unnatural base pair, is electroporated into cells. The cell culture medium is treated with a third unnatural deoxyribonucleotide and a fourth unnatural deoxyribonucleotide, the bases of which are configured to form a second unnatural base pair with each other, where the base of the first unnatural deoxyribonucleotide and the base of the third unnatural deoxyribonucleotide form a second unnatural base pair, and where the base of the second unnatural deoxyribonucleotide and the base of the fourth unnatural deoxyribonucleotide form a third unnatural base pair. In some cases, in vivo replication of the initially transformed double-stranded DNA plasmid results in subsequently replicated plasmids that include the third unnatural deoxyribonucleotide and the fourth unnatural deoxyribonucleotide. Alternatively or in combination, ribonucleotide variants of the third unnatural deoxyribonucleotide and the fourth unnatural deoxyribonucleotide are added to the cell culture medium. In some cases, these ribonucleotides are incorporated into RNAs such as mRNA or tRNA. In some cases, the first deoxynucleotide, the second deoxynucleotide, the third deoxynucleotide, and the fourth deoxynucleotide include different bases. In some cases, the first deoxynucleotide, the third deoxynucleotide, and the fourth deoxynucleotide include different bases. In some cases, the first deoxynucleotide and the third deoxynucleotide include the same base.
[0436] By practicing the methods of the present disclosure, one of ordinary skill in the art can obtain a population of viable proliferating cells that have at least one unnatural nucleotide and / or at least one unnatural base pair (UBP) within at least one nucleic acid maintained within at least some of the individual cells, where the at least one nucleic acid stably proliferates within the cells, and where the cells express a nucleotide triphosphate transporter adapted to provide a form of one or more unnatural nucleotides that is suitable for cellular uptake when in contact with (e.g., grown in the presence of) one or more unnatural nucleotides in a life support medium suitable for the growth and replication of the organism.
[0437] After being transported into the cell via a nucleotide triphosphate transporter, the unnatural base-pairing nucleotides are incorporated into nucleic acids within the cell by cellular machinery (e.g., the cell's own DNA and / or RNA polymerases, heterologous polymerases, or polymerases that have been evolved using directed evolution) (Chen T, Romesberg FE, FEBS Lett. Jan 21, 2014; 588(2):219-29; Betz K et al., J Am Chem Soc. Dec 11, 2013; 135(49):18637-43; the disclosures of each of these references are hereby incorporated by reference in their entirety). The unnatural nucleotides can be incorporated into cellular nucleic acids such as genomic DNA, genomic RNA, mRNA, tRNA, structural RNA, microRNA, and self-replicating nucleic acids (e.g., plasmids, viruses, or vectors).
[0438] In some instances, genetically engineered cells are produced by introducing nucleic acids (e.g., heterologous nucleic acids) into the cell. Any cell described herein can be a host cell and can contain an expression vector. In one embodiment, the host cell is a prokaryotic cell. In another embodiment, the host cell is Escherichia coli. In some embodiments, the cell contains one or more heterologous polynucleotides. A variety of techniques can be used to introduce nucleic acid reagents into microorganisms. Non-limiting examples of methods for introducing heterologous nucleic acids into various organisms include: transformation, transfection, transduction, electroporation, sonication-mediated transformation, conjugation, particle bombardment, etc. In some cases, the addition of a carrier molecule (e.g., a bisbenzimidazole-based compound, see, e.g., U.S. Patent No. US 5,595,899) can increase the uptake of DNA in cells that are generally considered difficult to transform by conventional methods. Conventional transformation methods are readily available to those skilled in the art and can be found in the following references: Maniatis, T., E.F. Fritsch, and J. Sambrook (1982) Molecular Cloning: a Laboratory Manual; Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, the disclosures of these references are hereby incorporated by reference in their entirety).
[0439] In some cases, genetic transformation is obtained using direct transfer of expression cassettes in, but not limited to, plasmids, viral vectors, viral nucleic acids, phage nucleic acids, phages, cosmids, and artificial chromosomes, or transfer of genetic material via cells or vectors such as cationic liposomes. Such methods are available in the art and are readily adaptable for use in the methods described herein. The transfer vector can be any nucleotide construct for delivering a gene to a cell (e.g., a plasmid), or as part of a general strategy for delivering genes, e.g., as part of a recombinant retrovirus or adenovirus (Ram et al. Cancer Res. 53:83-88, (1993)). Suitable transfection methods, including viral vectors, chemical transfectants, or physical-mechanical methods such as electroporation and direct diffusion of DNA, are described, for example, in: Wolff, J.A., et al., Science, 247, 1465-1468, (1990); and Wolff, J.A. Nature, 352, 815-818, (1991), the disclosures of each of which are hereby incorporated by reference in their entirety.
[0440] For example, DNA encoding a nucleoside triphosphate transporter or polymerase expression cassette and / or vector can be introduced into cells by any method, including but not limited to calcium-mediated transformation, electroporation, microinjection, lipofection, particle bombardment, etc.
[0441] In some instances, cells contain unnatural nucleoside triphosphates incorporated into one or more nucleic acids within the cell. For example, the cells can be live cells capable of incorporating at least one unnatural nucleotide into DNA or RNA maintained within the cell. The cells can also incorporate at least one unnatural base pair (UBP) comprising a pair of unnatural mutually base-pairing nucleotides into nucleic acids within the cell under in vivo conditions, wherein the unnatural mutually base-pairing nucleotides (e.g., their corresponding triphosphates) are taken up into the cell by the action of a nucleoside triphosphate transporter, the gene for which is delivered (e.g., introduced) into the cell by genetic transformation. For example, after incorporation into nucleic acids maintained within the cell, dTPT3 and dCNMO can form a stable unnatural base pair, which can be stably propagated by the DNA replication machinery of an organism, e.g., when grown in a life-supporting medium containing dTPT3TP and dCNMOTP.
[0442] In some cases, cells are capable of replicating nucleic acids containing unnatural nucleotides. Such methods can include genetically transforming cells with an expression cassette encoding a nucleoside triphosphate transporter that is capable of transporting one or more unnatural nucleotides as the corresponding triphosphates into the cell under in vivo conditions. Alternatively, cells that have been previously genetically transformed with an expression cassette that can express the encoded nucleoside triphosphate transporter can be employed. The method can further include contacting or exposing the genetically transformed cells to potassium phosphate and the corresponding triphosphate forms of at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides capable of forming an unnatural base pair (UBP)) in a life support medium suitable for the growth and replication of the cells, and maintaining the transformed cells in the life support medium under in vivo conditions in the presence of the corresponding triphosphate forms of at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides capable of forming an unnatural base pair (UBP)) for at least one replication cycle of the cells. The method can further include contacting or exposing the genetically transformed cells to potassium phosphate and the corresponding triphosphate forms of at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides configured to form an unnatural base pair (UBP)) in a life support medium suitable for the growth and replication of the cells, and maintaining the transformed cells in the life support medium under in vivo conditions in the presence of the corresponding triphosphate forms of at least one unnatural nucleotide (e.g., two mutually base-pairing nucleotides configured to form an unnatural base pair (UBP)) for at least one replication cycle of the cells.
[0443] In some embodiments, cells contain stably incorporated unnatural nucleic acids. Some embodiments include cells (e.g., Escherichia coli) in which nucleotides other than A, G, T, and C are stably incorporated into nucleic acids maintained within the cell. For example, nucleotides other than A, G, T, and C can be d5SICS, dCNMO, dNaM, and / or dTPT3, which can form stable unnatural base pairs within the nucleic acid after being incorporated into the nucleic acids of the cell. In one aspect, when an organism transformed with the gene for a triphosphate transporter is grown in a life support medium comprising potassium phosphate and the triphosphate forms of d5SICS, dNaM, dCNMO, and / or dTPT3, the unnatural nucleotides and unnatural base pairs can be stably propagated by the replication apparatus of the organism.
[0444] In some cases, cells contain an expanded genetic alphabet. The cells can contain non-natural nucleic acids stably incorporated. In some embodiments, cells having an expanded genetic alphabet contain non-natural nucleic acids that contain non-natural nucleotides that can pair with another non-natural nucleotide. In some embodiments, cells having an expanded genetic alphabet contain non-natural nucleic acids that hydrogen bond to another nucleic acid. In some embodiments, cells having an expanded genetic alphabet contain non-natural nucleic acids that hydrogen bond to another nucleic acid that is not base-paired. In some embodiments, cells having an expanded genetic alphabet contain non-natural nucleic acids that contain non-natural nucleotides having nucleobases that base pair with the nucleobase or another non-natural nucleotide by hydrophobic and / or stacking interactions. In some embodiments, cells having an expanded genetic alphabet contain non-natural nucleic acids that base pair with another nucleic acid via non-hydrogen bond interactions. Cells having an expanded genetic alphabet can be cells that can copy homologous nucleic acids to form nucleic acids that contain non-natural nucleic acids. Cells having an expanded genetic alphabet can be cells that contain non-natural nucleic acids (unnatural base pairs (UBPs)) that base pair with another non-natural nucleic acid.
[0445] In some embodiments, cells form unnatural DNA base pairs (UBPs) from input non-natural nucleotides under in vivo conditions. In some embodiments, inhibitors of potassium phosphate and / or phosphatase and / or nucleotide enzyme activity can facilitate the transport of non-natural nucleotides. The method includes using cells that express a heterologous nucleoside triphosphate transporter. When such cells are contacted with one or more nucleoside triphosphates, the nucleoside triphosphates are transported into the cells. The cells can be in the presence of inhibitors of potassium phosphate and / or phosphatase and nucleotide enzyme. Non-natural nucleoside triphosphates can be incorporated into nucleic acids within the cell by the natural machinery of the cell (i.e., polymerase) and can, for example, base pair with each other within the nucleic acids of the cell to form unnatural base pairs. In some embodiments, UBPs form between DNA with unnatural bases and RNA nucleotides.
[0446] In some embodiments, UBPs can be incorporated into a cell or population of cells when exposed to non-natural triphosphates. In some embodiments, UBPs can be incorporated into a cell or population of cells when substantially uniformly exposed to non-natural triphosphates.
[0447] In some embodiments, inducing the expression of a heterologous gene (e.g., a nucleoside triphosphate transporter (NTT)) in a cell can result in slower cell growth and increased uptake of non-native triphosphates compared to the growth of cells in which the expression of the heterologous gene is not induced and the uptake of one or more non-native triphosphates in said cells. Uptake differently includes transporting nucleotides into the cell, such as by diffusion, osmosis, or by the action of a transporter. In some embodiments, inducing the expression of a heterologous gene (e.g., NTT) in a cell can result in increased cell growth and increased non-native nucleic acid uptake compared to the growth and uptake of cells in which the expression of the heterologous gene is not induced.
[0448] In some embodiments, the UBP is incorporated during the log growth phase. In some embodiments, the UBP is incorporated during a non-log growth phase. In some embodiments, the UBP is incorporated during a substantially linear growth phase. In some embodiments, the UBP is stably incorporated into a cell or cell population after a period of growth. For example, the UBP can be stably incorporated into a cell or cell population after growing for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45 or 50 or more doublings. For example, the UBP can be stably incorporated into a cell or cell population after growing for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24 hours. For example, the UBP can be stably incorporated into a cell or cell population after growing for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or 31 days. For example, the UBP can be stably incorporated into a cell or cell population after growing for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months. For example, the UBP can be stably incorporated into a cell or cell population after growing for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or 50 years.
[0449] In some embodiments, the semi-synthetic organisms disclosed herein contain DNA comprising at least one unnatural nucleobase selected from the following:
[0450] In some embodiments, the DNA of the semi-synthetic organisms comprising at least one of the unnatural bases forms unnatural base pairs (UBPs). In some embodiments, the unnatural base pairs (UBPs) are dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1, or dNaM-dTAT1. In some embodiments, the DNA comprises at least one unnatural nucleobase selected from the following:
[0451] In some embodiments, the DNA comprises at least one unnatural nucleobase selected from the following: In some embodiments, the DNA comprises at least one unnatural nucleobase selected from the following:
[0452]
[0453] In some embodiments, the DNA comprises at least one unnatural nucleobase selected from the following: In some embodiments, the DNA comprises at least one unnatural nucleobase selected from In some embodiments, the DNA comprises at least two unnatural nucleobases selected from In some embodiments, the DNA comprises two strands, the first strand comprising at least one nucleobase that is and the second strand comprising at least one nucleobase that is In some embodiments, the DNA comprises at least one unnatural nucleobase that is
[0454] In some embodiments, the cell further utilizes an RNA polymerase to produce mRNA containing one or more unnatural nucleotides. In some embodiments, the RNA polymerase is a heterologous RNA polymerase. In some cases, the cell further utilizes a polymerase to produce tRNA containing an anticodon that comprises one or more unnatural nucleotides. In some embodiments, the tRNA is a heterologous tRNA. In some cases, the tRNA is loaded with an unnatural amino acid. In some cases, the unnatural anticodon of the tRNA pairs with the unnatural codon of the mRNA during translation to synthesize a protein containing an unnatural amino acid.
[0455] In some embodiments, the semi-synthetic organisms disclosed herein express a heterologous nucleoside triphosphate transporter. In some embodiments, the heterologous nucleoside triphosphate transporter is PtNTT2. In some embodiments, the semi-synthetic organisms further express a heterologous tRNA synthetase. In some embodiments, the heterologous tRNA synthetase is Methanosarcina barkeri pyrrolysyl-tRNA synthetase (Mb PylRS). In some embodiments, the semi-synthetic organisms express the nucleoside triphosphate transporter PtNTT2 and also express the tRNA synthetase Methanosarcina barkeri pyrrolysyl-tRNA synthetase (Mb PylRS). In some embodiments, the semi-synthetic organisms further express a heterologous RNA polymerase. In some embodiments, the heterologous RNA polymerase is T7 RNAP. In some embodiments, the semi-synthetic organisms do not express a protein with DNA recombination repair function. In some embodiments, the semi-synthetic organisms are Escherichia coli, and the organisms do not express RecA.
[0456] In some embodiments, the semi-synthetic organisms further comprise a heterologous mRNA. In some embodiments, the heterologous mRNA comprises at least one selected from of unnatural bases. In some embodiments, the heterologous mRNA comprises at least one that is of unnatural bases. In some embodiments, the heterologous mRNA comprises at least one that is of unnatural bases. In some embodiments, the heterologous mRNA comprises at least one that is of unnatural bases.
[0457] In some embodiments, the semi-synthetic organisms further comprise a heterologous tRNA. In some embodiments, the heterologous tRNA comprises at least one selected from of unnatural bases. In some embodiments, the heterologous tRNA comprises at least one that is of unnatural bases. In some embodiments, the heterologous tRNA comprises at least one that is of unnatural bases. In some embodiments, the heterologous tRNA comprises at least one that is of unnatural bases.
[0458] In some embodiments, the semi-synthetic organisms disclosed herein further comprise heterologous mRNA and heterologous tRNA. In some embodiments, the semi-synthetic organisms further comprise (a) a heterologous nucleoside triphosphate transporter, (b) heterologous mRNA, (c) heterologous tRNA, (d) a heterologous tRNA synthetase, and (e) a heterologous RNA polymerase, and wherein the organism does not express proteins with DNA recombination repair functions. In some embodiments, the nucleoside triphosphate transporter is PtNTT2, the tRNA synthetase is Methanosarcina barkeri pyrrolysyl-tRNA synthetase (Mb PylRS), and the RNA polymerase is T7 RNAP. In some embodiments, the semi-synthetic organism is Escherichia coli, and the organism does not express RecA. In some embodiments, the semi-synthetic organism overexpresses one or more DNA polymerases. In some embodiments, the organism overexpresses DNA Pol II.
[0459] Natural and unnatural amino acids
[0460] As used herein, an amino acid residue may refer to a molecule that contains both an amino group and a carboxyl group. Suitable amino acids include, but are not limited to, both D-isomers and L-isomers of naturally occurring amino acids, as well as non-naturally occurring amino acids prepared by organic synthesis or any other method. As used herein, the term amino acid includes, but is not limited to, α-amino acids, β-amino acids, naturally occurring amino acids, non-canonical amino acids, non-natural amino acids, and amino acid analogs.
[0461] The term "α-amino acid" may refer to a molecule that contains both an amino group and a carboxyl group attached to a carbon named the α-carbon. For example:
[0462]
[0463] The term "β-amino acid" may refer to a molecule that contains both an amino group and a carboxyl group in the β-configuration.
[0464] "Naturally occurring amino acids" may refer to any of the twenty amino acids generally found in peptides synthesized in nature, and are known by the single-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V.
[0465] The following table shows a summary of the properties of naturally occurring amino acids:
[0466]
[0467] "Hydrophobic amino acids" include small hydrophobic amino acids and large hydrophobic amino acids. "Small hydrophobic amino acids" may be glycine, alanine, proline, and their analogs. "Large hydrophobic amino acids" may be valine, leucine, isoleucine, phenylalanine, methionine, tryptophan, and their analogs. "Polar amino acids" may be serine, threonine, asparagine, glutamine, cysteine, tyrosine, and their analogs. "Charged amino acids" may be lysine, arginine, histidine, aspartic acid, glutamic acid, and their analogs.
[0468] "Amino acid analogs" may be molecules that are structurally similar to amino acids and can replace amino acids in the formation of peptidic macrocycles. Amino acid analogs include, but are not limited to, β-amino acids and amino acids in which the amino or carboxyl group is replaced by a similar reactive group (e.g., a primary amine is replaced by a secondary or tertiary amine, or a carboxyl group is replaced by an ester).
[0469] "Non-canonical amino acid (ncAA)" or "non-natural amino acid" or "non-naturally occurring amino acid" may be an amino acid that is not one of the twenty amino acids commonly found in naturally synthesized peptides and is known by the single-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. In some cases, non-natural amino acids are a subset of non-canonical amino acids.
[0470] Amino acid analogs can include β - amino acid analogs. Examples of β - amino acid analogs include, but are not limited to, the following: cyclic β - amino acid analogs; β - alanine; (R)-β - phenylalanine; (R)-1,2,3,4 - tetrahydro - isoquinoline - 3 - acetic acid; (R)-3 - amino - 4-(1 - naphthyl)-butyric acid; (R)-3 - amino - 4-(2,4 - dichlorophenyl)butyric acid; (R)-3 - amino - 4-(2 - chlorophenyl)-butyric acid; (R)-3 - amino - 4-(2 - cyanophenyl)-butyric acid; (R)-3 - amino - 4-(2 - fluorophenyl)-butyric acid; (R)-3 - amino - 4-(2 - furyl)-butyric acid; (R)-3 - amino - 4-(2 - methylphenyl)-butyric acid; (R)-3 - amino - 4-(2 - naphthyl)-butyric acid; (R)-3 - amino - 4-(2 - thienyl)-butyric acid; (R)-3 - amino - 4-(2 - trifluoromethylphenyl)-butyric acid; (R)-3 - amino - 4-(3,4 - dichlorophenyl)butyric acid; (R)-3 - amino - 4-(3,4 - difluorophenyl)butyric acid; (R)-3 - amino - 4-(3 - benzothienyl)-butyric acid; (R)-3 - amino - 4-(3 - chlorophenyl)-butyric acid; (R)-3 - amino - 4-(3 - cyanophenyl)-butyric acid; (R)-3 - amino - 4-(3 - fluorophenyl)-butyric acid; (R)-3 - amino - 4-(3 - methylphenyl)-butyric acid; (R)-3 - amino - 4-(3 - pyridyl)-butyric acid; (R)-3 - amino - 4-(3 - thienyl)-butyric acid; (R)-3 - amino - 4-(3 - trifluoromethylphenyl)-butyric acid; (R)-3 - amino - 4-(4 - bromophenyl)-butyric acid; (R)-3 - amino - 4-(4 - chlorophenyl)-butyric acid; (R)-3 - amino - 4-(4 - cyanophenyl)-butyric acid; (R)-3 - amino - 4-(4 - fluorophenyl)-butyric acid; (R)-3 - amino - 4-(4 - iodophenyl)-butyric acid; (R)-3 - amino - 4-(4 - methylphenyl)-butyric acid; (R)-3 - amino - 4-(4 - nitrophenyl)-butyric acid; (R)-3 - amino - 4-(4 - pyridyl)-butyric acid; (R)-3 - amino - 4-(4 - trifluoromethylphenyl)-butyric acid; (R)-3 - amino - 4 - pentafluoro - phenylbutyric acid; (R)-3 - amino - 5 - hexenoic acid; (R)-3 - amino - 5 - hexyneic acid; (R)-3 - amino - 5 - phenylpentanoic acid; (R)-3 - amino - 6 - phenyl - 5 - hexenoic acid; (S)-1,2,3,4 - tetrahydro - isoquinoline - 3 - acetic acid; (S)-3 - amino - 4-(1 - naphthyl)-butyric acid; (S)-3 - amino - 4-(2,4 - dichlorophenyl)butyric acid; (S)-3 - amino - 4-(2 - chlorophenyl)-butyric acid; (S)-3 - amino - 4-(2 - cyanophenyl)-butyric acid; (S)-3 - amino - 4-(2 - fluorophenyl)-butyric acid; (S)-3 - amino - 4-(2 - furyl)-butyric acid;(S)-3-Amino-4-(2-methylphenyl)-butyric acid; (S)-3-Amino-4-(2-naphthyl)-butyric acid; (S)-3-Amino-4-(2-thienyl)-butyric acid; (S)-3-Amino-4-(2-trifluoromethylphenyl)-butyric acid; (S)-3-Amino-4-(3,4-dichlorophenyl)butyric acid; (S)-3-Amino-4-(3,4-difluorophenyl)butyric acid; (S)-3-Amino-4-(3-benzothienyl)-butyric acid; (S)-3-Amino-4-(3-chlorophenyl)-butyric acid; (S)-3-Amino-4-(3-cyanophenyl)-butyric acid; (S)-3-Amino-4-(3-fluorophenyl)-butyric acid; (S)-3-Amino-4-(3-methylphenyl)-butyric acid; (S)-3-Amino-4-(3-pyridyl)-butyric acid; (S)-3-Amino-4-(3-thienyl)-butyric acid; (S)-3-Amino-4-(3-trifluoromethylphenyl)-butyric acid; (S)-3-Amino-4-(4-bromophenyl)-butyric acid; (S)-3-Amino-4-(4-chlorophenyl)butyric acid; (S)-3-Amino-4-(4-cyanophenyl)-butyric acid; (S)-3-Amino-4-(4-fluorophenyl)butyric acid; (S)-3-Amino-4-(4-iodophenyl)-butyric acid; (S)-3-Amino-4-(4-methylphenyl)-butyric acid; (S)-3-Amino-4-(4-nitrophenyl)-butyric acid; (S)-3-Amino-4-(4-pyridyl)-butyric acid; (S)-3-Amino-4-(4-trifluoromethylphenyl)-butyric acid; (S)-3-Amino-4-pentafluoro-phenylbutyric acid; (S)-3-Amino-5-hexenoic acid; (S)-3-Amino-5-hexynoic acid; (S)-3-Amino-5-phenylpentanoic acid; (S)-3-Amino-6-phenyl-5-hexenoic acid; 1,2,5,6-Tetrahydropyridine-3-carboxylic acid; 1,2,5,6-Tetrahydropyridine-4-carboxylic acid; 3-Amino-3-(2-chlorophenyl)-propionic acid; 3-Amino-3-(2-thienyl)-propionic acid; 3-Amino-3-(3-bromophenyl)-propionic acid; 3-Amino-3-(4-chlorophenyl)-propionic acid; 3-Amino-3-(4-methoxyphenyl)-propionic acid; 3-Amino-4,4,4-trifluoro-butyric acid; 3-Aminoadipic acid; D-β-Phenylalanine; β-Leucine; L-β-Homoserine; L-β-Homoserine γ-benzyl ester; L-β-Homoglutamic acid δ-benzyl ester; L-β-Homoisoleucine; L-β-Homoleucine; L-β-Homomethionine; L-β-Homophenylalanine; L-β-Homoproline; L-β-Homotrypthophan; L-β-Homovaline; L-Nω-Carbobenzyloxy-β-homolysine; Nω-L-β-Homoarginine; O-Benzyl-L-β-Homohydroxyproline; O-Benzyl-L-β-Homoserine; O-Benzyl-L-β-Homothreonine; O-Benzyl-L-β-Homotyrosine; γ-Trityl-L-β-Homothreonine; (R)-β-Phenylalanine;L-β-homol-aspartic acid γ-tert-butyl ester; L-β-homol-glutamic acid δ-tert-butyl ester; L-Nω-β-homol-lysine; Nδ-trityl-L-β-homol-glutamine; Nω-2,2,4,6,7-pentamethyl-dihydrobenzofuran-5-sulfonyl-L-β-homol-arginine; O-tert-butyl-L-β-homol-hydroxyproline; O-tert-butyl-L-β-homol-serine; O-tert-butyl-L-β-homol-threonine; O-tert-butyl-L-β-homol-tyrosine; 2-aminocyclopentanecarboxylic acid; and 2-aminocyclohexanecarboxylic acid.
[0471] Amino acid analogs can include analogs of alanine, valine, glycine, or leucine. Examples of amino acid analogs of alanine, valine, glycine, and leucine include, but are not limited to, the following: α-methoxy glycine; α-allyl-L-alanine; α-aminoisobutyric acid; α-methyl-leucine; β-(1-naphthyl)-D-alanine; β-(1-naphthyl)-L-alanine; β-(2-naphthyl)-D-alanine; β-(2-naphthyl)-L-alanine; β-(2-pyridyl)-D-alanine; β-(2-pyridyl)-L-alanine; β-(2-thienyl)-D-alanine; β-(2-thienyl)-L-alanine; β-(3-benzothienyl)-D-alanine; β-(3-benzothienyl)-L-alanine; β-(3-pyridyl)-D-alanine; β-(3-pyridyl)-L-alanine; β-(4-pyridyl)-D-alanine; β-(4-pyridyl)-L-alanine; β-chloro-L-alanine; β-cyano-L-alanine; β-cyclohexyl-D-alanine; β-cyclohexyl-L-alanine; β-cyclopenten-1-yl-alanine; β-cyclopentyl-alanine; β-cyclopropyl-L-Ala-OH. Dicyclohexylammonium salt; β-tert-butyl-D-alanine; β-tert-butyl-L-alanine; γ-aminobutyric acid; L-α,β-diaminopropionic acid; 2,4-dinitro-phenyl glycine; 2,5-dihydro-D-phenyl glycine; 2-amino-4,4,4-trifluorobutyric acid; 2-fluoro-phenyl glycine; 3-amino-4,4,4-trifluoro-butyric acid; 3-fluoro-valine; 4,4,4-trifluoro-valine; 4,5-dehydro-L-leu-OH. Dicyclohexylammonium salt; 4-fluoro-D-phenyl glycine; 4-fluoro-L-phenyl glycine; 4-hydroxy-D-phenyl glycine; 5,5,5-trifluoro-leucine; 6-aminohexanoic acid; cyclopentyl-D-Gly-OH. Dicyclohexylammonium salt; cyclopentyl-Gly-OH.Dicyclohexylammonium salts; D-α,β-diaminopropionic acid; D-α-aminobutyric acid; D-α-tert-butylglycine; D-(2-thienyl)glycine; D-(3-thienyl)glycine; D-2-aminohexanoic acid; D-2-indanylglycine; D-allylglycine-dicyclohexylammonium salt; D-cyclohexylglycine; D-norvaline; D-phenylglycine; β-aminobutyric acid; β-aminoisobutyric acid; (2-bromophenyl)glycine; (2-methoxyphenyl)glycine; (2-methylphenyl)glycine; (2-thiazolyl)glycine; (2-thienyl)glycine; 2-amino-3-(dimethylamino)-propanoic acid; L-α,β-diaminopropionic acid; L-α-aminobutyric acid; L-α-tert-butylglycine; L-(3-thienyl)glycine; L-2-amino-3-(dimethylamino)-propanoic acid; L-2-aminohexanoic acid dicyclohexyl-ammonium salt; L-2-indanylglycine; L-allylglycine dicyclohexylammonium salt; L-cyclohexylglycine; L-phenylglycine; L-propargylglycine; L-norvaline; N-α-aminomethyl-L-alanine; D-α,γ-diaminobutyric acid; L-α,γ-diaminobutyric acid; β-cyclopropyl-L-alanine; (N-β-(2,4-dinitrophenyl))-L-α,β-diaminopropionic acid; (N-β-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-D-α,β-diaminopropionic acid; (N-β-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-L-α,β-diaminopropionic acid; (N-β-4-methyltrityl)-L-α,β-diaminopropionic acid; (N-β-allyloxycarbonyl)-L-α,β-diaminopropionic acid; (N-γ-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-D-α,γ-diaminobutyric acid; (N-γ-1-(4,4-dimethyl-2,6-dioxocyclohex-1-ylidene)ethyl)-L-α,γ-diaminobutyric acid; (N-γ-4-methyltrityl)-D-α,γ-diaminobutyric acid; (N-γ-4-methyltrityl)-L-α,γ-diaminobutyric acid; (N-γ-allyloxycarbonyl)-L-α,γ-diaminobutyric acid; D-α,γ-diaminobutyric acid; 4,5-dehydro-L-leucine; cyclopentyl-D-Gly-OH; cyclopentyl-Gly-OH; D-allylglycine; D-homocyclohexylalanine; L-1-pyrenylalanine; L-2-aminohexanoic acid; L-allylglycine; L-homocyclohexylalanine; and N-(2-hydroxy-4-methoxy-Bzl)-Gly-OH.
[0472] Amino acid analogs can include analogs of arginine or lysine. Examples of amino acid analogs of arginine and lysine include, but are not limited to, the following: citrulline; L-2-amino-3-guanidinopropionic acid; L-2-amino-3-ureidopropionic acid; L-citrulline; Lys(Me)2-OH; Lys(N3)-OH; Nδ-benzyloxycarbonyl-L-ornithine; Nω-nitro-D-arginine; Nω-nitro-L-arginine; α-methyl-ornithine; 2,6-diaminopimelic acid; L-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxo-cyclohex-1-ylidene)ethyl)-D-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxo-cyclohex-1-ylidene)ethyl)-L-ornithine; (Nδ-4-methyltrityl)-D-ornithine; (Nδ-4-methyltrityl)-L-ornithine; D-ornithine; L-ornithine; Arg(Me)(Pbf)-OH; Arg(Me)2-OH (asymmetric); Arg(Me)2-OH (symmetric); Lys(ivDde)-OH; Lys(Me)2-OH.HCl; Lys(Me3)-OH chloride; Nω-nitro-D-arginine; and Nω-nitro-L-arginine.
[0473] Amino acid analogs can include analogs of aspartic acid or glutamic acid. Examples of amino acid analogs of aspartic acid and glutamic acid include, but are not limited to, the following: α-methyl-D-aspartic acid; α-methyl-glutamic acid; α-methyl-L-aspartic acid; γ-methylene-glutamic acid; (N-γ-ethyl)-L-glutamine; [N-α-(4-aminobenzoyl)]-L-glutamic acid; 2,6-diaminopimelic acid; L-α-aminoadipic acid; D-2-aminoadipic acid; D-α-aminoadipic acid; α-aminoheptanoic acid; iminodiacetic acid; L-2-aminoadipic acid; threo-β-methyl-aspartic acid; γ-carboxy-D-glutamic acid γ,γ-di-tert-butyl ester; γ-carboxy-L-glutamic acid γ,γ-di-tert-butyl ester; Glu(OAll)-OH; L-Asu(OtBu)-OH; and pyroglutamic acid.
[0474] Amino acid analogs can include analogs of cysteine and methionine. Examples of amino acid analogs of cysteine and methionine include, but are not limited to, the following: Cys(farnesyl)-OH, Cys(farnesyl)-OMe, α-methyl-methionine, Cys(2-hydroxyethyl)-OH, Cys(3-aminopropyl)-OH, 2-amino-4-(ethylthio)butyric acid, buthionine, buthionine sulfoximine, ethionine, methionine methylsulfonium chloride, selenomethionine, sulfopropylalanine, [2-(4-pyridyl)ethyl]-DL-penicillamine, [2-(4-pyridyl)ethyl]-L-cysteine, 4-methoxybenzyl-D-penicillamine, 4-methoxybenzyl-L-penicillamine, 4-methylbenzyl-D-penicillamine, 4-methylbenzyl-L-penicillamine, benzyl-D-cysteine, benzyl-L-cysteine, benzyl-DL-homocysteine, carbamoyl-L-cysteine, carboxyethyl-L-cysteine, carboxymethyl-L-cysteine, diphenylmethyl-L-cysteine, ethyl-L-cysteine, methyl-L-cysteine, tert-butyl-D-cysteine, trityl-L-homocysteine, trityl-D-penicillamine, cystathionine, homocystine, L-homocystine, (2-aminoethyl)-L-cysteine, seleno-L-cystine, cystathionine, Cys(StBu)-OH, and acetamidomethyl-D-penicillamine.
[0475] Amino acid analogs can include analogs of phenylalanine and tyrosine.Examples of amino acid analogs of phenylalanine and tyrosine include, but are not limited to, the following: β-methyl-phenylalanine, β-hydroxyphenylalanine, α-methyl-3-methoxy-DL-phenylalanine, α-methyl-D-phenylalanine, α-methyl-L-phenylalanine, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, 2,4-dichloro-phenylalanine, 2-(trifluoromethyl)-D-phenylalanine, 2-(trifluoromethyl)-L-phenylalanine, 2-bromo-D-phenylalanine, 2-bromo-L-phenylalanine, 2-chloro-D-phenylalanine, 2-chloro-L-phenylalanine, 2-cyano-D-phenylalanine, 2-cyano-L-phenylalanine, 2-fluoro-D-phenylalanine, 2-fluoro-L-phenylalanine, 2-methyl-D-phenylalanine, 2-methyl-L-phenylalanine, 2-nitro-D-phenylalanine, 2-nitro-L-phenylalanine, 2;4;5-trihydroxy-phenylalanine, 3,4,5-trifluoro-D-phenylalanine, 3,4,5-trifluoro-L-phenylalanine, 3,4-dichloro-D-phenylalanine, 3,4-dichloro-L-phenylalanine, 3,4-difluoro-D-phenylalanine, 3,4-difluoro-L-phenylalanine, 3,4-dihydroxy-L-phenylalanine, 3,4-dimethoxy-L-phenylalanine, 3,5,3'-triiodo-L-thyronine, 3,5-diiodo-D-tyrosine, 3,5-diiodo-L-tyrosine, 3,5-diiodo-L-thyronine, 3-(trifluoromethyl)-D-phenylalanine, 3-(trifluoromethyl)-L-phenylalanine, 3-amino-L-tyrosine, 3-bromo-D-phenylalanine, 3-bromo-L-phenylalanine, 3-chloro-D-phenylalanine, 3-chloro-L-phenylalanine, 3-chloro-L-tyrosine, 3-cyano-D-phenylalanine, 3-cyano-L-phenylalanine, 3-fluoro-D-phenylalanine, 3-fluoro-L-phenylalanine, 3-fluoro-tyrosine, 3-iodo-D-phenylalanine, 3-iodo-L-phenylalanine, 3-iodo-L-tyrosine, 3-methoxy-L-tyrosine, 3-methyl-D-phenylalanine, 3-methyl-L-phenylalanine, 3-nitro-D-phenylalanine, 3-nitro-L-phenylalanine, 3-nitro-L-tyrosine, 4-(trifluoromethyl)-D-phenylalanine, 4-(trifluoromethyl)-L-phenylalanine, 4-amino-D-phenylalanine, 4-amino-L-phenylalanine, 4-benzoyl-D-phenylalanine, 4-benzoyl-L-phenylalanine, 4-bis(2-chloroethyl)amino-L-phenylalanine, 4-bromo-D-phenylalanine, 4-bromo-L-phenylalanine, 4-chloro-D-phenylalanine, 4-chloro-L-phenylalanine, 4-cyano-D-phenylalanine, 4-cyano-L-phenylalanine, 4-fluoro-D-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-D-phenylalanine, 4-iodo-L-phenylalanine, homophenylalanine, thyroxine, 3,3-diphenylalanine, thyronine, ethyl-tyrosine, and methyl-tyrosine.
[0476] Amino acid analogs can include analogs of proline. Examples of amino acid analogs of proline include, but are not limited to, the following: 3,4-dehydro-proline, 4-fluoro-proline, cis-4-hydroxy-proline, thiazolidine-2-carboxylic acid, and trans-4-fluoro-proline.
[0477] Amino acid analogs can include analogs of serine and threonine. Examples of amino acid analogs of serine and threonine include, but are not limited to, the following: 3-amino-2-hydroxy-5-methylhexanoic acid, 2-amino-3-hydroxy-4-methylpentanoic acid, 2-amino-3-ethoxybutyric acid, 2-amino-3-methoxybutyric acid, 4-amino-3-hydroxy-6-methylheptanoic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-ethoxypropionic acid, 4-amino-3-hydroxybutyric acid, and α-methylserine.
[0478] Amino acid analogs can include analogs of tryptophan. Examples of amino acid analogs of tryptophan include, but are not limited to, the following: α-methyl-tryptophan; β-(3-benzothienyl)-D-alanine; β-(3-benzothienyl)-L-alanine; 1-methyl-tryptophan; 4-methyl-tryptophan; 5-benzyloxy-tryptophan; 5-bromo-tryptophan; 5-chloro-tryptophan; 5-fluoro-tryptophan; 5-hydroxy-tryptophan; 5-hydroxy-L-tryptophan; 5-methoxy-tryptophan; 5-methoxy-L-tryptophan; 5-methyl-tryptophan; 6-bromo-tryptophan; 6-chloro-D-tryptophan; 6-chloro-tryptophan; 6-fluoro-tryptophan; 6-methyl-tryptophan; 7-benzyloxy-tryptophan; 7-bromo-tryptophan; 7-methyl-tryptophan; D-1,2,3,4-tetrahydro-norharman-3-carboxylic acid; 6-methoxy-1,2,3,4-tetrahydro-norharman-1-carboxylic acid; 7-azatryptophan; L-1,2,3,4-tetrahydro-norharman-3-carboxylic acid; 5-methoxy-2-methyl-tryptophan; and 6-chloro-L-tryptophan.
[0479] Amino acid analogs can be racemic. In some cases, the D-isomer of the amino acid analog is used. In some instances, the L-isomer of the amino acid analog is used. In some cases, the amino acid analog contains a chiral center in the R or S configuration. Sometimes, one or more amino groups of a β-amino acid analog are replaced by a protecting group such as tert-butoxycarbonyl (BOC group), 9-fluorenylmethoxycarbonyl (FMOC), tosyl, etc. Sometimes, the carboxylic acid functional group of a β-amino acid analog is protected, for example, as its ester derivative. In some instances, salts of the amino acid analog are used.
[0480] In some embodiments, the unnatural amino acid is an unnatural amino acid described in Liu C.C., Schultz, P.G. Annu. Rev. Biochem. 2010, 79, 413, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, the unnatural amino acid includes N6(2-azidoethoxy)-carbonyl-L-lysine.
[0481] In some embodiments, the amino acid residues described herein (e.g., within a protein) are mutated to non-natural amino acids prior to conjugation to the conjugation moiety. In some cases, the mutation to a non-natural amino acid prevents or minimizes an autoimmune response of the immune system. As used herein, the term “non-natural amino acid” refers to an amino acid other than the 20 amino acids that are naturally present in proteins. Non-limiting examples of non-natural amino acids include: p-acetyl-L-phenylalanine, p-iodo-L-phenylalanine, p-methoxyphenylalanine, O-methyl-L-tyrosine, p-propargyloxy-phenylalanine, p-propargyl-phenylalanine, L-3-(2-naphthyl)alanine, 3-methyl-phenylalanine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, tri-O-acetyl-GlcNAcp-serine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-boronic acid phenylalanine, O-propargyl tyrosine, L-phosphoserine, phosphonoserine, phosphonotyrosine, p-bromophenylalanine, selenocysteine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, azido-lysine (N6-azidoethoxy-carbonyl-L-lysine, AzK), non-natural analogs of tyrosine amino acids; non-natural analogs of glutamine amino acids; non-natural analogs of phenylalanine amino acids; non-natural analogs of serine amino acids; non-natural analogs of threonine amino acids; alkyl, aryl, acyl, azido, cyano, halogen, hydrazine, hydrazide, hydroxy, alkenyl, alkynyl, ether, thiol, sulfonyl, seleno, ester, thioacid, borate, boronate, phosphoric acid, phosphonyl, phosphine, heterocycle, enone, imine, aldehyde, hydroxylamine, ketone or amino-substituted amino acids or combinations thereof; amino acids having a photoactivatable crosslinker; spin-labeled amino acids; fluorescent amino acids; metal-binding amino acids; metal-containing amino acids; radioactive amino acids; photocaged and / or photo-isomerizable amino acids; amino acids containing biotin or a biotin analog; amino acids containing a ketone; amino acids containing polyethylene glycol or polyether; heavy atom-substituted amino acids; chemically cleavable or photocleavable amino acids; amino acids having an extended side chain; amino acids containing a toxic group; sugar-substituted amino acids; carbon-linked sugar-containing amino acids; redox-active amino acids; α-hydroxy-containing acids; aminothio acids; α,α-disubstituted amino acids; β-amino acids; cyclic amino acids other than proline or histidine, and aromatic amino acids other than phenylalanine, tyrosine or tryptophan.
[0482] In some embodiments, the unnatural amino acid comprises a selective reactive group, or a reactive group for site-selectively labeling a target protein or polypeptide. In some cases, the chemistry is a bioorthogonal reaction (e.g., biocompatible and selectively reactive). In some cases, the chemistry is a Cu(I)-catalyzed or "copper-free" alkyne-azide cycloaddition reaction, Staudinger ligation, inverse-electron-demand Diels-Alder (IEDDA) reaction, "photo-click" chemistry, or a metal-mediated process (such as olefin metathesis and Suzuki-Miyaura or Sonogashira cross-coupling). In some embodiments, the unnatural amino acid comprises a photoreactive group that crosslinks upon irradiation with, for example, UV. In some embodiments, the unnatural amino acid includes a photocaged amino acid. In some cases, the unnatural amino acid is a para-substituted, meta-substituted, or ortho-substituted amino acid derivative.
[0483] In some cases, the unnatural amino acid includes p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, O-methyl-L-tyrosine, p-methoxyphenylalanine, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, L-3-(2-naphthyl)alanine, 3-methyl-phenylalanine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, tri-O-acetyl-GlcNAcp-serine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, L-phosphoserine, phosphonoserine, phosphonotyrosine, p-bromophenylalanine, p-amino-L-phenylalanine, or isopropyl-L-phenylalanine.
[0484] In some cases, the unnatural amino acid is 3 - aminotyrosine, 3 - nitrotyrosine, 3,4 - dihydroxy - phenylalanine or 3 - iodotyrosine. In some cases, the unnatural amino acid is phenylselenocysteine. In some cases, the unnatural amino acid is a phenylalanine derivative containing benzophenone, ketone, iodide, methoxy, acetyl, benzoyl or azide. In some cases, the unnatural amino acid is a lysine derivative containing benzophenone, ketone, iodide, methoxy, acetyl, benzoyl or azide. In some cases, the unnatural amino acid contains an aromatic side chain. In some cases, the unnatural amino acid does not contain an aromatic side chain. In some cases, the unnatural amino acid contains an azide group. In some cases, the unnatural amino acid contains a Michael receptor group. In some cases, the receptor group contains an unsaturated moiety capable of forming a covalent bond through a 1,2 - addition reaction. In some cases, the receptor group includes an electron - deficient alkene or alkyne. In some cases, the receptor group includes, but is not limited to, α,β - unsaturated: ketone, aldehyde, sulfoxide, sulfone, nitrile, imine or aromatic compound. In some cases, the unnatural amino acid is dehydroalanine. In some cases, the unnatural amino acid contains an aldehyde or ketone group. In some cases, the unnatural amino acid is a lysine derivative containing an aldehyde or ketone group. In some cases, the unnatural amino acid is a lysine derivative containing one or more O, N, Se or S atoms at the β, γ or δ position. In some cases, the unnatural amino acid is a lysine derivative containing an O, N, Se or S atom at the γ position. In some cases, the unnatural amino acid is a lysine derivative in which the εN atom is replaced by an oxygen atom. In some cases, the unnatural amino acid is a lysine derivative that is not a naturally occurring post - translationally modified lysine.
[0485] In some cases, the unnatural amino acid is an amino acid containing a side chain in which the sixth atom from the α position contains a carbonyl group. In some cases, the unnatural amino acid is an amino acid containing a side chain in which the sixth atom from the α position contains a carbonyl group and the fifth atom from the α position is nitrogen. In some cases, the unnatural amino acid is an amino acid containing a side chain in which the seventh atom from the α position is an oxygen atom.
[0486] In some cases, the unnatural amino acid is a serine derivative containing selenium. In some cases, the unnatural amino acid is selenoserine (2 - amino - 3 - hydrogenselenopropionic acid). In some cases, the unnatural amino acid is 2 - amino - 3 - ((2 - ((3 - (benzyloxy) - 3 - oxopropyl)amino)ethyl)seleno)propionic acid. In some cases, the unnatural amino acid is 2 - amino - 3 - (phenylseleno)propionic acid. In some cases, the unnatural amino acid contains selenium, and the oxidation of selenium results in the formation of an unnatural amino acid containing an alkene.
[0487] In some cases, the unnatural amino acid comprises a cyclooctynyl group. In some cases, the unnatural amino acid comprises a trans-cyclooctenyl group. In some cases, the unnatural amino acid comprises a norbornenyl group. In some cases, the unnatural amino acid comprises a cyclopropenyl group. In some cases, the unnatural amino acid comprises a diazacyclopropenyl group. In some cases, the unnatural amino acid comprises a tetrazine group.
[0488] In some cases, the unnatural amino acid is a lysine derivative in which the side-chain nitrogen is carbamylated. In some cases, the unnatural amino acid is a lysine derivative in which the side-chain nitrogen is acylated. In some cases, the unnatural amino acid is 2-amino-6-{[(tert-butoxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(tert-butoxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-Boc-N6-methyllysine. In some cases, the unnatural amino acid is N6-acetyllysine. In some cases, the unnatural amino acid is pyrrolysine. In some cases, the unnatural amino acid is N6-trifluoroacetyllysine. In some cases, the unnatural amino acid is 2-amino-6-{[(benzyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(p-iodobenzyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(p-nitrobenzyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-prolyllysine. In some cases, the unnatural amino acid is 2-amino-6-{[(cyclopentyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-(cyclopentanecarbonyl)lysine. In some cases, the unnatural amino acid is N6-(tetrahydrofuran-2-carbonyl)lysine. In some cases, the unnatural amino acid is N6-(3-ethynyltetrahydrofuran-2-carbonyl)lysine. In some cases, the unnatural amino acid is N6-((prop-2-yn-1-yloxy)carbonyl)lysine. In some cases, the unnatural amino acid is 2-amino-6-{[(2-azidocyclopentyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-((2-azidoethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is 2-amino-6-{[(2-nitrobenzyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(2-cyclooctynyl)oxycarbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-(2-aminobut-3-ynoyl)lysine. In some cases, the unnatural amino acid is 2-amino-6-((2-aminobut-3-ynoyl)oxy)hexanoic acid. In some cases, the unnatural amino acid is N6-(allyloxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-(butenyl-4-oxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-(pentenyl-5-oxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-((but-3-yn-1-yloxy)carbonyl)-lysine. In some cases, the unnatural amino acid is N6-((pent-4-yn-1-yloxy)carbonyl)-lysine. In some cases, the unnatural amino acid is N6-(thiazolidine-4-carbonyl)lysine.In some cases, the unnatural amino acid is 2-amino-8-oxononanoic acid. In some cases, the unnatural amino acid is 2-amino-8-oxooctanoic acid. In some cases, the unnatural amino acid is N6-(2-oxoacetyl)lysine.
[0489] In some cases, the unnatural amino acid is N6-propionyllysine. In some cases, the unnatural amino acid is N6-butyrylllysine. In some cases, the unnatural amino acid is N6-(but-2-enoyl)lysine. In some cases, the unnatural amino acid is N6-((bicyclo[2.2.1]hept-5-ene-2-yloxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((spiro[2.3]hex-1-ene-5-ylmethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((((4-(1-(trifluoromethyl)cycloprop-2-en-1-yl)benzyl)oxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((bicyclo[2.2.1]hept-5-ene-2-ylmethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is cysteine lysine. In some cases, the unnatural amino acid is N6-(((1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((3-(3-methyl-3H-diazirin-3-yl)propoxy)carbonyl)lysine. In some cases, the unnatural amino acid is N6-((m-nitrobenzyloxy)N6-methylcarbonyl)lysine. In some cases, the unnatural amino acid is N6-((bicyclo[6.1.0]non-4-yn-9-ylmethoxy)carbonyl)-lysine. In some cases, the unnatural amino acid is N6-((cyclohept-3-en-1-yloxy)carbonyl)-L-lysine.
[0490] In some cases, the unnatural amino acid is 2-amino-3-(((((benzyloxy)carbonyl)amino)methyl)seleno)propanoic acid. In some embodiments, the unnatural amino acid is incorporated into the protein through a repurposed amber, opal, or ochre stop codon. In some embodiments, the unnatural amino acid is incorporated into the protein through a 4-base codon. In some embodiments, the unnatural amino acid is incorporated into the protein through a repurposed rare sense codon.
[0491] In some embodiments, the unnatural amino acid is incorporated into the protein through an unnatural codon containing unnatural nucleotides.
[0492] In some embodiments, the protein comprises at least two unnatural amino acids. In some embodiments, the protein comprises at least three unnatural amino acids. In some embodiments, the protein comprises at least two different unnatural amino acids. In some embodiments, the protein comprises at least three different unnatural amino acids. The at least one unnatural amino acid: is a lysine analogue; comprises an aromatic side chain; comprises an azide group; comprises an alkyne group; or comprises an aldehyde or ketone group. In some embodiments, the at least one unnatural amino acid does not comprise an aromatic side chain. In some embodiments, the at least one unnatural amino acid comprises N6-azidoethoxycarbonyl-L-lysine (AzK) or N6-propynyl-ethoxycarbonyl-L-lysine (PraK). In some embodiments, the at least one unnatural amino acid comprises N6-azidoethoxycarbonyl-L-lysine (AzK). In some embodiments, the at least one unnatural amino acid comprises N6-propynyl-ethoxycarbonyl-L-lysine (PraK).
[0493] In some cases, incorporation of non-natural amino acids into proteins is mediated by orthogonal, modified synthetase / tRNA pairs. Such orthogonal pairs comprise a native or mutant synthetase that is capable of loading a non-natural tRNA with a specific non-natural amino acid, typically while minimizing: a) loading of other endogenous amino acids or alternative non-natural amino acids onto the non-natural tRNA and b) loading of any other (including endogenous) tRNAs. Such orthogonal pairs comprise a tRNA that is capable of being loaded by the synthetase while avoiding loading of other endogenous amino acids by endogenous synthetases. In some embodiments, such pairs are identified from a variety of organisms such as bacterial, yeast, archaeal or human sources. In some embodiments, the orthogonal synthetase / tRNA pair comprises components from a single organism. In some embodiments, the orthogonal synthetase / tRNA pair comprises components from two different organisms. In some embodiments, the orthogonal synthetase / tRNA pair comprises components that promote the translation of different amino acids prior to modification. In some embodiments, the orthogonal synthetase is a modified alanine synthetase. In some embodiments, the orthogonal synthetase is a modified arginine synthetase. In some embodiments, the orthogonal synthetase is a modified asparagine synthetase. In some embodiments, the orthogonal synthetase is a modified aspartate synthetase. In some embodiments, the orthogonal synthetase is a modified cysteine synthetase. In some embodiments, the orthogonal synthetase is a modified glutamine synthetase. In some embodiments, the orthogonal synthetase is a modified glutamate synthetase. In some embodiments, the orthogonal synthetase is a modified alanine glycine. In some embodiments, the orthogonal synthetase is a modified histidine synthetase. In some embodiments, the orthogonal synthetase is a modified leucine synthetase. In some embodiments, the orthogonal synthetase is a modified isoleucine synthetase. In some embodiments, the orthogonal synthetase is a modified lysine synthetase. In some embodiments, the orthogonal synthetase is a modified methionine synthetase. In some embodiments, the orthogonal synthetase is a modified phenylalanine synthetase. In some embodiments, the orthogonal synthetase is a modified proline synthetase. In some embodiments, the orthogonal synthetase is a modified serine synthetase. In some embodiments, the orthogonal synthetase is a modified threonine synthetase. In some embodiments, the orthogonal synthetase is a modified tryptophan synthetase. In some embodiments, the orthogonal synthetase is a modified tyrosine synthetase. In some embodiments, the orthogonal synthetase is a modified valine synthetase. In some embodiments, the orthogonal synthetase is a modified phosphoserine synthetase. In some embodiments, the orthogonal tRNA is a modified alanine tRNA. In some embodiments, the orthogonal tRNA is a modified arginine tRNA. In some embodiments, the orthogonal tRNA is a modified asparagine tRNA. In some embodiments, the orthogonal tRNA is a modified aspartate tRNA.In some embodiments, the orthogonal tRNA is a modified cysteine tRNA. In some embodiments, the orthogonal tRNA is a modified glutamine tRNA. In some embodiments, the orthogonal tRNA is a modified glutamate tRNA. In some embodiments, the orthogonal tRNA is a modified alanine glycine. In some embodiments, the orthogonal tRNA is a modified histidine tRNA. In some embodiments, the orthogonal tRNA is a modified leucine tRNA. In some embodiments, the orthogonal tRNA is a modified isoleucine tRNA. In some embodiments, the orthogonal tRNA is a modified lysine tRNA. In some embodiments, the orthogonal tRNA is a modified methionine tRNA. In some embodiments, the orthogonal tRNA is a modified phenylalanine tRNA. In some embodiments, the orthogonal tRNA is a modified proline tRNA. In some embodiments, the orthogonal tRNA is a modified serine tRNA. In some embodiments, the orthogonal tRNA is a modified threonine tRNA. In some embodiments, the orthogonal tRNA is a modified tryptophan tRNA. In some embodiments, the orthogonal tRNA is a modified tyrosine tRNA. In some embodiments, the orthogonal tRNA is a modified valine tRNA. In some embodiments, the orthogonal tRNA is a modified phosphoserine tRNA. In any of these embodiments, the tRNA can be a heterologous tRNA.
[0494] In some embodiments, non-natural amino acids are incorporated into proteins via an aminoacyl (aaRS or RS)-tRNA synthetase-tRNA pair. Exemplary aaRS-tRNA pairs include, but are not limited to, the Methanococcus jannaschii (Mj-Tyr) aaRS / tRNA pair, the Escherichia coli TyrRS (Ec-Tyr) / Bacillus stearothermophilus tRNA CUA pair, the Escherichia coli LeuRS (Ec-Leu) / Bacillus stearothermophilus tRNA CUA pair, and the pyrrolysyl-tRNA pair. In some cases, non-natural amino acids are incorporated into proteins via the Mj-TyrRS / tRNA pair. Exemplary unnatural amino acids (UAAs) that can be incorporated via the Mj-TyrRS / tRNA pair include, but are not limited to, para-substituted phenylalanine derivatives such as p-aminophenylalanine and p-methoxyphenylalanine; meta-substituted tyrosine derivatives such as 3-aminotyrosine, 3-nitrotyrosine, 3,4-dihydroxyphenylalanine, and 3-iodotyrosine; phenylselenocysteine; p-boronophenylalanine; and o-nitrobenzyltyrosine.
[0495] In some cases, non-natural amino acids are incorporated via Ec-Tyr / tRNACUA or Ec-Leu / tRNA CUA into the protein. It can be through Ec-Tyr / tRNA CUA or Ec-Leu / tRNA CUA Exemplary UAAs that can be incorporated include, but are not limited to, phenylalanine derivatives containing benzophenone, ketone, iodide, or azide substituents; O-propargyl tyrosine; α-aminooctanoic acid, O-methyl tyrosine, O-nitrobenzyl cysteine; and 3-(naphthalen-2-ylamino)-2-aminopropanoic acid.
[0496] In some cases, unnatural amino acids are incorporated into the proteins herein through pyrrolysyl-tRNA pairs. In some instances, PylRS is obtained from archaeal species, such as those that produce methane. In some instances, PylRS is obtained from Methanosarcina barkeri, Methanosarcina mazei, or Methanosarcina acetivorans. Exemplary UAAs that can be incorporated through pyrrolysyl-tRNA pairs include, but are not limited to, amide and carbamate-substituted lysines, such as 2-amino-6-((R)-tetrahydrofuran-2-carboxamido)hexanoic acid, N-ε- D -prolyl- L -lysine and N-ε-cyclopentyloxycarbonyl- L -lysine; N-ε-acryloyl- L -lysine; N-ε-[(1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl]- L -lysine; and N-ε-(1-methylcycloprop-2-enecarboxamido)lysine.
[0497] In some cases, unnatural amino acids are incorporated into the proteins described herein through the synthetases disclosed in US 9,988,619 and US 9,938,516, the disclosures of each of which are hereby incorporated by reference in their entirety. Exemplary UAAs that can be incorporated through such synthetases include p-methylazido-L-phenylalanine, aralkyl, heterocyclic, heteroaralkyl unnatural amino acids, etc. In some embodiments, such UAAs contain pyridyl, pyrazinyl, pyrazolyl, triazolyl, oxazolyl, thiazolyl, thiophenyl, or other heterocycles. In some embodiments, such amino acids contain azide, tetrazine, or other chemical groups capable of conjugating with a coupling partner (such as a water-soluble moiety). In some embodiments, such synthetases are expressed and used to incorporate UAAs into proteins in vivo. In some embodiments, a cell-free translation system is used to incorporate UAAs into proteins using such synthetases.
[0498] In some cases, non-natural amino acids are incorporated into the proteins described herein by naturally occurring synthetases. In some embodiments, non-natural amino acids are incorporated into proteins by organisms auxotrophic for one or more amino acids. In some embodiments, the synthetase corresponding to the auxotrophic amino acid is capable of loading the non-natural amino acid onto the corresponding tRNA. In some embodiments, the non-natural amino acid is selenocysteine or a derivative thereof. In some embodiments, the non-natural amino acid is selenomethionine or a derivative thereof. In some embodiments, the non-natural amino acid is an aromatic amino acid, wherein the aromatic amino acid contains an aryl halide, such as iodide. In an embodiment, the non-natural amino acid is structurally similar to the auxotrophic amino acid.
[0499] In some cases, the non-natural amino acids include Figure 8A the non-natural amino acids shown.
[0500] In some cases, the non-natural amino acids include lysine or phenylalanine derivatives or analogs. In some cases, the non-natural amino acids include lysine derivatives or lysine analogs. In some cases, the non-natural amino acid includes pyrrolysine (Pyl). In some cases, the non-natural amino acids include phenylalanine derivatives or phenylalanine analogs. In some cases, the non-natural amino acid is the non-natural amino acid described by Wan et al., "Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool," Biocheim Biophys Aceta 1844(6):1059 - 4070(2014). In some cases, the non-natural amino acids include Figure 8B and Figure 8C the non-natural amino acids shown.
[0501] In some embodiments, the non-natural amino acids include Figure 8D - Figure 8G the non-natural amino acids shown (obtained from Table 1 of Dumas et al., Chemical Science 2015, 6, 50 - 69).
[0502] In some embodiments, the unnatural amino acids incorporated into proteins as described herein are disclosed in US9,840,493; US 9,682,934; US 2017 / 0260137; US 9,938,516; or US 2018 / 0086734; the disclosures of each of which are hereby incorporated by reference in their entirety. Exemplary UAAs that can be incorporated by such synthetases include p-methylazido-L-phenylalanine, aralkyl, heterocyclic, and heteroaralkyl, as well as lysine derivative unnatural amino acids. In some embodiments, such UAAs contain pyridyl, pyrazinyl, pyrazolyl, triazolyl, oxazolyl, thiazolyl, thienyl, or other heterocycles. In some embodiments, such amino acids contain azide, tetrazine, or other chemical groups capable of conjugating with a coupling partner (such as a water-soluble moiety). In some embodiments, the UAA contains an azide attached to an aromatic moiety via an alkyl linker. In some embodiments, the alkyl linker is C1-C 10 linker. In some embodiments, the UAA contains a tetrazine attached to an aromatic moiety via an alkyl linker. In some embodiments, the UAA contains a tetrazine attached to an aromatic moiety via an amino group. In some embodiments, the UAA contains a tetrazine attached to an aromatic moiety via an alkylamino group. In some embodiments, the UAA contains an azide attached to the terminal nitrogen of an amino acid side chain (e.g., N6 of a lysine derivative, or N5, N4, or N3 of a derivative containing a shorter alkyl side chain) via an alkyl chain. In some embodiments, the UAA contains a tetrazine attached to the terminal nitrogen of an amino acid side chain via an alkyl chain. In some embodiments, the UAA contains an azide or tetrazine attached to an amide via an alkyl linker. In some embodiments, the UAA is a carbamate or amide of 3-aminopropionic acid, serine, lysine, or a derivative thereof containing an azide or tetrazine. In some embodiments, such UAAs are incorporated into proteins in vivo. In some embodiments, such UAAs are incorporated into proteins in a cell-free system.
[0503] Cell type
[0504] In some embodiments, many types of cells / microorganisms are used, e.g., for transformation or genetic engineering. In some embodiments, the cells are prokaryotic or eukaryotic cells. In some embodiments, the prokaryotic cells are bacterial cells. In some embodiments, the eukaryotic cells are fungal cells or single-celled protozoa. In some embodiments, the fungal cells are yeast cells. In other cases, the eukaryotic cells are cultured animal, plant, or human cells. In additional cases, the eukaryotic cells are present in an organism such as a plant, multicellular fungus, or animal.
[0505] In some embodiments, the engineered microorganism is a single-celled organism and is generally capable of dividing and proliferating. As used herein, an "engineered microorganism" is a microorganism whose genetic material has been altered using genetic engineering techniques (i.e., recombinant DNA technology). The microorganism can include one or more of the following characteristics: aerobic, anaerobic, filamentous, non-filamentous, haploid, diploid, auxotrophic, and / or non-auxotrophic. In certain embodiments, the engineered microorganism is a prokaryotic microorganism (e.g., a bacterium), and in certain embodiments, the engineered microorganism is a non-prokaryotic microorganism, such as a eukaryotic microorganism. In some embodiments, the engineered microorganism is a eukaryotic microorganism (e.g., yeast, other fungi, amoeba). In some embodiments, the engineered microorganism is a fungus. In some embodiments, the engineered organism is yeast.
[0506] Any suitable yeast can be selected as the host microorganism, engineered microorganism, genetically modified organism, or source of heterologous or modified polynucleotides. Yeasts include, but are not limited to, yeasts of the genus Yarrowia (e.g., Yarrowia lipolytica (formerly classified as Candida lipolytica)), yeasts of the genus Candida (e.g., C. revkaufi, C. viswanathii, C. pulcherrima, C. tropicalis, C. utilis), yeasts of the genus Rhodotorula (e.g., Rhodotorula glutinus, Rhodotorula graminis), yeasts of the genus Rhodosporidium (e.g., Rhodosporidium toruloides), yeasts of the genus Saccharomyces (e.g., Saccharomyces cerevisiae, Saccharomyces bayanus, Saccharomyces pastorianus, Saccharomyces carlsbergensis), yeasts of the genus Cryptococcus, yeasts of the genus Trichosporon (e.g., Trichosporon pullans, Trichosporon cutaneum), yeasts of the genus Pichia (e.g., Pichia pastoris, Komagataella phaffii), and yeasts of the genus Lipomyces (e.g., Lipomyces starkeyii, Lipomyces lipoferus). In some embodiments, suitable yeasts belong to the following genera: Arachniotus, Aspergillus, Aureobasidium, Auxarthron, Blastomyces, Candida, Chrysosporium, Debaryomyces, Coccidiodes, Cryptococcus, Gymnoascus, Hansenula, Histoplasma, Issatchenkia, Kluyveromyces, Lipomyces, Lssatchenkia, Microsporum, Myxotrichum,Myxozyma, Oidiodendron, Pachysolen, Penicillium, Pichia, Rhodosporidium, Rhodotorula, Rhodotorula, Saccharomyces, Schizosaccharomyces, Scopulariopsis, Sepedonium, Trichosporon or Yarrowia. In some embodiments, suitable yeasts belong to the following species: Arachniotus flavoluteus, Aspergillus flavus, Aspergillus fumigatus, Aspergillus niger, Aureobasidium pullulans, Auxarthron thaxteri, Blastomyces dermatitidis, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xestobii, Chrysosporuim keratinophilum, Coccidiodes immitis, Cryptococcus albidus var. diffluens, Cryptococcus laurentii, Cryptococcus neofomans, Debaryomyces hansenii, Gymnoascus dugwayensis, Hansenula anomala,Histoplasma capsulatum, Issatchenkia occidentalis, Issatchenkia orientalis, Kluyveromyces lactis, Kluyveromyces marxianus, Kluyveromyces thermotolerans, Kluyveromyces waltii, Lipomyces lipofer, Lipomyces starkeyi, Microsporum gypseum, Myxotrichum deflexum, Oidiodendron echinuulatum, Pachysolen tannophilis, Penicillium notatum, Pichia anomala, Pichia pastoris, Pichia stipitis, Rhodosporidium toruloides, Rhodotorula glutinis, Rhodotorula graminis, Saccharomyces cerevisiae, Saccharomyces kluyveri, Schizosaccharomyces pombe, Scopulariopsis acremonium, Sepedonium chrysospermum, Trichosporon cutaneum, Trichosporon pullulans, Yarrowia lipolytica, or Yarrowia lipolytica (formerly classified as Candida lipolytica). In some embodiments, the yeast is a strain of Yarrowia lipolytica, including but not limited to ATCC 20362, ATCC 8862, ATCC 18944, ATCC 20228, ATCC 76982, and LGAM S(7)1 strain (Papanikolaou S. and Aggelis G., Bioresour. Technol. 82(1):43-9 (2002)). In certain embodiments, the yeast is a Candida species (i.e., a Candida species) yeast. Any suitable Candida species can be used to produce aliphatic dicarboxylic acids (e.g., suberic acid, sebacic acid, dodecanedioic acid, tetradecanedioic acid, hexadecanedioic acid, octadecanedioic acid, eicosanedioic acid), and / or any suitable Candida species can be genetically modified for the production of aliphatic dicarboxylic acids (e.g., suberic acid, sebacic acid, dodecanedioic acid, tetradecanedioic acid, hexadecanedioic acid, octadecanedioic acid, eicosanedioic acid). In some embodiments, suitable Candida species include but are not limited to Candida albicans, Candida dubliniensis,Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xestobii, and any other Candida species yeast described herein. Non-limiting examples of Candida species strains include, but are not limited to, sAA001 (ATCC 20336), sAA002 (ATCC 20913), sAA003 (ATCC 20962), sAA496 (US 2012 / 0077252), sAA106 (US 2012 / 0077252), SU-2 (ura3- / ura3-), H5343 (β-oxidation blocked; U.S. Patent No. 5,648,247) strains. Any suitable strain from a Candida species yeast can be utilized as a parent strain for genetic modification.,
[0507] The genetic content of yeast genera, species, and strains is often closely related, making it difficult to distinguish, classify, and / or name them. In some cases, strains of Candida lipolytica and Yarrowia lipolytica may be difficult to distinguish, classify, and / or name and, in some cases, may be considered the same organism. In some cases, various strains of Candida tropicalis and Candida viswanathii may be difficult to distinguish, classify, and / or name (e.g., see Arie et al., J. Gen. Appl. Microbiol., 46, 257-262 (2000)). Some Candida tropicalis and Candida viswanathii strains obtained from the ATCC and from other commercial or academic sources can be considered equivalent and equally suitable for the embodiments described herein. In some embodiments, some parent strains of Candida tropicalis and Candida viswanathii are considered to differ only in name.
[0508] Any suitable fungus can be selected as a source of host microorganisms, engineered microorganisms, or heterologous polynucleotides. Non-limiting examples of fungi include, but are not limited to, fungi of the genus Aspergillus (e.g., Aspergillus parasiticus, Aspergillus nidulans), fungi of the genus Thraustochytrium, fungi of the genus Schizochytrium, and fungi of the genus Rhizopus (e.g., Rhizopus arrhizus, Rhizopus oryzae, Rhizopus nigricans). In some embodiments, the fungus is an Aspergillus parasiticus strain, including but not limited to strain ATCC 24690, and in certain embodiments, the fungus is an Aspergillus nidulans strain, including but not limited to strain ATCC 38163.
[0509] Any suitable prokaryote can be selected as a source of host microorganisms, engineered microorganisms, or heterologous polynucleotides. Gram-negative or Gram-positive bacteria can be selected. Examples of bacteria include, but are not limited to, bacteria of the genus Bacillus (e.g., Bacillus subtilis, Bacillus megaterium), bacteria of the genus Acinetobacter, bacteria of the genus Norcardia, bacteria of the genus Xanthobacter, bacteria of the genus Escherichia (e.g., Escherichia coli (e.g., strains DH10B, Stbl2, DH5-α, DB3, DB3.1), DB4, DB5, JDP682, and ccdA-over (e.g., U.S. Application No. 09 / 518,188)), bacteria of the genus Streptomyces, bacteria of the genus Erwinia, bacteria of the genus Klebsiella, bacteria of the genus Serratia (e.g., Serratia marcescens), bacteria of the genus Pseudomonas (e.g., Pseudomonas aeruginosa), bacteria of the genus Salmonella (e.g., Salmonella typhimurium, Salmonella typhi), bacteria of the genus Megasphaera (e.g., Megasphaera elsdenii). Bacteria also include, but are not limited to, photosynthetic bacteria (e.g., green non-sulfur bacteria (e.g., bacteria of the genus Chloroflexus (e.g., Chloroflexus aurantiacus), bacteria of the genus Chloronema (e.g., Chloronema gigateum)), green sulfur bacteria (e.g., bacteria of the genus Chlorobium (e.g., Chlorobium limicola), bacteria of the genus Pelodictyon (e.g., Pelodictyon luteolum)), purple sulfur bacteria (e.g., bacteria of the genus Chromatium (e.g., Chromatium okenii)), and purple non-sulfur bacteria (e.g., bacteria of the genus Rhodospirillum (e.g., Rhodospirillum rubrum), bacteria of the genus Rhodobacter (e.g., Rhodobacter sphaeroides, Rhodobacter capsulatus), and bacteria of the genus Rhodomicrobium (e.g., Rhodomicrobium vanellii))).
[0510] Cells from non-microbial organisms can be utilized as a source of host microorganisms, engineered microorganisms, or heterologous polynucleotides. Examples of such cells include, but are not limited to, insect cells (e.g., Drosophila (e.g., D. melanogaster), Spodoptera (e.g., S. frugiperda Sf9 or Sf21 cells), and Trichoplusa (e.g., High-Five cells)); nematode cells (e.g., C. elegans cells); avian cells; amphibian cells (e.g., Xenopus laevis cells); reptilian cells; mammalian cells (e.g., NIH3T3, 293, CHO, COS, VERO, C127, BHK, Per-C6, Bowes melanoma, and HeLa cells); and plant cells (e.g., Arabidopsis thaliana, Nicotania tabacum, Cuphea acinifolia, Cuphea aequipetala, Cuphea angustifolia, Cuphea appendiculata, Cuphea avigera, Cuphea avigera var. pulcherrima, Cuphea axilliflora, Cuphea bahiensis, Cuphea baillonis, Cuphea brachypoda, Cuphea bustamanta, Cuphea calcarata, Cuphea calophylla, Cuphea calophylla subsp. mesostemon, Cuphea carthagenensis, Cuphea circaeoides, Cuphea confertiflora, Cuphea cordata, Cuphea crassiflora, Cuphea cyanea, Cuphea decandra, Cuphea denticulata, Cuphea disperma, Cuphea epilobiifolia, Cuphea ericoides, Cuphea flava, Cuphea flavisetula, Cuphea fuchsiifolia, CupheaGaumeri, Cuphea glutinosa, Cuphea heterophylla, Cuphea hookeriana, Cuphea hyssopifolia (Mexican heather), Cuphea hyssopoides, Cuphea ignea, Cuphea ingrata, Cuphea jorullensis, Cuphea lanceolata, Cuphea linarioides, Cuphea llavea, Cuphea lophostoma, Cuphea lutea, Cuphea lutescens, Cuphea melanium, Cuphea melvilla, Cuphea micrantha, Cuphea micropetala, Cuphea mimuloides, Cuphea nitidula, Cuphea palustris, Cuphea parsonsia, Cuphea pascuorum, Cuphea paucipetala, Cuphea procumbens, Cuphea pseudosilene, Cuphea pseudovaccinium, Cuphea pulchra, Cuphea racemosa, Cuphea repens, Cuphea salicifolia, Cuphea salvadorensis, Cuphea schumannii, Cuphea sessiliflora, Cuphea sessilifolia, Cuphea setosa, Cuphea spectabilis, Cuphea spermacoce, Cuphea splendida, Cuphea splendida var. viridiflava, Cuphea strigulosa, Cuphea subuligera, Cuphea teleandra, Cuphea thymoides, Cuphea tolucana, Cuphea urens, Cuphea utriculosa, Cuphea with blue petiolesviscosissima), Cuphea watsoniana, Cuphea wrightii, Cuphea lanceolata
[0511] Microorganisms or cells used as sources of host organisms or heterologous polynucleotides are commercially available. The microorganisms and cells described herein, as well as other suitable microorganisms, can be obtained from, for example: Invitrogen Corporation (Carlsbad, California), American Type Culture Collection (Manassas, Virginia), and Agricultural Research Culture Collection (NRRL; Peoria, Illinois). Host microorganisms and engineered microorganisms can be provided in any suitable form. For example, such microorganisms can be provided in liquid culture or solid culture (e.g., agar-based medium), which can be a primary culture or can have been passaged (e.g., diluted and cultured) one or more times. Microorganisms can also be provided in frozen form or dried form (e.g., lyophilized). Microorganisms can be provided at any suitable concentration.
[0512] Polymerase
[0513] A particularly useful function of polymerase is to catalyze the polymerization of nucleic acid chains using an existing nucleic acid as a template. Other useful functions are described elsewhere herein. Examples of useful polymerases include DNA polymerase and RNA polymerase.
[0514] The ability of unnatural nucleic acids to improve the specificity, processivity, or other characteristics of polymerase is highly desirable in a variety of situations where unnatural nucleic acid incorporation is required, including amplification, sequencing, labeling, detection, cloning, and many other situations
[0515] In some cases, the disclosure herein includes, for example, polymerase that incorporates unnatural nucleic acids into growing template copies during DNA amplification. In some embodiments, the polymerase can be modified such that the active site of the polymerase is modified to reduce steric entry inhibition of unnatural nucleic acids into the active site. In some embodiments, the polymerase can be modified to provide complementarity to one or more unnatural features of the unnatural nucleic acid. Such polymerases can be expressed in cells or engineered for stable incorporation of UBP into cells. Accordingly, the invention includes compositions comprising heterologous or recombinant polymerase and methods of using the same.
[0516] Polymerases can be modified using methods of protein engineering. For example, molecular modeling can be performed based on crystal structures to identify positions in the polymerase where mutations can be made to modify the target activity. Residues identified as alternative targets can be replaced with residues selected using energy minimization modeling, homology modeling, and / or conservative amino acid substitutions, as described in the following references: Bordo, et al. J Mol Biol 217:721-729 (1991) and Hayes, et al. Proc Natl Acad Sci, USA 99:15926-15931 (2002), the disclosures of each of which are hereby incorporated by reference in their entirety.
[0517] Any of a variety of polymerases can be used in the methods or compositions described herein, including, for example, protein-based enzymes isolated from biological systems and functional variants thereof. When a particular polymerase (such as those exemplified below) is mentioned, it will be understood to include its functional variants unless otherwise indicated. In some embodiments, the polymerase is a wild-type polymerase. In some embodiments, the polymerase is a modified or mutant polymerase. In some embodiments, the polymerase can be a heterologous polymerase.
[0518] Polymerases having characteristics of improved entry of non-natural nucleic acids into the active site region and coordination with non-natural nucleotides in the active site region can also be used. In some embodiments, the modified polymerase has a modified nucleotide binding site.
[0519] In some embodiments, the specificity of the modified polymerase for non-natural nucleic acids is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for non-natural nucleic acids. In some embodiments, the specificity of the modified or wild-type polymerase for non-natural nucleic acids containing a modified sugar is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for natural nucleic acids and / or non-natural nucleic acids without a modified sugar. In some embodiments, the specificity of the modified or wild-type polymerase for non-natural nucleic acids containing a modified base is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for natural nucleic acids and / or non-natural nucleic acids without a modified base. In some embodiments, the specificity of the modified or wild-type polymerase for non-natural nucleic acids containing a triphosphate is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for nucleic acids containing a triphosphate and / or non-natural nucleic acids without a triphosphate. For example, the specificity of the modified or wild-type polymerase for non-natural nucleic acids containing a triphosphate can be at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for non-natural nucleic acids having a diphosphate or monophosphate, or no phosphate or a combination thereof.
[0520] In some embodiments, the modified or wild-type polymerase has relaxed specificity for non-natural nucleic acids. In some embodiments, the specificity of the modified or wild-type polymerase for non-natural nucleic acids and for natural nucleic acids is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for natural nucleic acids. In some embodiments, the specificity of the modified or wild-type polymerase for non-natural nucleic acids containing modified sugars and for natural nucleic acids is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for natural nucleic acids. In some embodiments, the specificity of the modified or wild-type polymerase for non-natural nucleic acids containing modified bases and for natural nucleic acids is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the specificity of the wild-type polymerase for natural nucleic acids.
[0521] The absence of exonuclease activity can be a wild-type characteristic or a characteristic conferred by a variant or engineered polymerase. For example, the exo-Klenow fragment is a mutant form of the Klenow fragment that lacks 3' to 5' proofreading exonuclease activity.
[0522] The methods of the invention can be used to expand the substrate range of any DNA polymerase that lacks intrinsic 3 to 5' exonuclease proofreading activity or in which the 3 to 5' exonuclease proofreading activity has been disabled, e.g., by mutation. Examples of DNA polymerases include polA, polB (see, e.g., Parrel and Loeb, Nature Struc Biol 2001) polC, polD, polY, polX, and reverse transcriptase (RT), but preferably processive high-fidelity polymerases (PCT / GB2004 / 004643). In some embodiments, the modified or wild-type polymerase substantially lacks 3' to 5' proofreading exonuclease activity. In some embodiments, the modified or wild-type polymerase substantially lacks 3' to 5' proofreading exonuclease activity for non-natural nucleic acids. In some embodiments, the modified or wild-type polymerase has 3' to 5' proofreading exonuclease activity. In some embodiments, the modified or wild-type polymerase has 3' to 5' proofreading exonuclease activity for natural nucleic acids and substantially lacks 3' to 5' proofreading exonuclease activity for non-natural nucleic acids.
[0523] In some embodiments, the 3'-to-5' proofreading exonuclease activity of the modified polymerase is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the proofreading exonuclease activity of the wild-type polymerase. In some embodiments, the 3'-to-5' proofreading exonuclease activity of the modified polymerase on non-natural nucleic acids is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, 99.99% of the proofreading exonuclease activity of the wild-type polymerase on natural nucleic acids. In some embodiments, the 3'-to-5' proofreading exonuclease activity of the modified polymerase on non-natural nucleic acids and the 3'-to-5' proofreading exonuclease activity on natural nucleic acids are at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the proofreading exonuclease activity of the wild-type polymerase on natural nucleic acids. In some embodiments, the 3'-to-5' proofreading exonuclease activity of the modified polymerase on natural nucleic acids is at least about 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.99% of the proofreading exonuclease activity of the wild-type polymerase on natural nucleic acids.
[0524] In some embodiments, a polymerase is characterized by its rate of dissociation from a nucleic acid. In some embodiments, the polymerase has a relatively low rate of dissociation with respect to one or more natural and non-natural nucleic acids. In some embodiments, the polymerase has a relatively high rate of dissociation with respect to one or more natural and non-natural nucleic acids. The dissociation rate is a polymerase activity that can be adjusted in the methods described herein to tune the reaction rate.
[0525] In some embodiments, a polymerase is characterized by its fidelity when used with a particular natural and / or non-natural nucleic acid or collection of natural and / or non-natural nucleic acids. Fidelity generally refers to the accuracy with which a polymerase incorporates the correct nucleic acid into a growing nucleic acid chain when preparing a copy of a nucleic acid template. DNA polymerase fidelity can be measured as the ratio of correct to incorrect natural and non-natural nucleic acid incorporation when natural and non-natural nucleic acids are present, for example, at equal concentrations to compete for the same site in the polymerase-strand-template nucleic acid ternary complex for strand synthesis. DNA polymerase fidelity can be calculated as the ratio of (k cat / K m ) of correct natural and non-natural nucleic acids to (k cat / K m ) of incorrect natural and non-natural nucleic acids; where k cat and K mis a Michaelis-Menten parameter in steady-state enzyme kinetics (Fersht, A.R. (1985) Enzyme Structure and Mechanism, 2nd ed., p. 350, W.H. Freeman & Co., New York, incorporated herein by reference). In some embodiments, the polymerase has a fidelity value of at least about 100, 1000, 10,000, 100,000, or 1x10 6 , with or without proofreading activity.
[0526] In some embodiments, assays that detect the incorporation of unnatural nucleic acids having a specific structure are used to screen polymerases or variants thereof from natural sources. In one example, polymerases can be screened for the ability to incorporate unnatural nucleic acids or UBPs (e.g., d5SICSTP, dCNMOTP, dTPT3TP, dNaMTP, dCNMOTP-dTPT3TP, or d5SICSTP-dNaMTP UBPs). Polymerases that exhibit properties such as modifications to unnatural nucleic acids compared to wild-type polymerases (e.g., heterologous polymerases) can be used. For example, the modified properties can be, for example, K m , k cat , V max , the ability of the polymerase to continuously synthesize in the presence of unnatural nucleic acids (or naturally occurring nucleotides), the average template read-length of the polymerase in the presence of unnatural nucleic acids, the specificity of the polymerase for unnatural nucleic acids, the binding rate of unnatural nucleic acids, the rate of product (pyrophosphate, triphosphate, etc.) release, the branching rate, or any combination thereof. In one embodiment, the modified properties are a reduced K m for unnatural nucleic acids and / or an increased k cat / K m or V max / K m for unnatural nucleic acids. Similarly, the polymerase optionally has an increased binding rate of unnatural nucleic acids, an increased product release rate, and / or a reduced branching rate compared to the wild-type polymerase.
[0527] Meanwhile, the polymerase can incorporate natural nucleic acids (e.g., A, C, G, and T) into the growing nucleic acid copy. For example, the polymerase optionally exhibits a specific activity for natural nucleic acids up to at least about 5% of the corresponding wild-type polymerase (e.g., 5%, 10%, 25%, 50%, 75%, 100%, or higher), and the ability to continuously synthesize using natural nucleic acids in the presence of a template up to at least 5% of the wild-type polymerase in the presence of natural nucleic acids (e.g., 5%, 10%, 25%, 50%, 75%, 100%, or higher). Optionally, the polymerase exhibits a k cat / Km or V max / K m At least about 5% of the wild-type polymerase (e.g., about 5%, 10%, 25%, 50%, 75% or 100% or higher).
[0528] The polymerase used herein that can have the ability to incorporate unnatural nucleic acids with a specific structure can also be generated using directed evolution methods. Nucleic acid synthesis assays can be used to screen polymerase variants with specificity for any of a variety of unnatural nucleic acids. For example, polymerase variants can be screened for the ability to incorporate unnatural nucleoside triphosphates opposite unnatural nucleotides in a DNA template (e.g., dTPT3TP opposite dCNMO, dCNMOTP opposite dTPT3, NaMTP opposite dTPT3, or TAT1TP opposite dCNMO or dNaM). In some embodiments, such an assay is an in vitro assay, e.g., using recombinant polymerase variants. In some embodiments, such an assay is an in vivo assay, e.g., expressing the polymerase variant in cells. Such directed evolution techniques can be used to screen variants of any suitable polymerase for activity against any of the unnatural nucleic acids described herein. In some cases, the polymerase used herein has the ability to incorporate unnatural ribonucleotides into nucleic acids such as RNA. For example, NaM or TAT1 ribonucleotides are incorporated into nucleic acids using the polymerase described herein.
[0529] The modified polymerase of the composition can optionally be a modified and / or recombinant Φ29-type DNA polymerase. Optionally, the polymerase can be a modified and / or recombinant Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722 or L17 polymerase.
[0530] The modified polymerase of the composition can optionally be a modified and / or recombinant prokaryotic DNA polymerase, e.g., DNA polymerase II (Pol II), DNA polymerase III (Pol III), DNA polymerase IV (Pol IV), DNA polymerase V (Pol V). In some embodiments, the modified polymerase includes a polymerase that mediates DNA synthesis of nucleotides across non-instructive lesions. In some embodiments, the genes encoding Pol I, Pol II (polB), Poll IV (dinB) and / or Pol V (umuCD) are constitutively expressed or overexpressed in engineered cells or SSOs. In some embodiments, increased expression or overexpression of Pol II promotes increased retention of unnatural base pairs (UBPs) in engineered cells or SSOs.
[0531] Nucleic acid polymerases that are generally useful in the present invention include DNA polymerases, RNA polymerases, reverse transcriptases, and mutants or altered forms thereof. DNA polymerases and their properties are described in particular detail in the following: DNA Replication, 2nd Edition, Kornberg and Baker, W.H. Freeman, New York, NY (1991). Known conventional DNA polymerases that can be used in the present invention include, but are not limited to, Pyrococcus furiosus (Pfu) DNA polymerase (Lundberg et al., 1991, Gene, 108:1, Stratagene), Pyrococcus woesei (Pwo) DNA polymerase (Hinnisdaels et al., 1996, Biotechniques, 20:186-8, Boehringer Mannheim), Thermus thermophilus (Tth) DNA polymerase (Myers and Gelfand 1991, Biochemistry 30:7661), Bacillus stearothermophilus DNA polymerase (Stenesh and McGowan, 1977, Biochim Biophys Acta 475:32), Thermococcus litoralis (Tli) DNA polymerase (also known as Vent TM DNA polymerase, Cariello et al., 1991, Polynucleotides Res, 19:4193, New England Biolabs), 9°Nm TM DNA polymerase (New England Biolabs), Stoffel fragment, Thermo (Amersham Pharmacia Biotech UK), Therminator TM(New England Biolabs), Thermotoga maritima (Tma) DNA polymerase (Diaz and Sabino, 1998 Braz J Med.Res, 31:1239), Thermus aquaticus (Taq) DNA polymerase (Chien et al., 1976, J.Bacteoriol, 127:1550), DNA polymerase, Pyrococcus kodakaraensis KOD DNA polymerase (Takagi et al., 1997, Appl.Environ.Microbiol. 63:4504), JDF-3 DNA polymerase (from the hyperthermococcal species JDF-3, patent application WO 0132887), Thermococcus GB-D (PGB-D) DNA polymerase (also known as Deep Vent TM DNA polymerase, Juncosa-Ginesta et al., 1994, Biotechniques, 16:820, New England Biolabs), UlTma DNA polymerase (from the thermophilic organism Thermotoga maritima; Diaz and Sabino, 1998 Braz J.Med.Res, 31:1239; PE Applied Biosystems), Tgo DNA polymerase (from thermococcus gorgonarius, Roche Molecular Biochemicals), Escherichia coli DNA polymerase I (Lecomte and Doubleday, 1983, Polynucleotides Res. 11:7505), T7 DNA polymerase (Nordstrom et al., 1981, J Biol.Chem. 256:3112), and archaeal DP1I / DP2 DNA polymerase II (Cann et al., 1998, Proc.Natl.Acad.Sci.USA 95:14250). Consider both mesophilic and thermophilic polymerases. Thermophilic DNA polymerases include, but are not limited to 9°Nm TM 、Therminator TM 、Taq、Tne、Tma、Pfu、TfI、Tth、TIi、Stoffel fragment, Vent TM and Deep Vent TMDNA polymerases, KOD DNA polymerase, Tgo, JDF-3 and its mutants, variants and derivatives. Also contemplated are polymerases as 3'-exonuclease-deficient mutants. Reverse transcriptases useful in the present invention include, but are not limited to, reverse transcriptases from the following: HIV, HTLV-I, HTLV-II, FeLV, FIV, SIV, AMV, MMTV, MoMuLV and other retroviruses (see Levin, Cell 88:5-8 (1997); Verma, Biochim Biophys Acta. 473:1-38 (1977); Wu et al., CRC Crit Rev Biochem. 3:289-347 (1975)). Additional examples of polymerases include, but are not limited to, 9°N DNA polymerase, Taq DNA polymerase, DNA polymerase, Pfu DNA polymerase, RB69 DNA polymerase, KOD DNA polymerase, and DNA polymerase (Gardner et al. (2004) “Comparative Kinetics of Nucleotide Analog Incorporation by Vent DNA Polymerase” J. Biol. Chem., 279(12), 11834-11842; Gardner and Jack “Determinants of nucleotide sugar recognition in an archaeon DNA polymerase” Nucleic Acids Research, 27(12)2545-2553). Polymerases isolated from non-thermophilic organisms may be thermally inactivatable. An example is the DNA polymerase from bacteriophage. It will be understood that polymerases from any of a variety of sources may be modified to increase or decrease their tolerance to high temperature conditions. In some embodiments, the polymerase may be thermophilic. In some embodiments, the thermophilic polymerase may be thermally inactivatable. Thermophilic polymerases are generally useful in high temperature conditions or thermal cycling conditions, such as those used in polymerase chain reaction (PCR) techniques.
[0532] In some embodiments, the polymerase includes Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, L17, 9°Nm TM 、Therminator TMDNA polymerases, Tne, Tma, TfI, Tth, Tli, Stoffel fragment, Vent TM and Deep Vent TM DNA polymerases, KOD DNA polymerase, Tgo, JDF-3, Pfu, Taq, T7 DNA polymerase, T7 RNA polymerase, PGB-D, UlTma DNA polymerase, Escherichia coli DNA polymerase I, Escherichia coli DNA polymerase III, archaeal DP1I / DP2 DNA polymerase II, 9°N DNA polymerase, Taq DNA polymerase, DNA polymerases, Pfu DNA polymerase, SP6 RNA polymerase, RB69 DNA polymerase, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, II reverse transcriptase or III reverse transcriptase.
[0533] In some embodiments, the polymerase is DNA polymerase I (or Klenow fragment), Vent polymerase, DNA polymerases, KOD DNA polymerase, Taq polymerase, T7 DNA polymerase, T7 RNA polymerase, Therminator TM DNA polymerases, POLB polymerase, SP6 RNA polymerase, Escherichia coli DNA polymerase I, Escherichia coli DNA polymerase III, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, II reverse transcriptase or III reverse transcriptase.
[0534] Nucleotide transporters
[0535] Nucleotide transporters (NTs) are a group of membrane transporters that facilitate the transfer of nucleotide substrates across cell membranes and vesicles. In some embodiments, there are two types of NTs, namely concentrative nucleoside transporters and equilibrative nucleoside transporters. In some cases, NTs also encompass organic anion transporters ('OAT) and organic cation transporters (OCT). In some cases, the nucleotide transporter is a nucleoside triphosphate transporter (NTT).
[0536] In some embodiments, the nucleoside triphosphate transporter (NTT) is from bacteria, plants, or algae. In some embodiments, the nucleotide nucleoside triphosphate transporter is TpNTT1, TpNTT2, TpNTT3, TpNTT4, TpNTT5, TpNTT6, TpNTT7, TpNTT8 (Thalassiosira pseudonana), PtNTT1, PtNTT2, PtNTT3, PtNTT4, PtNTT5, PtNTT6 (Phaeodactylum tricornutum), GsNTT (Galdieria sulphuraria), AtNTT1, AtNTT2 (Arabidopsis thaliana), CtNTT1, CtNTT2 (Chlamydia trachomatis), PamNTT1, PamNTT2 (Protochlamydia amoebophila), CcNTT (Caedibacter caryophilus), or RpNTT1 (Rickettsia prowazekii).
[0537] In some embodiments, the NTT is CNT1, CNT2, CNT3, ENT1, ENT2, OAT1, OAT3, or OCT1.
[0538] In some embodiments, the NTT imports unnatural nucleic acids into an organism (e.g., a cell). In some embodiments, the NTT can be modified such that the nucleotide binding site of the NTT is modified to reduce steric entry inhibition of the unnatural nucleic acid into the nucleotide binding site. In some embodiments, the NTT can be modified to provide increased interaction with one or more natural or unnatural features of the unnatural nucleic acid. Such NTTs can be expressed or engineered in cells for stable import of UBP into the cell. Accordingly, the present invention includes compositions comprising a heterologous or recombinant NTT and methods of using the same.
[0539] The NTT can be modified using methods regarding protein engineering. For example, molecular modeling can be performed based on crystal structures to identify positions in the NTT that can be mutated to modify the target activity or binding site. Residues identified as substitution targets can be replaced with residues selected using energy minimization modeling, homology modeling, and / or conservative amino acid substitutions, as described in Bordo, et al. J Mol Biol 217:721-729 (1991) and Hayes, et al. Proc Natl Acad Sci, USA 99:15926-15931 (2002), the disclosures of each of which are hereby incorporated by reference in their entirety.
[0540] Any of a variety of NTTs can be used in the methods or compositions described herein, including, for example, protein-based enzymes isolated from biological systems and functional variants thereof. When a specific NTT (such as those exemplified below) is mentioned, it will be understood to include its functional variants, unless otherwise indicated. In some embodiments, the NTT is a wild-type NTT. In some embodiments, the NTT is a modified or mutant NTT.
[0541] NTTs with characteristics of improved entry of unnatural nucleic acids into cells and coordination with unnatural nucleotides in the nucleotide-binding region can also be used. In some embodiments, the modified NTT has a modified nucleotide-binding site. In some embodiments, the modified or wild-type NTT has a relaxed specificity for unnatural nucleic acids. For example, the NTT optionally exhibits specific import activity for unnatural nucleotides up to at least about 0.1% of the corresponding wild-type NTT (e.g., about 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.1%, 1.2%, 1.5%, 1.8%, 2%, 3%, 4%, 5%, 10%, 25%, 50%, 75%, 100% or higher). Optionally, the NTT exhibits a k cat / K m or V max / K m up to at least about 0.1% of the wild-type NTT (e.g., about 0.1%, 0.2%, 0.5%, 0.8%, 1%, 1.1%, 1.2%, 1.5%, 1.8%, 2%, 3%, 4%, 5%, 10%, 25%, 50%, 75% or 100% or higher).
[0542] NTTs can be characterized according to their affinity for triphosphates (i.e., Km) and / or import rate (i.e., Vmax). In some embodiments, the NTT has a relative Km or Vmax for one or more natural and unnatural triphosphates. In some embodiments, the NTT has a relatively high Km or Vmax for one or more natural and unnatural triphosphates.
[0543] NTTs from natural sources or their variants can be screened using assays that detect the amount of triphosphate (using mass spectrometry or radioactivity if the triphosphate is appropriately labeled). In one example, NTTs can be screened for the ability to import unnatural triphosphates (e.g., dTPT3TP, dCNMOTP, d5SICSTP, dNaMTP, NaMTP, and / or TPT1TP). NTTs (e.g., heterologous NTTs) that exhibit modified characteristics for unnatural nucleic acids compared to wild-type NTTs can be used. For example, the modified characteristics can be, for example, K m , k cat , Vmax 。In one embodiment, the modified property is a reduced K for non-natural triphosphates m and / or an increased k for non-natural triphosphates cat / K m or V max / K m 。Similarly, the NTT optionally has an increased binding rate of non-natural triphosphates, an increased intracellular release rate, and / or an increased cellular import rate as compared to the wild-type NTT.
[0544] Meanwhile, the NTT can import natural triphosphates, such as dATP, dCTP, dGTP, dTTP, ATP, CTP, GTP, and / or TTP, into cells. In some cases, the NTT optionally exhibits specific import activity for natural nucleic acids capable of supporting replication and transcription. In some embodiments, the NTT optionally exhibits a k cat / K m or V max / K m 。
[0545] The NTTs that can have the ability to import non-natural triphosphates of a specific structure as used herein can also be generated using directed evolution methods. Nucleic acid synthesis assays can be used to screen for NTT variants with specificity for any of a variety of non-natural triphosphates. For example, NTT variants can be screened for the ability to import non-natural triphosphates (such as d5SICSTP, dNaMTP, dCNMOTP, dTPT3TP, NaMTP, and / or TPT1TP). In some embodiments, such an assay is an in vitro assay, for example, using recombinant NTT variants. In some embodiments, such an assay is an in vivo assay, for example, expressing NTT variants in cells. Such techniques can be used to screen for variants of any suitable NTT for activity against any of the non-natural triphosphates described herein.
[0546] Nucleic Acid Reagents and Tools
[0547] The nucleotide and / or nucleic acid reagents (or polynucleotides) for the methods, cells, or engineered microorganisms described herein comprise one or more ORFs with or without unnatural nucleotides. The ORF can be from any suitable source, sometimes from genomic DNA, mRNA, reverse transcribed RNA, or complementary DNA (cDNA), or a nucleic acid library comprising one or more of the foregoing, and from any organism species containing the nucleic acid sequence of interest, the protein of interest, or the activity of interest. Non-limiting examples of organisms from which the ORF can be obtained include, for example, bacteria, yeast, fungi, humans, insects, nematodes, bovines, equines, canines, felines, rats, or mice. In some embodiments, the nucleotide and / or nucleic acid reagents or other reagents described herein are isolated or purified. The ORF containing unnatural nucleotides can be created by published in vitro methods. In some cases, the nucleotide or nucleic acid reagent comprises unnatural nucleobases.
[0548] The nucleic acid reagent sometimes comprises a nucleotide sequence adjacent to the ORF that binds and translates to encode an amino acid tag. The nucleotide sequence encoding the tag is located 3' and / or 5' of the ORF in the nucleic acid reagent, thereby encoding a tag at the C-terminus or N-terminus of the protein or peptide encoded by the ORF. Any tag that does not eliminate in vitro transcription and / or translation can be utilized and can be appropriately selected by one skilled in the art. The tag can facilitate the isolation and / or purification of the desired ORF product from the culture or fermentation medium. In some cases, a nucleic acid reagent library is used with the methods and compositions described herein. For example, a library with at least 100, 1000, 2000, 5000, 10,000, or more than 50,000 unique polynucleotides, wherein each polynucleotide comprises at least one unnatural nucleobase.
[0549] Nucleic acids or nucleic acid reagents with or without non-natural nucleotides can contain certain elements, typically selected according to the intended use of the nucleic acid, such as regulatory elements. Any of the following elements can be included or excluded in the nucleic acid reagent. For example, the nucleic acid reagent can include one or more or all of the following nucleotide elements: one or more promoter elements, one or more 5' untranslated regions (5' UTRs), one or more regions into which a target nucleotide sequence can be inserted ("insertion elements"), one or more target nucleotide sequences, one or more 3' untranslated regions (3' UTRs), and one or more selection elements. The nucleic acid reagent can be provided with one or more such elements, and additional elements can be inserted into the nucleic acid prior to introducing the nucleic acid into the desired organism. In some embodiments, the provided nucleic acid reagent contains a promoter, a 5' UTR, an optional 3' UTR, and one or more insertion elements through which the target nucleotide sequence is inserted (i.e., cloned) into the nucleic acid reagent. In certain embodiments, the provided nucleic acid reagent contains a promoter, one or more insertion elements, and an optional 3' UTR, and the 5' UTR / target nucleotide sequence is inserted with the optional 3' UTR. The elements can be arranged in any order suitable for expression in the selected expression system (e.g., expression in the selected organism or, for example, expression in a cell-free system), and in some embodiments, the nucleic acid reagent contains the following elements in the 5' to 3' direction: (1) a promoter element, a 5' UTR, and one or more insertion elements; (2) a promoter element, a 5' UTR, and a target nucleotide sequence; (3) a promoter element, a 5' UTR, one or more insertion elements, and a 3' UTR; and (4) a promoter element, a 5' UTR, a target nucleotide sequence, and a 3' UTR. In some embodiments, the UTRs can be optimized to alter or increase the transcription or translation of a fully native or ORF containing non-natural nucleotides.
[0550] Nucleic acid reagents (e.g., expression cassettes and / or expression vectors) can include a variety of regulatory elements, including promoters, enhancers, translation initiation sequences, transcription termination sequences, and other elements. A "promoter" is generally one or more DNA sequences that function when located at a relatively fixed position with respect to the transcription start site. For example, a promoter can be located upstream of the nucleic acid segment of a nucleotide triphosphate transporter. A "promoter" contains the core elements required for the basal interaction of RNA polymerase with transcription factors and can contain upstream elements and response elements. An "enhancer" generally refers to a DNA sequence that does not function at a fixed distance from the transcription start site and can be located 5' or 3' of the transcription unit. In addition, enhancers can be within introns as well as within the coding sequence itself. Enhancers are generally between 10 and 300 in length and are cis-acting. Enhancers function to increase transcription from a nearby promoter. Like promoters, enhancers generally also contain response elements that mediate transcriptional regulation. Enhancers generally determine the regulation of expression and can be used to alter or optimize the expression of an ORF (including a fully native or an ORF containing non-native nucleotides).
[0551] As described above, nucleic acid reagents can also contain one or more 5' UTRs and one or more 3' UTRs. For example, expression vectors used in eukaryotic host cells (e.g., yeast, fungi, insects, plants, animals, humans, or nucleated cells) and prokaryotic host cells (e.g., viruses, bacteria) can contain sequences that signal for transcription termination, which may affect mRNA expression. These regions can be transcribed as polyadenylation segments in the untranslated portion of the mRNA encoding the tissue factor protein. The 3' untranslated region also includes the transcription termination site. In some preferred embodiments, the transcription unit contains a polyadenylation region. One benefit of this region is that it increases the likelihood that the transcribed unit is processed and transported like mRNA. The identification and use of polyadenylation signals in expression constructs are well known. In some preferred embodiments, homologous polyadenylation signals can be used in transgenic constructs.
[0552] The 5’UTR can contain one or more elements endogenous to the nucleotide sequence from which it is derived and sometimes includes one or more exogenous elements. The 5’UTR can be derived from any suitable nucleic acid, such as genomic DNA, plasmid DNA, RNA, or mRNA, for example, from any suitable organism (e.g., virus, bacterium, yeast, fungus, plant, insect, or mammal). One skilled in the art can select appropriate elements for the 5’UTR based on the selected expression system (e.g., expression in a selected organism or, for example, expression in a cell-free system). The 5’UTR sometimes contains one or more of the following elements known to one skilled in the art: enhancer sequences (e.g., for transcription or translation), transcription start sites, transcription factor binding sites, translation regulatory sites, translation start sites, translation factor binding sites, accessory protein binding sites, feedback regulator binding sites, Pribnow box, TATA box, -35 element, E-box (helix-loop-helix binding element), ribosome binding sites, replicons, internal ribosome entry sites (IRES), silencer elements, etc. In some embodiments, promoter elements can be isolated such that all 5’UTR elements required for appropriate conditional regulation are contained within a promoter element fragment or within a functional subsequence of a promoter element fragment.
[0553] The 5’UTR in a nucleic acid reagent can contain a translation enhancer nucleotide sequence. The translation enhancer nucleotide sequence is typically located between a promoter and a target nucleotide sequence in the nucleic acid reagent. The translation enhancer sequence typically binds to ribosomes, sometimes to an 18S rRNA-binding ribonucleotide sequence (i.e., a 40S ribosome-binding sequence), and sometimes to an internal ribosome entry sequence (IRES). The IRES typically forms an RNA scaffold with precisely positioned RNA tertiary structure that contacts the 40S ribosomal subunit via multiple specific intermolecular interactions. Examples of ribosome enhancer sequences are known and can be identified by one of ordinary skill in the art (e.g., Mignone et al., Nucleic Acids Research 33:D141-D146 (2005); Paulous et al., Nucleic Acids Research 31:722-733 (2003); Akbergenov et al., Nucleic Acids Research 32:239-247 (2004); Mignone et al., Genome Biology 3(3):reviews0004.1-0001.10 (2002); Gallie, Nucleic Acids Research 30:3401-3411 (2002); Shaloiko et al., DOI:10.1002 / bit.20267; and Gallie et al., Nucleic Acids Research 15:3257-3273 (1987); the disclosures of each of which are hereby incorporated by reference in their entirety).
[0554] The translation enhancer sequence is sometimes a eukaryotic sequence, such as a Kozak consensus sequence or other sequences (e.g., a hydra sequence, GenBank accession number U07128). The translation enhancer sequence is sometimes a prokaryotic sequence, such as a Shine-Dalgarno consensus sequence. In certain embodiments, the translation enhancer sequence is a viral nucleotide sequence. The translation enhancer sequence is sometimes from the 5’UTR of a plant virus, such as, for example, tobacco mosaic virus (TMV), alfalfa mosaic virus (AMV); tobacco etch virus (ETV); potato virus Y (PVY); turnip mosaic (poty) virus; and pea seed-borne mosaic virus. In certain embodiments, an ω sequence of about 67 bases in length from TMV is included as a translation enhancer sequence in the nucleic acid reagent (e.g., lacking guanosine nucleotides and including a poly(CAA) central region of 25 nucleotides in length).
[0555] The 3’UTR can contain one or more elements endogenous to the nucleotide sequence from which it is derived and sometimes includes one or more exogenous elements. The 3’UTR can be derived from any suitable nucleic acid, such as genomic DNA, plasmid DNA, RNA or mRNA, for example, from any suitable organism (e.g., virus, bacterium, yeast, fungus, plant, insect or mammal). One skilled in the art can select appropriate elements for the 3’UTR based on the selected expression system (e.g., expression in the selected organism). The 3’UTR sometimes contains one or more of the following elements known to those skilled in the art: transcriptional regulatory sites, transcriptional start sites, transcriptional termination sites, transcription factor binding sites, translational regulatory sites, translational termination sites, translational start sites, translation factor binding sites, ribosome binding sites, replicons, enhancer elements, silencer elements and polyadenylation tails. The 3’UTR typically includes a polyadenylation tail and sometimes does not, and if a polyadenylation tail is present, one or more adenosine moieties can be added or deleted therein (e.g., about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45 or about 50 adenosine moieties can be added or subtracted).
[0556] In some embodiments, modifications of the 5’UTR and / or 3’UTR are used to alter (e.g., increase, add, decrease or substantially eliminate) the activity of a promoter. The alteration of promoter activity, in turn, can alter the activity (e.g., enzyme activity) of a peptide, polypeptide or protein by alteration of transcription of one or more target nucleotide sequences from a promoter element operably linked to a modified 5’ or 3’UTR. For example, in certain embodiments, a microorganism can be engineered by genetic modification to express a nucleic acid reagent comprising a modified 5’ or 3’UTR, which modified 5’ or 3’UTR can add novel activity (e.g., an activity not normally found in the host organism), or increase the expression of an existing activity by increasing transcription from a homologous or heterologous promoter operably linked to a target nucleotide sequence (e.g., a target homologous or heterologous nucleotide sequence). In some embodiments, in certain embodiments, a microorganism can be engineered by genetic modification to express a nucleic acid reagent comprising a modified 5’ or 3’UTR, which modified 5’ or 3’UTR can decrease the expression of an activity by decreasing or substantially eliminating transcription from a homologous or heterologous promoter operably linked to a target nucleotide sequence.
[0557] The expression of a nucleoside triphosphate transporter from an expression cassette or expression vector can be controlled by any promoter capable of expressing in a prokaryotic or eukaryotic cell. Promoter elements are generally required for DNA synthesis and / or RNA synthesis. A promoter element generally comprises a DNA region that can promote transcription of a particular gene, by providing a starting site for RNA synthesis corresponding to the gene. In some embodiments, the promoter is generally located near the gene it regulates, upstream of the gene (e.g., 5' of the gene), and on the same DNA strand as the sense strand of the gene. In some embodiments, the promoter element can be isolated from a gene or organism and inserted to be operably linked to a polynucleotide sequence to allow alteration and / or regulation of expression. A non-native promoter for nucleic acid expression (e.g., a promoter not normally associated with a given nucleic acid sequence) is generally referred to as a heterologous promoter. In certain embodiments, a heterologous promoter and / or 5' UTR can be inserted to be operably linked to a polynucleotide encoding a polypeptide having the desired activity as described herein. As used herein, the terms "operably linked" and "functionally linked to" with respect to a promoter refer to the relationship between a coding sequence and a promoter element. A promoter is operably linked or functionally linked to a coding sequence when the promoter element regulates or controls the expression of the coding sequence via transcription. The terms "operably linked" and "functionally linked to" are used interchangeably herein with respect to a promoter element.
[0558] A promoter generally interacts with an RNA polymerase. A polymerase is an enzyme that catalyzes the synthesis of nucleic acids using pre-existing nucleic acid reagents. When the template is a DNA template, proteins are synthesized after transcription of RNA molecules. Enzymes having polymerase activity suitable for use in the present method include any polymerase that is active in the selected system for synthesizing proteins using the selected template. In some embodiments, a promoter (e.g., a heterologous promoter), also referred to herein as a promoter element, can be operably linked to a nucleotide sequence or open reading frame (ORF). Transcription from the promoter element can catalyze the synthesis of an RNA corresponding to the nucleotide sequence or ORF sequence operably linked to the promoter, which in turn results in the synthesis of the desired peptide, polypeptide, or protein.
[0559] Promoter elements sometimes exhibit responsiveness to regulatory control. Promoter elements can sometimes also be regulated by a selection agent. That is, transcription from a promoter element can sometimes be turned on, off, upregulated, or downregulated in response to changes in environmental, nutritional, or internal conditions or signals (e.g., heat-inducible promoters, light-regulated promoters, feedback-regulated promoters, hormone-influenced promoters, tissue-specific promoters, oxygen- and pH-influenced promoters, promoters responsive to a selection agent (e.g., kanamycin), etc.). Promoters that are influenced by environmental, nutritional, or internal signals are often affected by signals (direct or indirect) that bind at or near the promoter and increase or decrease the expression of the target sequence under certain conditions. In all cases of the methods disclosed herein, promoters, including natural or modified promoters, can be used to alter or optimize the expression of a completely natural ORF (e.g., NTT or aaRS) or an ORF containing unnatural nucleotides (e.g., mRNA or tRNA).
[0560] Non-limiting examples of select agents or modulators used in the embodiments described herein that affect transcription from a promoter element include, but are not limited to: (1) nucleic acid segments encoding products that confer resistance to a compound that is otherwise toxic (e.g., an antibiotic); (2) nucleic acid segments encoding products that are otherwise absent in the recipient cell (e.g., an essential product, a tRNA gene, a auxotrophic marker); (3) nucleic acid segments encoding products that inhibit the activity of a gene product; (4) nucleic acid segments encoding products that may be readily identifiable (e.g., phenotypic markers such as an antibiotic (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and cell surface proteins); (5) nucleic acid segments that bind to products that are otherwise harmful to cell survival and / or function; (6) nucleic acid segments that otherwise inhibit the activity of any of the nucleic acid segments described in items 1-5 above (e.g., antisense oligonucleotides); (7) nucleic acid segments that bind to products that modify substrates (e.g., restriction endonucleases); (8) nucleic acid segments that can be used to isolate or identify a desired molecule (e.g., a specific protein binding site); (9) nucleic acid segments encoding a specific nucleotide sequence that may otherwise be non-functional (e.g., for PCR amplification of a subpopulation of molecules); (10) nucleic acid segments that directly or indirectly confer resistance or sensitivity to a specific compound in the absence thereof; (11) nucleic acid segments encoding products that are toxic in the recipient cell or convert a relatively non-toxic compound into a toxic compound (e.g., herpes simplex thymidine kinase, cytosine deaminase); (12) nucleic acid segments that inhibit the replication, segregation, or heritability of a nucleic acid molecule that contains such nucleic acid segments; (13) nucleic acid segments encoding a conditional replication function (e.g., replication in certain hosts or host cell lines or under certain environmental conditions (e.g., temperature, nutrient conditions, etc.)); and / or (14) nucleic acid encoding one or more mRNAs or tRNAs that contain non-natural nucleotides. In some embodiments, a modulating or selecting agent may be added to alter the existing growth conditions to which the organism is subjected (e.g., growing in liquid culture, growing in a fermenter, growing on a solid nutrient plate, etc.).
[0561] In some embodiments, the regulation of promoter elements can be used to alter (e.g., increase, add, decrease, or substantially eliminate) the activity of a peptide, polypeptide, or protein (e.g., enzymatic activity). For example, in certain embodiments, a microorganism can be engineered by genetic modification to express a nucleic acid reagent that can add novel activity (e.g., an activity not normally found in the host organism), or increase the expression of an existing activity by increasing transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest (e.g., a homologous or heterologous nucleotide sequence of interest). In some embodiments, in certain embodiments, a microorganism can be engineered by genetic modification to express a nucleic acid reagent that can decrease the expression of an activity by decreasing or substantially eliminating transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest.
[0562] A nucleic acid encoding a heterologous protein (e.g., a nucleotide triphosphate transporter) can be inserted or used in any suitable expression system. In some embodiments, in certain embodiments, the nucleic acid reagent is sometimes stably integrated into the chromosome of the host organism, or the nucleic acid reagent can be a deletion of a portion of the host chromosome (e.g., a genetically modified organism where an alteration of the host genome confers the ability to selectively or preferentially maintain the desired organism carrying the genetic modification). Such nucleic acid reagents (e.g., a nucleic acid or a genetically modified organism whose altered genome confers an optional trait on the organism) can be selected for their ability to direct the production of a desired protein or nucleic acid molecule. When desired, the nucleic acid reagent can be altered such that the codons encode: (i) the same amino acid, using a different tRNA than that specified in the native sequence, or (ii) an amino acid different from normal, including non-conventional or non-natural amino acids (including detectably labeled amino acids).
[0563] Recombinant expression is effectively accomplished using an expression cassette that can be part of a vector such as a plasmid. The vector can include a promoter operably linked to the nucleic acid encoding the nucleotide triphosphate transporter. The vector can also include other elements required for transcription and translation as described herein. The expression cassette, expression vector, and the sequences in the cassette or vector can be heterologous to the cell that contacts the non-natural nucleotide. For example, the nucleotide triphosphate transporter sequence can be heterologous to the cell.
[0564] A variety of prokaryotic and eukaryotic expression vectors can be generated that are suitable for carrying, encoding, and / or expressing nucleotide triphosphate transporters. Such expression vectors include, for example, pET, pET3d, pCR2.1, pBAD, pUC, and yeast vectors. The vectors can be used, for example, in a variety of in vivo and in vitro situations. Non-limiting examples of prokaryotic promoters that can be used include the SP6, T7, T5, tac, bla, trp, gal, lac, or maltose promoters. Non-limiting examples of eukaryotic promoters that can be used include constitutive promoters, such as viral promoters, e.g., CMV, SV40, and RSV promoters; and regulatable promoters, such as inducible or repressible promoters, e.g., the tet promoter, the hsp70 promoter, and synthetic promoters regulated by CRE. Vectors for bacterial expression include pGEX-5X-3, and vectors for eukaryotic expression include pCIneo-CMV. Viral vectors that can be employed include those related to: lentiviruses, adenoviruses, adeno-associated viruses, herpesviruses, vaccinia viruses, polioviruses, AIDS viruses, neurotrophic viruses, Sindbis viruses, and other viruses. Also useful are any viral families that share the characteristics of these viruses and are thus suitable for use as vectors. Retroviral vectors that can be employed include those described in: Verma, American Society for Microbiology, pages 229-232, Washington, (1985). For example, such retroviral vectors can include Moloney murine leukemia virus, MMLV, and other retroviruses that express the desired characteristics. Generally, viral vectors contain non-structural early genes, structural late genes, RNA polymerase III transcripts, inverted terminal repeats required for replication and encapsidation, and promoters that control the transcription and replication of the viral genome. When engineered as vectors, the virus typically has one or more early genes removed, and a gene or gene / promoter cassette is inserted into the viral genome in place of the removed viral nucleic acid.
[0565] Clone
[0566] Components such as ORFs can be incorporated into nucleic acid reagents using any convenient cloning strategy known in the art. Components can be inserted into a template that is unrelated to the inserted component using known methods such as: (1) cutting the template at one or more existing restriction enzyme sites and ligating the component of interest, and (2) adding a restriction enzyme site to the template by hybridizing an oligonucleotide primer that includes one or more appropriate restriction enzyme sites and amplifying by polymerase chain reaction (described in more detail herein). Other cloning strategies utilize one or more insertion sites present in or inserted into the nucleic acid reagent, such as, for example, oligonucleotide primer hybridization sites for PCR, and other sites described herein. In some embodiments, the cloning strategy can be combined with genetic manipulation such as recombination (e.g., recombining a nucleic acid reagent having a nucleic acid sequence of interest into the genome of an organism to be modified, as further described herein). In some embodiments, one or more cloned ORFs can be produced by engineering a microorganism with one or more ORFs of interest to produce (directly or indirectly) a modified or wild-type nucleotide triphosphate transporter and / or polymerase, the microorganism having an activity that alters the nucleotide triphosphate transporter activity or polymerase activity.
[0567] The nucleic acid can be specifically cleaved by contacting the nucleic acid with one or more specific cleavage agents. Specific cleavage agents generally will specifically cleave at specific sites according to specific nucleotide sequences. Examples of enzyme specific cleavage agents include, but are not limited to, endonucleases (e.g., DNases (e.g., DNase I, II); RNases (e.g., RNase E, F, H, P); Cleavase TMEnzymes; Taq DNA polymerase; Escherichia coli DNA polymerase I and eukaryotic structure-specific endonucleases; murine FEN-1 endonuclease; type I, II, or III restriction endonucleases, such as Acc I, Afl III, Alu I, Alw44 I, Apa I, AsnI, Ava I, Ava II, BamH I, Ban II, Bcl I, Bgl I, Bgl II, Bln I, BsaI, Bsm I, BsmBI, BssHII, BstE II, Cfo I, CIa I, Dde I, Dpn I, Dra I, EcIX I, EcoR I, EcoR I, EcoR II, EcoR V, Hae II, Hae II, Hind II, Hind III, Hpa I, Hpa II, Kpn I, Ksp I, Mlu I, MIuN I, Msp I, Nci I, Nco I, Nde I, Nde II, Nhe I, Not I, Nru I, Nsi I, Pst I, Pvu I, Pvu II, Rsa I, SacI, Sal I, Sau3A I, Sca I, ScrF I, Sfi I, Sma I, Spe I, Sph I, Ssp I, Stu I, Sty I, Swa I, Taq I, Xba I, Xho I); glycosylases (e.g., uracil-DNA glycosylase (UDG), 3-methyladenine DNA glycosylase, 3-methyladenine DNA glycosylase II, pyrimidine hydrate-DNA glycosylase, FaPy-DNA glycosylase, thymine mismatch-DNA glycosylase, hypoxanthine-DNA glycosylase, 5-hydroxymethyluracil DNA glycosylase (HmUDG), 5-hydroxymethylcytosine DNA glycosylase, or 1,N6-etheno-adenine DNA glycosylase); exonucleases (e.g., exonuclease III); ribozymes; and DNases. The sample nucleic acid can be treated with chemical agents or synthesized using modified nucleotides and the modified nucleic acid can be cleaved. In non-limiting examples, the sample nucleic acid can be treated with: (i) an alkylating agent, such as methyl nitrosourea, which generates several alkylated bases, including N3-methyladenine and N3-methylguanine, which are recognized and cleaved by alkylpurine DNA-glycosylase; (ii) sodium bisulfite, which causes deamination of cytosine residues in DNA to form uracil residues, which can be cleaved by uracil N-glycosylase; and (iii) a chemical agent that converts guanine to its oxidized form 8-hydroxyguanine, which can be cleaved by formamidopyrimidine DNA N-glycosylase.Examples of chemical cleavage processes include, but are not limited to, alkylation (e.g., alkylation of phosphorothioate-modified nucleic acids); acid-labile cleavage of nucleic acids containing P3'-N5'-aminophosphates; and osmium tetroxide and piperidine treatment of nucleic acids.
[0568] In some embodiments, the nucleic acid reagent includes one or more recombinase insertion sites. A recombinase insertion site is a recognition sequence on a nucleic acid molecule that participates in an integration / recombination reaction of a recombinase protein. For example, the recombination site for Cre recombinase is loxP, which is a 34 base pair sequence consisting of two 13 base pair inverted repeats (serving as recombinase binding sites) flanking an 8 base pair core sequence (e.g., Sauer, Curr. Opin. Biotech. 5:521-527 (1994)). Other examples of recombination sites include attB, attP, attL, and attR sequences and their mutants, fragments, variants, and derivatives, which are recognized by the recombinase protein λInt and by the accessory proteins integration host factor (IHF), FIS, and excisionase (Xis) (e.g., U.S. Patent Nos. 5,888,732; 6,143,557; 6,171,861; 6,270,969; 6,277,608; and 6,720,140; U.S. Patent Application Nos. 09 / 517,466 and 09 / 732,914; U.S. Patent Publication No. US2002 / 0007051; and Landy, Curr. Opin. Biotech. 3:699-707 (1993); the disclosures of each of which are hereby incorporated by reference in their entirety).
[0569] Examples of recombinases for cloning nucleic acids are in the system (Invitrogen, California), which system includes at least one recombination site for cloning a desired nucleic acid molecule in vivo or in vitro. In some embodiments, the system utilizes a vector containing at least two different site-specific recombination sites, which are generally based on the bacteriophage λ system (e.g., att1 and att2) and are mutated from the wild-type (att0) site. Each mutated site has a unique specificity for its homologous partner att site of the same type (i.e., its binding partner recombination site) (e.g., attB1 for attP1, or attL1 for attR1), and does not cross-react with other mutated types of recombination sites or with the wild-type att0 site. The different site specificities allow for the directional cloning or ligation of the desired molecule, thereby providing the desired orientation of the cloned molecule. Using The system clones and subclones nucleic acid fragments flanked by recombination sites by replacing an optional marker (e.g., ccdB) flanked by att sites on the recipient plasmid molecule, which recipient plasmid molecule is sometimes referred to as the Destination Vector. The desired clone is then selected by transformation of a ccdB-sensitive host strain and positive selection for the marker on the recipient molecule. Similar strategies for negative selection (e.g., using a toxic gene) can be used in other organisms, such as thymidine kinase (TK) in mammals and insects.
[0570] Nucleic acid reagents sometimes contain one or more origin of replication (ORI) elements. In some embodiments, the template contains two or more ORIs, where one ORI functions efficiently in one organism (e.g., bacteria), and the other ORI functions efficiently in another organism (e.g., eukaryotes, such as yeast for example). In some embodiments, one ORI can function efficiently in one species (e.g., Saccharomyces cerevisiae), and another ORI can function efficiently in a different species (e.g., Schizosaccharomyces pombe). Nucleic acid reagents sometimes also include one or more transcriptional regulatory sites.
[0571] Nucleic acid reagents (e.g., expression cassettes or vectors) can include nucleic acid sequences encoding a marker product. The marker product is used to determine whether a gene has been delivered to a cell, and, once delivered, whether the gene is expressed. Examples of marker genes include the Escherichia coli lacZ gene encoding β-galactosidase and green fluorescent protein. In some embodiments, the marker can be an optional marker. When such an optional marker is successfully transferred into a host cell, the transformed host cell can survive when placed under a selection pressure. There are two different broad classes of selection schemes. The first class is based on the metabolism of the cell and the use of mutant cell lines that lack the ability to grow in medium without supplementation. The second class is dominant selection, which refers to selection schemes that can be used for any cell type and do not require the use of mutant cell lines. These schemes typically use drugs to block the growth of the host cell. Those cells with the novel gene will express a protein conferring drug resistance and will survive the selection. Examples of such dominant selection use the following drugs: neomycin (Southern et al., J. Molec. Appl. Genet. 1:327 (1982)), mycophenolic acid (Mulligan et al., Science 209:1422 (1980)), or hygromycin (Sugden, et al., Mol. Cell. Biol. 5:410 - 413 (1985); the disclosures of each of the said references are hereby incorporated by reference in their entirety).
[0572] The nucleic acid reagent can include one or more selection elements (e.g., an element for selecting the presence of the nucleic acid reagent and not for activating a promoter element that can be selectively regulated). Selection elements are typically used in known processes to determine whether a nucleic acid reagent is included in a cell. In some embodiments, the nucleic acid reagent includes two or more selection elements, where one selection element functions efficiently in one organism and another selection element functions efficiently in another organism. Examples of selection elements include, but are not limited to: (1) a nucleic acid segment encoding a product that confers resistance to a compound that is otherwise toxic (e.g., an antibiotic); (2) a nucleic acid segment encoding a product that is otherwise absent in the recipient cell (e.g., an essential product, a tRNA gene, a auxotrophic marker); (3) a nucleic acid segment encoding a product that inhibits the activity of a gene product; (4) a nucleic acid segment encoding a product that may be readily identifiable (e.g., a phenotypic marker such as an antibiotic (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and a cell surface protein); (5) a nucleic acid segment that binds to a product that is otherwise harmful to cell survival and / or function; (6) a nucleic acid segment that otherwise inhibits the activity of any of the nucleic acid segments described in items 1-5 above (e.g., an antisense oligonucleotide); (7) a nucleic acid segment that binds to a product that modifies a substrate (e.g., a restriction endonuclease); (8) a nucleic acid segment that can be used to isolate or identify a desired molecule (e.g., a specific protein binding site); (9) a nucleic acid segment encoding a specific nucleotide sequence that may otherwise be non-functional (e.g., for PCR amplification of a subpopulation of molecules); (10) a nucleic acid segment that directly or indirectly confers resistance or sensitivity to a specific compound in the absence thereof; (11) a nucleic acid segment encoding a product that is toxic in the recipient cell or converts a relatively non-toxic compound into a toxic compound (e.g., herpes simplex thymidine kinase, cytosine deaminase); (12) a nucleic acid segment that inhibits the replication, segregation, or heritability of a nucleic acid molecule that contains the nucleic acid segment; and / or (13) a nucleic acid segment encoding a conditional replication function (e.g., replication in certain hosts or host cell lines or under certain environmental conditions (e.g., temperature, nutritional conditions, etc.)).
[0573] The nucleic acid reagent can be in any form for in vivo transcription and / or translation. The nucleic acid is sometimes a plasmid such as a supercoiled plasmid, sometimes a yeast artificial chromosome (e.g., YAC), sometimes a linear nucleic acid (e.g., a linear nucleic acid generated by PCR or by restriction digestion), sometimes single-stranded and sometimes double-stranded. The nucleic acid reagent is sometimes prepared by an amplification process such as the polymerase chain reaction (PCR) process or the transcription-mediated amplification process (TMA). In TMA, two enzymes are used in an isothermal reaction to generate an amplification product detected by light emission (e.g., Biochemistry, June 25, 1996; 35(25):8429-38). Standard PCR processes are known (e.g., U.S. Patent Nos. 4,683,202; 4,683,195; 4,965,188; and 5,656,493), and are typically carried out in cycles. Each cycle includes heat denaturation, where hybrid nucleic acids dissociate; cooling, where primer oligonucleotides hybridize; and extension of the oligonucleotides by a polymerase (i.e., Taq polymerase). An example of a PCR cycling process is to treat the sample at 95°C for 5 minutes; repeat forty-five cycles of 95°C for 1 minute, 59°C for 1 minute 10 seconds, and 72°C for 1 minute 30 seconds; and then treat the sample at 72°C for 5 minutes. Multiple cycles are typically carried out using a commercially available thermal cycler. Sometimes the PCR amplification product is stored at a lower temperature (e.g., at 4°C) for a period of time and sometimes it is frozen (e.g., at -20°C) before analysis.
[0574] Cloning strategies similar to those described above can be employed to generate DNA containing unnatural nucleotides. For example, oligonucleotides containing unnatural nucleotides at desired positions are synthesized using standard solid-phase synthesis methods and purified by HPLC. The oligonucleotides are then inserted into a plasmid having a cloning site such as a BsaI site (but other sites discussed above can be used) containing the desired sequence context (i.e., UTR and coding sequence) using cloning methods such as Golden Gate Assembly.
[0575] Kit / article
[0576] In certain embodiments, kits and articles for use with one or more of the methods described herein are disclosed. Such kits include a carrier, package, or container that is compartmentalized to receive one or more containers such as vials, tubes, etc., each of the one or more containers containing one of the individual elements to be used in the methods described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. In one embodiment, the containers are formed from a variety of materials such as glass or plastic.
[0577] In some embodiments, the kit includes suitable packaging materials to contain the contents of the kit. In some cases, the packaging materials are constructed by well-known methods, preferably to provide a sterile and contamination-free environment. The packaging materials used herein may include, for example, those commonly used in commercially available kits sold for use with nucleic acid sequencing systems. Exemplary packaging materials include, but are not limited to, glass, plastic, paper, foil, etc. that can hold the components described herein within fixed boundaries.
[0578] The packaging materials may include labels indicating the specific uses of the components. The uses of the kit indicated by the labels may be one or more of the methods described herein that are appropriate for a particular combination of components present in the kit. For example, the label may indicate that the kit is for use in a method for synthesizing polynucleotides or in a method for determining nucleic acid sequences.
[0579] The kit may also include instructions for using the packaged reagents or components. The instructions will generally include a tangible expression describing reaction parameters such as the relative amounts of kit components and samples to be mixed, the duration of maintenance of the reagent / sample mixture, temperature, buffer conditions, etc.
[0580] It will be understood that not all components required for a particular reaction must be present in a particular kit. Instead, one or more additional components may be provided from other sources. The instructions provided with the kit may identify one or more additional components to be provided and where the components can be obtained.
[0581] In some embodiments, a kit is provided that is for stably incorporating non-natural nucleic acids into cellular nucleic acids, e.g., using the methods for preparing genetically engineered cells provided by the present invention. In one embodiment, the kit described herein includes genetically engineered cells and one or more non-natural nucleic acids. In another embodiment, the kit described herein includes an isolated and purified plasmid that contains a sequence selected from SEQ ID NO: 1-2. In additional embodiments, the kit described herein includes primers that contain a sequence selected from SEQ ID NO: 3-20.
[0582] In additional embodiments, the kit described herein provides cells and a nucleic acid molecule containing a heterologous gene for introduction into the cells to thereby provide genetically engineered cells, such as an expression vector containing the nucleic acid of any of the embodiments described previously in this paragraph.
[0583] Exemplary Embodiments
[0584] The present disclosure is further described by the following embodiments. The features of each embodiment may be combined with any other embodiment where appropriate and practical.
[0585] Embodiment 1. A method for producing a protein containing unnatural amino acids in vivo, the method comprising:
[0586] Transcribing a DNA template containing a first unnatural base and a complementary second unnatural base to incorporate a third unnatural base into mRNA, the third unnatural base being configured to form a first unnatural base pair with the first unnatural base;
[0587] Transcribing the DNA template to incorporate a fourth unnatural base into tRNA, wherein the fourth unnatural base is configured to form a second unnatural base pair with the second unnatural base, wherein the first unnatural base pair and the second unnatural base pair are different; and
[0588] Translating a protein from the mRNA and the tRNA, wherein the protein contains unnatural amino acids.
[0589] Embodiment 2. The method according to Embodiment 1, wherein the in vivo method comprises using a semi-synthetic organism.
[0590] Embodiment 3. The method according to Embodiment 2, wherein the organism comprises a microorganism.
[0591] Embodiment 4. The method according to Embodiment 3, wherein the organism comprises a bacterium.
[0592] Embodiment 5. The method according to Embodiment 4, wherein the organism comprises a Gram-positive bacterium.
[0593] Embodiment 6. The method according to Embodiment 4, wherein the organism comprises a Gram-negative bacterium.
[0594] Embodiment 7. The method according to any one of Embodiments 2-4, wherein the organism comprises Escherichia coli.
[0595] Embodiment 8. The method according to any one of Embodiments 1-7, wherein at least one unnatural base is selected from
[0596] (i) 2-Thiouracil, 2-thiothymine, 2'-deoxyuridine, 4-thiouracil, 4-thiothymine, uracil-5-yl, hypoxanthin-9-yl (I), 5-halouracil; 5-propynyluracil, 6-azathymine, 6-azauracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, methyl uracil-5-oxyacetate, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, uracil-5-oxyacetic acid, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiothymidine, 5-carboxymethylaminomethyluracil or dihydrouracil;
[0597] (ii) 5-Hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytarabine, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azacytosine, azacytidine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one) or pyridoindole cytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one);
[0598] (iii) 2-Aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2'-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxy substituted adenines, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-aza-adenine;
[0599] (iv) 2-methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thioguanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxy substituted guanines, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-aza-guanine; and
[0600] (v) hypoxanthine, xanthine, 1-methylinosine, queosine, β-D-galactosyl queosine, inosine, β-D-mannosyl queosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone.
[0601] Embodiment 9. The method according to any one of Embodiments 1-7, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is selected from:
[0602]
[0603] Embodiment 10. The method according to any one of Embodiments 1-7, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is selected from:
[0604]
[0605] Embodiment 11. The method according to Embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is selected from:
[0606]
[0607] Embodiment 12. The method according to Embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is selected from:
[0608]
[0609] Embodiment 13. The method according to Embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base or the fourth unnatural base is selected from:
[0610]
[0611] Embodiment 14. The method according to embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0612]
[0613] Embodiment 15. The method according to embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0614]
[0615] Embodiment 16. The method according to embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from
[0616]
[0617] Embodiment 17. The method according to embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0618]
[0619] Embodiment 18. The method according to embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is
[0620] Embodiment 19. The method according to embodiment 9, wherein at least one of the first unnatural base, the second unnatural base, the third unnatural base, or the fourth unnatural base is selected from:
[0621]
[0622] Embodiment 20. The method according to embodiment 9, wherein the first unnatural base or the second unnatural base is
[0623] Embodiment 21. The method according to embodiment 9, wherein the first unnatural base or the second unnatural base is
[0624] Embodiment 22. The method according to embodiment 9, wherein the first unnatural base is and the second unnatural base is
[0625] Embodiment 23. The method according to Embodiment 9, wherein the first unnatural base is and the second unnatural base is
[0626] Embodiment 24. The method according to Embodiment 9, wherein the third unnatural base or the fourth unnatural base is
[0627] Embodiment 25. The method according to Embodiment 9, wherein the third unnatural base is
[0628] Embodiment 26. The method according to Embodiment 9, wherein the fourth unnatural base is
[0629] Embodiment 27. The method according to Embodiment 9, wherein the third unnatural base or the fourth unnatural base is
[0630] Embodiment 28. The method according to Embodiment 9, wherein the third unnatural base is
[0631] Embodiment 29. The method according to Embodiment 9, wherein the fourth unnatural base is
[0632] Embodiment 30. The method according to Embodiment 9, wherein the first unnatural base is the second unnatural base is the third unnatural base is and the fourth unnatural base is
[0633] Embodiment 31. The method according to Embodiment 9, wherein the first unnatural base is the second unnatural base is the third unnatural base is and the fourth unnatural base is
[0634] Embodiment 32. The method according to Embodiment 9, wherein the first unnatural base is the second unnatural base is the third unnatural base is and the fourth unnatural base is
[0635] Embodiment 33. The method according to embodiment 9, wherein the third unnatural base is
[0636] Embodiment 34. The method according to embodiment 9, wherein the fourth unnatural base is
[0637] Embodiment 35. The method according to embodiment 9, wherein the first unnatural base is The second unnatural base is The third unnatural base is And the fourth unnatural base is
[0638] Embodiment 36. The method according to any one of embodiments 9 to 35, wherein the third unnatural base and the fourth unnatural base comprise ribose.
[0639] Embodiment 37. The method according to any one of embodiments 9 to 35, wherein the third unnatural base and the fourth unnatural base comprise deoxyribose.
[0640] Embodiment 38. The method according to any one of embodiments 9 to 35, wherein the first unnatural base and the second unnatural base comprise deoxyribose.
[0641] Embodiment 39. The method according to any one of embodiments 9 to 35, wherein the first unnatural base and the second unnatural base comprise deoxyribose, and the third unnatural base and the fourth unnatural base comprise ribose.
[0642] Embodiment 40. The method according to embodiment 9, wherein the DNA template comprises at least one unnatural base pair (UBP) selected from the following:
[0643]
[0644]
[0645] Embodiment 41. The method according to embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) of dNaM-d5SICS.
[0646] Embodiment 42. The method according to embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) of dCNMO-dTPT3.
[0647] Embodiment 43. The method according to embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3.
[0648] Embodiment 44. The method according to embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dPTMO-dTPT3.
[0649] Embodiment 45. The method according to embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1.
[0650] Embodiment 46. The method according to embodiment 40, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.
[0651] Embodiment 47. The method according to embodiment 1, wherein the DNA template comprises at least one unnatural base pair (UBP) selected from the following:
[0652]
[0653] And
[0654] wherein the mRNA and the tRNA comprise at least one unnatural base selected from the following:
[0655]
[0656] Embodiment 48. The method according to embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-d5SICS.
[0657] Embodiment 49. The method according to embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTPT3.
[0658] Embodiment 50. The method according to embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTPT3.
[0659] Embodiment 51. The method according to embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dPTMO-dTPT3.
[0660] Embodiment 52. The method according to embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dNaM-dTAT1.
[0661] Embodiment 53. The method according to embodiment 47, wherein the DNA template comprises at least one unnatural base pair (UBP) that is dCNMO-dTAT1.
[0662] Embodiment 54. The method according to any one of embodiments 47 to 53, wherein the mRNA and the tRNA comprise unnatural bases selected from of unnatural bases.
[0663] Embodiment 55. The method according to embodiment 54, wherein the mRNA and the tRNA comprise unnatural bases selected from of unnatural bases.
[0664] Embodiment 56. The method according to embodiment 54, wherein the mRNA comprises an unnatural base that is of unnatural bases.
[0665] Embodiment 57. The method according to embodiment 54, wherein the mRNA comprises an unnatural base that is of unnatural bases.
[0666] Embodiment 58. The method according to embodiment 54, wherein the mRNA comprises an unnatural base that is of unnatural bases.
[0667] Embodiment 59. The method according to embodiment 54, wherein the tRNA comprises unnatural bases selected from of unnatural bases.
[0668] Embodiment 60. The method according to embodiment 54, wherein the tRNA comprises an unnatural base that is of unnatural bases.
[0669] Embodiment 61. The method according to embodiment 54, wherein the tRNA comprises an unnatural base that is of unnatural bases.
[0670] Embodiment 62. The method according to embodiment 54, wherein the tRNA comprises an unnatural base that is of unnatural bases.
[0671] Embodiment 63. The method according to any one of embodiments 1 - 7, wherein the first unnatural base comprises dCNMO, and the second unnatural base comprises dTPT3.
[0672] Embodiment 64. The method according to any one of Embodiments 1-7 or 41, wherein the third unnatural base comprises NaM, and the second unnatural base comprises TAT1.
[0673] Embodiment 65. The method according to any one of Embodiments 1-64, wherein the first unnatural base or the second unnatural base is recognized by a DNA polymerase.
[0674] Embodiment 66. The method according to any one of Embodiments 1-65, wherein the third unnatural base or the fourth unnatural base is recognized by an RNA polymerase.
[0675] Embodiment 67. The method according to any one of Embodiments 1-66, wherein the protein comprises at least two unnatural amino acids.
[0676] Embodiment 68. The method according to any one of Embodiments 1-66, wherein the protein comprises at least three unnatural amino acids.
[0677] Embodiment 69. The method according to any one of Embodiments 1-66, wherein the protein comprises at least two different unnatural amino acids.
[0678] Embodiment 70. The method according to any one of Embodiments 1-66, wherein the protein comprises at least three different unnatural amino acids.
[0679] Embodiment 71. The method according to any one of Embodiments 1-70, wherein the at least one unnatural amino acid:
[0680] is a lysine analogue;
[0681] comprises an aromatic side chain;
[0682] comprises an azide group;
[0683] comprises an alkynyl group; or
[0684] comprises an aldehyde group or a ketone group.
[0685] Embodiment 72. The method according to any one of Embodiments 1-70, wherein the at least one unnatural amino acid does not comprise an aromatic side chain.
[0686] Embodiment 73. The method of any one of Embodiments 1-70, wherein the at least one unnatural amino acid comprises N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propynylethoxy-carbonyl-L-lysine (PraK), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propynyloxy phenylalanine, p-propynyl-phenylalanine, 3-methyl-phenylalanine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)seleno)propanoic acid, 2-amino-3-(phenylseleno)propanoic acid, or selenocysteine.
[0687] Embodiment 74. The method according to Embodiment 73, wherein the at least one unnatural amino acid comprises N6-azidoethoxy-carbonyl-L-lysine (AzK) and N6-propynylethoxy-carbonyl-L-lysine (PraK).
[0688] Embodiment 75. The method according to Embodiment 72, wherein the at least one unnatural amino acid comprises N6-azidoethoxy-carbonyl-L-lysine (AzK).
[0689] Embodiment 76. The method according to Embodiment 72, wherein the at least one unnatural amino acid comprises N6-propynylethoxy-carbonyl-L-lysine (PraK).
[0690] Embodiment 77. A semi-synthetic organism, the semi-synthetic organism comprising an expanded genetic alphabet, wherein the genetic alphabet comprises at least two unique unnatural bases.
[0691] Embodiment 78. The semi-synthetic organism according to Embodiment 77, wherein the organism comprises a microorganism.
[0692] Embodiment 79. The semi-synthetic organism according to any one of embodiments 77-78, wherein the organism comprises bacteria.
[0693] Embodiment 80. The semi-synthetic organism according to embodiment 79, wherein the organism comprises Gram-positive bacteria.
[0694] Embodiment 81. The semi-synthetic organism according to embodiment 79, wherein the organism comprises Gram-positive bacteria.
[0695] Embodiment 82. The semi-synthetic organism according to any one of embodiments 77-79, wherein the organism comprises Escherichia coli.
[0696] Embodiment 83. The semi-synthetic organism according to any one of Embodiments 77-82, wherein at least one of the unnatural bases is selected from: 2-aminoadenin-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-amino-adenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine 3-deazaguanine, 3-deazaadenine, 4-thiouracil, 4-thiothymine, uracil-5-yl, hypoxanthin-9-yl (I), 5-methyl-cytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo and 5-trifluoromethyluracil and cytosine; 5-halouracil, 5-halocytosine, 5-propynyl-uracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halo, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorocytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitro cytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azo-uracil, 6-azocytosine, azacytosine, 6-azo-thymine, 6-thioguanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxy substituted adenine and guanine;N4-ethylcytosine, N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, those that increase the stability of duplex formation, general nucleic acids, hydrophobic nucleic acids, chimeric nucleic acids, size-expanded nucleic acids, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), G-clamp, phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d]pyrimidin-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylqueuosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylqueuosine, 5’-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, wybutoxosine, pseudouracil, queuosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxaacetic acid methyl ester, uracil-5-oxaacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w and 2,6-diaminopurine and those in which the purine or pyrimidine base is replaced by a heterocycle.;
[0697] Embodiment 84. The semi-synthetic organism according to any one of Embodiments 77-82, wherein the organism comprises DNA containing at least one unnatural nucleobase selected from:
[0698]
[0699] Embodiment 85. The semi-synthetic organism according to any one of Embodiments 77-84, wherein the DNA containing at least one of the unnatural bases forms an unnatural base pair (UBP).
[0700] Embodiment 86. The semi-synthetic organism according to Embodiment 85, wherein the unnatural base pair (UBP) is dCNMO-dTPT3, dNaM-dTPT3, dCNMO-dTAT1 or dNaM-dTAT1.
[0701] Embodiment 87. The semi-synthetic organism according to Embodiment 84, wherein the DNA comprises at least one unnatural nucleobase selected from:
[0702]
[0703] Embodiment 88. The semi-synthetic organism according to Embodiment 87, wherein the DNA comprises at least one unnatural nucleobase selected from:
[0704]
[0705] Embodiment 89. The semi-synthetic organism according to Embodiment 88, wherein the DNA comprises at least one unnatural nucleobase selected from:
[0706]
[0707] Embodiment 90. The semi-synthetic organism according to Embodiment 88, wherein the DNA comprises at least one unnatural nucleobase selected from:
[0708]
[0709] Embodiment 91. The semi-synthetic organism according to Embodiment 88, wherein the DNA comprises at least one unnatural nucleobase selected from:
[0710]
[0711] Embodiment 92. The semi-synthetic organism according to Embodiment 88, wherein ...
Claims
1. A nucleobase having the following structure: wherein the wavy line represents the attachment of the nucleobase to a ribosyl, deoxyribosyl, or dideoxyribosyl moiety, and wherein the ribosyl, deoxyribosyl, or dideoxyribosyl moiety is in free form, linked to a monophosphate, diphosphate, triphosphate, α-thiotriphosphate, β-thiotriphosphate, or γ-thiotriphosphate group, or is incorporated in a polynucleotide.
2. The nucleobase according to claim 1, wherein the nucleobase binds to a complementary base-pairing nucleobase to form an unnatural base pair (UBP).
3. The nucleobase according to claim 2, wherein the complementary base-pairing nucleobase is selected from:
4. A double-stranded oligonucleotide comprising a first oligonucleotide strand that comprises the nucleobase according to claim 1, and a second oligonucleotide strand that is complementary to the first oligonucleotide strand, wherein the second oligonucleotide strand comprises a complementary base-pairing nucleobase at its complementary base-pairing site.
5. The double-stranded oligonucleotide according to claim 4, wherein the second oligonucleotide strand comprises a complementary base-pairing nucleobase selected from the following at its complementary base-pairing site:
6. The double-stranded oligonucleotide according to claim 5, wherein the second oligonucleotide strand comprises complementary base-pairing nucleobases 7. The double-stranded oligonucleotide according to claim 5, wherein the second oligonucleotide strand comprises complementary base-pairing nucleobases 8. A plasmid comprising a gene encoding a transfer RNA (tRNA) and / or a gene encoding a protein of interest, wherein the gene comprises at least one nucleobase according to claim 1 and at least one complementary base-pairing nucleobase according to claim 3, and wherein the complementary base-pairing nucleobase is at a complementary base-pairing site.
9. A tRNA that comprises the nucleobase according to claim 1.
10. An mRNA that comprises the nucleobase according to claim 1, or comprises a codon that contains the nucleobase according to claim 1.
11. A transfer RNA (tRNA) that comprises: an anticodon that comprises the nucleobase according to claim 1, and wherein the anticodon pairs with an unnatural codon encoding an unnatural amino acid; and an identification element that promotes selective loading of the unnatural amino acid by the tRNA by an aminoacyl-tRNA synthetase.
12. The tRNA according to claim 11, wherein the aminoacyl-tRNA synthetase is derived from the genus Methanosarcina, or Methanocaldococcus.
13. The tRNA according to claim 11 or 12, wherein the unnatural amino acid comprises an aromatic moiety.
14. The tRNA according to claim 11 or 12, wherein the unnatural amino acid is a lysine derivative selected from: or a phenylalanine derivative selected from:
15. A nucleic acid molecule having the following structure: N1-Zx-N2 wherein: each Z is independently a nucleobase according to claim 1, which is bonded to a ribosyl or deoxyribosyl; N1 is one or more nucleotides or a terminal phosphate group attached at the 5'-end of the ribosyl or deoxyribosyl of Z; N2 is one or more nucleotide or terminal hydroxyl groups attached to the 3'-end of the ribosyl or deoxyribosyl of Z; and x is an integer from 1 to 20.
16. The nucleic acid molecule according to claim 15, wherein the structure encodes a gene.
17. The nucleic acid molecule according to claim 16, wherein Zx is located in the translation region of the gene.
18. The nucleic acid molecule according to claim 16, wherein Zx is located in the untranslated region of the gene.
19. A polynucleotide library, wherein the library comprises at least 5000 different polynucleotides, and wherein each polynucleotide comprises at least one nucleobase according to claim 1.
20. A nucleoside triphosphate comprising a nucleobase, wherein the nucleobase is:
21. The nucleoside triphosphate according to claim 20, wherein the nucleoside comprises ribose or deoxyribose.
22. A DNA, said DNA comprising nucleobases having the structure and complementary base-pairing nucleobases having the structure .
23. A DNA, said DNA comprising nucleobases having a structure and complementary base-pairing nucleobases having a structure .
24. A semi-synthetic organism, the semi-synthetic organism comprising DNA, the DNA comprising at least three different unnatural bases, and wherein a first unnatural base of the at least three different unnatural bases comprises the nucleobase according to claim 1.
25. The semi-synthetic organism according to claim 24, wherein the organism comprises a microorganism.
26. The semi-synthetic organism according to claim 25, wherein the microorganism is Escherichia coli.
27. The semi-synthetic organism according to claim 24, wherein the semi-synthetic organism is a cell.
28. The semi-synthetic organism according to claim 24, 25, 26 or 27, wherein a second unnatural nucleobase of the at least three different unnatural bases is selected from:
29. The semi-synthetic organism according to claim 24, 25, 26 or 27, wherein the DNA comprises at least one unnatural base pair (UBP), wherein the unnatural base pair (UBP) is 30. The semi-synthetic organism according to claim 24, 25, 26 or 27, which further comprises a heterologous nucleoside triphosphate transporter.
31. The semi-synthetic organism according to claim 30, wherein the heterologous nucleoside triphosphate transporter is PtNTT2.
32. The semi-synthetic organism according to claim 24, 25, 26 or 27, which further comprises a heterologous tRNA synthetase.
33. The semi-synthetic organism according to claim 32, wherein the heterologous tRNA synthetase is Methanosarcina barkeri pyrrolysyl-tRNA synthetase (Mb PylRS).
34. The semi-synthetic organism according to claim 24, 25, 26 or 27, which further comprises a heterologous RNA polymerase.
35. The semi-synthetic organism according to claim 34, wherein the heterologous RNA polymerase is T7 RNAP.
36. The semi-synthetic organism according to claim 24, 25, 26 or 27, wherein the organism does not express a protein having a DNA recombination repair function.
37. The semi-synthetic organism according to claim 36, wherein the organism does not express RecA.
38. The semi-synthetic organism according to claim 24, 25, 26 or 27, wherein the semi-synthetic organism further comprises mRNA, and the mRNA comprises at least one unnatural base.
39. The semi-synthetic organism according to claim 38, wherein the mRNA comprises at least one unnatural base selected from 40. The semi-synthetic organism according to claim 24, 25, 26 or 27, wherein the semi-synthetic organism further comprises tRNA, and the tRNA comprises at least one unnatural base.
41. The semi-synthetic organism according to claim 40, wherein the tRNA comprises at least one unnatural base selected from 42. The nucleobase according to any one of claims 1-3, wherein the polynucleotide is RNA, DNA, peptide nucleic acid (PNA), locked nucleic acid (LNA) or a nucleic acid containing phosphorothioate.
43. The nucleobase according to any one of claims 1-3, wherein the polynucleotide is a bicyclic nucleic acid comprising a bridge between the 4' and 2' ribosyl ring atoms.
44. The nucleobase according to any one of claims 1-3, wherein the polynucleotide is a linked nucleic acid comprising nucleic acids linked together by internucleic linkages.
Citation Information
Patent Citations
Dinucleotide and oligonucleotide analogues
EP0614907A1
Dinucleotide analogues, intermediates therefor and oligonucleotides derived therefrom
EP0629633A2
System for positioning an intubation tube
US11571179B2
Atomizing device
US2001817A
Use of multiple recombination sites with unique specificity in recombinational cloning
US20020007051A1