Semi-synthetic organisms of eukaryotes
Eukaryotic semi-synthetic organisms using unnatural base pairs overcome the limitations of existing methods by directly incorporating non-standard amino acids into proteins, enabling the generation of unnatural polypeptides with improved properties.
Patent Information
- Application Number
- JP2022519674
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-30
- Filing Date
- 2020-09-29
- Publication Date
- 2025-06-16
- Estimated Expiration
- 2040-09-29
AI Technical Summary
Current methods for incorporating non-standard amino acids into proteins are limited by the need to compete with natural codon functions, and genome synthesis is impractical for large eukaryotic genomes.
The development of eukaryotic semi-synthetic organisms (SSOs) that utilize unnatural base pairs (UBPs) to translate mRNA containing unnatural codons, allowing for the direct incorporation of non-standard amino acids into proteins.
This approach enables the generation of unnatural polypeptides with improved properties, overcoming the limitations of existing methods by avoiding interference from natural codon functions and facilitating the production of novel therapeutics and materials.
Smart Images

Figure 0007692900000606 
Figure 0007692900000607 
Figure 0007692900000608
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 62 / 908,421, filed September 30, 2019.
[0002] Sequence Listing This application includes a sequence listing that was electronically submitted in ASCII format and is hereby incorporated by reference in its entirety. The ASCII copy was created on September 24, 2020, has the name 36271 - 810_601_SL.txt, and is 19,000 bytes in size.
[0003] Statement Regarding Federally Sponsored Research This invention was made with government support under Grant No. GM118178 awarded by the National Institutes of Health (NIH). The U.S. government has certain rights in this invention.
Background Art
[0004] All proteins produced to date in cells are encoded by a four-letter two-base pair genetic alphabet. This generally restricts the amino acids that can build proteins to the standard 20, the proteinogenic amino acids. This has enabled the diversity of life, but expanding to include non-standard amino acids (ncAAs), including those selected to provide desired activities, could potentially enable the creation of novel proteins with improved properties for applications ranging from materials to therapeutics. Efforts to incorporate ncAAs mainly rely on expanding the genetic alphabet by nonsense codon (UAG) or four-letter codon (quadruplet codon) suppression, but in these cases, ncAA incorporation must compete with the natural function of the codon. To avoid this limitation, efforts have focused on synthesizing genomes depleted of natural stop or rare codons, thus freeing them for reassignment to ncAAs. However, rare codons can play important roles in regulating translation and protein folding, and genome synthesis is generally impractical as a general strategy, especially when using large eukaryotic genomes.
[0005] An alternative approach relies on the use of unnatural base pairs (UBPs), which, in principle, enable the creation of a virtually unlimited number of novel and completely new codons that are not interfered with by natural functions from a practical perspective. By pursuing pharmaceutical chemical mimics, a family of UBPs, such as dNaM-dTPT3, which has been used as the basis for Escherichia coli (E. coli) semi-synthetic organisms (SSOs) (Figure 1B), has been developed. E. coli SSOs store UBPs in their genomes or plasmids, transcribe them into mRNA and tRNA, and translate proteins containing ncAAs using tRNAs charged with ncAAs by orthogonal synthetases. E. coli SSOs have important practical applications as they are currently being used to generate novel therapeutics.
[0006] The breadth of ncAAs and the resulting unnatural polypeptides that can be generated is at least in part determined by the SSOs used. To date, the use of UBPs such as dNAM-dTPT3 has not been shown in eukaryotic SSOs or systems. The proof-of-concept of the approach summarized herein in eukaryotic cells will enable the generation of unnatural polypeptides that could be useful for important practical applications such as a broader range of ncAAs and the generation of novel therapeutics. Summary of the Invention Means for Solving the Problems
[0007] In some embodiments, eukaryotic semi-synthetic organisms (SSOs) generated by exploring the translation of unnatural codons are provided herein. Protein production was characterized after direct, transient, triple transfection using messenger RNA (mRNA) containing an unnatural codon, transfer RNA (tRNA) containing a cognate unnatural codon, and DNA encoding a synthetase suitable for charging the tRNA with a non-standard amino acid (ncAA).
[0008] Aspects disclosed herein provide eukaryotic cells comprising (a) a messenger RNA (mRNA) having a codon comprising a first unnatural base and (b) a transfer RNA (tRNA) having an anticodon comprising a second unnatural base, wherein the first and second unnatural bases form an unnatural base pair (UBP) in the eukaryotic cell and the mRNA is translated in the cell to produce a polypeptide comprising at least one unnatural amino acid. In some embodiments, the tRNA is charged with an unnatural amino acid. In some embodiments, the eukaryotic cell further comprises a polypeptide translated from the mRNA, the polypeptide comprising at least one unnatural amino acid. In some embodiments, the eukaryotic cell further comprises ribosomes capable of translating, using the tRNA, an mRNA into a polypeptide comprising at least one unnatural amino acid.
[0009] Aspects disclosed herein also provide eukaryotic cells comprising (a) a first unnatural ribonucleotide comprising a first unnatural base; and (b) a second unnatural ribonucleotide comprising a second unnatural base, wherein the first and second unnatural bases comprise an unnatural base pair (UBP) that forms an unnatural base pair (UBP) in the eukaryotic cell.
[0010] In some embodiments, the first unnatural base or the second unnatural base is (i) 2-thiouracil, 2-thio-thymine, 2'-deoxyuridine, 4-thio-uracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl (I), 5-halouracil, 5-propynyl-uracil, 6-azothymine, 6-azouracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, methyl ester of uracil-5-oxoacetic acid, uracil-5-oxoacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, uracil-5-oxoacetic acid, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil or dihydrouracil; (ii) 5-hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytosine arabinoside, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azocytosine, azacytosine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one) or pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d]pyrimidin-2-one);(iii) 2-Aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2'-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted adenine, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-aza-adenine; (iv) 2-methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thio-guanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted guanine, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-aza-guanine; and (v) hypoxanthine, xanthine, 1-methylinosine, queuosine, beta-D-galactosylqueuosine, inosine, beta-D-mannosylqueuosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone, selected from the group consisting of. In some embodiments, the first unnatural base and the second unnatural base are each independently;
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
[0011] In some embodiments, the eukaryotic cell further comprises: (a) a transfer RNA (tRNA) having an anticodon that includes a first unnatural base; and (b) a messenger RNA (mRNA) having a codon that includes a second unnatural base, wherein the first and second unnatural bases are capable of forming an unnatural base pair (UBP) in the eukaryotic cell. In some embodiments, the eukaryotic cell further comprises: (a) a transfer RNA (tRNA) having an anticodon that includes a second unnatural base; and (b) a messenger RNA (mRNA) having a codon that includes a first unnatural base, wherein the first and second unnatural bases are capable of forming an unnatural base pair (UBP) in the eukaryotic cell. In some embodiments, the codon of the mRNA includes three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the first position (X-N-N) in the codon of the mRNA. In some embodiments, the codon of the mRNA includes three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the middle position (N-X-N) in the codon of the mRNA. In some embodiments, the codon of the mRNA includes three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the last position (N-N-X) in the codon of the mRNA. In some embodiments, the eukaryote further comprises a polypeptide translated from the mRNA, and the polypeptide includes at least one unnatural amino acid. In some embodiments, the at least one unnatural amino acid is: (a) a lysine analog; (b) includes an aromatic side chain; (c) includes an azide group; (d) includes an alkyne group; or (e) includes an aldehyde or ketone group.In some embodiments, one or more non-natural amino acids are selected from the group consisting of N6-((azidoethoxy)-carbonyl)-L-lysine (AzK), N6-((propargylethoxy)-carbonyl)-L-lysine (PraK), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine. In some embodiments, at least one non-natural amino acid is N6-((azidoethoxy)-carbonyl)-L-lysine (AzK). In some embodiments, at least one non-natural amino acid is N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine. In some embodiments, at least one non-natural amino acid is N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine.In some embodiments, at least one non-natural amino acid is N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine. In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the human cell is a HEK293T cell. In some embodiments, the cell is a hamster cell. In some embodiments, the hamster cell is a Chinese hamster ovary (CHO) cell. In some embodiments, the cell is isolated and purified. In some embodiments, the mRNA and tRNA are stabilized to degrade in eukaryotic cells.
[0012] Aspects disclosed herein provide a semi-synthetic organism comprising a eukaryotic cell described herein.
[0013] Aspects disclosed herein provide a eukaryotic cell line comprising a plurality of eukaryotic cells of the present disclosure.
[0014] Aspects disclosed herein are methods for producing a polypeptide comprising one or more non-natural amino acids in a eukaryotic cell, comprising: (a) introducing into the cell (i) a messenger RNA (mRNA) having a codon comprising a first unnatural base and (ii) a transfer RNA (tRNA) having an anticodon comprising a second unnatural base in the eukaryotic cell, wherein the first and second unnatural bases form an unnatural base pair (UBP) in the eukaryotic cell; and (b) translating, using the tRNA, the mRNA into a polypeptide comprising one or more non-natural amino acids. In some embodiments, the tRNA is charged with a non-natural amino acid.
[0015] Aspects disclosed herein also provide a method for producing a polypeptide comprising one or more unnatural amino acids in a eukaryotic cell, the method comprising: (a) providing a eukaryotic cell comprising: (i) a messenger RNA (mRNA) having a codon comprising a first unnatural base; and (ii) a transfer RNA (tRNA) having an anticodon comprising a second unnatural base, wherein the first and second unnatural bases form an unnatural base pair (UBP) in the eukaryotic cell; and (b) translating, using the tRNA by a ribosome endogenous to the eukaryotic cell, an mRNA into a polypeptide comprising one or more unnatural amino acids. In some embodiments, the polypeptide comprises a eukaryotic glycosylation pattern. The glycosylation pattern can correspond to the cell in which it is produced (e.g., mammalian glycosylation pattern if the cell is mammalian, human glycosylation pattern if the cell is human, etc.).
[0016] Aspects disclosed herein also provide a method for producing a polypeptide in a eukaryotic cell, the polypeptide comprising one or more unnatural amino acids, the method comprising: (a) providing a eukaryotic cell comprising: (i) an mRNA comprising a codon comprising a first unnatural base; (ii) a tRNA comprising an anticodon comprising a second unnatural base, wherein the first and second unnatural bases form a complementary base pair; and (iii) a tRNA synthetase that preferentially aminoacylates the tRNA with one or more unnatural amino acids as compared to natural amino acids; and (b) providing one or more unnatural amino acids to the eukaryotic cell, whereby the eukaryotic cell produces a polypeptide comprising one or more unnatural amino acids.
[0017] Aspects disclosed herein also provide a method for producing a polypeptide comprising one or more non-natural amino acids in a eukaryotic cell, the method comprising: (a) providing a eukaryotic cell comprising: (i) a transfer RNA (tRNA) having an anticodon comprising a first unnatural base; and (ii) a messenger RNA (mRNA) having a codon comprising a second unnatural base, wherein the first and second unnatural bases form an unnatural base pair (UBP) in the eukaryotic cell; and (c) translating, using the tRNA by ribosomes endogenous to the eukaryotic cell, an mRNA into a polypeptide comprising one or more non-natural amino acids.
[0018] In some embodiments, the codon of the mRNA comprises three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the first position (X-N-N) in the codon of the mRNA. In some embodiments, the codon of the mRNA comprises three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the central position (N-X-N) in the codon of the mRNA. In some embodiments, the codon of the mRNA comprises three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the last position (N-N-X) in the codon of the mRNA. In some embodiments, the first unnatural base or the second unnatural base is (a) 2-thiouracil, 2-thio-thymine, 2'-deoxyuridine, 4-thio-uracil, 4-thio-thymine, uracil-5-yl, hypoxanthin-9-yl (I), 5-halouracil; 5-propynyl-uracil, 6-azothymine, 6-azouracil, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, pseudouracil, methyl ester of uracil-5-oxoacetic acid, uracil-5-oxoacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 5-methyl-2-thiouracil, 4-thiouracil, 5-methyluracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, uracil-5-oxoacetic acid, 5-(carboxyhydroxylmethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil or dihydrouracil;(b) 5-Hydroxymethylcytosine, 5-trifluoromethylcytosine, 5-halocytosine, 5-propynylcytosine, 5-hydroxycytosine, cyclocytosine, cytarabine, 5,6-dihydrocytosine, 5-nitrocytosine, 6-azacytosine, azacytidine, N4-ethylcytosine, 3-methylcytosine, 5-methylcytosine, 4-acetylcytosine, 2-thiocytosine, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), phenoxazine cytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one) or pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d]pyrimidin-2-one); (c) 2-aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2’-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted adenine, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-aza-adenine; (d) 2-methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thio-guanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl substituted guanine, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-aza-guanine;and (e) selected from the group consisting of hypoxanthine, xanthine, 1-methylinosine, queuosine, beta-D-galactosylqueuosine, inosine, beta-D-mannosylqueuosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone. In some embodiments, the first unnatural base or the second unnatural base is;
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
Chemical Structure
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
[0019] In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the human cell is a HEK293T cell. In some embodiments, the cell is a hamster cell. In some embodiments, the hamster cell is a Chinese hamster ovary (CHO) cell. In some embodiments, the unnatural amino acid is (a) a lysine analog; (b) contains an aromatic side chain; (c) contains an azide group; (d) contains an alkyne group; or (e) contains an aldehyde or ketone group. In some embodiments, the unnatural amino acid is N6-((azidoethoxy)-carbonyl)-L-lysine (AzK), N6-((propargylethoxy)-carbonyl)-L-lysine (PraK), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine, selected from the group consisting of.In some embodiments, the unnatural amino acid is N6-((azidoethoxy)-carbonyl)-L-lysine (AzK). In some embodiments, one or more unnatural amino acids are N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine. In some embodiments, one or more unnatural amino acids are N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine. In some embodiments, one or more unnatural amino acids are N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine.
[0020] Aspects disclosed herein are methods of producing a polypeptide in a eukaryotic cell, wherein the polypeptide comprises one or more unnatural amino acids, the method comprising: (a) providing a eukaryotic cell, the eukaryotic cell comprising: (i) an mRNA comprising a codon comprising one or more unnatural bases; (ii) a tRNA comprising an anticodon comprising one or more unnatural bases, wherein one or more unnatural bases comprising the codon in the mRNA and one or more unnatural bases comprising the anticodon in the tRNA form complementary base pairs; and (iii) a tRNA synthetase that preferentially aminoacylates the tRNA with one or more unnatural amino acids as compared to natural amino acids; and (b) providing one or more unnatural amino acids to the eukaryotic cell, whereby the eukaryotic cell produces a polypeptide comprising one or more unnatural amino acids. In some embodiments, the codon of the mRNA comprises three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the first position (X-N-N) in the codon of the mRNA. In some embodiments, the codon of the mRNA comprises three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the middle position (N-X-N) in the codon of the mRNA. In some embodiments, the codon of the mRNA comprises three consecutive nucleobases (N-N-N), and the first unnatural base (X) is located at the last position (N-N-X) in the codon of the mRNA. In some embodiments, one or more unnatural bases comprising the codon in the mRNA are of the formula [Chemistry] or [Chemistry] and is selected from the group consisting of hydrogen, alkyl, alkenyl, alkynyl, methoxy, methanethiol, methaneseleno, halogen, cyano and azide, and the wavy line indicates the bond to the ribosyl moiety. In some embodiments, the first unnatural base or the second unnatural base is
[0021] [Chemistry] selected from the group consisting of, and the wavy line indicates the bond to the ribosyl moiety. In some embodiments, when the first unnatural base is
[0022] [Chemistry] then the second unnatural base is [Chemistry] and when the first unnatural base is [Chemistry] then the second unnatural base is [Chemistry] and the wavy line indicates the bond to the ribosyl moiety. In some embodiments, when the first unnatural base is [Chemistry] then the second unnatural base is [Chemistry] and when the first unnatural base is [Chemistry] If so, the second unnatural base is
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chem.
Chem.
Chem.
Chem.
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chem.
Chem.
Chem.
Chem.
Chem.
[0023] Aspects disclosed herein are systems for the expression of non-natural polypeptides, comprising: (a) at least one non-natural amino acid; (b) an mRNA encoding a non-natural polypeptide, the mRNA comprising at least one codon containing one or more first non-natural bases; (c) a tRNA comprising at least one anticodon containing one or more second non-natural bases, wherein the one or more first non-natural bases and the one or more second non-natural bases form one or more complementary base pairs; and (d) a eukaryotic ribosome capable of translating the mRNA into a polypeptide containing the non-natural amino acid using the tRNA and a tRNA synthetase. The tRNA is charged with the non-natural amino acid, and / or the system may further comprise one or more nucleic acid constructs comprising a tRNA synthetase and / or a nucleic acid sequence encoding a tRNA synthetase, wherein the tRNA synthetase preferentially aminoacylates the tRNA with at least one non-natural amino acid. The system can be in vitro (e.g., a cell-free or reconstituted system of purified components such as a cell lysate) or in a eukaryotic cell. In some embodiments, at least one codon of the mRNA comprises three consecutive nucleobases (N-N-N), and one or more first non-natural bases (X) are located at the first position (X-N-N) in at least one codon of the mRNA. In some embodiments, at least one codon of the mRNA comprises three consecutive nucleobases (N-N-N), and one or more first non-natural bases (X) are located at the middle position (N-X-N) in the codon of the mRNA. In some embodiments, at least one codon of the mRNA comprises three consecutive nucleobases (N-N-N), and one or more first non-natural bases (X) are located at the last position (N-N-X) in at least one codon of the mRNA. In some embodiments, the one or more non-natural bases have the formula [Chemical formula] wherein R2 is selected from the group consisting of hydrogen, alkyl, alkenyl, alkynyl, methoxy, methanethiol, methaneseleno, halogen, cyano and azide, and the wavy line indicates the bond to the ribosyl moiety. In some embodiments, one or more first unnatural bases or one or more second unnatural bases are
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chemical formula
Chemical formula
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chemical formula
Chemical formula
[0024] In one embodiment, the eukaryotic cell comprises mRNA encoding a hypersensitive green fluorescent protein (EGFP) having an unnatural codon at position 151 (EGFP151(NXN), where N refers to one of the natural nucleobases and X refers to NaM), Methanosarcina mazei tRNAPyl (tRNAPyl(NYN), where Y refers to TPT3) recoded with a cognate unnatural anticodon, and a chimeric Methanosarcina barkeri pyrrolysyl-tRNA synthetase (ChPylRS) that can charge N6-(2-azidoethoxy)-carbonyl-L-lysine (AzK) to the unnatural tRNAPyl.
[0025] The various aspects of the invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by reference to the following detailed description which illustrates exemplary embodiments in which the principles of the invention are utilized, and to its accompanying drawings.
Brief Description of the Drawings
[0026]
Figure 1A
Figure 1B
Figure 1C
Figure 2
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 4E
Figure 4F
Figure 4G
Figure 5
Figure 6
Figure 7A
Figure 7B
Figure 7C
Figure 7D
Modes for Carrying Out the Invention
[0027] Specific Terms Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claimed subject matter. In this application, unless specifically stated otherwise, the use of the singular includes the plural. As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless specifically stated otherwise. Further, the term "including" and the use of other forms such as "include", "includes", and "included" are not limiting.
[0028] As used herein, ranges and amounts may be expressed as being “about” a particular value or range. About includes the exact amount. Thus, “about 5 μL” means “about 5 μL” as well as “5 μL”. In general, the term “about” includes amounts that are expected to be within the range of experimental error.
[0029] As used herein, phrases such as “under conditions suitable to provide” or “under conditions sufficient to produce” in the context of a synthetic method refer to reaction conditions such as time, temperature, solvent, reactant concentration, etc., that are within the scope of ordinary skill that an experimenter can vary to provide a useful amount or yield of the reaction product. The desired reaction product need not be the only reaction product, nor does the starting material need to be completely consumed, provided that the desired reaction product can be isolated or otherwise used further.
[0030] “Chemically feasible” means a bond arrangement or compound that does not violate the generally understood rules of organic structures. For example, in certain situations, structures within the scope of a claim that contain pentavalent carbon atoms, which do not exist in nature, are understood to be outside the scope of the claim. The structures disclosed herein are intended to include only “chemically feasible” structures in all of their embodiments. For example, in structures represented by variable atoms or groups, any recited structure that is not chemically feasible is not intended to be disclosed or claimed herein.
[0031] A “chemical analog” of a chemical structure, as used herein, refers to a chemical structure that may not be readily derivable synthetically from the parent structure, but retains a substantial similarity to the parent structure. In some embodiments, a nucleotide analog is a non-natural nucleotide. In some embodiments, a nucleoside analog is a non-natural nucleoside. Related chemical structures that are readily derivable synthetically from the parent chemical structure are referred to as “derivatives”.
[0032] Accordingly, a polynucleotide, as the term is used herein, refers to DNA-like or RNA-like polymers such as DNA, RNA, peptide nucleic acid (PNA), locked nucleic acid (LNA), phosphorothioate, unnatural bases, etc., which are well known in the art. Polynucleotides can be synthesized on an automated synthesizer using, for example, phosphoramidite chemistry reactions or other chemical approaches adapted for use with a synthesizer.
[0033] Examples of DNA include, but are not limited to, cDNA and genomic DNA. DNA can be attached to other biomolecules, including but not limited to RNA and peptides, by covalent or non-covalent means. Examples of RNA also include coding RNA, such as messenger RNA (mRNA). In some embodiments, the RNA is rRNA, RNAi, snoRNA, microRNA, siRNA, snRNA, exRNA, piRNA, long non-coding RNA, or any combination or hybrid thereof. In some cases, the RNA is a component of a ribozyme. DNA and RNA can be in any form, including but not limited to linear, circular, supercoiled, single-stranded, and double-stranded.
[0034] Peptide nucleic acid (PNA) is a synthetic DNA / RNA analog in which a peptide-like backbone replaces the sugar-phosphate backbone of DNA or RNA. PNA oligomers exhibit higher binding strength and greater specificity in binding to complementary DNA, and PNA / DNA base mismatches are less stable than similar mismatches in DNA / DNA duplexes. This binding strength and specificity also apply to PNA / RNA duplexes. PNA is resistant to enzymatic degradation because it is not readily recognized by either nucleases or proteases. PNA is also stable over a wide pH range. See Nielsen PE, Egholm M, Berg RH, Buchardt O (December 1991), "Sequence-selective recognition of DNA by strand displacement with a thymine-substituted polyamide", Science 254(5037):1497-500. doi:10.1126 / science.1962210. PMID 1962210; and Egholm M, Buchardt O, Christensen L, Behrens C, Freier SM, Driver DA, Berg RH, Kim SK, Norden B, and Nielsen PE (1993), "PNA Hybridizes to Complementary Oligonucleotides Obeying the Watson-Crick Hydrogen Bonding Rules". Nature 365(6446):566-8. doi:10.1038 / 365566a0. PMID 7692304.
[0035] Locked nucleic acids (LNAs) are modified RNA nucleotides, where the ribose moiety of the LNA nucleotide is modified with an additional bridge connecting the 2'-oxygen and 4'-carbon. This bridge "locks" the ribose in the 3'-endo (North) conformation, which is commonly seen in A-form duplexes. LNA nucleotides can be mixed with DNA or RNA residues of oligonucleotides as needed. Such oligomers may be chemically synthesized and are commercially available. The locked ribose conformation enhances base stacking and pre-organization of the backbone. For example, Kaur, H; Arora, A; Wengel, J; Maiti, S (2006), "Thermodynamic, Counterion, and Hydration Effects for the Incorporation of Locked Nucleic Acid Nucleotides into DNA Duplexes", Biochemistry 45(23):7347-55.doi:10.1021 / bi060307w.PMID 16752924; Owczarzy R.; You Y., Groth C.L., Tataurov A.V. (2011), "Stability and mismatch discrimination of locked nucleic acid-DNA duplexes", Biochem.50(43):9352-9367.doi:10.1021 / bi200904e.PMC3201676.PMID21928795; Alexei A. Koshkin; Sanjay K. Singh, Poul Nielsen, Vivek K.Rajwanshi, Ravindra Kumar, Michael Meldgaard, Carl Erik Olsen, Jesper Wengel (1998), "LNA (Locked Nucleic Acids): Synthesis of the adenine, cytosine, guanine, 5-methylcytosine, thymine and uracil bicyclonucleoside monomers, oligomerisation, and unprecedented nucleic acid recognition", Tetrahedron 54(14):3607-30. doi:10.1016 / S0040-4020(98)00094-5; and, Satoshi Obika; Daishu Nanbu, Yoshiyuki Hari, Ken-ichiro Morio, Yasuko In, Toshimasa Ishida, Takeshi Imanishi (1997), "Synthesis of 2’-O,4’-C-methyleneuridine and -cytidine. Novel bicyclic nucleosides having a fixed C3’-endo sugar puckering", Tetrahedron Lett. 38(50):8735-8. doi:10.1016 / S0040-4039(97)10322-7. See also.
[0036] Molecular beacons or molecular beacon probes are oligonucleotide hybridization probes that can detect the presence of specific nucleic acid sequences in a homogeneous solution. A molecular beacon is a hairpin-shaped molecule with an internally quenched fluorophore, and fluorescence is restored when it binds to the target nucleic acid sequence. See, for example, Tyagi S, Kramer FR (1996), "Molecular beacons: probes that fluoresce upon hybridization", Nat Biotechnol. 14(3):303-8. PMID 9630890; Tapp I, Malmberg L, Rennel E, Wik M, Syvanen AC (April 2000), "Homogeneous scoring of single-nucleotide polymorphisms: comparison of the 5’-nuclease TaqMan assay and Molecular Beacon probes", Biotechniques 28(4):732-8. PMID 10769752; and Akimitsu Okamoto (2011), "ECHO probes: a concept of fluorescence control for practical nucleic acid sensing", Chem. Soc. Rev. 40:5815-5828.
[0037] In some embodiments, a nucleobase is generally the heterocyclic base portion of a nucleoside. The nucleobase may be naturally occurring, may be modified, may have no similarity to natural bases, and may be synthesized, for example, by organic synthesis. In certain embodiments, a nucleobase includes any atom or group of atoms that can interact with the base of another nucleic acid, with or without the use of hydrogen bonding. In certain embodiments, an unnatural nucleobase is not derived from a natural nucleobase. It should be noted that an unnatural nucleobase does not necessarily possess basic properties, but is called a nucleobase for simplicity. In some embodiments, when referring to a nucleobase, "(d)" indicates that the nucleobase can be attached to deoxyribose or ribose.
[0038] In some embodiments, a nucleoside is a compound that includes a nucleobase portion and a sugar portion. Nucleosides include, but are not limited to, naturally occurring nucleosides (found in DNA and RNA), abasic nucleosides, modified nucleosides, and nucleosides having mimetic bases and / or sugar groups. Nucleosides include nucleosides having any of a variety of substituents. A nucleoside can be a glycoside compound formed by a glycosidic linkage between a nucleobase and a reducing group of a sugar.
[0039] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.
[0040] Methods, systems, and compositions comprising unnatural base pairs in eukaryotic cells Disclosed herein are in vivo methods and compositions for generating nucleic acids using an expanded genetic alphabet in eukaryotic cells (FIGS. 1A - 3B). In some examples, the nucleic acids encode non - natural proteins, where the non - natural proteins contain at least one non - natural amino acid. In some cases, the in vivo methods or compositions described herein utilize or comprise semi - synthetic organisms. In some examples, the method includes the step of incorporating at least one unnatural base pair (UBP) into one or more nucleic acids. Such base pairs are formed by base pairing between the nucleobases of two nucleosides. In the exemplary workflow provided in FIG. 1B, DNA 101 encoding protein 102 and tRNA 103 containing unnatural nucleobases (X, Y) that are each complementary is transcribed 104 to yield tRNA 106 and mRNA 107. After charging the tRNA with non - natural amino acid 105, mRNA 107 is translated 108 to yield protein 110 containing one or more non - natural amino acids 109. The methods and compositions described herein enable site - specific incorporation of non - natural amino acids with high fidelity and yield in some examples. Also described herein are methods for using semi - synthetic organisms containing an expanded genetic alphabet, semi - synthetic organisms for generating protein products including those containing at least one non - natural amino acid residue.
[0041] The selection of unnatural nucleobases enables the optimization of one or more steps in the methods described herein. For example, the nucleobases are selected for high-efficiency replication, transcription, and / or translation. In some examples, one or more unnatural nucleobase pairs are utilized for the methods described herein. For example, a first set of nucleobases that include a deoxyribo moiety are used for DNA replication (e.g., a first nucleobase and a second nucleobase configured to form a first base pair), and a second set of nucleobases (e.g., a third and a fourth nucleobase that are attached to ribose and are configured to form a second base pair) are used for transcription / translation. Complementary base pairing between the first set of nucleobases and the second set of nucleobases enables, in some examples, the transcription of a gene to yield tRNA or protein from a DNA template that includes nucleobases from the first set. Complementary base pairing between the second set of nucleobases (the second base pair) enables, in some examples, translation by matching tRNA and mRNA that include unnatural nucleic acids. In some cases, the nucleobases in the first set bind to a deoxyribose moiety. In some cases, the nucleobases in the first set bind to a ribose moiety. In some examples, the nucleobases of both sets are unique. In some examples, at least one nucleobase is the same in both sets. In some examples, the first nucleobase and the third nucleobase are the same. In some embodiments, the first base pair and the second base pair are not the same. In some cases, the first base pair, the second base pair, and the third base pair are not the same.
[0042] Engineered eukaryotic organisms In some embodiments, the methods and plasmids disclosed herein are further used to generate engineered eukaryotic organisms, such as organisms that incorporate and replicate unnatural nucleotides or unnatural base pairs (UBPs), and to transcribe mRNAs and tRNAs used to translate proteins containing unnatural amino acid residues, and nucleic acids containing unnatural nucleotides can also be used. In some examples, the organism is a semi-synthetic organism (SSO). In some examples, the SSO is not a prokaryote. In some examples, the SSO is a mammal. In some examples, the mammalian SSO is a human. In some examples, the mammalian SSO is a hamster. In some examples, the human SSO is derived from HEK293T cells. In some examples, the human SSO is derived from Chinese hamster ovary (CHO) cells.
[0043] In some examples, the cells used are genetically transformed with an expression cassette encoding a heterologous protein, such as a tRNA synthetase. In some embodiments, the tRNA synthetase preferentially aminoacylates a tRNA containing an anticodon with an unnatural base with an unnatural amino acid. In some embodiments, the cell comprises a tRNA synthetase that preferentially aminoacylates a tRNA containing an anticodon with an unnatural base with an unnatural amino acid.
[0044] The cell may be a eukaryotic cell, and the pair of unnatural base-pairing nucleotides can be TPT3 and NaM or CNMO.
[0045] Compositions and methods are described herein that include the use of two or more non-natural base-pairing nucleotides. Such base-pairing nucleotides enter cells in some cases by standard nucleic acid transformation methods known in the art (e.g., electroporation, chemical transformation, or other methods capable of introducing nucleic acids containing non-natural nucleotides into cells). In some cases, three or more non-natural base-pairing nucleotides are used. In some cases, the base-pairing non-natural nucleotides enter cells as part of a polynucleotide such as mRNA and / or tRNA. One or more base-pairing non-natural nucleotides that enter cells as part of a polynucleotide (RNA) need not themselves be replicated in vivo.
[0046] In some cases, genetically engineered cells are produced by introduction of nucleic acids, such as heterologous nucleic acids, into cells. Any cell described herein can be a host cell and can contain an expression vector. In some embodiments, the cell is a mammalian cell. In some embodiments, the mammalian cell is a human cell (e.g., HEK293T cells). In some embodiments, the mammalian cell is a hamster cell (e.g., CHO cells). In some embodiments, the cell contains one or more heterologous polynucleotides. Nucleic acid reagents can be introduced into microorganisms using a variety of techniques. Non-limiting examples of methods used to introduce heterologous nucleic acids into various organisms include transformation, transfection, transduction, electroporation, sonication-mediated transformation, conjugation, particle gun, and the like. In some examples, the addition of carrier molecules (e.g., bis-benzoimidazolyl compounds, see, e.g., U.S. Patent No. 5,595,899) can typically increase the uptake of DNA in cells that are thought to be difficult to transform by conventional methods. Conventional methods of transformation are readily available to those of skill in the art and can be found in Maniatis, T., E. F. Fritsch and J. Sambrook (1982) Molecular Cloning: a Laboratory Manual; Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y.
[0047] In some instances, genetic transformation can be obtained, but not limited to, by direct introduction of an expression cassette in plasmids, viral vectors, viral nucleic acids, phage nucleic acids, phages, cosmids and artificial chromosomes, or through introduction of genetic material or a carrier such as cationic liposomes in cells. Such methods are available in the art and can be readily adapted for use in the methods described herein. The introduction vector can be any nucleotide construct (e.g., plasmid) used to deliver a gene to a cell, or can be part of a general strategy for delivering a gene, e.g., as part of a recombinant retrovirus or adenovirus (Ram et al., Cancer Res. 53:83-88, (1993)). Suitable means for transfection, including viral vectors, chemical transfectants, or physical and mechanical methods such as electroporation and direct diffusion of DNA, are described, for example, in Wolff, J.A. et al., Science, 247, 1465-1468, (1990); and, Wolff, J.A., Nature, 352, 815-818, (1991).
[0048] Nucleic acid molecule In some embodiments, the nucleic acid (e.g., also referred to herein as the nucleic acid molecule of interest) is from any origin or composition, e.g., from RNA, siRNA (short interfering RNA), RNAi, tRNA, mRNA or rRNA (ribosomal RNA), and can be, for example, in any form (e.g., linear, circular, highly coiled, single-stranded, double-stranded, etc.). In some embodiments, the nucleic acid comprises nucleotides, nucleosides or polynucleotides. In some cases, the nucleic acid includes natural nucleic acids and non-natural nucleic acids. In some cases, the nucleic acid also includes non-natural nucleic acids such as RNA analogs (e.g., containing base analogs, sugar analogs and / or non-native backbones). It is understood that the term "nucleic acid" does not refer to, or imply, a polynucleotide chain of a specific length, and thus polynucleotides and oligonucleotides are also included within its definition. Exemplary natural nucleotides include, but are not limited to, ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, GMP, dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP and dGMP. Exemplary natural deoxyribonucleotides include dATP, dTTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP and dGMP. Exemplary natural ribonucleotides include ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP and GMP. For natural RNA, the uracil base is uridine. The nucleic acid can be a vector, plasmid, phagemid, autonomously replicating sequence (ARS), centromere, artificial chromosome, yeast artificial chromosome (e.g., YAC), or other nucleic acid that is replicable in a host cell or replicated in a host cell. In some cases, the non-natural nucleic acid is a nucleic acid analog. In additional cases, the non-natural nucleic acid is of extracellular origin. In other cases, the non-natural nucleic acid is available in the intracellular space of an organism provided herein, e.g., a genetically modified organism.In some embodiments, the unnatural nucleotide is not a natural nucleotide. In some embodiments, a nucleotide that does not contain a natural base contains an unnatural nucleic acid base.
[0049] Unnatural nucleic acid Nucleotide analogs or unnatural nucleotides include nucleotides containing some type of modification in any of the base, sugar, or phosphate moieties. In some embodiments, the modification includes a chemical modification. In some cases, the modification occurs at the 3’OH or 5’OH group, the backbone, the sugar component, or the nucleotide base. In some examples, the modification optionally includes a linker molecule not naturally present and / or inter- or intra-strand cross-links. In one aspect, the modified nucleic acid includes one or more modifications of the 3’OH or 5’OH group, the backbone, the sugar component, or the nucleotide base, and / or the addition of a linker molecule not naturally present. In one aspect, the modified backbone includes a backbone other than a phosphodiester backbone. In one aspect, the modified sugar includes a sugar other than deoxyribose (in modified DNA) or other than ribose (in modified RNA). In one aspect, the modified base includes a base other than adenine, guanine, cytosine, or thymine (in modified DNA), or a base other than adenine, guanine, cytosine, or uracil (in modified RNA).
[0050] In some embodiments, the nucleic acid includes at least one modified base. In some examples, the nucleic acid includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more modified bases. In some cases, the modification to the base moiety includes not only natural and synthetic modifications of A, C, G, and T / U, but also different purine or pyrimidine bases. In some embodiments, the modification is a modified form of adenine, guanine, cytosine, or thymine (in modified DNA), or a modified form of adenine, guanine, cytosine, or uracil (in modified RNA).
[0051] Modified bases of unnatural nucleic acids include, but are not limited to, uracil-5-yl, hypoxanthin-9-yl (I), 2-aminoadenin-9-yl, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azouracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine. Certain unnatural nucleic acids, such as 5-substituted pyrimidines, 6-azapyrimidines and N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, 2-aminopropyladenine, 5-propynyluracil, 5-propynylcytosine, 5-methylcytosine, those that increase the stability of double-strand formation, universal nucleic acids, hydrophobic nucleic acids, random nucleic acids, nucleic acids with an enlarged size, fluorinated nucleic acids, 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6 and O-6 substituted purines are 2-aminopropyladenine, 5-propynyluracil and 5-propynylcytosine, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl of adenine and guanine, other alkyl derivatives, 2-propyl of adenine and guanine and other alkyl derivatives, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyl (-C≡C-CH3) uracil, 5-propynylcytosine, other alkynyl derivatives of pyrimidine nucleic acids, 6-azouracil, 6-azacytosine, 6-azathymine,5-Uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially, 5-bromo, 5-trifluoromethyl, other 5-substituted uracils and cytosines, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine, 8-azadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, tricyclic pyrimidines, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), G-clamps, phenoxazine cytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one), pyridoindole cytidine (H-pyrido[3’,2’:4,5]pyrrolo[2,3-d]pyrimidin-2-one), purine or pyrimidine bases substituted with other heterocycles, 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, 2-pyridone, azacytosine, 5-bromocytosine, bromouracil, 5-chlorocytosine, chlorinated cytosines, cyclocytosine, cytidine arabinoside, 5-fluorocytosine, fluoropyrimidines, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil and 5-iodouracil, 2-amino-adenine, 6-thio-guanine, 2-thio-thymine, 4-thio-thymine, 5-propynyl-uracil, 4-thio-uracil, N4-ethylcytosine, 7-deazaguanine, 7-deaza-8-azaguanine, 5-hydroxycytosine, 2’-deoxyuridine, 2-amino-2’-deoxyadenosine,and U.S. Patent Nos. 3,687,808; 4,845,205; 4,910,300; 4,948,882; 5,093,232; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121; 5,596,091; 5,614,617; 5,645,985; 5,681,941; 5,750,692; 5,763,588; 5,830,653 and 6,005,096; International Publication No. 99 / 62923; Kandimalla et al., (2001) Bioorg. Med. Chem. 9:807-813; The Concise Encyclopedia of Polymer Science and Engineering, Kroschwitz, J.I., ed., John Wiley & Sons, 1990, pp. 858-859; Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613; and Sanghvi, Chapter 15, Antisense Research and Applications, Crooke and Lebleu, eds., CRC Press, 1993, pp. 273-288. Additional base modifications can be found, for example, in U.S. Patent No. 3,687,808; Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613. In some examples, the unnatural nucleic acid contains the nucleobases of FIG. 2. In some examples, the unnatural nucleic acid contains the nucleobases of FIG. 3A. In some examples, the unnatural nucleic acid contains the nucleobases of FIG. 3B.,
[0052] Unnatural nucleic acids containing various heterocyclic bases and various sugar moieties (and sugar analogs) are available in the art, and in some cases, the nucleic acid contains one or several heterocyclic bases other than the five major base components of naturally occurring nucleic acids. For example, as heterocyclic bases, in some cases, uracil-5-yl, cytosine-5-yl, adenine-7-yl, adenine-8-yl, guanine-7-yl, guanine-8-yl, 4-aminopyrrolo[2.3-d]pyrimidin-5-yl, 2-amino-4-oxopyrrolo[2,3-d]pyrimidin-5-yl, 2-amino-4-oxopyrrolo[2.3-d]pyrimidin-3-yl groups can be mentioned, where purine is bonded to the sugar moiety of the nucleic acid via the 9-position, to pyrimidine via the 1-position, to pyrrolopyrimidine via the 7-position, and to pyrazolopyrimidine via the 1-position.
[0053] In some embodiments, the modified bases of the unnatural nucleic acid are represented below, where the wavy line identifies the point of attachment to deoxyribose or ribose.
Chemical formula
Chemical formula
Chemical formula
Chemical formula
[0054] In some embodiments, the nucleotide analog is also modified at the phosphate moiety. Modified phosphate moieties include, but are not limited to, those having a modification in the linkage between two nucleotides, such as phosphorothioate, chiral phosphorothioate, phosphorodithioate, phosphotriester, aminoalkyl phosphotriester, 3'-alkylene phosphonate and methyl and other alkyl phosphonates including chiral phosphonate, phosphinate, phosphoramidate including 3'-aminophosphoramidate and aminoalkyl phosphoramidate, thionophosphoramidate, thionoalkyl phosphonate, thionoalkyl phosphotriester, and boranophosphate. The linkage of these phosphates or modified phosphates between two nucleotides is by a 3'-5' linkage or a 2'-5' linkage, and it is understood that the linkage includes reverse orientations such as from 3'-5' to 5'-3' or from 2'-5' to 5'-2'. Also included are various salts, mixed salts, and free acid forms. Numerous U.S. patents teach methods of making and using nucleotides containing modified phosphates, including, but not limited to, U.S. Pat. Nos. 3,687,808; 4,469,863; 4,476,301; 5,023,243; 5,177,196; 5,188,897; 5,264,423; 5,276,019; 5,278,302; 5,286,717; 5,321,131; 5,399,676; 5,405,939; 5,453,496; 5,455,233; 5,466,677; 5,476,925; 5,519,126; 5,536,821; 5,541,306; 5,550,111; 5,563,253; 5,571,799; 5,587,361; and 5,625,050.
[0055] In some embodiments, the unnatural nucleic acids include 2′,3′-dideoxy-2′,3′-didehydro-nucleosides (PCT / US2002 / 006460), 5′-substituted DNA and RNA derivatives (PCT / US2011 / 033961; Saha et al., J. Org Chem., 1995, vol. 60, pp. 788-789; Wang et al., Bioorganic & Medicinal Chemistry Letters, 1999, vol. 9, pp. 885-890; and Mikhailov et al., Nucleosides & Nucleotides, 1991, vol. 10(1-3), pp. 339-343; Leonid et al., 1995, vol. 14(3-5), pp. 901-905; and Eppacher et al., Helvetica Chimica Acta, 2004, vol. 87, pp. 3004-3020; PCT / JP2000 / 004720; PCT / JP2003 / 002342; PCT / JP2004 / 013216; PCT / JP2005 / 020435; PCT / JP2006 / 315479; PCT / JP2006 / 324484; PCT / JP2009 / 056718; PCT / JP2010 / 067560), or 5′-substituted monomers made as monophosphates using modified bases (Wang et al., Nucleosides Nucleotides & Nucleic Acids, 2004, vol. 23(1 and 2), pp. 317-337).
[0056] In some embodiments, the unnatural nucleic acids include modifications at the 5' and 2' positions of the sugar ring (PCT / US94 / 02993), e.g., 5'-CH2 substituted 2'-O-protected nucleosides (Wu et al., Helvetica Chimica Acta, 2000, Vol. 83, pp. 1127-1143 and Wu et al., Bioconjugate Chem. 1999, Vol. 10, pp. 921-924). In some cases, the unnatural nucleic acids include amide-linked nucleoside dimers prepared for incorporation into oligonucleotides, where the 3'-linked nucleoside (5' to 3') in the dimer includes 2'-OCH3 and 5'-(S)-CH3 (Mesmaeker et al., Synlett, 1997, pp. 1287-1290). The unnatural nucleic acids can include 2'-substituted 5'-CH2 (or O) modified nucleosides (PCT / US92 / 01020). The unnatural nucleic acids can include 5'-methylene phosphonate DNA and RNA monomers and dimers (Bohringer et al., Tet. Lett., 1993, Vol. 34, pp. 2723-2726; Collingwood et al., Synlett, 1995, Vol. 7, pp. 703-705; and Hutter et al., Helvetica Chimica Acta, 2002, Vol. 85, pp. 2777-2806). The unnatural nucleic acids can include 5'-phosphonate monomers with 2'-substitutions (U.S. Patent Application Publication No. 2006 / 0074035) and other modified 5'-phosphonate monomers (International Publication No. 1997 / 35869). The unnatural nucleic acids can include 5'-modified methylene phosphonate monomers (European Patent Application Publication No. 614907 and European Patent Application Publication No. 629633).Non-natural nucleic acids can include analogs of 5'- or 6'-phosphonate ribonucleosides that contain a hydroxyl group at the 5'- and / or 6'-position (Chen et al., Phosphorus, Sulfur and Silicon, 2002, Vol. 777, pp. 1783-1786; Jung et al., Bioorg. Med. Chem., 2000, Vol. 8, pp. 2501-2509; Gallier et al., Eur. J. Org. Chem., 2007, pp. 925-933; and Hampton et al., J. Med. Chem., 1976, Vol. 19(8), pp. 1029-1033). Non-natural nucleic acids can include 5'-phosphonate deoxyribonucleoside monomers and dimers having a 5'-phosphate group (Nawrot et al., Oligonucleotides, 2006, Vol. 16(1), pp. 68-82). Non-natural nucleic acids can include nucleosides having a 6'-phosphonate group, where the 5'- or / and 6'-position is unsubstituted or substituted with a thio-tert-butyl group (SC(CH3)3) (and its analogs); a methyleneamino group (CH2NH2) (and its analogs) or a cyano group (CN) (and its analogs) (Fairhurst et al., Synlett, 2001, Vol. 4, pp. 467-472; Kappler et al., J. Med. Chem., 1986, Vol. 29, pp. 1030-1038; Kappler et al., J. Med. Chem., 1982, Vol. 25, pp. 1179-1184; Vrudhula et al., J. Med. Chem., 1987, Vol. 30, pp. 888-894; Hampton et al., J. Med. Chem., 1976, Vol. 19, pp. 1371-1377; Geze et al., J. Am. Chem. Soc, 1983, Vol. 105(26), pp. 7638-7640; and Hampton et al., J. Am. Chem. Soc, 1973, Vol. 95(13), pp. 4404-4414).
[0057] In some embodiments, the unnatural nucleic acids also include modifications to the sugar moiety. In some cases, the nucleic acid contains one or more nucleosides with modified sugar groups. Such sugar-modified nucleosides may confer enhanced nuclease stability, increased binding affinity, or some other advantageous biological property. In certain embodiments, the nucleic acid includes a chemically modified ribofuranose ring moiety. Examples of chemically modified ribofuranose rings include, but are not limited to, addition of substituents (5' and / or 2' substituents; cross-linking of two ring atoms to form bicyclic nucleic acids (BNAs); replacement of the oxygen atom of the ribosyl ring with S, N(R) or C(R1)(R2) (R = H, C1-C 12 alkyl or protecting group); and combinations thereof). Examples of chemically modified sugars can be found in WO 2008 / 101157, US 2005 / 0130923 A1, and WO 2007 / 134181.
[0058] In some instances, the modified nucleic acid includes a modified sugar or sugar analog. Thus, in addition to ribose and deoxyribose, the sugar moiety can be a pentose, deoxypentose, hexose, deoxyhexose, glucose, arabinose, xylose, lyxose, or a cyclopentyl group of a sugar “analog”. The sugar can be in the form of pyranosyl or furanosyl. The sugar moiety can be a furanoside of ribose, deoxyribose, arabinose or 2′-O-alkyl ribose, and the sugar can be attached to each heterocyclic base in either the [alpha] or [beta] anomeric configuration. Sugar modifications include, but are not limited to, 2′-alkoxy-RNA analogs, 2′-amino-RNA analogs, 2′-fluoro-DNA and 2′-alkoxy- or amino-RNA / DNA chimeras. For example, sugar modifications can include 2′-O-methyl-uridine or 2′-O-methyl-cytidine. Sugar modifications include 2′-O-alkyl substituted deoxyribonucleosides and 2′-O-ethylene glycol-like ribonucleosides. The production of these sugars or sugar analogs and their respective “nucleosides” in which such sugars or sugar analogs are attached to heterocyclic bases (nucleic acid bases) is known. Sugar modifications can also be made and combined with other modifications.
[0059] Modifications to the sugar moiety include natural and non-natural modifications of ribose and deoxyribose. Sugar modifications include, but are not limited to, modifications at the 2′ position: OH; F; O-, S- or N-alkyl; O-, S- or N-alkenyl; O-, S- or N-alkynyl; or O-alkyl-O-alkyl, where the alkyl, alkenyl and alkynyl are substituted or unsubstituted C1-C 10 alkyl or C2-C10 alkenyl and alkynyl. 2′ sugar modifications include, but are not limited to, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2) nNH2, -O(CH2) n CH3, -O(CH2) n ONH2 and -O(CH2) n ON[(CH2) n CH3)]2 (wherein n and m are from 1 to about 10) may also be mentioned.
[0060] Other modifications at the 2'-position include, but are not limited to, C1-C 10Lower alkyl, substituted lower alkyl, aralkyl, aralkyl, O-aralkyl, O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloaralkyl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, a group for improving the pharmacokinetic properties of an oligonucleotide, or a group for improving the pharmacodynamic properties of an oligonucleotide, and other substituents having similar properties. Similar modifications can also be made at other positions of the sugar, particularly at the 3'-position of the sugar in the 3'-terminal nucleotide or in a 2'-5' linked oligonucleotide, and at the 5'-position of the 5'-terminal nucleotide. Modified sugars also include those containing modifications at the bridging ring oxygen such as CH2 and S. Sugar analogs of nucleotides can also have a sugar mimetic such as a cyclobutyl moiety instead of a pentofuranosyl sugar.U.S. Patent Nos. 4,981,957; 5,118,800; 5,319,080; 5,359,044; 5,393,878; 5,446,137; 5,466,786; 5,514,785; 5,519,134; 5,567,811; 5,576,427; 5,591,722; 5,597,909; 5,610,300; 5,627,053; 5,639,873; 5,646,265; 5,658,873; 5,670,633; 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121; 5,596,091; 5,614,617; 5,681,941; and 5,700,920, etc., teach the manufacture of such modified sugar structures and numerous U.S. patents detail and describe the scope of base modifications, each of which is hereby incorporated by reference in its entirety.
[0061] Examples of nucleic acids having modified sugar moieties include, but are not limited to, nucleic acids containing 5'-vinyl, 5'-methyl (R or S), 4'-S, 2'-F, 2'-OCH3 and 2'-O(CH2)2OCH3 substituents. Substituents at the 2'-position are allyl, amino, azido, thio, O-allyl, O-(C1-C 1O alkyl), OCF3, O(CH2)2SCH3, O(CH2)2-O-N(R m )(R n ) and O-CH2-C(=O)-N(R m )(R n )(wherein each R m and R n is independently H, or substituted or unsubstituted C1-C 10 alkyl).
[0062] In certain embodiments, the nucleic acids described herein include one or more bicyclic nucleic acids. In certain such embodiments, the bicyclic nucleic acid includes a bridge between the 4’ and 2’ ribosyl ring atoms. In certain embodiments, the nucleic acids provided herein include one or more bicyclic nucleic acids, wherein the bridge includes a 4’-2’ bicyclic nucleic acid. Examples of such 4’-2’ bicyclic nucleic acids include, but are not limited to, the formulae: 4’-(CH2)-O-2’ (LNA); 4’-(CH2)-S-2’; 4’-(CH2)2-O-2’ (ENA); 4’-CH(CH3)-O-2’ and 4’-CH(CH2OCH3)-O-2’, and analogs thereof (see U.S. Patent No. 7,399,845); 4’-C(CH3)(CH3)-O-2’ and analogs thereof (see International Publication No. 2009 / 006478, International Publication No. 2008 / 150729, U.S. Patent Application Publication No. 2004 / 0171570, U.S. Patent No. 7,427,672, Chattopadhyaya et al., J. Org. Chem., Vol. 209, No. 74, pp. 118-134, and International Publication No. 2008 / 154401).For example, Singh et al., Chem. Commun., 1998, Vol. 4, pp. 455-456; Koshkin et al., Tetrahedron, 1998, Vol. 54, pp. 3607-3630; Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 2000, Vol. 97, pp. 5633-5638; Kumar et al., Bioorg. Med. Chem. Lett., 1998, Vol. 8, pp. 2219-2222; Singh et al., J. Org. Chem., 1998, Vol. 63, pp. 10035-10039; Srivastava et al., J. Am. Chem. Soc., 2007, Vol. 129 (No. 26), pp. 8362-8379; Elayadi et al., Curr. Opinion Invens. Drugs, 2001, Vol. 2, pp. 558-561; Braasch et al., Chem. Biol, 2001, Vol. 8, pp. 1-7; Oram et al., Curr. Opinion Mol. Ther., 2001, Vol. 3, pp. 239-243; U.S. Patent Nos. 4,849,513; 5,015,733; 5,118,800; 5,118,802; 7,053,207; 6,268,490; 6,770,748; 6,794,499; 7,034,133; 6,525,191; 6,670,461; and 7,399,845; International Publication Nos. WO 2004 / 106356, WO 1994 / 14226, WO 2005 / 021570, WO 2007 / 090071 and WO 2007 / 134181; U.S. Patent Application Publication Nos. 2004 / 0171570, 2007 / 0287831 and 2008 / 0039618; U.S. Provisional Application Nos. 60 / 989,574, 61 / 026,995, 61 / 026,998, 61 / 056,564, 61 / 086,231, 61 / 097,787 and 61 / 099,844; and International Application Nos. PCT / US2008 / 064591, PCT US2008 / 066154, PCT US2008 / 068922 and PCT / DK98 / 00393 are also referred to.
[0063] In certain embodiments, the nucleic acid comprises linked nucleic acids. The nucleic acids can be linked together using any inter-nucleic acid linkage. Two main classes of inter-nucleic acid linkers are defined by the presence or absence of a phosphorus atom. Representative phosphorus-containing inter-nucleic acid linkages include, but are not limited to, phosphodiester, phosphotriester, methylphosphonate, phosphoramidate, and phosphorothioate (P=S). Representative phosphorus-free inter-nucleic acid linker groups include, but are not limited to, methylene methylimino (-CH2-N(CH3)-O-CH2-), thiodiester (-O-C(O)-S-), thiocarbamate (-O-C(O)(NH)-S-); siloxane (-O-Si(H)2-O-); and N,N * -dimethylhydrazine (-CH2-N(CH3)-N(CH3)). In certain embodiments, inter-nucleic acid linkages having chiral atoms can be produced as racemic mixtures, as separate enantiomers, for example, as alkylphosphonates and phosphorothioates. The unnatural nucleic acids can contain a single modification. The unnatural nucleic acids can contain multiple modifications within one moiety or between different moieties.
[0064] Modifications of the phosphate of the backbone to the nucleic acid include, but are not limited to, methylphosphonate, phosphorothioate, phosphoramidate (bridged or unbridged), phosphotriester, phosphorodithioate, phosphodithioate, and boranophosphate, and can be used in any combination. Other non-phosphate linkages can also be used.
[0065] In some embodiments, backbone modifications (e.g., internucleotide linkages of methylphosphonate, phosphorothioate, phosphoramidate, and phosphorodithioate) can confer immunomodulatory activity in the modified nucleic acids and / or enhance their stability in vivo.
[0066] In some instances, the phosphorus derivative (or modified phosphate group) can be bonded to a sugar or sugar analog moiety and can be a monophosphate, diphosphate, triphosphate, alkylphosphonate, phosphorothioate, phosphorodithioate, phosphoramidate, etc. Exemplary polynucleotides containing modified phosphate linkages or non-phosphate linkages can be found in Peyrottes et al., 1996, Nucleic Acids Res. 24:1841-1848; Chaturvedi et al., 1996, Nucleic Acids Res. 24:2318-2323; and Schultz et al., (1996) Nucleic Acids Res. 24:2966-2973; Matteucci, 1997, "Oligonucleotide Analogs: an Overview" in Oligonucleotides as Therapeutic Agents, (Chadwick and Cardew eds) John Wiley and Sons, New York, NY; Zon, 1993, "Oligonucleoside Phosphorothioates" in Protocols for Oligonucleotides and Analogs, Synthesis and Properties, Humana Press, 165-190; Miller et al., 1971, JACS 93:6657-6665; Jager et al., 1988, Biochem. 27:7247-7246; Nelson et al., 1997, JOC 62:7278-7287; U.S. Patent No. 5,453,496; and Micklefield, 2001, Curr. Med. Chem. 8:1157-1179.
[0067] In some cases, modification of the backbone involves replacing the phosphodiester linkage with an alternative moiety such as an anionic, neutral or cationic group. Examples of such modifications include anionic internucleoside linkages; N3’-P5’ phosphoramidate modifications; boranophosphate DNA; prooligonucleotides; neutral internucleoside linkages, such as methylphosphonates; amide-linked DNA; methylene(methylimino) linkages; formacetal and thioformacetal linkages; backbones containing sulfonyl groups; morpholino oligos; peptide nucleic acids (PNAs); and positively charged deoxyribonucleic acid guanidine (DNG) oligos (Micklefield, 2001, Current Medicinal Chemistry 8:1157-1179). Modified nucleic acids can include chimeric or mixed backbones containing one or more combinations of modifications, such as combinations of phosphate linkages, for example, combinations of phosphodiester and phosphorothioate linkages.
[0068] Examples of alternatives to phosphate include, for example, short-chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short-chain heteroatom or heterocyclic internucleoside linkages. These include those having a morpholino linkage (in part formed from the sugar moiety of the nucleoside); a siloxane backbone; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; alkene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S, and CH2 components. Numerous U.S. patents disclose methods of making and using these types of phosphate replacements, including, but not limited to, U.S. Patent Nos. 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; 5,434,257; 5,466,677; 5,470,967; 5,489,677; 5,541,307; 5,561,225; 5,596,086; 5,602,240; 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439. It is also understood that in nucleotide alternatives, both the sugar and phosphate moieties of the nucleotide can be replaced, for example, by an amide-type linkage (aminoethylglycine) (PNA). U.S. Patent Nos. 5,539,082; 5,714,331; and 5,719,262 teach methods of making and using PNA molecules, each of which is incorporated herein by reference.See also Nielsen et al., Science, 1991, Vol. 254, pp. 1497-1500. For example, it is also possible to link (conjugate) other types of molecules to nucleotides or nucleotide analogs in order to enhance cellular uptake. The conjugate can be chemically linked to the nucleotide or nucleotide analog.Such conjugates include, but are not limited to, lipid moieties such as cholesterol moieties (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, 86, pp. 6553-6556), cholic acid (Manoharan et al., Bioorg. Med. Chem. Let., 1994, 4, pp. 1053-1060), thioethers such as hexyl-S-tritylthiol (Manoharan et al., Ann. KY. Acad. Sci., 1992, 660, pp. 306-309; Manoharan et al., Bioorg. Med. Chem. Let., 1993, 3, pp. 2765-2770), thiocolesterol (Oberhauser et al., Nucl. Acids Res., 1992, 20, pp. 533-538), aliphatic chains such as dodecanediol or undecyl residues (Saison-Behmoaras et al., EM5OJ, 1991, 10, pp. 1111-1118; Kabanov et al., FEBS Lett., 1990, 259, pp. 327-330; Svinarchuk et al., Biochimie, 1993, 75, pp. 49-54), phospholipids such as di-hexadecyl-rac-glycerol or triethylammonium l-di-O-hexadecyl-rac-glycero-S-H-phosphonate (Manoharan et al., Tetrahedron Lett., 1995, 36, pp. 3651-3654; Shea et al., Nucl. Acids Res., 1990, 18, pp. 3777-3783), polyamine or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides, 1995, 14, pp. 969-973), or adamantane acetic acid (Manoharan et al., Tetrahedron Lett., 1995, 36, pp. 3651-3654), palmitoyl moieties (Mishra et al., Biochem. Biophys. Acta, 1995, 1264, pp. 229-237), or octadecylamine or hexylamino-carbonyl-oxy cholesterol moieties (Crooke et al., J. Pharmacol. Exp. Ther., 1996, 277, pp. 923-937).Numerous U.S. patents teach the manufacture of such conjugates, including but not limited to U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241; 5,391,723; 5,416,203; 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928; and 5,688,941.
[0069] Nucleobases for use in compositions and methods for replication, transcription, translation, and incorporation of unnatural amino acids into proteins are described herein. In some embodiments, the nucleobases described herein have the structure:
Chemical formula
[0070] Non-natural deoxyribonucleic acid (DNA) is, in some cases, transcribed into messenger RNA (mRNA) containing non-natural bases described herein (e.g., d5SICS, dNaM, dTPT3, dMTMO, dCNMO, dTAT1). Exemplary mRNA codons are encoded by an exemplary region of non-natural DNA containing three consecutive deoxyribonucleotides (NNN) including TTX, TGX, CGX, AGX, GAX, CAX, GXT, CXT, GXG, AXG, GXC, AXC, GXA, CXC, TXC, ATX, CTX, TTX, GTX, TAX or GGX (where X is a non-natural base attached to the 2'-deoxyribosyl moiety). Exemplary mRNA codons resulting from the transcription of exemplary non-natural DNA each contain three consecutive ribonucleotides (NNN) including UUX, UGX, CGX, AGX, GAX, CAX, GXU, CXU, GXG, AXG, GXC, AXC, GXA, CXC, UXC, AUX, CUX, UUX, GUX, UAX or GGX (where X is a non-natural base attached to the ribosyl moiety). In some embodiments, the non-natural base is at the first position (X-N-N) in the codon sequence. In some embodiments, the non-natural base is at the second (or middle) position (N-X-N) in the codon sequence. In some embodiments, the non-natural base is at the third (last) position (N-N-X) in the codon sequence.
[0071] The mRNA containing the codons described herein is, in some cases, translated in vivo in cells (e.g., eukaryotic cells). The translation of the mRNA containing the unnatural bases described herein is mediated by transfer RNAs (tRNAs) containing anticodon sequences that are the reverse complements of the mRNA codon sequences described herein. In some embodiments, the tRNA anticodon contains an unnatural base including YAA, XAA, YCA, XCA, YCG, XCG, YCU, XCU, YUC, XUC, YUG, XUG, AYC, AYG, CYC, CYU, GYC, GYU, UYC, GYG, GYA, YAU, XAU, XAG, YAG, XAC, YAC, XUA, YUA, XCC or YCC (wherein X and Y each represent an unnatural base and X and Y are not the same). In some embodiments, the unnatural base is at the first position (X / Y-N-N) in the anticodon sequence. In some embodiments, the unnatural base is at the second (or middle) position (N-X / Y-N) in the anticodon sequence. In some embodiments, the unnatural base is at the third (last) position (N-N-X / Y) in the anticodon sequence.
[0072] Nucleobase pair-forming In some embodiments, non-natural nucleotides form base pairs (unnatural base pairs; UBPs) with another non-natural nucleotide, for example, during translation. For example, a first non-natural nucleic acid can form a base pair with a second non-natural nucleic acid. For example, as one pair of unnatural nucleoside triphosphates that can form base pairs during translation, nucleotides comprising (d)5SICS and nucleotides comprising (d)NaM are included. As other examples, but not limited to: nucleotides comprising (d)CNMO and nucleotides comprising (d)TPT3 are included. Such non-natural nucleotides can have a ribose or deoxyribose sugar moiety (indicated by "(d)"). For example, as one pair of unnatural nucleoside triphosphates that can form base pairs when incorporated into a nucleic acid, nucleotides comprising TAT1 and nucleotides comprising NaM are included. In some embodiments, as one pair of unnatural nucleoside triphosphates that can form base pairs when incorporated into a nucleic acid, nucleotides comprising dCNMO and nucleotides comprising TAT1 are included. In some embodiments, as one pair of unnatural nucleoside triphosphates that can form base pairs when incorporated into a nucleic acid, nucleotides comprising dTPT3 and nucleotides comprising NaM are included. In some embodiments, non-natural nucleic acids do not substantially form base pairs with natural nucleic acids (A, T, G, C). In some embodiments, non-natural nucleic acids can form base pairs with natural nucleic acids.
[0073] In some embodiments, the unnatural (deoxy)ribonucleotides are unnatural (deoxy)ribonucleotides that can form UBP, but do not substantially form base pairs with any of the respective natural (deoxy)ribonucleotides. In some embodiments, the unnatural (deoxy)ribonucleotides are unnatural (deoxy)ribonucleotides that can form UBP, but do not substantially form base pairs with one or more natural nucleic acids. For example, the unnatural nucleic acid cannot substantially form base pairs with A, T, and C, but can form base pairs with G. For example, the unnatural nucleic acid cannot substantially form base pairs with A, T, and G, but can form base pairs with C. For example, the unnatural nucleic acid cannot substantially form base pairs with C, G, and A, but can form base pairs with T. For example, the unnatural nucleic acid cannot substantially form base pairs with C, G, and T, but can form base pairs with A. For example, the unnatural nucleic acid cannot substantially form base pairs with A and T, but can form base pairs with C and G. For example, the unnatural nucleic acid cannot substantially form base pairs with A and C, but can form base pairs with T and G. For example, the unnatural nucleic acid cannot substantially form base pairs with A and G, but can form base pairs with C and T. For example, the unnatural nucleic acid cannot substantially form base pairs with C and T, but can form base pairs with A and G. For example, the unnatural nucleic acid cannot substantially form base pairs with C and G, but can form base pairs with T and G. For example, the unnatural nucleic acid cannot substantially form base pairs with T and G, but can form base pairs with A and G. For example, the unnatural nucleic acid cannot substantially form base pairs with G, but can form base pairs with A, T, and C. For example, the unnatural nucleic acid cannot substantially form base pairs with A, but can form base pairs with G, T, and C. For example, the unnatural nucleic acid cannot substantially form base pairs with T, but can form base pairs with G, A, and C.For example, non-natural nucleic acids are substantially unable to form base pairs with C, but can form base pairs with G, T, and A.
[0074] Exemplary non-natural nucleotides capable of forming non-natural base pairs (UBPs) under in vivo conditions (e.g., in RNA such as between tRNA and mRNA) include, but are not limited to, 5SICS, d5SICS, NaM, dNaM, dTPT3, dMTMO, dCNMO, TAT1, and combinations thereof. In some embodiments, non-natural nucleotide base pairs include, but are not limited to:
Chemical formula
[0075] Unnatural base pairs (UBPs) are formed between the codon sequence of mRNA and the anticodon sequence of tRNA and promote the translation of mRNA into unnatural polypeptides. The codon-anticodon UBPs, in some examples, include a codon sequence comprising three consecutive nucleic acids (e.g., UUX) read 5’ to 3’ of the mRNA and an anticodon sequence comprising three consecutive nucleic acids (e.g., YAA or XAA) read 5’ to 3’ of the tRNA. In some embodiments, when the mRNA codon is UUX, the tRNA anticodon is YAA or XAA. In some embodiments, when the mRNA codon is UGX, the tRNA anticodon is YCA or XCA. In some embodiments, when the mRNA codon is CGX, the tRNA anticodon is YCG or XCG. In some embodiments, when the mRNA codon is AGX, the tRNA anticodon is YCU or XCU. In some embodiments, when the mRNA codon is GAX, the tRNA anticodon is YUC or XUC. In some embodiments, when the mRNA codon is CAX, the tRNA anticodon is YUG or XUG. In some embodiments, when the mRNA codon is GXU, the tRNA anticodon is AYC. In some embodiments, when the mRNA codon is CXU, the tRNA anticodon is AYG. In some embodiments, when the mRNA codon is GXG, the tRNA anticodon is CYC. In some embodiments, when the mRNA codon is AXG, the tRNA anticodon is CYU. In some embodiments, when the mRNA codon is GXC, the tRNA anticodon is GYC. In some embodiments, when the mRNA codon is AXC, the tRNA anticodon is GYU. In some embodiments, when the mRNA codon is GXA, the tRNA anticodon is UYC. In some embodiments, when the mRNA codon is CXC, the tRNA anticodon is GYG. In some embodiments, when the mRNA codon is UXC, the tRNA anticodon is GYA.In some embodiments, when the mRNA codon is AUX, the tRNA anticodon is YAU or XAU. In some embodiments, when the mRNA codon is CUX, the tRNA anticodon is XAG or YAG. In some embodiments, when the mRNA codon is UUX, the tRNA anticodon is XAA or YAA. In some embodiments, when the mRNA codon is GUX, the tRNA anticodon is XAC or YAC. In some embodiments, when the mRNA codon is UAX, the tRNA anticodon is XUA or YUA. In some embodiments, when the mRNA codon is GGX, the tRNA anticodon is XCC or YCC.
[0076] Natural and unnatural amino acids As used herein, an amino acid residue may refer to a molecule that includes both an amino group and a carboxyl group. Suitable amino acids include, but are not limited to, both D and L isomers of naturally occurring amino acids, as well as amino acids not found in nature that are prepared by organic synthesis or any other method. The term amino acid as used herein includes, but is not limited to, α-amino acids, natural amino acids, unnatural amino acids, and amino acid analogs.
[0077] The term "α-amino acid" may refer to a molecule that includes both an amino group and a carboxyl group attached to a carbon, called the α-carbon. For example:
Chemical formula
[0078] The term "β-amino acid" can refer to a molecule that includes both an amino group and a carboxyl group in the β configuration.
[0079] "Naturally occurring amino acid" may refer to any one of the 20 amino acids commonly found in peptides synthesized in nature, known by the single-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V.
[0080] The following table shows a summary of the properties of natural amino acids.
[0081]
Table 1
[0082] "Hydrophobic amino acids" include small hydrophobic amino acids and large hydrophobic amino acids. "Small hydrophobic amino acids" may be glycine, alanine, proline, and their analogs. "Large hydrophobic amino acids" may be valine, leucine, isoleucine, phenylalanine, methionine, tryptophan, and their analogs. "Polar amino acids" may be serine, threonine, asparagine, glutamine, cysteine, tyrosine, and their analogs. "Charged amino acids" may be lysine, arginine, histidine, aspartic acid, glutamate, and their analogs.
[0083] "Amino acid analog" is a molecule that is structurally similar to an amino acid and can substitute for an amino acid in the formation of peptide-mimicking macrocyclic molecules. Amino acid analogs include, but are not limited to, β-amino acids and amino acids in which the amino or carboxy group is replaced by a similarly reactive group (e.g., substitution of a primary amine by a secondary or tertiary amine, or substitution of a carboxy group by an ester).
[0084] "Non-standard amino acid (ncAA)" or "non-natural amino acid" can be an amino acid that is not one of the 20 amino acids commonly found in peptides synthesized in nature, and is known by the one-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. In some cases, non-natural amino acids are a subset of non-standard amino acids.
[0085] Amino acid analogs may include β-amino acid analogs. Examples of β-amino acid analogs include, but are not limited to, the following: cyclic β-amino acid analogs; β-alanine; (R)-β-phenylalanine; (R)-1,2,3,4-tetrahydro-isoquinoline-3-acetic acid; (R)-3-amino-4-(1-naphthyl)-butyric acid; (R)-3-amino-4-(2,4-dichlorophenyl)butyric acid; (R)-3-amino-4-(2-chlorophenyl)-butyric acid; (R)-3-amino-4-(2-cyanophenyl)-butyric acid; (R)-3-amino-4-(2-fluorophenyl)-butyric acid; (R)-3-amino-4-(2-furyl)-butyric acid; (R)-3-amino-4-(2-methylphenyl)-butyric acid; (R)-3-amino-4-(2-naphthyl)-butyric acid; (R)-3-amino-4-(2-thienyl)-butyric acid; (R)-3-amino-4-(2-trifluoromethylphenyl)-butyric acid; (R)-3-amino-4-(3,4-dichlorophenyl)butyric acid; (R)-3-amino-4-(3,4-difluorophenyl)butyric acid; (R)-3-amino-4-(3-benzothienyl)-butyric acid; (R)-3-amino-4-(3-chlorophenyl)-butyric acid; (R)-3-amino-4-(3-cyanophenyl)-butyric acid; (R)-3-amino-4-(3-fluorophenyl)-butyric acid; (R)-3-amino-4-(3-methylphenyl)-butyric acid; (R)-3-amino-4-(3-pyridyl)-butyric acid; (R)-3-amino-4-(3-thienyl)-butyric acid; (R)-3-amino-4-(3-trifluoromethylphenyl)-butyric acid; (R)-3-amino-4-(4-bromophenyl)-butyric acid; (R)-3-amino-4-(4-chlorophenyl)-butyric acid; (R)-3-amino-4-(4-cyanophenyl)-butyric acid; (R)-3-amino-4-(4-fluorophenyl)-butyric acid; (R)-3-amino-4-(4-iodophenyl)-butyric acid; (R)-3-amino-4-(4-methylphenyl)-butyric acid; (R)-3-amino-4-(4-nitrophenyl)-butyric acid; (R)-3-amino-4-(4-pyridyl)-butyric acid; (R)-3-amino-4-(4-trifluoromethylphenyl)-butyric acid; (R)-3-amino-4-pentafluoro-phenylbutyric acid; (R)-3-amino-5-hexenoic acid; (R)-3-amino-5-hexynoic acid;(R)-3-Amino-5-phenylpentanoic acid; (R)-3-Amino-6-phenyl-5-hexenoic acid; (S)-1,2,3,4-Tetrahydroisoquinoline-3-acetic acid; (S)-3-Amino-4-(1-naphthyl)butyric acid; (S)-3-Amino-4-(2,4-dichlorophenyl)butyric acid; (S)-3-Amino-4-(2-chlorophenyl)butyric acid; (S)-3-Amino-4-(2-cyanophenyl)butyric acid; (S)-3-Amino-4-(2-fluorophenyl)butyric acid; (S)-3-Amino-4-(2-furyl)butyric acid; (S)-3-Amino-4-(2-methylphenyl)butyric acid; (S)-3-Amino-4-(2-naphthyl)butyric acid; (S)-3-Amino-4-(2-thienyl)butyric acid; (S)-3-Amino-4-(2-trifluoromethylphenyl)butyric acid; (S)-3-Amino-4-(3,4-dichlorophenyl)butyric acid; (S)-3-Amino-4-(3,4-difluorophenyl)butyric acid; (S)-3-Amino-4-(3-benzothienyl)butyric acid; (S)-3-Amino-4-(3-chlorophenyl)butyric acid; (S)-3-Amino-4-(3-cyanophenyl)butyric acid; (S)-3-Amino-4-(3-fluorophenyl)butyric acid; (S)-3-Amino-4-(3-methylphenyl)butyric acid; (S)-3-Amino-4-(3-pyridyl)butyric acid; (S)-3-Amino-4-(3-thienyl)butyric acid; (S)-3-Amino-4-(3-trifluoromethylphenyl)butyric acid; (S)-3-Amino-4-(4-bromophenyl)butyric acid; (S)-3-Amino-4-(4-chlorophenyl)butyric acid; (S)-3-Amino-4-(4-cyanophenyl)butyric acid; (S)-3-Amino-4-(4-fluorophenyl)butyric acid; (S)-3-Amino-4-(4-iodophenyl)butyric acid; (S)-3-Amino-4-(4-methylphenyl)butyric acid; (S)-3-Amino-4-(4-nitrophenyl)butyric acid; (S)-3-Amino-4-(4-pyridyl)butyric acid; (S)-3-Amino-4-(4-trifluoromethylphenyl)butyric acid; (S)-3-Amino-4-pentafluoro-phenylbutyric acid; (S)-3-Amino-5-hexenoic acid; (S)-3-Amino-5-hexynoic acid; (S)-3-Amino-5-phenylpentanoic acid; (S)-3-Amino-6-phenyl-5-hexenoic acid;1,2,5,6 - Tetrahydropyridine - 3 - carboxylic acid; 1,2,5,6 - Tetrahydropyridine - 4 - carboxylic acid; 3 - Amino - 3 - (2 - chlorophenyl) - propionic acid; 3 - Amino - 3 - (2 - thienyl) - propionic acid; 3 - Amino - 3 - (3 - bromophenyl) - propionic acid; 3 - Amino - 3 - (4 - chlorophenyl) - propionic acid; 3 - Amino - 3 - (4 - methoxyphenyl) - propionic acid; 3 - Amino - 4,4,4 - trifluoro - butyric acid; 3 - Aminoadipic acid; D - β - phenylalanine; β - leucine; L - β - homoalanine; L - β - homoaspartic acid γ - benzyl ester; L - β - homoglutamic acid δ - benzyl ester; L - β - homoisoleucine; L - β - homoleucine; L - β - homomethionine; L - β - homophenylalanine; L - β - homoproline; L - β - homotryptophan; L - β - homovaline; L - Nω - benzyloxycarbonyl - β - homolysine; Nω - L - β - homoarginine; O - benzyl - L - β - homohydroxyproline; O - benzyl - L - β - homoserine; O - benzyl - L - β - homothreonine; O - benzyl - L - β - homotyrosine; γ - trityl - L - β - homoaspartic acid; (R) - β - phenylalanine; L - β - homoaspartic acid γ - t - butyl ester; L - β - homoglutamic acid δ - t - butyl ester; L - Nω - β - homolysine; Nδ - trityl - L - β - homoglutamic acid; Nω - 2,2,4,6,7 - pentamethyl - dihydrobenzofuran - 5 - sulfonyl - L - β - homoarginine; O - t - butyl - L - β - homohydroxy - proline; O - t - butyl - L - β - homoserine; O - t - butyl - L - β - homothreonine; O - t - butyl - L - β - homotyrosine; 2 - Aminocyclopentanecarboxylic acid; and 2 - Aminocyclohexanecarboxylic acid.;
[0086] Examples of amino acid analogs may include alanine, valine, glycine, or leucine analogs. Examples of amino acid analogs of alanine, valine, glycine, and leucine include, but are not limited to, the following: α-methoxy glycine; α-allyl-L-alanine; α-aminoisobutyric acid; α-methyl-leucine; β-(1-naphthyl)-D-alanine; β-(1-naphthyl)-L-alanine; β-(2-naphthyl)-D-alanine; β-(2-naphthyl)-L-alanine; β-(2-pyridyl)-D-alanine; β-(2-pyridyl)-L-alanine; β-(2-thienyl)-D-alanine; β-(2-thienyl)-L-alanine; β-(3-benzothienyl)-D-alanine; β-(3-benzothienyl)-L-alanine; β-(3-pyridyl)-D-alanine; β-(3-pyridyl)-L-alanine; β-(4-pyridyl)-D-alanine; β-(4-pyridyl)-L-alanine; β-chloro-L-alanine; β-cyano-L-alanine; β-cyclohexyl-D-alanine; β-cyclohexyl-L-alanine; β-cyclopentene-1-yl-alanine; β-cyclopentyl-alanine; β-cyclopropyl-L-Ala-OH. dicyclohexylammonium salt; β-t-butyl-D-alanine; β-t-butyl-L-alanine; γ-aminobutyric acid; L-α,β-diaminopropionic acid; 2,4-dinitrophenylglycine; 2,5-dihydro-D-phenylglycine; 2-amino-4,4,4-trifluorobutyric acid; 2-fluoro-phenylglycine; 3-amino-4,4,4-trifluoro-butanoic acid; 3-fluorovaline; 4,4,4-trifluorovaline; 4,5-dehydro-L-leu-OH. dicyclohexylammonium salt; 4-fluoro-D-phenylglycine; 4-fluoro-L-phenylglycine; 4-hydroxy-D-phenylglycine; 5,5,5-trifluoroleucine; 6-aminohexanoic acid; cyclopentyl-D-Gly-OH. dicyclohexylammonium salt; cyclopentyl-Gly-OH. dicyclohexylammonium salt; D-α,β-diaminopropionic acid; D-α-aminobutyric acid; D-α-t-butylglycine; D-(2-thienyl)glycine; D-(3-thienyl)glycine; D-2-aminocaproic acid; D-2-indanylglycine;D-allylglycine-dicyclohexylammonium salt; D-cyclohexylglycine; D-norvaline; D-phenylglycine; β-aminobutyric acid; β-aminoisobutyric acid; (2-bromophenyl)glycine; (2-methoxyphenyl)glycine; (2-methylphenyl)glycine; (2-thiazolyl)glycine; (2-thienyl)glycine; 2-amino-3-(dimethylamino)-propionic acid; L-α,β-diaminopropionic acid; L-α-aminobutyric acid; L-α-t-butylglycine; L-(3-thienyl)glycine; L-2-amino-3-(dimethylamino)-propionic acid; L-2-aminocaproic acid dicyclohexyl-ammonium salt; L-2-indanylglycine; L-allylglycine dicyclohexylammonium salt; L-cyclohexylglycine; L-phenylglycine; L-propargylglycine; L-norvaline; N-α-aminomethyl-L-alanine; D-α,γ-diaminobutyric acid; L-α,γ-diaminobutyric acid; β-cyclopropyl-L-alanine; (N-β-(2,4-dinitrophenyl))-L-α,β-diaminopropionic acid; (N-β-1-(4,4-dimethyl-2,6-dioxocyclohexa-1-ylidene)ethyl)-D-α,β-diaminopropionic acid; (N-β-1-(4,4-dimethyl-2,6-dioxocyclohexa-1-ylidene)ethyl)-L-α,β-diaminopropionic acid; (N-β-4-methyltrityl)-L-α,β-diaminopropionic acid; (N-β-allyloxycarbonyl)-L-α,β-diaminopropionic acid; (N-γ-1-(4,4-dimethyl-2,6-dioxocyclohexa-1-ylidene)ethyl)-D-α,γ-diaminobutyric acid; (N-γ-1-(4,4-dimethyl-2,6-dioxocyclohexa-1-ylidene)ethyl)-L-α,γ-diaminobutyric acid; (N-γ-4-methyltrityl)-D-α,γ-diaminobutyric acid; (N-γ-4-methyltrityl)-L-α,γ-diaminobutyric acid; (N-γ-allyloxycarbonyl)-L-α,γ-diaminobutyric acid; D-α,γ-diaminobutyric acid; 4,5-dehydro-L-leucine; cyclopentyl-D-Gly-OH; cyclopentyl-Gly-OH; D-allylglycine; D-homocyclohexylalanine; L-1-pyrenylalanine; L-2-aminocaproic acid; L-allylglycine;L-Homocyclohexylalanine; and N-(2-hydroxy-4-methoxy-Bzl)-Gly-OH;
[0087] Examples of amino acid analogs can include analogs of arginine or lysine. Examples of amino acid analogs of arginine and lysine include, but are not limited to, the following: citrulline; L-2-amino-3-guanidinopropionic acid; L-2-amino-3-ureidopropionic acid; L-citrulline; Lys(Me)2-OH; Lys(N3)-OH; Nδ-benzyloxycarbonyl-L-ornithine; Nω-nitro-D-arginine; Nω-nitro-L-arginine; α-methylornithine; 2,6-diaminoheptanedioic acid; L-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxo-cyclohex-1-ylidene)ethyl)-D-ornithine; (Nδ-1-(4,4-dimethyl-2,6-dioxo-cyclohexen-1-ylidene)ethyl)-L-ornithine; (Nδ-4-methyltrityl)-D-ornithine; (Nδ-4-methyltrityl)-L-ornithine; D-ornithine; L-ornithine; Arg(Me)(Pbf)-OH; Arg(Me)2-OH (asymmetric); Arg(Me)2-OH (symmetric); Lys(ivDde)-OH; Lys(Me)2-OH.HCl; Lys(Me3)-OH chloride; Nω-nitro-D-arginine; and Nω-nitro-L-arginine.
[0088] Examples of amino acid analogs can include analogs of aspartic acid or glutamic acid. Examples of amino acid analogs of aspartic acid and glutamic acid include, but are not limited to, the following: α-methyl-D-aspartic acid; α-methyl-glutamic acid; α-methyl-L-aspartic acid; γ-methylene-glutamic acid; (N-γ-ethyl)-L-glutamic acid; [N-α-(4-aminobenzoyl)]-L-glutamic acid; 2,6-diaminopimelic acid; L-α-amino-suberic acid; D-2-aminoadipic acid; D-α-amino-suberic acid; α-aminopimelic acid; iminodiacetic acid; L-2-aminoadipic acid; threo-β-methyl-aspartic acid; γ-carboxy-D-glutamic acid γ,γ-di-t-butyl ester; γ-carboxy-L-glutamic acid γ,γ-di-t-butyl ester; Glu(OAll)-OH; L-Asu(OtBu)-OH; and pyroglutamic acid.
[0089] Amino acid analogs can include analogs of cysteine and methionine. Examples of amino acid analogs of cysteine and methionine include, but are not limited to: Cys(farnesyl)-OH, Cys(farnesyl)-OMe, α-methyl-methionine, Cys(2-hydroxyethyl)-OH, Cys(3-aminopropyl)-OH, 2-amino-4-(ethylthio)butyric acid, butionine, butionine sulfoximine, ethionine, methionine methylsulfonium chloride, selenomethionine, cysteic acid, [2-(4-pyridyl)ethyl]-DL-penicillamine, [2-(4-pyridyl)ethyl]-L-cysteine, 4-methoxybenzyl-D-penicillamine, 4-methoxybenzyl-L-penicillamine, 4-methylbenzyl-D-penicillamine, 4-methylbenzyl-L-penicillamine, benzyl-D-cysteine, benzyl-L-cysteine, benzyl-DL-homocysteine, carbamoyl-L-cysteine, carboxyethyl-L-cysteine, carboxymethyl-L-cysteine, diphenylmethyl-L-cysteine, ethyl-L-cysteine, methyl-L-cysteine, t-butyl-D-cysteine, trityl-L-homocysteine, trityl-D-penicillamine, cystathionine, homocystine, L-homocystine, (2-aminoethyl)-L-cysteine, seleno-L-cystine, cystathionine, Cys(StBu)-OH, and acetamidomethyl-D-penicillamine.
[0090] Examples of amino acid analogs can include analogs of phenylalanine and tyrosine. Examples of amino acid analogs of phenylalanine and tyrosine include: β-methyl-phenylalanine, β-hydroxyphenylalanine, α-methyl-3-methoxy-DL-phenylalanine, α-methyl-D-phenylalanine, α-methyl-L-phenylalanine, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, 2,4-dichloro-phenylalanine, 2-(trifluoromethyl)-D-phenylalanine, 2-(trifluoromethyl)-L-phenylalanine, 2-bromo-D-phenylalanine, 2-bromo-L-phenylalanine, 2-chloro-D-phenylalanine, 2-chloro-L-phenylalanine, 2-cyano-D-phenylalanine, 2-cyano-L-phenylalanine, 2-fluoro-D-phenylalanine, 2-fluoro-L-phenylalanine, 2-methyl-D-phenylalanine, 2-methyl-L-phenylalanine, 2-nitro-D-phenylalanine, 2-nitro-L-phenylalanine, 2,4,5-trihydroxy-phenylalanine, 3,4,5-trifluoro-D-phenylalanine, 3,4,5-trifluoro-L-phenylalanine, 3,4-dichloro-D-phenylalanine, 3,4-dichloro-L-phenylalanine, 3,4-difluoro-D-phenylalanine, 3,4-difluoro-L-phenylalanine, 3,4-dihydroxy-L-phenylalanine, 3,4-dimethoxy-L-phenylalanine, 3,5,3’-triiodo-L-thyronine, 3,5-diiodo-D-tyrosine, 3,5-diiodo-L-tyrosine, 3,5-Iodo-L-thyronine, 3-(trifluoromethyl)-D-phenylalanine, 3-(trifluoromethyl)-L-phenylalanine, 3-amino-L-tyrosine, 3-bromo-D-phenylalanine, 3-bromo-L-phenylalanine, 3-chloro-D-phenylalanine, 3-chloro-L-phenylalanine, 3-chloro-L-tyrosine, 3-cyano-D-phenylalanine, 3-cyano-L-phenylalanine, 3-fluoro-D-phenylalanine, 3-fluoro-L-phenylalanine, 3-fluoro-thyrosine, 3-iodo-D-phenylalanine, 3-iodo-L-phenylalanine, 3-iodo-L-thyronine, 3-methoxy-L-thyronine, 3-methyl-D-phenylalanine, 3-methyl-L-phenylalanine, 3-nitro-D-phenylalanine, 3-nitro-L-phenylalanine, 3-nitro-L-thyronine, 4-(trifluoromethyl)-D-phenylalanine, 4-(trifluoromethyl)-L-phenylalanine, 4-amino-D-phenylalanine, 4-amino-L-phenylalanine, 4-benzoyl-D-phenylalanine, 4-benzoyl-L-phenylalanine, 4-bis(2-chloroethyl)amino-L-phenylalanine, 4-bromo-D-phenylalanine, 4-bromo-L-phenylalanine, 4-chloro-D-phenylalanine, 4-chloro-L-phenylalanine, 4-cyano-D-phenylalanine, 4-cyano-L-phenylalanine, 4-fluoro-D-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-D-phenylalanine, 4-iodo-L-phenylalanine, homophenylalanine, thyroxine, 3,3-diphenylalanine, thyronine, ethyl-thyrosine, and methylthyrosine, may be mentioned.
[0091] Examples of amino acid analogs may include analogs of proline. Examples of proline amino acid analogs include, but are not limited to: 3,4-dehydro-proline, 4-fluoro-proline, cis-4-hydroxy-proline, thiazolidine-2-carboxylic acid, and trans-4-fluoro-proline.
[0092] Examples of amino acid analogs include analogs of serine and threonine. Examples of amino acid analogs of serine and threonine include, but are not limited to: 3-amino-2-hydroxy-5-methylhexanoic acid, 2-amino-3-hydroxy-4-methylpentanoic acid, 2-amino-3-ethoxybutanoic acid, 2-amino-3-methoxybutanoic acid, 4-amino-3-hydroxy-6-methylheptanoic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-benzyloxypropionic acid, 2-amino-3-ethoxypropionic acid, 4-amino-3-hydroxybutanoic acid, and α-methylserine.
[0093] Examples of amino acid analogs include analogs of tryptophan. Examples of amino acid analogs of tryptophan include, but are not limited to, the following: α-methyl-tryptophan; β-(3-benzothienyl)-D-alanine; β-(3-benzothienyl)-L-alanine; 1-methyl-tryptophan; 4-methyl-tryptophan; 5-benzyloxy-tryptophan; 5-bromo-tryptophan; 5-chloro-tryptophan; 5-fluoro-tryptophan; 5-hydroxytryptophan; 5-hydroxy-L-tryptophan; 5-methoxy-tryptophan; 5-methoxy-L-tryptophan; 5-methyl-tryptophan; 6-bromo-tryptophan; 6-chloro-D-tryptophan; 6-chloro-tryptophan; 6-fluoro-tryptophan; 6-methyl-tryptophan; 7-benzyloxy-tryptophan; 7-bromo-tryptophan; 7-methyl-tryptophan; D-1,2,3,4-tetrahydro-norharman-3-carboxylic acid; 6-methoxy-1,2,3,4-tetrahydronorharman-1-carboxylic acid; 7-azatryptophan; L-1,2,3,4-tetrahydro-norharman-3-carboxylic acid; 5-methoxy-2-methyl-tryptophan; and 6-chloro-L-tryptophan.
[0094] The amino acid analog may be a racemate. In some cases, the D-isomer of the amino acid analog is used. In some cases, the L-isomer of the amino acid analog is used. In some cases, the amino acid analog contains a chiral center in the R or S configuration. Sometimes, the amino group of the β-amino acid analog is substituted with a protecting group such as tert-butyloxycarbonyl (BOC group), 9-fluorenylmethyloxycarbonyl (FMOC), tosyl, etc. Sometimes, the carboxylic acid functional group of the β-amino acid analog is protected, for example, as its ester derivative. In some cases, salts of the amino acid analog are used.
[0095] In some embodiments, the unnatural amino acid is an unnatural amino acid described in Liu C.C., Schultz, P.G. Annu. Rev. Biochem. 2010, 79, 413. In some embodiments, the unnatural amino acid includes N6((2-azidoethoxy)-carbonyl)-L-lysine.
[0096] In some embodiments, the amino acid residues described herein (e.g., within a protein) are mutated to non-natural amino acids prior to attachment to the conjugate moiety. In some cases, the mutation to a non-natural amino acid prevents or minimizes an autoimmune response of the immune system. As used herein, the term "non-natural amino acid" refers to an amino acid other than the 20 amino acids that naturally occur in proteins. Non-limiting examples of non-natural amino acids include: p-acetyl-L-phenylalanine, p-iodo-L-phenylalanine, p-methoxyphenylalanine, O-methyl-L-tyrosine, p-propynyloxyphenylalanine, p-propynyl-phenylalanine, L-3-(2-naphthyl)alanine, 3-methyl-phenylalanine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, tri-O-acetyl-GlcNAcp-serine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-boronophenylalanine, O-propynyl tyrosine, L-phosphoserine, phosphonoserine, phosphonotyrosine, p-bromophenylalanine, selenocysteine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, N6-((azidoethoxy)-carbonyl)-L-lysine, (AzK), N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, or N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine, non-natural analogs of tyrosine amino acids; non-natural analogs of glutamine amino acids; non-natural analogs of phenylalanine amino acids; non-natural analogs of serine amino acids; non-natural analogs of threonine amino acids; alkyl, aryl, acyl, azide, cyano, halo, hydrazine, hydrazide, hydroxyl, alkenyl, alkynyl, ether, thiol, sulfonyl, seleno, ester, thioacid, borate, boronate, phospho, phosphono, phosphine, heterocyclic, enone, imine, aldehyde, hydroxylamine, keto, or amino-substituted amino acids, or combinations thereof; amino acids having a photoactivatable crosslinker;Spin-labeled amino acids; fluorescent amino acids; metal-binding amino acids; metal-containing amino acids; radioactive amino acids; photocaged and / or photo-isomerizable amino acids; biotin or biotin analogs containing amino acids; keto containing amino acids; amino acids containing polyethylene glycol or polyethers; heavy atom-substituted amino acids; chemically or photocleavable amino acids; amino acids having an elongated side chain; amino acids containing a toxic group; sugar-substituted amino acids; carbon-bonded sugar-containing amino acids; redox-active amino acids; α-hydroxy-containing acids; aminothiols; α,α-disubstituted amino acids; β-amino acids; cyclic amino acids other than proline or histidine, and aromatic amino acids other than phenylalanine, tyrosine or tryptophan, are included.;
[0097] In some embodiments, the unnatural amino acid comprises a selectively reactive group or a reactive group for site-selective labeling of a target protein or polypeptide. Optionally, the chemical reaction is a bioorthogonal reaction (e.g., biocompatible and selective reaction). Optionally, the chemical reaction is a Cu(I)-catalyzed or "copper-free" alkyne-azide triazole formation reaction, Staudinger ligation, inverse electron demand Diels-Alder (IEDDA) reaction, "photoclick" chemistry, or a metal-mediated process such as olefin metathesis and Suzuki-Miyaura or Sonogashira cross-coupling. In some embodiments, the unnatural amino acid comprises a photoreactive group that crosslinks upon irradiation with, for example, UV light. In some embodiments, the unnatural amino acid comprises a photocaged amino acid. Optionally, the unnatural amino acid is a para-substituted, meta-substituted, or ortho-substituted amino acid derivative.;
[0098] In some cases, the unnatural amino acids include p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, O-methyl-L-tyrosine, p-methoxyphenylalanine, p-propynyloxyphenylalanine, p-propynyl-phenylalanine, L-3-(2-naphthyl)alanine, 3-methyl-phenylalanine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, tri-O-acetyl-GlcNAcp-serine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, L-phosphoserine, phosphonoserine, phosphonotyrosine, p-bromophenylalanine, p-amino-L-phenylalanine, or isopropyl-L-phenylalanine.
[0099] In some cases, the unnatural amino acid is 3-aminotyrosine, 3-nitrotyrosine, 3,4-dihydroxy-phenylalanine, or 3-iodotyrosine. In some cases, the unnatural amino acid is phenylselenocysteine. In some cases, the unnatural amino acid is benzophenone, ketone, iodide, methoxy, acetyl, benzoyl, or azide (including phenylalanine derivatives). In some cases, the unnatural amino acid is a benzophenone, ketone, iodide, methoxy, acetyl, benzoyl, or azide-containing lysine derivative. In some cases, the unnatural amino acid contains an aromatic side chain. In some cases, the unnatural amino acid does not contain an aromatic side chain. In some cases, the unnatural amino acid contains an azide group. In some cases, the unnatural amino acid contains a Michael acceptor group. In some cases, the Michael acceptor group contains an unsaturated moiety capable of forming a covalent bond via a 1,2-addition reaction. In some cases, the Michael acceptor group contains an electron-deficient alkene or alkyne. In some cases, examples of the Michael acceptor group include, but are not limited to, alpha, beta-unsaturated: ketone, aldehyde, sulfoxide, sulfone, nitrile, imine, or aromatic. In some cases, the unnatural amino acid is dehydroalanine. In some cases, the unnatural amino acid contains an aldehyde or ketone group. In some cases, the unnatural amino acid is a lysine derivative containing an aldehyde or ketone group. In some cases, the unnatural amino acid is a lysine derivative containing one or more O, N, Se, or S atoms at the beta, gamma, or delta position. In some cases, the unnatural amino acid is a lysine derivative containing an O, N, Se, or S atom at the gamma position. In some cases, the unnatural amino acid is a lysine derivative in which the epsilon N atom is replaced by an oxygen atom. In some cases, the unnatural amino acid is a lysine derivative that is a post-translationally modified lysine not naturally occurring.
[0100] In some cases, the unnatural amino acid is an amino acid containing a side chain, and the sixth atom from the alpha position contains a carbonyl group. In some cases, the unnatural amino acid is an amino acid containing a side chain, the sixth atom from the alpha position contains a carbonyl group, and the fifth atom from the alpha position is nitrogen. In some cases, the unnatural amino acid is an amino acid containing a side chain, and the seventh atom from the alpha position is an oxygen atom.
[0101] In some cases, the unnatural amino acid is a serine derivative containing selenium. In some cases, the unnatural amino acid is selenoserine (2-amino-3-hydroxyselenopropanoic acid). In some cases, the unnatural amino acid is 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid. In some cases, the unnatural amino acid is 2-amino-3-(phenylselanyl)propanoic acid. In some cases, the unnatural amino acid contains selenium, and the oxidation of selenium results in the formation of an unnatural amino acid containing an alkene.
[0102] In some cases, the unnatural amino acid contains a cyclooctynyl group. In some cases, the unnatural amino acid contains a trans-cyclooctenyl group. In some cases, the unnatural amino acid contains a norbornenyl group. In some cases, the unnatural amino acid contains a cyclopropenyl group. In some cases, the unnatural amino acid contains a diazirinyl group. In some cases, the unnatural amino acid contains a tetrazinyl group.
[0103] In some cases, the unnatural amino acid is a lysine derivative in which the side-chain nitrogen is carbamylated. In some cases, the unnatural amino acid is a lysine derivative in which the side-chain nitrogen is acylated. In some cases, the unnatural amino acid is 2-amino-6-{[(tert-butoxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(tert-butoxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-Boc-N6-methyllysine. In some cases, the unnatural amino acid is N6-acetyllysine. In some cases, the unnatural amino acid is pyrrolysine. In some cases, the unnatural amino acid is N6-trifluoroacetyllysine. In some cases, the unnatural amino acid is 2-amino-6-{[(benzyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(p-iodobenzyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is 2-amino-6-{[(p-nitrobenzyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-prolyllysine. In some cases, the unnatural amino acid is 2-amino-6-{[(cyclopentyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-(cyclopentanecarbonyl)lysine. In some cases, the unnatural amino acid is N6-(tetrahydrofuran-2-carbonyl)lysine. In some cases, the unnatural amino acid is N6-(3-ethynyltetrahydrofuran-2-carbonyl)lysine. In some cases, the unnatural amino acid is N6-((prop-2-yn-1-yloxy)carbonyl)lysine. In some cases, the unnatural amino acid is 2-amino-6-{[(2-azidocyclopentyloxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-((2-azidoethoxy)carbonyl)lysine. In some cases, the unnatural amino acid is 2-amino-6-{[(2-nitrobenzyloxy)carbonyl]amino}hexanoic acid.In some cases, the unnatural amino acid is 2-amino-6-{[(2-cyclooctynyl oxy)carbonyl]amino}hexanoic acid. In some cases, the unnatural amino acid is N6-(2-aminobut-3-enoyl)lysine. In some cases, the unnatural amino acid is 2-amino-6-((2-aminobut-3-enoyl)oxy)hexanoic acid. In some cases, the unnatural amino acid is N6-(allyloxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-(butenyl-4-oxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-(pentenyl-5-oxycarbonyl)lysine. In some cases, the unnatural amino acid is N6-((but-3-yn-1-yl oxy)carbonyl)-lysine. In some cases, the unnatural amino acid is N6-((penta-4-yn-1-yl oxy)carbonyl)-lysine. In some cases, the unnatural amino acid is N6-(thiazolidine-4-carbonyl)lysine. In some cases, the unnatural amino acid is 2-amino-8-oxononanoic acid. In some cases, the unnatural amino acid is 2-amino-8-oxooctanoic acid. In some cases, the unnatural amino acid is N6-(2-oxoacetyl)lysine.
[0104] In some cases, the unnatural amino acid is N6-propionyl lysine. In some cases, the unnatural amino acid is N6-butyryl lysine. In some cases, the unnatural amino acid is N6-(but-2-enoyl) lysine. In some cases, the unnatural amino acid is N6-((bicyclo[2.2.1]hept-5-en-2-yloxy)carbonyl) lysine. In some cases, the unnatural amino acid is N6-((spiro[2.3]hex-1-en-5-ylmethoxy)carbonyl) lysine. In some cases, the unnatural amino acid is N6-((((4-(1-(trifluoromethyl)cycloprop-2-en-1-yl)benzyl)oxy)carbonyl) lysine. In some cases, the unnatural amino acid is N6-((bicyclo[2.2.1]hept-5-en-2-ylmethoxy)carbonyl) lysine. In some cases, the unnatural amino acid is cysteinyl lysine. In some cases, the unnatural amino acid is N6-((1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl) lysine. In some cases, the unnatural amino acid is N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl) lysine. In some cases, the unnatural amino acid is N6-((3-(3-methyl-3H-diazirin-3-yl)propoxy)carbonyl) lysine. In some cases, the unnatural amino acid is N6-((methanitrobenyloxy)N6-methylcarbonyl) lysine. In some cases, the unnatural amino acid is N6-((bicyclo[6.1.0]non-4-in-9-ylmethoxy)carbonyl)-lysine. In some cases, the unnatural amino acid is N6-((cyclohept-3-en-1-yloxy)carbonyl)-L-lysine.
[0105] In some embodiments, the unnatural amino acid is incorporated into the protein by an unnatural codon that includes an unnatural nucleotide.
[0106] In some cases, the incorporation of non-natural amino acids into proteins is mediated by pairs of orthogonal modified synthetases / tRNAs. Such orthogonal pairs include natural or mutant synthetases that can charge a specific non-natural amino acid to a non-natural tRNA, while often minimizing a) the charging of other endogenous amino acids or alternative non-natural amino acids to the non-natural tRNA and b) the charging of any other (including endogenous) tRNAs. Such orthogonal pairs include tRNAs that can be charged by a synthetase while avoiding the charging of other endogenous amino acids by endogenous synthetases. In some embodiments, such pairs are identified from various organisms such as bacteria, yeast, archaea, or human sources. In some embodiments, the orthogonal synthetase / tRNA pair comprises components from a single organism. In some embodiments, the orthogonal synthetase / tRNA pair comprises components from two different organisms. In some embodiments, the orthogonal synthetase / tRNA pair comprises components that promote the translation of different amino acids prior to modification. In some embodiments, the orthogonal synthetase is a modified alanine synthetase. In some embodiments, the orthogonal synthetase is a modified arginine synthetase. In some embodiments, the orthogonal synthetase is a modified asparagine synthetase. In some embodiments, the orthogonal synthetase is a modified aspartic acid synthetase. In some embodiments, the orthogonal synthetase is a modified cysteine synthetase. In some embodiments, the orthogonal synthetase is a modified glutamine synthetase. In some embodiments, the orthogonal synthetase is a modified glutamic acid synthetase. In some embodiments, the orthogonal synthetase is a modified alanine glycine. In some embodiments, the orthogonal synthetase is a modified histidine synthetase. In some embodiments, the orthogonal synthetase is a modified leucine synthetase. In some embodiments, the orthogonal synthetase is a modified isoleucine synthetase. In some embodiments, the orthogonal synthetase is a modified lysine synthetase.In some embodiments, the orthogonal synthetase is a modified methionine synthetase. In some embodiments, the orthogonal synthetase is a modified phenylalanine synthetase. In some embodiments, the orthogonal synthetase is a modified proline synthetase. In some embodiments, the orthogonal synthetase is a modified serine synthetase. In some embodiments, the orthogonal synthetase is a modified threonine synthetase. In some embodiments, the orthogonal synthetase is a modified tryptophan synthetase. In some embodiments, the orthogonal synthetase is a modified tyrosine synthetase. In some embodiments, the orthogonal synthetase is a modified valine synthetase. In some embodiments, the orthogonal synthetase is a modified phosphoserine synthetase. In some embodiments, the orthogonal tRNA is a modified alanine tRNA. In some embodiments, the orthogonal tRNA is a modified arginine tRNA. In some embodiments, the orthogonal tRNA is a modified asparagine tRNA. In some embodiments, the orthogonal tRNA is a modified aspartic acid tRNA. In some embodiments, the orthogonal tRNA is a modified cysteine tRNA. In some embodiments, the orthogonal tRNA is a modified glutamine tRNA. In some embodiments, the orthogonal tRNA is a modified glutamic acid tRNA. In some embodiments, the orthogonal tRNA is a modified alanine glycine. In some embodiments, the orthogonal tRNA is a modified histidine tRNA. In some embodiments, the orthogonal tRNA is a modified leucine tRNA. In some embodiments, the orthogonal tRNA is a modified isoleucine tRNA. In some embodiments, the orthogonal tRNA is a modified lysine tRNA. In some embodiments, the orthogonal tRNA is a modified methionine tRNA. In some embodiments, the orthogonal tRNA is a modified phenylalanine tRNA. In some embodiments, the orthogonal tRNA is a modified proline tRNA.In some embodiments, the orthogonal tRNA is a modified serine tRNA. In some embodiments, the orthogonal tRNA is a modified threonine tRNA. In some embodiments, the orthogonal tRNA is a modified tryptophan tRNA. In some embodiments, the orthogonal tRNA is a modified tyrosine tRNA. In some embodiments, the orthogonal tRNA is a modified valine tRNA. In some embodiments, the orthogonal tRNA is a modified phosphoserine tRNA.
[0107] In some embodiments, the unnatural amino acid is incorporated into the protein by an aminoacyl (aaRS or RS)-tRNA synthetase-tRNA pair. Exemplary aaRS-tRNA pairs include, but are not limited to, the Methanococcus jannaschii (Mj-Tyr) aaRS / tRNA pair, the E. coli TyrRS (Ec-Tyr) / B. stearothermophilus tRNA CUA pair, the E. coli LeuRS (Ec-Leu) / B. stearothermophilus tRNA CUA pair, and the pyrrolysyl-tRNA pair. In some cases, the unnatural amino acid is incorporated into the protein by the Mj-TyrRS / tRNA pair. Exemplary unnatural amino acids (UAAs) that can be incorporated by the Mj-TyrRS / tRNA pair include, but are not limited to, para-substituted phenylalanine derivatives such as p-aminophenylalanine and p-methoxyphenylalanine; meta-substituted tyrosine derivatives such as 3-aminotyrosine, 3-nitrotyrosine, 3,4-dihydroxyphenylalanine, and 3-iodotyrosine; phenylselenocysteine; p-boronophenylalanine; and o-nitrobenzyltyrosine.
[0108] In some cases, non-natural amino acids are incorporated into proteins by the Ec-Tyr / tRNACUA or Ec-Leu / tRNACUA pair. Exemplary UAAs that can be incorporated by the Ec-Tyr / tRNACUA or Ec-Leu / tRNACUA pair include, but are not limited to, phenylalanine derivatives containing benzophenone, ketone, iodide, or azide substitutions; O-propargyl tyrosine; α-aminocaprylic acid, O-methyl tyrosine, O-nitrobenzyl cysteine; and 3-(naphthalen-2-ylamino)-2-amino-propanoic acid.
[0109] In some cases, non-natural amino acids are incorporated into proteins by the pyrrolysyl-tRNA pair. In some cases, PylRS is obtained from archaeal species such as methanogenic bacteria. In some cases, PylRS is obtained from Methanosarcina barkeri, Methanosarcina mazei, or Methanosarcina acetivorans. Exemplary UAAs that can be incorporated by the pyrrolysyl-tRNA pair include, but are not limited to, amide and carbamate substituted lysines, such as 2-amino-6-((R)-tetrahydrofuran-2-carboxamide)hexanoic acid, N-ε-D-prolyl-L-lysine, and N-ε-cyclopentyloxycarbonyl-L-lysine; N-ε-acryloyl-L-lysine; N-ε-[(1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl]-L-lysine; and N-ε-(1-methylcycloprop-2-ene-carboxamide)lysine.
[0110] In some cases, non-natural amino acids are incorporated into the proteins described herein by the synthetases disclosed in U.S. Patent No. 9,988,619 and U.S. Patent No. 9,938,516. Exemplary UAAs that can be incorporated by such synthetases include para-methylazido-L-phenylalanine, aralkyl, heterocyclyl, heteroaralkyl non-natural amino acids, and the like. In some embodiments, such UAAs include pyridyl, pyrazinyl, pyrazolyl, triazolyl, oxazolyl, thiazolyl, thiophenyl, or other heterocycles. Such amino acids in some embodiments include other chemical groups that can conjugate to coupling partners such as azide, tetrazine, or a water-soluble moiety. In some embodiments, such synthetases are expressed and used to incorporate UAAs into proteins in vivo. In some embodiments, such synthetases are used to incorporate UAAs into proteins using a cell-free translation system such as a reconstituted system of cell lysates or purified components. In a cell-free system or in a separate reaction beforehand, tRNA can be charged with a non-natural amino acid (such that the charged tRNA is added directly to a system containing ribosomes, mRNA, and other components, obviating the need to add a synthetase or a construct encoding a synthetase to the system).
[0111] Systems for in vitro translation are described, for example, in Zeenko et al., RNA 14:593-602 (2008); Spirin, Trends Biotechnol. 2004:538-545 (2004); and Endo et al., Curr. Opin. Biotechnol. 17:373-380 (2006). The systems can be made from cell lysates (e.g., extracts) or reconstituted from purified components. The systems can include one or more translation initiation factors; ATP; and one or more translation termination factors, in addition to ribosomes, tRNA, and other components described herein. In some embodiments, the systems further include one or more molecular chaperones, which can assist in the folding of nascent polypeptides during and / or after translation.
[0112] In some cases, non-natural amino acids are incorporated into the proteins described herein by naturally occurring synthetases. In some embodiments, non-natural amino acids are incorporated into proteins by organisms that are auxotrophic for one or more amino acids. In some embodiments, the synthetase corresponding to the auxotrophic amino acid can charge the corresponding tRNA with the non-natural amino acid. In some embodiments, the non-natural amino acid is selenocysteine or a derivative thereof. In some embodiments, the non-natural amino acid is selenomethionine or a derivative thereof. In some embodiments, the non-natural amino acid is an aromatic amino acid, and the aromatic amino acid includes aryl halides such as iodide. In embodiments, the non-natural amino acid is structurally similar to an auxotrophic amino acid.
[0113] In some cases, the non-natural amino acid includes the non-natural amino acid shown in FIG. 4A.
[0114] In some cases, the unnatural amino acid includes a lysine or phenylalanine derivative or analog. In some cases, the unnatural amino acid includes a lysine derivative or lysine analog. In some cases, the unnatural amino acid includes pyrrolysine (Pyl). In some cases, the unnatural amino acid includes a phenylalanine derivative or phenylalanine analog. In some cases, the unnatural amino acid is an unnatural amino acid described in "Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool" by Wan et al., Biocheim Biophys Aceta 1844(6):1059 - 4070(2014). In some cases, the unnatural amino acid includes the unnatural amino acids shown in FIGS. 4B and 4C.
[0115] In some embodiments, the unnatural amino acid includes the unnatural amino acids shown in FIGS. 4D - 4G. (Adopted from Table 1 of Dumas et al., Chemical Science 2015, 6, 50 - 69).
[0116] In some embodiments, the unnatural amino acids incorporated into the proteins described herein are disclosed in U.S. Patent No. 9,840,493; U.S. Patent No. 9,682,934; U.S. Patent Application Publication No. 2017 / 0260137; U.S. Patent No. 9,938,516; or U.S. Patent Application Publication No. 2018 / 0086734. Exemplary UAAs that can be incorporated by such synthetases include para-methylazido-L-phenylalanine, aralkyl, heterocyclyl, and heteroaralkyl, as well as unnatural amino acids of lysine derivatives. In some embodiments, such UAAs include pyridyl, pyrazinyl, pyrazolyl, triazolyl, oxazolyl, thiazolyl, thiophenyl, or other heterocycles. Such amino acids in some embodiments include other chemical groups that can bind to coupling partners such as azide, tetrazine, or a water-soluble moiety. In some embodiments, the UAA includes an azide bonded to an aromatic moiety via an alkyl linker. In some embodiments, the alkyl linker is a C1-C10 linker. In some embodiments, the UAA includes a tetrazine bonded to an aromatic moiety via an alkyl linker. In some embodiments, the UAA includes a tetrazine bonded to an aromatic moiety via an amino group. In some embodiments, the UAA includes a tetrazine bonded to an aromatic moiety via an alkylamino group. In some embodiments, the UAA includes an azide bonded to the terminal nitrogen of an amino acid side chain (e.g., N6 of a lysine derivative, or N5, N4, or N3 of a derivative containing a shorter alkyl side chain) via an alkyl chain. In some embodiments, the UAA includes a tetrazine bonded to the terminal nitrogen of an amino acid side chain via an alkyl chain. In some embodiments, the UAA includes an azide or tetrazine bonded to an amide via an alkyl linker. In some embodiments, the UAA is an azide- or tetrazine-containing carbamate or an amide of 3-aminoalanine, serine, lysine, or derivatives thereof. In some embodiments, such UAAs are incorporated into proteins in vivo. In some embodiments, such UAAs are incorporated into cell-free system proteins.
[0117] Cell type In some embodiments, many types of cells / microorganisms are used for, e.g., transformation or genetic engineering. In some embodiments, the cells are eukaryotic cells. In some cases, the cells are eukaryotic cells such as cultured animal, plant, or human cells. In additional cases, the cells are present in organisms such as plants or animals.
[0118] In some embodiments, the engineered microorganisms are single-celled organisms and can often divide and proliferate. The microorganisms can include one or more of the following characteristics: aerobic, anaerobic, filamentous, non-filamentous, haploid, diploid, auxotrophic and / or non-auxotrophic. In certain embodiments, the engineered microorganisms are non-prokaryotic microorganisms. In some embodiments, the engineered microorganisms are eukaryotic microorganisms (e.g., yeast, fungi, amoeba). In some embodiments, the engineered microorganisms are fungi. In some embodiments, the engineered organisms are yeast.
[0119] Any suitable yeast can be selected as a source of host microorganisms, engineered microorganisms, genetically modified organisms, or heterologous or modified polynucleotides. Yeasts include, but are not limited to, Yarrowia yeasts (e.g., Y. lipolytica (formerly classified as Candida lipolytica)), Candida yeasts (e.g., C. revkaufi, C. viswanathii, C. pulcherrima, C. tropicalis, C. utilis), Rhodotorula yeasts (e.g., R. glutinus, R. graminis), Rhodosporidium yeasts (e.g., R. toruloides), Saccharomyces yeasts (e.g., S. cerevisiae, S. bayanus, S. pastorianus, S. carlsbergensis), Cryptococcus yeasts, Trichosporon yeasts (e.g., T. pullans, T. cutaneum), Pichia yeasts (e.g., P. pastoris), and Lipomyces yeasts (e.g., L. starkeyii, L. lipoferus). In some embodiments, suitable yeasts are yeasts of the genus Arachniotus, Aspergillus, Aureobasidium, Auxarthron, Blastomyces, Candida, Chrysosporuim, Chrysosporuim, Debaryomyces, Coccidiodes, Cryptococcus, Gymnoascus, Hansenula, Histoplasma, Issatchenkia, Kluyveromyces, Lipomyces, Lssatchenkia, Microsporum, Myxotrichum, Myxozyma, Oidiodendron, Pachysolen, Penicillium, Pichia, Rhodosporidium, Rhodotorula, Rhodotorula, Saccharomyces, Schizosaccharomyces, Scopulariopsis, Sepedonium, Trichosporon, or Yarrowia. In some embodiments, suitable yeasts are Arachniotus flavoluteus, Aspergillusflavus, Aspergillus fumigatus, Aspergillus niger, Aureobasidium pullulans, Auxarthron thaxteri, Blastomyces dermatitidis, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, Candida viswanathii, Candida xestobii, Chrysosporuim keratinophilum, Coccidiodes immitis, Cryptococcus albidus var. diffluens, Cryptococcus laurentii, Cryptococcus neofomans, Debaryomyces hansenii, Gymnoascus dugwayensis, Hansenula anomala, Histoplasma capsulatum, Issatchenkia occidentalis, Isstachenkia orientalis, Kluyveromyces lactis, Kluyveromyces marxianus, Kluyveromyces thermotolerans, Kluyveromyces waltii, Lipomyces lipoferus, Lipomyces starkeyii, Microsporum gypseum, Myxotrichum deflexum, Oidiodendron echinulatum, Pachysolen tannophilis, Penicillium notatum, Pichia anomala, Pichia pastoris, PichiaYeast of the species of Pichia stipitis, Rhodosporidium toruloides, Rhodotorula glutinus, Rhodotorula graminis, Saccharomyces cerevisiae, Saccharomyces kluyveri, Schizosaccharomyces pombe, Scopulariopsis acremonium, Sepedonium chrysospermum, Trichosporon cutaneum, Trichosporon pullans, Yarrowia lipolytica, or Yarrowia lipolytica (formerly classified as Candida lipolytica). In some embodiments, the yeast is a Y. lipolytica strain including, but not limited to, ATCC20362, ATCC8862, ATCC18944, ATCC20228, ATCC76982, and LGAM S(7)1 strain (Papanikolaou S. and Aggelis G., Bioresor.Technol. 82(1):43-9(2002)). In certain embodiments, the yeast is a Candida species (i.e., Candida genus) yeast. Any suitable Candida species may be used and / or may be genetically modified for the production of fatty dicarboxylic acids (e.g., octanedioic acid, decanedioic acid, dodecanedioic acid, tetradecanedioic acid, hexadecanedioic acid, octadecanedioic acid, eicosanedioic acid). In some embodiments, suitable Candida species include, but are not limited to, Candida albicans, Candida dubliniensis, Candida famata, Candida glabrata, Candida guilliermondii, Candida kefyr, Candida krusei, Candida lambica, Candida lipolytica, Candida lustitaniae, Candida parapsilosis, Candida pulcherrima, Candida revkaufi, Candida rugosa, Candida tropicalis, Candida utilis, CandidaExamples include viswanathii, Candida xestobii, and any other Candida species yeast described herein. Non-limiting examples of Candida species strains include, but are not limited to, sAA001 (ATCC20336), sAA002 (ATCC20913), sAA003 (ATCC20962), sAA496 (U.S. Patent Application Publication No. 2012 / 0077252), sAA106 (U.S. Patent Application Publication No. 2012 / 0077252), SU-2 (ura3- / ura3-), H5343 (beta oxidation block; U.S. Patent No. 5648247). Any suitable strain yeast derived from Candida species can be used as a parent strain for genetic recombination.
[0120] Genus, species, and strains of yeast are often very closely related genetically, which can make it difficult to distinguish, classify, and / or name them. In some cases, it may be difficult to distinguish, classify, and / or name strains of C. lipolytica and Y. lipolytica, and in some cases, they may be considered the same organism. In some cases, it may be difficult to distinguish, classify, and / or name various strains of C. tropicalis and C. viswanathii (see, e.g., Arie et al., J. Gen. Appl. Microbiol., 46, 257-262 (2000)). Some C. tropicalis and C. viswanathii strains obtained from ATCC and other commercial or academic sources may be considered equivalent and equally suitable for the embodiments described herein. In some embodiments, some parent strains of C. tropicalis and C. viswanathii are considered to differ only in name.
[0121] Any suitable fungus can be selected as a source of host microorganisms, engineered microorganisms, or heterologous polynucleotides. Non-limiting examples of fungi include, but are not limited to, Aspergillus fungi (e.g., A. parasiticus, A. nidulans), Thraustochytrium fungi, Schizochytrium fungi, and Rhizopus fungi (e.g., R. arrhizus, R. oryzae, R. nigricans). In some embodiments, the fungus is an A. parasiticus strain, including, but not limited to, the ATCC 24690 strain, and in certain embodiments, the fungus is an A. nidulans strain, including, but not limited to, the ATCC 38163 strain.
[0122] Non-microbiologically derived cells can be used as a source of host microorganisms, engineered microorganisms, or heterologous polynucleotides. Examples of such cells include, but are not limited to, insect cells (e.g., Drosophila (e.g., D. melanogaster), Spodoptera (e.g., S. frugiperda Sf9 or Sf21 cells) and Trichoplusa (e.g., High-Five cells); nematode cells (e.g., C. elegans cells); avian cells; amphibian cells (e.g., Xenopus laevis cells); reptilian cells; mammalian cells (e.g., NIH3T3, 293, CHO, COS, VERO, C127, BHK, Per-C6, Bowes melanoma and HeLa cells); and plant cells (e.g., Arabidopsis thaliana, Nicotania tabacum, Cuphea acinifolia, Cuphea aequipetala, Cuphea angustifolia, Cuphea appendiculata, Cuphea avigera, Cuphea avigera var. pulcherrima, Cuphea axilliflora, Cuphea bahiensis, Cuphea baillonis, Cuphea brachypoda, Cuphea bustamanta, Cuphea calcarata, Cuphea calophylla, Cuphea calophylla subsp.mesostemon, Cuphea carthagenensis, Cuphea circaeoides, Cuphea confertiflora, Cuphea cordata, Cuphea crassiflora, Cuphea cyanea, Cuphea decandra, Cuphea denticulata, Cuphea disperma, Cuphea epilobiifolia, Cuphea ericoides, Cuphea flava, Cuphea flavisetula, Cuphea fuchsiifolia, Cuphea gaumeri, Cuphea glutinosa, Cuphea heterophylla, Cuphea hookeriana, Cuphea hyssopifolia (Mexican - heather), Cuphea hyssopoides, Cuphea ignea, Cuphea ingrata, Cuphea jorullensis, Cuphea lanceolata, Cuphea linarioides, Cuphea llavea, Cuphea lophostoma, Cuphea lutea, Cuphea lutescens, Cuphea melanium, Cuphea melvilla, Cuphea micrantha, Cuphea micropetala, Cuphea mimuloides, Cuphea nitidula, Cuphea palustris, Cuphea parsonsia, Cuphea pascuorum, Cuphea paucipetala, Cuphea procumbens, Cuphea pseudosilene, Cuphea pseudovaccinium, Cuphea pulchra, Cuphea racemosa, Cuphea repens, Cuphea salicifolia, Cuphea salvadorensis, Cuphea schumannii, Cuphea sessiliflora, Cuphea sessilifolia, Cuphea setosa, Cuphea spectabilis, Cuphea spermacoce, Cuphea splendida, Cuphea splendida var.Examples include viridiflava, Cuphea strigulosa, Cuphea subuligera, Cuphea teleandra, Cuphea thymoides, Cuphea tolucana, Cuphea urens, Cuphea utriculosa, Cuphea viscosissima, Cuphea watsoniana, Cuphea wrightii, Cuphea lanceolata).
[0123] Microorganisms or cells used as a source of a host organism or heterologous polynucleotide are commercially available. The microorganisms and cells described herein, as well as other suitable microorganisms and cells, are available, for example, from Invitrogen Corporation (Carlsbad, CA), American Type Culture Collection (Manassas, Virginia), and Agricultural Research Culture Collection (NRRL; Peoria, Illinois). The host microorganism and the engineered microorganism can be provided in any suitable form. For example, such microorganisms can be provided in liquid culture or solid culture (e.g., agar-based medium), which can be a primary culture or can have been passaged one or more times (e.g., diluted and cultured). The microorganisms can also be provided in frozen form or dried form (e.g., lyophilized). The microorganisms can be provided at any suitable concentration.
[0124] Nucleic Acid Reagents and Tools The nucleotide and / or nucleic acid reagents (or polynucleotides) for use with the methods, cells, or engineered microorganisms described herein include one or more ORFs, with or without unnatural nucleotides. The ORFs can be derived from any suitable source, sometimes genomic DNA, mRNA, reverse transcribed RNA or complementary DNA (cDNA), or a nucleic acid library containing one or more of the foregoing, and can be from any species of organism containing the nucleic acid sequence of interest, the protein of interest, or the activity of interest. Non-limiting examples of organisms from which ORFs can be obtained include, for example, bacteria, yeast, fungi, humans, insects, nematodes, cows, horses, dogs, cats, rats or mice. In some embodiments, the nucleotides and / or nucleic acid reagents or other reagents described herein are isolated or purified. ORFs containing unnatural nucleotides can be made by published in vitro methods. In some cases, the nucleotide or nucleic acid reagent contains unnatural nucleobases.
[0125] The nucleic acid reagent may, in combination with the ORF, contain a nucleotide sequence adjacent to the ORF that encodes an amino acid tag. The nucleotide sequence encoding the tag is located 3' and / or 5' of the ORF in the nucleic acid reagent, thereby encoding the tag at the C-terminus or N-terminus of the protein or peptide encoded by the ORF. Any tag that does not inactivate in vitro transcription and / or translation may be utilized and may be appropriately selected by the skilled person. The tag may facilitate the isolation and / or purification of the desired ORF product from the culture or fermentation medium. In some cases, a library of nucleic acid reagents is used with the methods and compositions described herein. For example, a library of at least 100, 1000, 2000, 5000, 10,000, or more than 50,000 unique polynucleotides is present in the library, and each polynucleotide contains at least one unnatural nucleobase.
[0126] Nucleic acids or nucleic acid reagents, whether containing or not containing unnatural nucleotides, may contain certain elements, such as regulatory elements often selected according to the purpose of use of the nucleic acid. Any of the following elements may be included in or excluded from the nucleic acid reagent. For example, the nucleic acid reagent may contain one or more or all of the following nucleotide elements: one or more promoter elements, one or more 5' untranslated regions (5' UTRs), one or more regions ( "insertion elements") into which the target nucleotide sequence can be inserted, one or more target nucleotide sequences, one or more 3' untranslated regions (3' UTRs), and one or more selection elements. The nucleic acid reagent may be provided with one or more of such elements, and other elements may be inserted into the nucleic acid before the nucleic acid is introduced into the desired organism. In some embodiments, the provided nucleic acid reagent contains a promoter, 5' UTR, any 3' UTR, and an insertion element into which the target nucleotide sequence is inserted (i.e., cloned) into the nucleic acid reagent. In certain embodiments, the provided nucleic acid reagent contains a promoter, an insertion element, and any 3' UTR, and the 5' UTR / target nucleotide sequence is inserted together with any 3' UTR. The elements may be arranged in any order suitable for expression in the selected expression system (e.g., expression in the selected organism, or expression in a cell-free system, for example), and in some embodiments, the nucleic acid reagent contains the following elements in the 5'→3' direction: (1) a promoter element, 5' UTR, and insertion element; (2) a promoter element, 5' UTR, and target nucleotide sequence; (3) a promoter element, 5' UTR, insertion element and 3' UTR; and (4) a promoter element, 5' UTR, target nucleotide sequence and 3' UTR. In some embodiments, the UTR may be fully natural or optimized to alter or increase the transcription or translation of the ORF, including unnatural nucleotides.
[0127] The nucleic acids (e.g., mRNA) containing nucleobases described herein include 5’UTRs and / or 3’UTRs that enhance mRNA stability in some cases in vivo (e.g., in eukaryotic cells or eukaryotic SSOs). In some examples, the 5’ or 3’UTR or both are engineered to reduce mRNA degradation or decay in vivo. Non-limiting examples of 5’ and 3’UTRs that enhance mRNA stability in the eukaryotic systems disclosed herein are the CS2 3’ and 5’UTRs. In some embodiments, the mRNA is modified to reduce the rate of removal of the poly(A) tail of the mRNA as compared to the mRNA containing nucleobases described herein that are otherwise unmodified. In some embodiments, cis-acting AU-rich elements (AREs) are blocked from intracellular and extracellular signaling that promotes mRNA decay. In some embodiments, premature stop codons in the mRNA are removed to reduce nonsense-mediated decay (NMD) of the mRNA.
[0128] In some cases, the 5’ and / or 3’UTRs directly or indirectly increase the translation of the mRNA into a polypeptide. Non-limiting examples of ways in which the 5’UTR or 3’UTR directly affects the translation of the mRNA into a polypeptide include the recruitment of RNA-binding proteins that bind to 5’ or 3’ cis-elements and result in the recruitment of ribosomes or effector proteins (e.g., mRNA deadenylase, decapping enzyme). Non-limiting examples of ways in which the 5’UTR or 3’UTR indirectly affects the translation of the mRNA into a polypeptide include the formation of 5’ and 3’UTR secondary structures and mRNA intracellular localization that block or enhance the binding of RNA-binding proteins to the 5’ or 3’UTR region.
[0129] In some embodiments, the 5’UTR and / or 3’UTR increases the translation efficiency of the mRNA in vitro or in vivo as compared to the translation efficiency of an mRNA containing unmodified nucleobases. In some embodiments, the translation efficiency is increased by engineering the mRNA to reduce skipping of the selected AUG (start codon) by the ribosome during scanning. In some embodiments, the mRNA includes a sequence element that improves start codon recognition, such as a Kozak sequence or a variant thereof. In some embodiments, the 5’UTR of the mRNA is engineered to reduce the overall guanine-cytosine (GC) content.
[0130] In some embodiments, the formation of secondary structure (e.g., RNA G-quadruplex structure, RG4) in the mRNA involving an AUG start codon within the 5’UTR is reduced, thereby increasing the efficiency of translation therefrom. In some embodiments, the 5’UTR is engineered to have a negative folding free energy (ΔG) as compared to an unengineered mRNA. In some embodiments, the ΔG is up to -40, -41, -42, -43, -44, -45, -46, -47, -48, -49, -50, -51, -52, -53, -54, -55, -56, -57, -58, -59 or -60. In some embodiments, the mRNA is chemically modified with the 5’UTR or 3’UTR to promote translation efficiency. In some embodiments, the chemical modification is N 6-Methyladenosine. In in vitro systems (e.g., engineered eukaryotic cells or semi-synthetic organisms), overexpression of eIF4A, a subunit of the eIF4F complex that cooperates with eIF3B and eIF4H to promote unwinding of RNA secondary structure, increases the translation efficiency of mRNA. In some embodiments, knockout or knockdown of a stabilizing protein that promotes mRNA secondary structure formation (e.g., fragile X mental retardation protein (FMRP)) reduces secondary structure formation, thereby increasing the translation efficiency of mRNA. In some embodiments, a trans-acting agent (e.g., a small molecule of RNA, a protein) is introduced into a cell (e.g., a eukaryotic cell) to promote translation of mRNA.
[0131] In some examples, the 5’UTR and / or 3’UTR promote the intracellular localization of mRNA, thereby promoting translation of mRNA in vivo. In some embodiments, 3’ or 5’UTR cis-acting elements such as mRNA zip codes are modified such that binding of the mRNA zip code by a zip code-binding protein (e.g., Staufen) is inhibited or enhanced, thereby increasing the translation efficiency of mRNA.
[0132] Nucleic acid reagents, such as expression cassettes and / or expression vectors (e.g., for expressing a heterologous tRNA synthetase), can contain various regulatory elements including promoters, enhancers, translation initiation sequences, transcription termination sequences, and other elements. A "promoter" is generally a DNA sequence that functions when located at a relatively fixed position with respect to the transcription start site. For example, a promoter can be upstream of a nucleoside triphosphate transporter nucleic acid segment. A "promoter" contains the core elements necessary for the basic interaction of RNA polymerase and transcription factors, and may also include upstream elements and response elements. An "enhancer" generally refers to a DNA sequence that functions at a non-fixed distance from the transcription start site and can be either 5' or 3' to the transcription unit. Additionally, enhancers can be present within introns and within the coding sequences themselves. They are usually between 10 and 300 in length and function in cis. Enhancers function to increase transcription from nearby promoters. Similar to promoters, enhancers often contain response elements that mediate the regulation of transcription. Enhancers often determine the regulation of expression and can be used to alter or optimize ORF expression, such as those that are completely natural or contain non-natural nucleotides in an ORF.
[0133] As described above, the nucleic acid reagent may also include one or more 5’UTRs and one or more 3’UTRs. For example, expression vectors used in eukaryotic host cells (e.g., yeast, fungi, insects, plants, animals, humans or nucleated cells) and prokaryotic host cells (e.g., viruses, bacteria) may include sequences that signal for the termination of transcription, which can affect the expression of mRNA. These regions can be transcribed as polyadenylation segments of the untranslated portion of the mRNA encoding the tissue factor protein. The 3’ untranslated region also includes the transcription termination site. In some preferred embodiments, the transcription unit includes a polyadenylation region. One advantage of this region is that it increases the likelihood that the transcribed unit will be processed and transported like mRNA. The identification and use of polyadenylation signals in expression constructs are well established. In some preferred embodiments, homologous polyadenylation signals can be used in transgene constructs.
[0134] The 5’UTR may contain one or more elements endogenous to the nucleotide sequence from which it is derived and sometimes one or more exogenous elements. The 5’UTR may be derived from any suitable nucleic acid such as genomic DNA, plasmid DNA, RNA or mRNA, for example, from any suitable organism (e.g., virus, bacterium, yeast, fungus, plant, insect or mammal). A skilled person may select suitable elements for the 5’UTR based on the selected expression system (e.g., expression in a selected organism or, for example, expression in a cell-free system). The 5’UTR may sometimes contain one or more of the following elements known to a skilled person: enhancer sequences (e.g., for transcription or translation), transcription start sites, transcription factor binding sites, translation regulatory sites, translation start sites, translation factor binding sites, accessory protein binding sites, feedback regulator binding sites, Pribnow box, TATA box, -35 element, E-box (helix-loop-helix binding element), ribosome binding sites, replicons, internal ribosome entry sites (IRES), silencer elements, etc. In some embodiments, the promoter element may be separated such that all 5’UTR elements necessary for appropriate conditional regulation are contained within the promoter element fragment or within a functional partial sequence of the promoter element fragment.
[0135] The 5’UTR in the nucleic acid reagent may contain a translational enhancer nucleotide sequence. The translational enhancer nucleotide sequence is often located between the promoter of the nucleic acid reagent and the target nucleotide sequence. The translational enhancer sequence often binds to the ribosome and may be an 18S rRNA-binding ribonucleotide sequence (i.e., a 40S ribosome-binding sequence) or may be an internal ribosome entry site (IRES). The IRES generally forms an RNA scaffold with an accurately arranged RNA tertiary structure that contacts the 40S ribosome subunit through a number of specific intermolecular interactions. Examples of ribosome enhancer sequences are known and can be identified by those skilled in the art (e.g., Mignone et al., Nucleic Acids Research 33:D141-D146 (2005); Paulous et al., Nucleic Acids Research 31:722-733 (2003); Akbergenov et al., Nucleic Acids Research 32:239-247 (2004); Mignone et al., Genome Biology 3(3):reviews0004.1-0001.10 (2002); Gallie, Nucleic Acids Research 30:3401-3411 (2002); Shaloiko et al., DOI:10.1002 / bit.20267; and Gallie et al., Nucleic Acids Research 15:3257-3273 (1987)).
[0136] The translation enhancer sequence may be a eukaryotic sequence such as a Kozak consensus sequence or other sequences (e.g., the hydra polyp sequence, GenBank accession number U07128). The translation enhancer sequence may be a prokaryotic sequence such as a Shine-Dalgarno consensus sequence. In certain embodiments, the translation enhancer sequence is a viral nucleotide sequence. The translation enhancer sequence may be derived from, for example, the 5’ UTR of plant viruses such as tobacco mosaic virus (TMV), alfalfa mosaic virus (AMV); tobacco etch virus (ETV); potato virus Y (PVY); turnip mosaic (poty) virus and bean pod mottle virus. In certain embodiments, an omega sequence of approximately 67 bases in length from TMV is included in the nucleic acid reagent as a translation enhancer sequence (e.g., lacking guanosine nucleotides and including a 25 nucleotide long poly(CAA) central region).
[0137] The 3’UTR may contain one or more elements endogenous to the nucleotide sequence from which it is derived and may also contain one or more exogenous elements. The 3’UTR may be derived from any suitable nucleic acid such as genomic DNA, plasmid DNA, RNA or mRNA, for example, from any suitable organism (e.g., virus, bacterium, yeast, fungus, plant, insect or mammal). A technician may select suitable elements for the 3’UTR based on the selected expression system (e.g., expression in the selected organism, etc.). The 3’UTR may contain one or more of the following elements known to the technician: transcriptional regulatory site, transcription start site, transcription termination site, transcription factor binding site, translational regulatory site, translation termination site, translation start site, translation factor binding site, ribosome binding site, replicon, enhancer element, silencer element and polyadenosine tail. The 3’UTR often contains a polyadenosine tail and may not contain it. When the polyadenosine tail is present, one or more adenosine moieties may be added or deleted (e.g., about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45 or about 50 adenosine moieties may be added or deleted).
[0138] In some embodiments, modifications of the 5’UTR and / or 3’UTR are used to alter (e.g., increase, add, decrease, or substantially eliminate) the activity of a promoter. The alteration of promoter activity can then result in a change in the activity (e.g., enzymatic activity) of a peptide, polypeptide, or protein by a change in the transcription of the nucleotide sequence of interest from an operably linked promoter element that includes the modified 5’ or 3’UTR. For example, a microorganism can be engineered by genetic modification to express a nucleic acid reagent that includes a modified 5’ or 3’UTR that can add a new activity (e.g., an activity not normally found in the host organism), or in certain embodiments, increase the expression of an existing activity by increasing transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest (e.g., a homologous or heterologous nucleotide sequence of interest). In some embodiments, a microorganism can be engineered by genetic modification to express a nucleic acid reagent that includes a modified 5’ or 3’UTR that can decrease the expression of an activity by decreasing or substantially eliminating transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest in certain embodiments.
[0139] Expression of a heterologous polypeptide such as a tRNA synthetase from an expression cassette or expression vector is controlled by any promoter that is expressible in a prokaryotic or eukaryotic cell. Promoter elements are usually required for DNA synthesis and / or RNA synthesis. A promoter element often includes a region of DNA that can promote transcription of a particular gene by providing a starting site for synthesis of RNA corresponding to that gene. Promoters are generally located near the genes they regulate, upstream of the gene (e.g., 5' of the gene), and in some embodiments, on the same DNA strand as the sense strand of the gene. In some embodiments, a promoter element may be isolated from a gene or organism and inserted in operative association with a polynucleotide sequence to allow for altered and / or regulated expression. A non-native promoter used for expression of a nucleic acid (e.g., a promoter not normally associated with a given nucleic acid sequence) is often referred to as a heterologous promoter. In certain embodiments, a heterologous promoter and / or 5' UTR can be inserted in operative association with a polynucleotide encoding a polypeptide having a desired activity as described herein. The terms "operatively linked" and "functionally associated" as used herein with respect to a promoter refer to the relationship between a coding sequence and a promoter element. A promoter is operatively linked to, or functionally associated with, a coding sequence if expression from the coding sequence via transcription is regulated or controlled by the promoter element. The terms "operatively linked" and "functionally associated" are used interchangeably herein with respect to a promoter element.
[0140] Promoters often interact with RNA polymerase. Polymerase is an enzyme that catalyzes nucleic acid synthesis using existing nucleic acid reagents. When the template is a DNA template, an RNA molecule is transcribed before the protein is synthesized. Enzymes having polymerase activity suitable for use in the methods of the present invention include any polymerase that is active in a selected system using a template selected for protein synthesis. In some embodiments, a promoter (e.g., a heterologous promoter), also referred to herein as a promoter element, can be operably linked to a nucleotide sequence or open reading frame (ORF). Transcription from the promoter element can catalyze the synthesis of RNA corresponding to the nucleotide sequence or ORF sequence operably linked to the promoter, which in turn results in the synthesis of the desired peptide, polypeptide, or protein.
[0141] Promoter elements may exhibit responsiveness to regulatory control. Promoter elements may also be regulated by a selective agent. That is, transcription from a promoter element can be turned on, off, upregulated, or downregulated in response to changes in environmental, nutritional, or internal conditions or signals (e.g., heat-inducible promoters, light-regulated promoters, feedback-regulated promoters, hormone-responsive promoters, tissue-specific promoters, oxygen- and pH-responsive promoters, promoters responsive to a selective agent (e.g., kanamycin), etc.). Promoters that are affected by environmental, nutritional, or internal signals are often affected by the influence of a signal (direct or indirect) that binds to or near the promoter and increases or decreases the expression of the target sequence under specific conditions. As with all methods disclosed herein, the inclusion of natural or modified promoters can be used to alter or optimize the expression of an ORF containing a fully natural ORF (e.g., aaRS) or unnatural nucleotides (e.g., mRNA or tRNA).
[0142] Non-limiting examples of selective agents or modulators that affect transcription from promoter elements used in the embodiments described herein include, but are not limited to, the following: (1) nucleic acid segments encoding products that confer resistance to otherwise toxic compounds (e.g., antibiotics); (2) nucleic acid segments encoding products that are otherwise lacking in the recipient cell (e.g., essential products, tRNA genes, auxotrophic markers); (3) nucleic acid segments encoding products that suppress the activity of gene products; (4) nucleic acid segments encoding products that can be readily identified (e.g., phenotypic markers such as antibiotics (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and cell surface proteins); (5) nucleic acid segments that bind to products that are otherwise harmful to cell survival and / or function; (6) nucleic acid segments that inhibit the activity of any of the nucleic acid segments described in (1)-(5) above (e.g., antisense oligonucleotides); (7) nucleic acid segments that bind to products that modify substrates (e.g., restriction endonucleases); (8) nucleic acid segments that can be used to isolate or identify desired molecules (e.g., specific protein binding sites); (9) nucleic acid segments encoding specific nucleotide sequences that cannot otherwise function (e.g., for PCR amplification of subpopulations of molecules); (10) nucleic acid segments that, when absent, directly or indirectly confer resistance or sensitivity to specific compounds; (11) nucleic acid segments encoding products that convert a compound that is toxic or relatively non-toxic in the recipient cell into a toxic compound (e.g., herpes simplex thymidine kinase, cytosine deaminase); (12) nucleic acid segments that inhibit the replication, segregation, or heritability of the nucleic acid molecules containing them; (13) nucleic acid segments encoding conditional replication functions, e.g., replication in a specific host or host cell line or under specific environmental conditions (e.g., temperature, nutrient conditions, etc.); and / or (14) nucleic acids encoding one or more mRNAs or tRNAs containing non-natural nucleotides.In some embodiments, a modulator or a selective agent may be added to alter the existing growth conditions to which the organism is subjected (e.g., growth in liquid culture, growth in a fermenter, growth on a solid nutrient plate, etc.).
[0143] In some embodiments, regulation of promoter elements can be used to alter (e.g., increase, add, decrease, or substantially eliminate) the activity of a peptide, polypeptide, or protein (e.g., enzymatic activity, etc.). For example, a microorganism can be engineered by genetic modification to express a nucleic acid reagent that can add a new activity (e.g., an activity not normally found in the host organism), or in certain embodiments, increase the expression of an existing activity by increasing transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest (e.g., a homologous or heterologous nucleotide sequence of interest). In some embodiments, a microorganism can be engineered by genetic modification to express a nucleic acid reagent that can decrease or substantially eliminate transcription from a homologous or heterologous promoter operably linked to a nucleotide sequence of interest, thereby decreasing the expression of the activity, in certain embodiments.
[0144] Heterologous proteins, such as nucleic acids encoding tRNA synthetases, may be inserted into or used in any suitable expression system. In some embodiments, the nucleic acid reagent is sometimes stably integrated into the host organism's chromosome, or in certain embodiments, the nucleic acid reagent may be a deletion of a portion of the host chromosome (e.g., in a genetically modified organism, the alteration of the host genome confers the ability to selectively or preferentially maintain the desired organism with the genetic modification). Such nucleic acid reagents (e.g., nucleic acids whose modified genome confers a selectable trait on the organism or genetically modified organisms) can be selected for their ability to induce the production of the desired protein or nucleic acid molecule. Optionally, the nucleic acid reagent may be modified so that the codons (i) encode the same amino acid using a tRNA different from that specified by the native sequence, or (ii) encode a different amino acid than usual or a non-natural amino acid (including a detectable labeled amino acid).
[0145] Recombinant expression is usefully achieved using an expression cassette, which can be part of a vector such as a plasmid. The vector can contain a promoter operably linked to the nucleic acid. The vector may also contain other elements necessary for transcription and translation, as described herein. The expression cassette, expression vector, and the sequences in the cassette or vector can be heterologous to the cells in which the unnatural nucleotides are contacted.
[0146] It is possible to generate various prokaryotic and eukaryotic expression vectors suitable for encoding and / or expressing heterologous proteins such as tRNA synthetase. Such expression vectors include, for example, pET, pET3d, pCR2.1, pBAD, pUC, and yeast vectors. The vectors can be used, for example, in various in vivo and in vitro situations. Non-limiting examples of prokaryotic promoters that can be used include SP6, T7, T5, tac, bla, trp, gal, lac, or maltose promoters. Non-limiting examples of eukaryotic promoters that can be used include constitutive promoters, such as viral promoters like CMV, SV40, and RSV promoters, as well as regulatable promoters, such as inducible or repressible promoters like the tet promoter, hsp70 promoter, and synthetic promoters regulated by CRE. Vectors for bacterial expression include pGEX-5X-3, and vectors for eukaryotic expression include pCIneo-CMV. Viral vectors that can be used include those related to lentivirus, adenovirus, adeno-associated virus, herpesvirus, vaccinia virus, poliovirus, AIDS virus, neurotrophic virus, Sindbis, and other viruses. Any viral family that shares the characteristics of these viruses and is suitable for use as a vector is also useful. Retroviral vectors that can be used include those described in Verma, American Society for Microbiology, pp. 229-232, Washington, (1985). For example, such retroviral vectors include murine Moloney leukemia virus, MMLV, and other retroviruses that express desirable characteristics. Usually, viral vectors include non-structural early genes, structural late genes, RNA polymerase III transcripts, inverted terminal repeats necessary for replication and capsid formation, and promoters that control transcription and replication of the viral genome.When designed as a vector, the virus usually has one or more initial genes removed, and a gene or gene / promoter cassette is inserted into the viral genome in place of the removed viral nucleic acid.
[0147] Cloning Any convenient cloning strategy known in the art can be utilized to incorporate elements such as ORFs into nucleic acid reagents. Elements can be inserted into a template independently of the inserted element using known methods such as, for example, (1) cutting the template at one or more existing restriction enzyme sites and ligating the element of interest, and (2) hybridizing oligonucleotide primers containing one or more appropriate restriction enzyme sites and adding restriction enzyme sites to the template by amplification by polymerase chain reaction (described in more detail herein). Other cloning strategies utilize one or more insertion sites present in or inserted into nucleic acid reagents, such as, for example, oligonucleotide primer hybridization sites for PCR and others described herein. In some embodiments, the cloning strategy may be combined with genetic manipulation such as recombination (e.g., recombination of a nucleic acid reagent having a nucleic acid sequence of interest into the genome of the organism to be modified, as further described herein). In some embodiments, the cloned ORF can generate (directly or indirectly) a modified or wild-type polymerase by manipulating a microorganism having one or more ORFs of interest, and the microorganism contains an altered activity of polymerase activity.
[0148] A nucleic acid can be specifically cleaved by contacting it with one or more specific cleavage agents. Specific cleavage agents often specifically cleave according to a specific nucleotide sequence at a specific site.Examples of enzyme-specific cleavage agents include, but are not limited to, endonucleases (e.g., DNase (e.g., DNase I, II); RNase (e.g., RNase E, F, H, P); Cleavase™ enzyme; Taq DNA polymerase; E. coli DNA polymerase I and eukaryotic structure-specific endonucleases; mouse FEN-1 endonuclease; type I, II or III restriction endonucleases, e.g., Acc I, Afl III, Alu I, Alw44 I, Apa I, Asn I, Ava I, Ava II, BamH I, Ban II, Bcl I, Bgl I, Bgl II, Bln I, BsaI, Bsm I, BsmBI, BssH II, BstE II, Cfo I, CIa I, Dde I, Dpn I, Dra I, EcIX I, EcoR I, EcoR I, EcoR II, EcoR V, Hae II, Hae II, Hind II, Hind III, Hpa I, Hpa II, Kpn I, Ksp I, Mlu I, MIuN I, Msp I, Nci I, Nco I, Nde I, Nde II, Nhe I, Not I, Nru I, Nsi I, Pst I, Pvu I, Pvu II, Rsa I, Sac I, Sal I, Sau3A I, Sca I, ScrF I, Sfi I, Sma I, Spe I, Sph I, Ssp I, Stu I, Sty I, Swa I, Taq I, Xba I, Xho I); glycosylases (e.g., uracil-DNA glycosylase (UDG), 3-methyladenine DNA glycosylase, 3-methyladenine DNA glycosylase II, pyrimidine hydrate-DNA glycosylase, FaPy-DNA glycosylase, thymine mismatch-DNA glycosylase, hypoxanthine-DNA glycosylase, 5-hydroxymethyluracil DNA glycosylase (HmUDG), 5-hydroxymethylcytosine DNA glycosylase, or 1,N6-etheno-adenine DNA glycosylase); exonucleases (e.g., exonuclease III); ribozymes; and DNAzymes. The sample nucleic acid may be treated with chemicals or synthesized using modified nucleotides, and the modified nucleic acid may be cleaved.In a non-limiting example, the sample nucleic acid can be treated with (i) an alkylating agent such as methylnitrosourea that generates several alkylated bases, including N3-methyladenine and N3-methylguanine, which are recognized and cleaved by alkylpurine DNA glycosylase; (ii) sodium bisulfite, which causes deamination of cytosine residues in DNA to form uracil residues that can be cleaved by uracil N-glycosylase; and (iii) a chemical that converts guanine to its oxidized form, 8-hydroxyguanine, which can be cleaved by formamidopyrimidine DNA N-glycosylase. Examples of chemical cleavage processes include, but are not limited to, alkylation (e.g., alkylation of phosphorothioate-modified nucleic acids); acid-labile cleavage of P3’-N5’-phosphoramidate-containing nucleic acids; and treatment of nucleic acids with osmium tetroxide and piperidine.
[0149] In some embodiments, the nucleic acid reagent comprises one or more recombinase insertion sites. A recombinase insertion site is a recognition sequence on a nucleic acid molecule that participates in an integration / recombination reaction by a recombinant protein. For example, the recombination site for Cre recombinase is loxP, which is a 34-base pair sequence composed of two 13-base pair inverted repeat sequences (functioning as recombinase binding sites) flanking an 8-base pair core sequence (e.g., Sauer, Curr. Opin. Biotech. 5:521-527 (1994)). Other examples of recombination sites include the attB, attP, attL, and attR sequences, as well as variants, fragments, variants, and derivatives thereof recognized by the recombinant protein λInt and by the accessory proteins integration host factor (IHF), FIS, and excisionase (Xis) (e.g., U.S. Patent Nos. 5,888,732; 6,143,557; 6,171,861; 6,270,969; 6,277,608; and 6,720,140; U.S. Patent Application Publication Nos. 09 / 517,466 and 09 / 732,914; U.S. Patent Application Publication No. 2002 / 0007051; and Landy, Curr. Opin. Biotech. 3:699-707 (1993)).
[0150] Examples of recombinase cloning nucleic acids are found in the Gateway™ system (Invitrogen, California), which includes at least one recombination site for cloning a desired nucleic acid molecule in vivo or in vitro. In some embodiments, this system often includes at least two different site-specific recombination sites, based on the bacteriophage lambda system (e.g., att1 and att2), and utilizes vectors that are mutated from the wild-type (att0) site. Each mutant site has specificity unique to the cognate partner att site of the same type (i.e., its binding partner recombination site) (e.g., attB1 and attP1, or attL1 and attR1), and does not cross-react with other mutant-type recombination sites or with the wild-type att0 site. The different site specificities allow for directional cloning or ligation of the desired molecule, and thus provide the desired orientation of the cloned molecule. Nucleic acid fragments adjacent to the recombination sites are cloned and subcloned by replacing a selectable marker (e.g., ccdB) adjacent to the att site on a recipient plasmid molecule, sometimes called the Destination Vector, using the Gateway™ system. The desired clone is then selected by transformation of a ccdB-sensitive host strain and positive selection of the marker on the recipient molecule. Similar strategies for negative selection (e.g., use of a toxic gene) can be used in other organisms, such as mammalian and insect thymidine kinase (TK).
[0151] The nucleic acid reagent may contain one or more origin of replication (ORI) elements. In some embodiments, the template contains two or more ORIs, one of which functions efficiently in one organism (e.g., bacteria), and the other functions efficiently in another organism (e.g., eukaryotes such as yeast). In some embodiments, an ORI may function efficiently in one species (e.g., S. cerevisiae, etc.), and another ORI may function efficiently in a different species (e.g., S. pombe, etc.). The nucleic acid reagent may also contain one or more transcriptional regulatory sites.
[0152] The nucleic acid reagent, e.g., an expression cassette or vector, may contain a nucleic acid sequence encoding a marker product. The marker product is used to determine whether a certain gene has been delivered to a cell and is expressed after delivery. Examples of marker genes include the E. coli lacZ gene encoding β-galactosidase and green fluorescent protein. In some embodiments, the marker may be a selectable marker. When such a selectable marker is successfully transferred to a host cell, the transformed host cell can survive when placed under a selection pressure. There are two distinct categories of selection regimens that are widely used. The first category is based on the use of mutant cell lines that lack the ability to metabolize and grow independently of supplemented media. The second category is dominant selection, which refers to a selection scheme used in any cell type and does not require the use of mutant cell lines. These schemes usually use drugs to block the growth of host cells. Those cells with the new gene express a protein carrying drug resistance and survive the selection. Examples of such dominant selections are the use of the drug neomycin (Southern et al., J. Molec. Appl. Genet. 1:327 (1982)), mycophenolic acid (Mulligan et al., Science 209:1422 (1980)) or hygromycin (Sugden et al., Mol. Cell. Biol. 5:410 - 413 (1985)).
[0153] The nucleic acid reagent can include one or more selection elements (e.g., elements for selecting the presence of the nucleic acid reagent, elements that are not for activating the activity of a promoter element that can be selectively regulated). The selection elements are often utilized using known processes to determine whether the nucleic acid reagent is contained in a cell. In some embodiments, the nucleic acid reagent includes two or more selection elements, where one element functions efficiently in one organism and another element functions efficiently in another organism.Examples of selectable elements include, but are not limited to, the following: (1) nucleic acid segments encoding products that provide resistance to compounds that are toxic by other means (e.g., antibiotics); (2) nucleic acid segments encoding products that are lacking in recipient cells by other means (e.g., essential products, tRNA genes, auxotrophic markers); (3) nucleic acid segments encoding products that suppress the activity of gene products; (4) nucleic acid segments encoding products that can be easily identified (e.g., phenotypic markers such as antibiotics (e.g., β-lactamase), β-galactosidase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), and cell surface proteins); (5) nucleic acid segments that bind to products that are harmful to cell survival and / or function by other means; (6) nucleic acid segments that inhibit the activity of any of the nucleic acid segments described in (1) to (5) above (e.g., antisense oligonucleotides) by other means; (7) nucleic acid segments that bind to products that modify substrates (e.g., restriction endonucleases); (8) nucleic acid segments that can be used to isolate or identify desired molecules (e.g., specific protein binding sites); (9) nucleic acid segments encoding specific nucleotide sequences that may not function by other means (e.g., for PCR amplification of subpopulations of molecules); (10) nucleic acid segments that directly or indirectly confer resistance or sensitivity to specific compounds when absent; (11) nucleic acid segments encoding products that convert a compound that is toxic or relatively non-toxic in recipient cells into a toxic compound (e.g., herpes simplex thymidine kinase, cytosine deaminase); (12) nucleic acid segments that inhibit the replication, partitioning, or heritability of nucleic acid molecules containing them; and / or (13) nucleic acid segments encoding conditional replication functions, e.g., replication in specific hosts or host cell lines, or under specific environmental conditions (e.g., temperature, nutrient status, etc.).
[0154] The nucleic acid reagent can be in any form useful for in vivo transcription and / or translation. The nucleic acid can be a plasmid such as a supercoiled plasmid, a yeast artificial chromosome (e.g., YAC), a linear nucleic acid (e.g., a linear nucleic acid produced by PCR or restriction digestion), single-stranded, or sometimes double-stranded. The nucleic acid reagent can be prepared by an amplification process such as the polymerase chain reaction (PCR) process or the transcription-mediated amplification process (TMA). In TMA, two enzymes are used in an isothermal reaction to generate amplification products that are detected by luminescence (e.g., Biochemistry 1996 Jun 25;35(25):8429-38). The standard PCR process is known (e.g., U.S. Patent Nos. 4,683,202; 4,683,195; 4,965,188; and 5,565,493) and is generally carried out in cycles. Each cycle includes heat denaturation (where the hybrid nucleic acid dissociates), cooling (where the primer oligonucleotide hybridizes); and elongation of the oligonucleotide by a polymerase (i.e., Taq polymerase). An example of a PCR cycling process is to treat the sample at 95°C for 5 minutes; repeat 45 cycles of 1 minute at 95°C, 1 minute at 59°C, 10 seconds, and 1 minute 30 seconds at 72°C; then treat the sample at 72°C for 5 minutes. Multiple cycles are frequently performed using a commercially available thermal cycler. The PCR amplification products may be temporarily stored at a low temperature (e.g., 4°C) or frozen before analysis (e.g., -20°C).
[0155] Using a cloning strategy similar to the above, DNA containing unnatural nucleotides can be generated. For example, an oligonucleotide containing an unnatural nucleotide at a desired position is synthesized using standard solid-phase synthesis and purified by HPLC. The oligonucleotide is then inserted into a plasmid containing the necessary sequence context (i.e., UTR and coding sequence) using a cloning method (such as Golden Gate assembly) that uses a cloning site such as a BsaI site (although other ones described above may be used).
[0156] Kit / Manufactured article In certain embodiments, disclosed herein are kits and manufactured articles for use in one or more of the methods described herein. Such kits include a carrier, package, or container compartmentalized to receive one or more containers such as vials, tubes, etc., each of the containers containing one of the distinct elements used in the methods described herein. Suitable containers include, for example, bottles, vials, syringes, and test tubes. In one embodiment, the container is formed from various materials such as glass or plastic.
[0157] In some embodiments, the kit comprises a suitable packaging material for containing the contents of the kit. Optionally, the packaging material is constructed by well-known methods, preferably to provide a sterile and contaminant-free environment. Packaging materials used herein include, for example, those customarily utilized in commercially available kits sold for use in nucleic acid sequencing systems. Exemplary packaging materials include, but are not limited to, glass, plastic, paper, foil, etc. that can hold the components described herein within a defined range.
[0158] The packaging material may comprise a label indicating the specific use of the components. The use of the kit indicated by this label may be one or more of the methods described herein that are appropriate for the specific combination of components present in the kit. For example, the label may indicate that the kit is useful for a method of synthesizing polynucleotides or for a method of determining the sequence of nucleic acids.
[0159] Instructions for use of the packaged reagent or component may also be included in the kit. Such instructions typically include specific statements explaining the relative amounts of kit components and samples to be mixed, the maintenance period of the reagent / sample mixture, temperature, buffer conditions, and other reaction parameters.
[0160] It is understood that not all components necessary for a particular reaction need to be present in a particular kit. Instead, one or more additional components may be provided from other sources. The instructions attached to the kit may identify the additional components provided and where they can be obtained.
[0161] In some embodiments, for example, kits are provided that are useful for stably integrating unnatural nucleic acids into cellular nucleic acids using the methods provided by the present invention for producing genetically engineered mammalian cells (e.g., CHO or HEK293T cells). In one embodiment, the kits described herein comprise genetically engineered cells and one or more unnatural nucleic acids.
[0162] In additional embodiments, the kits described herein provide a cell and a nucleic acid molecule comprising a heterologous gene for introduction into the cell, thereby providing a genetically engineered cell such as an expression vector comprising the nucleic acid of any of the above embodiments described in this paragraph.
[0163] In some embodiments, the cells described herein are delivered to an organism that can be a multicellular organism, such as a mammal, e.g., a human. As such, eukaryotic cells containing polypeptides with non-natural amino acids can be introduced into the organism.
[0164] Numbered embodiments The present disclosure includes the following non-limiting numbered embodiments: Embodiment 1. A method of producing a polypeptide comprising one or more non-natural amino acids in a eukaryotic cell, comprising: (a) (i) providing a eukaryotic cell comprising a transfer RNA (tRNA) having an anticodon comprising a first unnatural base and (ii) a messenger RNA (mRNA) having a codon comprising a second unnatural base, wherein the first and second unnatural bases form an unnatural base pair (UBP) in the eukaryotic cell; (b) translating, using the tRNA by a ribosome endogenous to the eukaryotic cell, an mRNA into a polypeptide comprising one or more non-natural amino acids. A method comprising.
[0165] Embodiment 2. The method of embodiment 1, wherein the codon of the mRNA comprises three consecutive nucleobases (N-N-N); and the first unnatural base (X) is located at the first position (X-N-N) in the codon of the mRNA.
[0166] Embodiment 3. The method of embodiment 1, wherein the codon of the mRNA comprises three consecutive nucleobases (N-N-N); and the first unnatural base (X) is located at the middle position (N-X-N) in the codon of the mRNA.
[0167] Embodiment 4. The method of embodiment 1, wherein the codon of the mRNA comprises three consecutive nucleobases (N-N-N); and the first unnatural base (X) is located at the last position (N-N-X) in the codon of the mRNA.
[0168] Embodiment 5. The first unnatural base or the second unnatural base is (i) 2 - Thiouracil, 2 - Thio - thymine, 2’ - Deoxyuridine, 4 - Thio - uracil, 4 - Thio - thymine, Uracil - 5 - yl, Hypoxanthin - 9 - yl (I), 5 - Halouracil; 5 - Propynyl - uracil, 6 - Azothymine, 6 - Azouracil, 5 - Methylaminomethyluracil, 5 - Methoxyaminomethyl - 2 - thiouracil, Pseudouracil, Methyl Uracil - 5 - oxoacetate, Uracil - 5 - oxoacetic acid, 5 - Methyl - 2 - thiouracil, 3 - (3 - Amino - 3 - N - 2 - carboxypropyl)uracil, 5 - Methyl - 2 - thiouracil, 4 - Thiouracil, 5 - Methyluracil, 5’ - Methoxycarboxymethyluracil, 5 - Methoxyuracil, Uracil - 5 - oxyacetic acid, 5 - (Carboxyhydroxylmethyl)uracil, 5 - Carboxymethylaminomethyl - 2 - thiouridine, 5 - Carboxymethylaminomethyluracil or Dihydrouracil; (ii) 5 - Hydroxymethylcytosine, 5 - Trifluoromethylcytosine, 5 - Halocytosine, 5 - Propynylcytosine, 5 - Hydroxycytosine, Cyclocytosine, Cytosine Arabinoside, 5,6 - Dihydrocytosine, 5 - Nitrocytosine, 6 - Azocytosine, Azacytosine, N4 - Ethylcytosine, 3 - Methylcytosine, 5 - Methylcytosine, 4 - Acetylcytosine, 2 - Thiocytosine, Phenoxazine Cytidine ([5,4 - b][1,4]benzoxazin - 2(3H) - one), Phenothiazine Cytidine (1H - Pyrimido[5,4 - b][1,4]benzothiazin - 2(3H) - one), Phenoxazine Cytidine (9 - (2 - Aminoethoxy) - H - pyrimido[5,4 - b][1,4]benzoxazin - 2(3H) - one), Carbazole Cytidine (2H - Pyrimido[4,5 - b]indol - 2 - one) or Pyridoindole Cytidine (H - Pyrido[3’,2’:4,5]pyrrolo[2,3 - d]pyrimidin - 2 - one); (iii) Adenine substituted with 2-aminoadenine, 2-propyladenine, 2-amino-adenine, 2-F-adenine, 2-amino-propyl-adenine, 2-amino-2'-deoxyadenosine, 3-deazaadenine, 7-methyladenine, 7-deaza-adenine, 8-azaadenine, 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl, N6-isopentenyladenine, 2-methyladenine, 2,6-diaminopurine, 2-methylthio-N6-isopentenyladenine or 6-aza-adenine; (iv) 2-Methylguanine, 2-propyl and alkyl derivatives of guanine, 3-deazaguanine, 6-thio-guanine, 7-methylguanine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-8-azaguanine, 8-azaguanine, guanine substituted with 8-halo, 8-amino, 8-thiol, 8-thioalkyl and 8-hydroxyl, 1-methylguanine, 2,2-dimethylguanine, 7-methylguanine or 6-aza-guanine; and (v) Hypoxanthine, xanthine, 1-methylinosine, queuosine, beta-D-galactosylqueuosine, inosine, beta-D-mannosylqueuosine, wybutoxosine, hydroxyurea, (acp3)w, 2-aminopyridine or 2-pyridone The method according to any one of Embodiments 1 to 4 selected from the group consisting of.
[0169] Embodiment 6. The first unnatural base or the second unnatural base is
Chemical formula
[0170] Embodiment 7. The first unnatural base is
Chemical formula
Chemical formula
[0171] Embodiment 8. When the first unnatural base is [Chemical formula] when it is, the second unnatural base is [Chemical formula] and when the first unnatural base is [Chemical formula] when it is, the second unnatural base is [Chemical formula] where the wavy line indicates the bond to the ribosyl moiety, the method according to Embodiment 6.
[0172] Embodiment 9. When the first unnatural base is [Chemical formula] when it is, the second unnatural base is [Chemical formula] and when the first unnatural base is [Chemical formula] when it is, the second unnatural base is [Chemical formula] which is the method according to Embodiment 6, where the wavy line indicates the bond to the ribosyl moiety.
[0173] Embodiment 10. When the first unnatural base is [Chemical formula] then the second unnatural base is [Chemical formula] When the first unnatural base is [Chemical formula] then the second unnatural base is [Chemical formula] which is the method according to Embodiment 6, where the wavy line indicates the bond to the ribosyl moiety.
[0174] Embodiment 11. When the first unnatural base is [Chemical formula] then the second unnatural base is [Chemical formula] When the first unnatural base is [Chemical formula] then the second unnatural base is [Chemical formula] which is the method according to Embodiment 6, where the wavy line indicates the bond to the ribosyl moiety.
[0175] Embodiment 12. When the first unnatural base is [Chemical formula] In the case where it is, the second unnatural base is
Chemical formula
Chemical formula
Chemical formula
[0176] Embodiment 13. The first unnatural base or the second unnatural base is Modification at the 2'-position: OH, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl; O-alkenyl, S-alkenyl, N-alkenyl; O-alkynyl, S-alkynyl, N-alkynyl; O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, (where alkyl, alkenyl and alkynyl are substituted or unsubstituted C1-C 10 , alkyl, C2-C 10 alkenyl, C2-C 10 alkynyl, -O[(CH2) n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n -NH2 and -O(CH2) n ON[(CH2) n CH3)]2 may be, and n and m are from 1 to about 10); And / or modification at the 5'-position: 5'-vinyl, 5'-methyl (R or S); Modification at the 4-position: 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, a group for improving the pharmacokinetic properties of the oligonucleotide or a group for improving the pharmacodynamic properties of the oligonucleotide, and a modified sugar moiety selected from the group consisting of any combination thereof, the method according to any one of Embodiments 1 to 12.
[0177] Embodiment 14. The method is the method according to any one of Embodiments 1 to 13, which is a human cell.
[0178] Embodiment 15. The human cell is a HEK293T cell, the method according to Embodiment 14.
[0179] Embodiment 16. The cell is a hamster cell, the method according to any one of Embodiments 1 to 13.
[0180] Embodiment 17. The hamster cell is a Chinese hamster ovary (CHO) cell, the method according to Embodiment 16.
[0181] Embodiment 18. The unnatural amino acid is: A lysine analog; Containing an aromatic side chain; Containing an azide group; Containing an alkyne group; or Containing an aldehyde or ketone group, the method according to any one of Embodiments 1 to 17.
[0182] Embodiment 19. The non-natural amino acid is selected from the group consisting of N6-((azidoethoxy)-carbonyl)-L-lysine (AzK), N6-((propargylethoxy)-carbonyl)-L-lysine (PraK), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine or N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine, the method according to any one of Embodiments 1 to 17.
[0183] Embodiment 20. The non-natural amino acid is N6-((azidoethoxy)-carbonyl)-L-lysine (AzK), the method according to Embodiment 19.
[0184] Embodiment 21. A method for producing a polypeptide in a eukaryotic cell, wherein the polypeptide contains one or more unnatural amino acids, and the method comprises: (a) providing a eukaryote, wherein the eukaryote comprises: (i) an mRNA containing a codon that contains one or more unnatural bases; (ii) a tRNA containing an anticodon that contains one or more unnatural bases, wherein one or more unnatural bases contained in the codon in the mRNA and one or more unnatural bases contained in the anticodon in the tRNA form complementary base pairs; (iii) a tRNA synthetase that preferentially aminoacylates the tRNA with one or more unnatural amino acids as compared to natural amino acids and, (b) providing one or more unnatural amino acids to the eukaryotic cell, wherein the eukaryotic cell produces a polypeptide containing one or more unnatural amino acids.
[0185] Embodiment 22. The method according to Embodiment 21, wherein the codon of the mRNA contains three consecutive nucleobases (N-N-N); and the first unnatural base (X) is located at the first position (X-N-N) in the codon of the mRNA.
[0186] Embodiment 23. The method according to Embodiment 21, wherein the codon of the mRNA contains three consecutive nucleobases (N-N-N); and the first unnatural base (X) is located at the middle position (N-X-N) in the codon of the mRNA.
[0187] Embodiment 24. The method according to Embodiment 21, wherein the codon of the mRNA contains three consecutive nucleobases (N-N-N); and the first unnatural base (X) is located at the last position (N-N-X) in the codon of the mRNA.
[0188] Embodiment 25. One or more unnatural bases contained in the codon in the mRNA have the formula
Chemical formula
[0189] Embodiment 26. The first unnatural base or the second unnatural base is [Chem.] selected from the group consisting of, and the wavy line indicates a bond to the ribosyl moiety, the method according to any one of embodiments 21 to 24.
[0190] Embodiment 27. When the first unnatural base is [Chem.] then the second unnatural base is [Chem.] and when the first unnatural base is [Chem.] then the second unnatural base is [Chem.] and the wavy line indicates a bond to the ribosyl moiety, the method according to embodiment 26.
[0191] Embodiment 28. When the first unnatural base is [Chem.] then the second unnatural base is [Chem.] and when the first unnatural base is
Chem.
Chem.
[0192] Embodiment 29. When the first unnatural base is
Chem.
Chem.
Chem.
Chem.
[0193] Embodiment 30. When the first unnatural base is
Chem.
Chem.
Chem.
Chem.
[0194] Embodiment 31. When the first unnatural base is
Chemical formula
Chemical formula
Chemical formula
Chemical formula
[0195] Embodiment 32. When the first unnatural base is
Chemical formula
Chemical formula
Chemical formula
Chemical formula
[0196] Embodiment 33. When the first unnatural base is
Chemical formula
[0197] Embodiment 34. The unnatural nucleotide containing the codon in the mRNA is [Chemical formula] selected from, and the wavy line indicates the bond to the ribosyl moiety, the method according to any one of Embodiments 21 to 24.
[0198] Embodiment 35. The unnatural nucleotide containing the codon in the mRNA is [Chemical formula] which is, and the wavy line indicates the bond to the ribosyl moiety, the method according to Embodiment 34.
[0199] Embodiment 36. The unnatural nucleotide containing the codon in the mRNA is [Chemical formula] which is, and the wavy line indicates the bond to the ribosyl moiety, the method according to Embodiment 34.
[0200] Embodiment 37. The unnatural nucleotide containing the codon in the mRNA is [Chemical formula] which is, and the wavy line indicates the bond to the ribosyl moiety, the method according to Embodiment 34.
[0201] Embodiment 38. The codon of the mRNA contains three consecutive nucleobases (N-N-N), and the unnatural base (X) is located at the first position (X-N-N) in the codon of the mRNA, and the unnatural base is [Chemical formula] The method according to Embodiment 21, which is selected from [description of selection range] and the wavy line indicates the bond to the ribosyl moiety.
[0202] Embodiment 39. The unnatural base is
Chemical formula
[0203] Embodiment 40. The unnatural base is
Chemical formula
[0204] Embodiment 41. The unnatural base is
Chemical formula
[0205] Embodiment 42. The codon of the mRNA contains three consecutive nucleobases (N-N-N), the unnatural base (X) is located at the central position (N-X-N) in the codon of the mRNA, and the unnatural base is
Chemical formula
[0206] Embodiment 43. The unnatural base is
Chemical formula
[0207] Embodiment 44. The unnatural base is
Chemical formula
[0208] Embodiment 45. The unnatural base is
Chemical formula
[0209] Embodiment 46. The codon of the mRNA contains three consecutive nucleobases (N-N-N), the unnatural base (X) is located at the last position (N-N-X) in the codon of the mRNA, and the unnatural base is
Chemical formula
[0210] Embodiment 47. The unnatural base is
Chemical formula
[0211] Embodiment 48. The unnatural base is
Chemical formula
[0212] Embodiment 49. The unnatural base is
Chemical formula
[0213] Embodiment 50. The anticodon of the tRNA contains three consecutive nucleobases (N-N-N); the first unnatural base (X) is located at the first position (X-N-N) in the anticodon of the tRNA, the method according to Embodiment 21.
[0214] Embodiment 51. The unnatural base is
Chemical formula
[0215] Embodiment 52. The unnatural base is
Chemical formula
[0216] Embodiment 53. The unnatural base is
Chemical formula
[0217] Embodiment 54. The unnatural base is
Chemical formula
[0218] Embodiment 55. The anticodon of the tRNA contains three consecutive nucleobases (N-N-N); the first unnatural base (X) is located at the central position (N-X-N) in the anticodon of the tRNA, the method according to Embodiment 21.
[0219] Embodiment 56. The unnatural base is
Chemical formula
[0220] Embodiment 57. The unnatural base is [Chemical formula] and the wavy line indicates a bond to the ribosyl moiety, the method according to Embodiment 55.
[0221] Embodiment 58. The unnatural base is [Chemical formula] and the wavy line indicates a bond to the ribosyl moiety, the method according to Embodiment 55.
[0222] Embodiment 59. The unnatural base is [Chemical formula] and the wavy line indicates a bond to the ribosyl moiety, the method according to Embodiment 55.
[0223] Embodiment 60. The anticodon of tRNA contains three consecutive nucleobases (N-N-N); the first unnatural base (X) is located at the last position (N-N-X) in the anticodon of tRNA, the method according to Embodiment 21.
[0224] Embodiment 61. The unnatural base is [Chemical formula] selected from, and the wavy line indicates a bond to the ribosyl moiety, the method according to Embodiment 60.
[0225] Embodiment 62. The unnatural base is [Chemical formula] and the wavy line indicates a bond to the ribosyl moiety, the method according to Embodiment 61.
[0226] Embodiment 63. The unnatural base is [Chemical formula] as shown, and the wavy line indicates the bond to the ribosyl moiety, the method according to Embodiment 61.
[0227] Embodiment 64. The unnatural base is [Chemical formula] as shown, and the wavy line indicates the bond to the ribosyl moiety, the method according to Embodiment 61.
[0228] Embodiment 65. Each of the codon and the anticodon contains three consecutive nucleobases (N-N-N), the codon in the mRNA contains a first unnatural base (X) located at the first position (X-N-N) of the codon, and the anticodon in the tRNA contains a second unnatural base (Y) located at the last position (N-N-Y) of the anticodon, the method according to Embodiment 21.
[0229] Embodiment 66. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are the same or different, the method according to Embodiment 65.
[0230] Embodiment 67. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are the same, the method according to Embodiment 66.
[0231] Embodiment 68. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are different, the method according to Embodiment 66.
[0232] Embodiment 69. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are [Chemical formula] Selected from the group consisting of, the wavy line indicates the bond to the ribosyl moiety, the method according to any one of Embodiments 65 to 68.
[0233] Embodiment 70. The first unnatural base (X) located in the codon of mRNA and the second unnatural base (Y) located in the anticodon of tRNA are
Chemical formula
[0234] Embodiment 71. Both the first unnatural base (X) located in the codon of mRNA and the second unnatural base (Y) located in the anticodon of tRNA are
Chemical formula
[0235] Embodiment 72. Both the first unnatural base (X) located in the codon of mRNA and the second unnatural base (Y) located in the anticodon of tRNA are
Chemical formula
[0236] Embodiment 73. Both the first unnatural base (X) located in the codon of mRNA and the second unnatural base (Y) located in the anticodon of tRNA are
Chemical formula
[0237] Embodiment 74. The first unnatural base (X) located in the codon of mRNA is [Chemical formula] selected from, and the second unnatural base (Y) located in the anticodon of the tRNA is [Chemical formula] as described in embodiment 70, where in each case the wavy line indicates the bond to the ribosyl moiety.
[0238] Embodiment 75. The first unnatural base (X) located in the codon of the mRNA is [Chemical formula] as described in embodiment 74.
[0239] Embodiment 76. The first unnatural base (X) located in the codon of the mRNA is [Chemical formula] as described in embodiment 74.
[0240] Embodiment 77. Each of the codon and the anticodon contains three consecutive nucleobases (N-N-N), the codon in the mRNA contains a first unnatural base (X) located at the central position (N-X-N) of the codon, and the anticodon in the tRNA contains a second unnatural base (Y) located at the central position (N-Y-N) of the anticodon, as described in embodiment 21.
[0241] Embodiment 78. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are the same or different, as described in embodiment 77.
[0242] Embodiment 79. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are the same, as described in embodiment 78.
[0243] Embodiment 80. The method according to embodiment 78, wherein the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are different.
[0244] Embodiment 81. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are [Chemical formula] selected from the group consisting of, and the wavy line indicates the bond to the ribosyl moiety, the method according to any one of embodiments 77 to 79.
[0245] Embodiment 82. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are [Chemical formula] selected from the group consisting of, and the wavy line indicates the bond to the ribosyl moiety, the method according to embodiment 81.
[0246] Embodiment 83. Both the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are [Chemical formula] as shown, and the wavy line indicates the bond to the ribosyl moiety, the method according to embodiment 82.
[0247] Embodiment 84. Both the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are [Chemical formula] as shown, and the wavy line indicates the bond to the ribosyl moiety, the method according to embodiment 82.
[0248] Embodiment 85. Both the first unnatural base (X) located in the codon of mRNA and the second unnatural base (Y) located in the anticodon of tRNA are
Chemical formula
[0249] Embodiment 86. The first unnatural base (X) located in the codon of mRNA is
Chemical formula
Chemical formula
[0250] Embodiment 87. The first unnatural base (X) located in the codon of mRNA is
Chemical formula
[0251] Embodiment 88. The first unnatural base (X) located in the codon of mRNA is
Chemical formula
[0252] Embodiment 89. Each of the codon and the anticodon contains three consecutive nucleobases (N-N-N), the codon in mRNA contains the first unnatural base (X) located at the last position (N-N-X) of the codon, and the anticodon in tRNA contains the second unnatural base (Y) located at the first position (Y-N-N) of the anticodon, the method according to Embodiment 21.
[0253] Embodiment 90. The method according to embodiment 89, wherein the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are the same or different.
[0254] Embodiment 91. The method according to embodiment 89, wherein the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are the same.
[0255] Embodiment 92. The method according to embodiment 89, wherein the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are different.
[0256] Embodiment 93. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are
Chem.
[0257] Embodiment 94. The first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are
Chem.
[0258] Embodiment 95. Both the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are
Chem.
[0259] Embodiment 96. Both the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are
Chemical formula
[0260] Embodiment 97. Both the first unnatural base (X) located in the codon of the mRNA and the second unnatural base (Y) located in the anticodon of the tRNA are
Chemical formula
[0261] Embodiment 98. The first unnatural base (X) located in the codon of the mRNA is
Chemical formula
Chemical formula
[0262] Embodiment 99. The first unnatural base (X) located in the codon of the mRNA is
Chemical formula
[0263] Embodiment 100. The first unnatural base (X) located in the codon of the mRNA is [Chemistry] The method according to embodiment 98, which is as follows.
[0264] Embodiment 101. The method according to any one of embodiments 21, 23, 25 to 37, 42 to 45, 55 to 59, and 77 to 88, wherein the codon in the mRNA is selected from AXC, GXC, or GXU, and X is a non-natural base.
[0265] Embodiment 102. The method according to embodiment 101, wherein the codon in the mRNA is AXC and X is a non-natural base.
[0266] Embodiment 103. The method according to embodiment 101, wherein the codon in the mRNA is GXC and X is a non-natural base.
[0267] Embodiment 104. The method according to embodiment 101, wherein the codon in the mRNA is GXU and X is a non-natural base.
[0268] Embodiment 105. The method according to any one of embodiments 21, 23, 25 to 37, 42 to 45, 55 to 59, and 77 to 88, wherein the codon in the mRNA is selected from AXC, GXC, or GXU, the anticodon in the tRNA is selected from GYU, GYC, and AYC, X is a first non-natural base, and Y is a second non-natural base.
[0269] Embodiment 106. The method according to embodiment 105, wherein X and Y are the same or different.
[0270] Embodiment 107. The method according to embodiment 106, wherein X and Y are the same.
[0271] Embodiment 108. The method according to embodiment 106, wherein X and Y are different.
[0272] Embodiment 109. The method according to embodiment 105, wherein the codon in the mRNA is AXC and the anticodon in the tRNA is GYU.
[0273] Embodiment 110. The method according to Embodiment 109, wherein X and Y are the same or different.
[0274] Embodiment 111. The method according to Embodiment 109, wherein X and Y are the same.
[0275] Embodiment 112. The method according to Embodiment 109, wherein X and Y are different.
[0276] Embodiment 113. The method according to Embodiment 106, wherein the codon in the mRNA is GXC and the anticodon in the tRNA is GYC.
[0277] Embodiment 114. The method according to Embodiment 113, wherein X and Y are the same or different.
[0278] Embodiment 115. The method according to Embodiment 113, wherein X and Y are the same.
[0279] Embodiment 116. The method according to Embodiment 113, wherein X and Y are different.
[0280] Embodiment 117. The method according to Embodiment 106, wherein the codon in the mRNA is GXU and the anticodon is AYC.
[0281] Embodiment 118. The method according to Embodiment 117, wherein X and Y are the same or different.
[0282] Embodiment 119. The method according to Embodiment 117, wherein X and Y are the same.
[0283] Embodiment 120. The method according to Embodiment 117, wherein X and Y are different.
[0284] Embodiment 121. The tRNA is derived from Methanococcus jannaschii, Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans, and the method according to any one of Embodiments 21 to 120.
[0285] Embodiment 122. The tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans, and the method according to any one of Embodiments 21 to 120.
[0286] Embodiment 123. The tRNA and the tRNA synthetase are derived from Methanococcus jannaschii, and the method according to Embodiment 122.
[0287] Embodiment 124. The tRNA and the tRNA synthetase are derived from Methanosarcina barkeri, and the method according to Embodiment 122.
[0288] Embodiment 125. The tRNA and the tRNA synthetase are derived from Methanosarcina mazei, and the method according to Embodiment 122.
[0289] Embodiment 126. The tRNA and the tRNA synthetase are derived from Methanosarcina acetivorans, and the method according to Embodiment 122.
[0290] Embodiment 127. The tRNA is derived from Methanococcus jannaschii, and the tRNA synthetase is derived from Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans, and the method according to any one of Embodiments 21 to 120.
[0291] Embodiment 128. The tRNA is derived from Methanosarcina barkeri, and the tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina mazei or Methanosarcina acetivorans, and the method according to any one of Embodiments 21 to 120.
[0292] Embodiment 129. The method according to any one of Embodiments 21 to 120, wherein the tRNA is derived from Methanosarcina mazei, and the tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina barkeri or Methanosarcina acetivorans.
[0293] Embodiment 130. The method according to any one of Embodiments 21 to 120, wherein the tRNA is derived from Methanosarcina acetivorans, and the tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina barkeri or Methanosarcina mazei.
[0294] Embodiment 131. The method according to any one of Embodiments 21 to 120, wherein the tRNA is derived from Methanosarcina mazei, and the tRNA synthetase is derived from Methanosarcina barkeri.
[0295] Embodiment 132. The method according to any one of Embodiments 21 to 120, wherein the cell is a human cell.
[0296] Embodiment 133. The method according to Embodiment 132, wherein the human cell is a HEK293T cell.
[0297] Embodiment 134. The method according to any one of Embodiments 21 to 120, wherein the cell is a hamster cell.
[0298] Embodiment 135. The method according to Embodiment 134, wherein the hamster cell is a Chinese hamster ovary (CHO) cell.
[0299] Embodiment 136. The unnatural amino acid is: a lysine analog; comprising an aromatic side chain; comprising an azide group; comprising an alkyne group; or comprising an aldehyde or ketone group, according to any one of Embodiments 21 to 135.
[0300] Embodiment 137. The non-natural amino acid is selected from the group consisting of N6-((azidoethoxy)-carbonyl)-L-lysine (AzK), N6-((propargylethoxy)-carbonyl)-L-lysine (PraK), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine or N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine, the method according to any one of Embodiments 21 to 135.
[0301] Embodiment 138. The non-natural amino acid is N6-((azidoethoxy)-carbonyl)-L-lysine (AzK), the method according to Embodiment 137.
[0302] Embodiment 139. A system for the expression of a non-natural polypeptide in a eukaryotic cell, comprising: (a) At least one non-natural amino acid; (b) An mRNA encoding a non-natural polypeptide, comprising at least one codon containing one or more first non-natural bases; (c) A tRNA comprising at least one anticodon containing one or more second non-natural bases, wherein the one or more first non-natural bases and the one or more second non-natural bases form one or more complementary base pairs, the tRNA; (d) One or more nucleic acid constructs comprising a nucleic acid sequence encoding a tRNA synthetase that preferentially aminoacylates the tRNA with at least one non-natural amino acid; and (e) A eukaryotic cell capable of translating the mRNA into a polypeptide containing a non-natural amino acid using the tRNA and the tRNA synthetase A system comprising.
[0303] Embodiment 140. At least one codon of the mRNA comprises three consecutive nucleic acid bases (N-N-N); one or more first non-natural bases (X) are located at the first position (X-N-N) in at least one codon of the mRNA, the system according to embodiment 139.
[0304] Embodiment 141. At least one codon of the mRNA comprises three consecutive nucleic acid bases (N-N-N); one or more first non-natural bases (X) are located at the central position (N-X-N) in the codon of the mRNA, the system according to embodiment 139.
[0305] Embodiment 142. At least one codon of the mRNA comprises three consecutive nucleic acid bases (N-N-N); one or more first non-natural bases (X) are located at the last position (N-N-X) in at least one codon of the mRNA, the system according to embodiment 139.
[0306] Embodiment 143. One or more non-natural bases have the formula
Chemical formula
[0307] Embodiment 144. One or more first unnatural bases or one or more second unnatural bases are [Chemical formula] selected from the group consisting of, and the wavy line indicates a bond to the ribosyl moiety, the system according to any one of embodiments 139 to 142.
[0308] Embodiment 145. When one or more first unnatural bases are [Chemical formula] then one or more second unnatural bases are [Chemical formula] and when one or more first unnatural bases are [Chemical formula] then the second unnatural base is [Chemical formula] and the wavy line indicates a bond to the ribosyl moiety, the system according to embodiment 144.
[0309] Embodiment 146. When one or more first unnatural bases are [Chemical formula] then one or more second unnatural bases are [Chemical formula] and one or more first unnatural bases are [Chemical formula] in the case of, one or more second unnatural bases are [Chemical formula] as shown, where the wavy line indicates the bond to the ribosyl moiety, the system according to embodiment 144.
[0310] Embodiment 147. One or more first unnatural bases are [Chemical formula] in the case of, one or more second unnatural bases are [Chemical formula] and one or more first unnatural bases are [Chemical formula] in the case of, one or more second unnatural bases are [Chemical formula] as shown, where the wavy line indicates the bond to the ribosyl moiety, the system according to embodiment 144.
[0311] Embodiment 148. One or more first unnatural bases are [Chemical formula] in the case of, one or more second unnatural bases are [Chemical formula] and one or more first unnatural bases are [Chemical formula] In the case where it is, one or more second unnatural bases are [Chemical formula] as shown, and the wavy line indicates the bond to the ribosyl moiety, the system described in Embodiment 144.
[0312] Embodiment 149. One or more first unnatural bases are [Chemical formula] In the case where it is, one or more second unnatural bases are [Chemical formula] as shown, and one or more first unnatural bases are [Chemical formula] In the case where it is, one or more second unnatural bases are [Chemical formula] as shown, and the wavy line indicates the bond to the ribosyl moiety, the system described in Embodiment 144.
[0313] Embodiment 150. One or more first unnatural bases are [Chemical formula] In the case where it is, one or more second unnatural bases are [Chemical formula] as shown, and one or more first unnatural bases are [Chemical formula] In the case where it is, one or more second unnatural bases are [Chemical formula] The system according to Embodiment 144, wherein the wavy line indicates a bond to the ribosyl moiety.
[0314] Embodiment 151. When one or more first unnatural bases are
Chemical formula
Chemical formula
[0315] Embodiment 152. One or more first unnatural bases are
Chemical formula
[0316] Embodiment 153. One or more first unnatural bases are
Chemical formula
[0317] Embodiment 154. One or more first unnatural bases are
Chemical formula
[0318] Embodiment 155. One or more first unnatural bases are
Chemical formula
[0319] Embodiment 156. At least one codon of the mRNA contains three consecutive nucleobases (N-N-N), and one or more first unnatural bases (X) are located at the first position (X-N-N) in the codon of the mRNA, and one or more first unnatural bases are
Chemical formula
[0320] Embodiment 157. One or more first unnatural bases are
Chemical formula
[0321] Embodiment 158. One or more first unnatural bases are
Chemical formula
[0322] Embodiment 159. One or more first unnatural bases are
Chemical formula
[0323] Embodiment 160. At least one codon of the mRNA contains three consecutive nucleobases (N-N-N), and one or more first unnatural bases (X) are located at the central position (N-X-N) in the codon of the mRNA, and one or more first unnatural bases are
Chemical formula
[0324] Embodiment 161. One or more first unnatural bases are
Chemical formula
[0325] Embodiment 162. One or more first unnatural bases are
Chemical formula
[0326] Embodiment 163. One or more first unnatural bases are
Chemical formula
[0327] Embodiment 164. At least one codon of the mRNA contains three consecutive nucleobases (N-N-N), one or more first unnatural bases (X) are located at the last position (N-N-X) in the codon of the mRNA, and one or more first unnatural bases are
Chemical formula
[0328] Embodiment 165. One or more first unnatural bases are
Chemical formula
[0329] Embodiment 166. One or more first unnatural bases are
Chemical formula
[0330] Embodiment 167. One or more first unnatural bases are
Chemical formula
[0331] Embodiment 168. At least one anticodon of tRNA contains three consecutive nucleobases (N-N-N); one or more second unnatural bases (X) are located at the first position (X-N-N) in the anticodon of tRNA, the system according to Embodiment 139.
[0332] Embodiment 169. One or more second unnatural bases are
Chemical formula
[0333] Embodiment 170. One or more second unnatural bases are
Chemical formula
[0334] Embodiment 171. One or more second unnatural bases are
Chemical formula
[0335] Embodiment 172. One or more second non-natural bases are
Chemical formula
[0336] Embodiment 173. At least one anticodon of tRNA contains three consecutive nucleobases (N-N-N); one or more second non-natural bases (X) are located at the central position (N-X-N) in the anticodon of tRNA, which is the system according to Embodiment 139.
[0337] Embodiment 174. One or more second non-natural bases are
Chemical formula
[0338] Embodiment 175. One or more second non-natural bases are
Chemical formula
[0339] Embodiment 176. One or more second non-natural bases are
Chemical formula
[0340] Embodiment 177. One or more second non-natural bases are
Chemical formula
[0341] Embodiment 178. At least one anticodon of the tRNA contains three consecutive nucleobases (N-N-N); one or more second unnatural bases (X) are located at the last position (N-N-X) in the anticodon of the tRNA, which is the system according to Embodiment 139.
[0342] Embodiment 179. One or more second unnatural bases are
Chemical formula
[0343] Embodiment 180. One or more second unnatural bases are
Chemical formula
[0344] Embodiment 181. One or more second unnatural bases are
Chemical formula
[0345] Embodiment 182. One or more second unnatural bases are
Chemical formula
[0346] Embodiment 183. At least one codon and at least one anticodon each independently contain three consecutive nucleobases (N-N-N), and at least one codon contains one or more first unnatural bases (X) located at the first position (X-N-N) of the codon, and at least one anticodon in the tRNA contains one or more second unnatural bases (Y) located at the last position (N-N-Y) of the anticodon, the system according to Embodiment 139.
[0347] Embodiment 184. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are the same or different, the system according to Embodiment 183.
[0348] Embodiment 185. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are the same, the system according to Embodiment 184.
[0349] Embodiment 186. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are different, the system according to Embodiment 184.
[0350] Embodiment 187. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are
Chemical formula
[0351] Embodiment 188. One or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are
Chemical formula
[0352] Embodiment 189. Both one or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are
Chemical formula
[0353] Embodiment 190. Both one or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are
Chemical formula
[0354] Embodiment 191. Both one or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are
Chemical formula
[0355] Embodiment 192. One or more first unnatural bases (X) located in the codon of mRNA are [Chemical formula] selected from, and one or more second unnatural bases (Y) located in the anticodon of the tRNA are [Chemical formula] as described in Embodiment 188, where in each case the wavy line indicates the bond to the ribosyl moiety.
[0356] Embodiment 193. One or more first unnatural bases (X) located in the codon of the mRNA are [Chemical formula] the system according to Embodiment 192.
[0357] Embodiment 194. One or more first unnatural bases (X) located in the codon of the mRNA are [Chemical formula] the system according to Embodiment 192.
[0358] Embodiment 195. At least one codon and at least one anticodon each independently comprise three consecutive nucleobases (N-N-N), and at least one codon in the mRNA comprises one or more first unnatural bases (X) located at the central position (N-X-N) of at least one codon, and at least one anticodon in the tRNA comprises one or more second unnatural bases (Y) located at the central position (N-Y-N) of the anticodon, the system according to Embodiment 139.
[0359] Embodiment 196. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are the same or different, the system according to Embodiment 195.
[0360] Embodiment 197. The system according to Embodiment 195, wherein one or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are the same.
[0361] Embodiment 198. The system according to Embodiment 195, wherein one or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are different.
[0362] Embodiment 199. One or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are
Chemical formula
[0363] Embodiment 200. One or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are
Chemical formula
[0364] Embodiment 201. Both one or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are
Chemical formula
[0365] Embodiment 202. One or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are both
Chemical formula
[0366] Embodiment 203. One or more first unnatural bases (X) located in the codon of mRNA and one or more second unnatural bases (Y) located in the anticodon of tRNA are both
Chemical formula
[0367] Embodiment 204. One or more first unnatural bases (X) located in the codon of mRNA are
Chemical formula
Chemical formula
[0368] Embodiment 205. One or more first unnatural bases (X) located in the codon of mRNA are
Chemical formula
[0369] Embodiment 206. One or more first unnatural bases (X) located in the codon of mRNA are [Chemistry] The system according to Embodiment 204, which is.
[0370] Embodiment 207. At least one codon and at least one anticodon each independently contain three consecutive nucleobases (N-N-N), and at least one codon in the mRNA contains one or more first unnatural bases (X) located at the last position (N-N-X) of the at least one codon, and at least one anticodon in the tRNA contains one or more second unnatural bases (Y) located at the first position (Y-N-N) of the anticodon. The system according to Embodiment 139.
[0371] Embodiment 208. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are the same or different. The system according to Embodiment 207.
[0372] Embodiment 209. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are the same. The system according to Embodiment 208.
[0373] Embodiment 210. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are different. The system according to Embodiment 208.
[0374] Embodiment 211. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are [Chemistry] A system according to any one of Embodiments 207 to 210, selected from the group consisting of, wherein the wavy line indicates a bond to the ribosyl moiety.
[0375] Embodiment 212. One or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are
Chemical formula
[0376] Embodiment 213. Both one or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are
Chemical formula
[0377] Embodiment 214. Both one or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are
Chemical formula
[0378] Embodiment 215. Both one or more first unnatural bases (X) located in the codon of the mRNA and one or more second unnatural bases (Y) located in the anticodon of the tRNA are
Chemical formula
[0379] Embodiment 216. One or more first unnatural bases (X) located in the codon of the mRNA are [Chemical formula] selected from, and one or more second unnatural bases (Y) located in the anticodon of the tRNA are [Chemical formula] as shown, and in each case the wavy line indicates the bond to the ribosyl moiety, the system according to Embodiment 212.
[0380] Embodiment 217. One or more first unnatural bases (X) located in the codon of the mRNA are [Chemical formula] as shown, the system according to Embodiment 216.
[0381] Embodiment 218. One or more first unnatural bases (X) located in the codon of the mRNA are [Chemical formula] as shown, the system according to Embodiment 216.
[0382] Embodiment 219. At least one codon in the mRNA is selected from AXC, GXC or GXU, and X is an unnatural base, the system according to any one of Embodiments 139 to 218.
[0383] Embodiment 220. At least one codon in the mRNA is AXC, and X is an unnatural base, the system according to Embodiment 219.
[0384] Embodiment 221. At least one codon in the mRNA is GXC, and X is an unnatural base, the system according to Embodiment 219.
[0385] Embodiment 222. The system according to Embodiment 219, wherein at least one codon in the mRNA is GXU, and X is a non-natural base.
[0386] Embodiment 223. The system according to any one of Embodiments 139 to 218, wherein at least one codon in the mRNA is selected from AXC, GXC or GXU, at least one anticodon in the tRNA is selected from GYU, GYC and AYC, X is one or more first non-natural bases, and Y is one or more second non-natural bases.
[0387] Embodiment 224. The system according to Embodiment 223, wherein X and Y are the same or different.
[0388] Embodiment 225. The system according to Embodiment 224, wherein X and Y are the same.
[0389] Embodiment 226. The system according to Embodiment 224, wherein X and Y are different.
[0390] Embodiment 227. The system according to Embodiment 223, wherein at least one codon in the mRNA is AXC and at least one anticodon in the tRNA is GYU.
[0391] Embodiment 228. The system according to Embodiment 227, wherein X and Y are the same or different.
[0392] Embodiment 229. The system according to Embodiment 228, wherein X and Y are the same.
[0393] Embodiment 230. The system according to Embodiment 228, wherein X and Y are different.
[0394] Embodiment 231. The system according to Embodiment 223, wherein at least one codon in the mRNA is GXC and at least one anticodon in the tRNA is GYC.
[0395] Embodiment 232. The system according to Embodiment 231, wherein X and Y are the same or different.
[0396] Embodiment 233. The system according to Embodiment 232, wherein X and Y are the same.
[0397] Embodiment 234. The system according to Embodiment 232, wherein X and Y are different.
[0398] Embodiment 235. The system according to Embodiment 223, wherein at least one codon in the mRNA is GXU and at least one anticodon is AYC.
[0399] Embodiment 236. The system according to Embodiment 235, wherein X and Y are the same or different.
[0400] Embodiment 237. The system according to Embodiment 236, wherein X and Y are the same.
[0401] Embodiment 238. The system according to Embodiment 236, wherein X and Y are different.
[0402] Embodiment 239. The system according to any one of Embodiments 139 to 238, wherein the tRNA is derived from Methanococcus jannaschii, Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans.
[0403] Embodiment 240. The system according to any one of Embodiments 139 to 238, wherein the tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans.
[0404] Embodiment 241. The system according to Embodiment 240, wherein the tRNA and the tRNA synthetase are derived from Methanococcus jannaschii.
[0405] Embodiment 242. The tRNA and the tRNA synthetase are the system described in Embodiment 240, which is derived from Methanosarcina barkeri.
[0406] Embodiment 243. The tRNA and the tRNA synthetase are the system described in Embodiment 240, which is derived from Methanosarcina mazei.
[0407] Embodiment 244. The tRNA and the tRNA synthetase are the system described in Embodiment 240, which is derived from Methanosarcina acetivorans.
[0408] Embodiment 245. The tRNA is derived from Methanococcus jannaschii, and the tRNA synthetase is derived from Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans, and the system is as described in any one of Embodiments 139 to 239.
[0409] Embodiment 246. The tRNA is derived from Methanosarcina barkeri, and the tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina mazei or Methanosarcina acetivorans, and the system is as described in any one of Embodiments 139 to 239.
[0410] Embodiment 247. The tRNA is derived from Methanosarcina mazei, and the tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina barkeri or Methanosarcina acetivorans, and the system is as described in any one of Embodiments 139 to 239.
[0411] Embodiment 248. The tRNA is derived from Methanosarcina acetivorans, and the tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina barkeri or Methanosarcina mazei, and the system is as described in any one of Embodiments 139 to 239.
[0412] Embodiment 249. The tRNA is derived from Methanosarcina mazei, and the tRNA synthetase is derived from Methanosarcina barkeri, and the system is as described in any one of Embodiments 139 to 239.
[0413] Embodiment 250. The system according to any one of Embodiments 139 to 249, wherein the cell is a human cell.
[0414] Embodiment 251. The system according to Embodiment 250, wherein the human cell is a HEK293T cell.
[0415] Embodiment 252. The system according to any one of Embodiments 139 to 239, wherein the cell is a hamster cell.
[0416] Embodiment 253. The system according to Embodiment 252, wherein the hamster cell is a Chinese hamster ovary (CHO) cell.
[0417] Embodiment 254. The unnatural amino acid is: a lysine analog; including an aromatic side chain; including an azide group; including an alkyne group; or including an aldehyde or ketone group, the system according to any one of Embodiments 139 to 253.
[0418] Embodiment 255. The unnatural amino acid is selecte...
Claims
1. (a)A messenger RNA (mRNA) comprising a codon containing a first unnatural base, the mRNA encoding a polypeptide containing at least one unnatural amino acid encoded by the codon; (b)A pyrrolysyl or tyrosyl transfer RNA (tRNA) comprising an anticodon containing a second unnatural base; (c)A pyrrolysyl or tyrosyl tRNA synthetase, the tRNA synthetase being characterized by aminoacylating the tRNA with at least one unnatural amino acid, the tRNA synthetase A eukaryotic cell comprising The first and the second unnatural bases are capable of forming an unnatural base pair (UBP) in a eukaryotic cell, and the codon and the anticodon each independently comprise three consecutive nucleobases (N-N-N), The first unnatural base (X) is located at the central position (N-X-N) of the codon, and the second unnatural base (Y) is located at the central position (N-Y-N) of the anticodon, or The first unnatural base (X) is located at the last position (N-N-X) of the codon, and the second unnatural base (Y) is located at the first position (Y-N-N) of the anticodon, said eukaryotic cell.
2. The eukaryotic cell according to claim 1, wherein the tRNA is charged with at least one unnatural amino acid.
3. A polypeptide translated from the mRNA, the polypeptide containing at least one unnatural amino acid, the eukaryotic cell according to claim 1 or 2.
4. The eukaryotic cell according to any one of claims 1 to 3, wherein the polypeptide contains a eukaryotic glycosylation pattern.
5. The first unnatural base and the second unnatural base are each independently 【Chemical Formula 1】 selected from, the wavy line indicating the bond to the ribosyl moiety, the eukaryotic cell according to any one of claims 1 to 4.
6. The first unnatural base is of the formula 【Chemical Formula 2】 wherein R₂ is selected from hydrogen, alkyl, alkenyl, alkynyl, methoxy, methanethiol, methaneseleno, halogen, cyano and azide, and the wavy line indicates the bond to the ribosyl moiety, the eukaryotic cell according to any one of claims 1 to 5.
7. (a) (i) When the first unnatural base is 【Chemical Formula 3】 then the second unnatural base is 【Chemical Formula 4】 or (ii) when the first unnatural base is 【Chemical Formula 5】 then the second unnatural base is 【Chemical Formula 6】 and the wavy line indicates the bond to the ribosyl moiety; or (b) (i) When the first unnatural base is 【Chemical Formula 7】 then the second unnatural base is 【Chemical Formula 8】 or (ii) when the first unnatural base is 【Chemical Formula 9】 then the second unnatural base is 【Chemical Formula 10】 and the wavy line indicates the bond to the ribosyl moiety; or (c) (i) When the first unnatural base is 【Chemical Formula 11】 then the second unnatural base is 【Chemical Formula 12】 or (ii) when the first unnatural base is 【Chemical Formula 13】 then the second unnatural base is [Chemical Formula 14] and the wavy line indicates the bond to the ribosyl moiety; or (d) (i) When the first unnatural base is [Chemical Formula 15] then the second unnatural base is [Chemical Formula 16] or (ii) when the first unnatural base is [Chemical Formula 17] then the second unnatural base is [Chemical Formula 18] and the wavy line indicates the bond to the ribosyl moiety; or (e) (i) When the first unnatural base is [Chemical Formula 19] then the second unnatural base is [Chemical Formula 20] or (ii) when the first unnatural base is [Chemical Formula 21] then the second unnatural base is [Chemical Formula 22] and the wavy line indicates the bond to the ribosyl moiety; or (f) (i) When the first unnatural base is [Chemical Formula 23] then the second unnatural base is [Chemical Formula 24] or (ii) when the first unnatural base is [Chemical Formula 25] then the second unnatural base is [Chemical Formula 26] and the wavy line indicates the bond to the ribosyl moiety; or, (g) The first and second unnatural bases are each [Chemical Formula 27] and the wavy line indicates a bond to the ribosyl moiety, the eukaryotic cell according to any one of claims 1 to 6.
8. The first unnatural base (X) is located at the central position (N-X-N) in the codon, and the first unnatural base is 【Chemical formula 28】 selected from, and the wavy line indicates a bond to the ribosyl moiety; or The first unnatural base (X) is located at the last position (N-N-X) in the codon, and the first unnatural base is 【Chemical formula 29】 selected from, and the wavy line indicates a bond to the ribosyl moiety; or The second unnatural base (Y) is located at the first position (Y-N-N) of the anticodon, and the second unnatural base is 【Chemical formula 30】 selected from, and the wavy line indicates a bond to the ribosyl moiety; or The second unnatural base (Y) is located at the central position (N-Y-N) of the anticodon, and the second unnatural base is 【Chemical formula 31】 selected from, and the wavy line indicates a bond to the ribosyl moiety, the eukaryotic cell according to any one of claims 1 to 7.
9. The first unnatural base (X) and the second unnatural base (Y) are each independently 【Chemical formula 32】 selected from, and the wavy line indicates a bond to the ribosyl moiety; or The first unnatural base (X) and the second unnatural base (Y) are both 【Chemical formula 33】 and the wavy line indicates a bond to the ribosyl moiety; or The first unnatural base (X) is 【Chemical formula 34】 selected from, and the second unnatural base (Y) is 【Chemical Formula 35】 wherein in each case the wavy line indicates the bond to the ribosyl moiety, the eukaryotic cell according to any one of claims 1 to 8.
10. The eukaryotic cell according to any one of claims 1 to 9, wherein the first unnatural base (X) and the second unnatural base (Y) are the same.
11. The eukaryotic cell according to any one of claims 1 to 9, wherein the first unnatural base (X) and the second unnatural base (Y) are different.
12. The eukaryotic cell according to any one of claims 1 to 11, wherein the three consecutive nucleobases in the codon are selected from AXC, GXC or GXU, and X is the first unnatural base.
13. The eukaryotic cell according to claim 12, wherein the three consecutive nucleobases in the anticodon are selected from GYU, GYC and AYC, and Y is the second unnatural base.
14. The three consecutive nucleobases in the codon are AXC, and the three consecutive nucleobases in the anticodon are GYU; or the three consecutive nucleobases in the codon are GXC, and the three consecutive nucleobases in the anticodon are GYC; or the three consecutive nucleobases in the codon are GXU, and the three consecutive nucleobases in the anticodon are AYC, the eukaryotic cell according to claim 12.
15. The first unnatural base or the second unnatural base is OH, substituted lower alkyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F or a combination thereof; O-alkyl, S-alkyl, N-alkyl or a combination thereof; O-alkenyl, S-alkenyl, N-alkenyl or a combination thereof; O-alkynyl, S-alkynyl, N-alkynyl or a combination thereof; Modifications at the 2'-position including O-alkyl-O-alkyl, 2'-F, 2'-OCH₃, 2'-O(CH₂)₂OCH₃, or combinations thereof (wherein alkyl, alkenyl, and alkynyl are substituted or unsubstituted C₁-C₁₀ alkyl, C₂-C₁₀ alkenyl, C₂-C₁₀ alkynyl, -O[(CH₂)ₙO]ₘCH₃, -O(CH₂)ₙOCH₃, -O(CH₂)ₙNH₂, -O(CH₂)ₙCH₃, -O(CH₂)ₙ-NH₂, or -O(CH₂)ₙON[(CH₂)ₙCH₃)]₂, and n and m are from 1 to about 10); Modifications at the 5'-position including 5'-vinyl, 5'-methyl (R or S), or combinations thereof ; Modifications at the 4'-position including 4'-S, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of the oligonucleotide, or group for improving the pharmacodynamic properties of the oligonucleotide, or combinations thereof ; or combinations thereof A eukaryotic cell according to any one of claims 1 to 14, comprising a modified sugar moiety independently selected from these.
16. At least one non-natural amino acid is: A lysine analog; Comprising an aromatic side chain; Comprising an azide group; Comprising an alkyne group; or Comprising an aldehyde or ketone group, a eukaryotic cell according to any one of claims 1 to 15.
17. At least one non-natural amino acid comprises N6-((azidoethoxy)-carbonyl)-L-lysine (AzK), a eukaryotic cell according to claim 16.
18. The eukaryotic cell is a human cell, the eukaryotic cell according to any one of claims 1 to 17.
19. The cell is a mammalian cell, the eukaryotic cell according to any one of claims 1 to 18.
20. The cell is a HEK293T cell or a Chinese hamster ovary (CHO) cell, the eukaryotic cell according to any one of claims 1 to 17.
21. Further comprising a polypeptide translated from mRNA, the polypeptide comprising at least one non-natural amino acid and a mammalian glycosylation pattern, the eukaryotic cell according to any one of claims 1 to 20.
22. An isolated cell, the eukaryotic cell according to any one of claims 1 to 21.
23. The pyrrolysyl or tyrosyl tRNA is derived from Methanococcus jannaschii, Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans, and / or The pyrrolysyl or tyrosyl tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans, the eukaryotic cell according to any one of claims 1 to 22.
24. A semi-synthetic organism comprising the eukaryotic cell according to any one of claims 1 to 23.
25. A eukaryotic cell culture comprising a plurality of eukaryotic cells according to any one of claims 1 to 23.
26. A method of delivering a cell to an organism, the method comprising the step of contacting the organism with the cell according to any one of claims 1 to 23.
27. The organism is a mammal, the method according to claim 26.
28. The mammal is a human, the method according to claim 27.
29. A method for producing a polypeptide comprising at least one unnatural amino acid in a eukaryotic cell, comprising: (i) a messenger RNA (mRNA) comprising a codon containing a first unnatural base, the mRNA encoding a polypeptide comprising at least one unnatural amino acid encoded by the codon; (ii) introducing into the eukaryotic cell a pyrrolysyl or tyrosyl transfer RNA (tRNA) comprising an anticodon containing a second unnatural base, the first and the second unnatural bases are capable of forming an unnatural base pair (UBP) in the eukaryotic cell, and the codon and the anticodon each independently comprise three consecutive nucleobases (N-N-N), the eukaryotic cell comprises an expression vector encoding a pyrrolysyl or tyrosyl tRNA synthetase that aminoacylates the tRNA with at least one unnatural amino acid, the first unnatural base (X) is located at the central position (N-X-N) of the codon, and the second unnatural base (Y) is located at the central position (N-Y-N) of the anticodon, or the first unnatural base (X) is located at the last position (N-N-X) of the codon, and the second unnatural base (Y) is located at the first position (Y-N-N) of the anticodon, the eukaryotic cell is capable of translating the mRNA into a polypeptide comprising at least one unnatural amino acid using the tRNA, said method.
30. The method according to claim 29, wherein the tRNA is charged with an unnatural amino acid.
31. A method for producing a polypeptide comprising at least one unnatural amino acid in a eukaryotic cell, comprising: culturing the eukaryotic cell according to any one of claims 1 to 23 under conditions such that in the eukaryotic cell, the mRNA is translated into a polypeptide comprising at least one unnatural amino acid, using the tRNA and ribosomes endogenous to the eukaryotic cell. The method as described above, which comprises **Claim 32** A method for producing a polypeptide in a eukaryotic cell, wherein the polypeptide comprises at least one unnatural amino acid, the method comprising: Contacting the eukaryotic cell according to any one of claims 1 to 23 with at least one unnatural amino acid under conditions such that the mRNA is translated in the eukaryotic cell into a polypeptide comprising at least one unnatural amino acid. **Claim 33** A system for producing an unnatural polypeptide, comprising: (a) at least one unnatural amino acid; (b) an mRNA encoding an unnatural polypeptide comprising at least one unnatural amino acid, the mRNA comprising a codon containing a first unnatural base; (c) a pyrrolysyl or tyrosyl tRNA comprising an anticodon containing a second unnatural base; and (d) a eukaryotic ribosome capable of translating the mRNA into a polypeptide comprising at least one unnatural amino acid using the pyrrolysyl or tyrosyl tRNA wherein (i) the pyrrolysyl or tyrosyl tRNA is charged with the at least one unnatural amino acid, or (ii) the system further comprises one or more nucleic acid constructs comprising a pyrrolysyl or tyrosyl tRNA synthetase or a nucleic acid sequence encoding a pyrrolysyl or tyrosyl tRNA synthetase, and the pyrrolysyl or tyrosyl tRNA synthetase is characterized by aminoacylating the tRNA with the at least one unnatural amino acid. The first unnatural base and the second unnatural base are capable of forming an unnatural base pair (UBP), and the codon and the anticodon each independently comprise three consecutive nucleobases (N-N-N). The first unnatural base (X) is located at the central position (N-X-N) of the codon, and the second unnatural base (Y) is located at the central position (N-Y-N) of the anticodon, or The first unnatural base (X) is located at the last position of the codon (N-N-X), and the second unnatural base (Y) is located at the first position of the anticodon (Y-N-N), said system. **Claim 34**: The first unnatural base and the second unnatural base are each independently, **Chemical Formula 36** selected from, wherein R 2 is selected from hydrogen, alkyl, alkenyl, alkynyl, methoxy, methanethiol, methaneseleno, halogen, cyano and azide, and the wavy line indicates the bond to the ribosyl moiety, the system according to claim 33. **Claim 35**: The first unnatural base or the second unnatural base is **Chemical Formula 37** independently selected from, and the wavy line indicates the bond to the ribosyl moiety, the system according to claim 33 or 34. **Claim 36**: (a) (i) When the first unnatural base is **Chemical Formula 38** then the second unnatural base is **Chemical Formula 39** or (ii) when the first unnatural base is **Chemical Formula 40** then the second unnatural base is **Chemical Formula 41** and the wavy line indicates the bond to the ribosyl moiety; or (b) (i) When the first unnatural base is **Chemical Formula 42** then the second unnatural base is **Chemical Formula 43** or (ii) when the first unnatural base is **Chemical Formula 44** then the second unnatural base is **Chemical Formula 45** and the wavy line indicates the bond to the ribosyl moiety; or (c) (i) When the first unnatural base is [Chemical Formula 46] then the second unnatural base is [Chemical Formula 47] or (ii) when the first unnatural base is [Chemical Formula 48] then the second unnatural base is [Chemical Formula 49] and the wavy line indicates the bond to the ribosyl moiety; or (d) (i) When the first unnatural base is [Chemical Formula 50] then the second unnatural base is [Chemical Formula 51] or (ii) when the first unnatural base is [Chemical Formula 52] then the second unnatural base is [Chemical Formula 53] and the wavy line indicates the bond to the ribosyl moiety; or (e) (i) When the first unnatural base is [Chemical Formula 54] then the second unnatural base is [Chemical Formula 55] or (ii) when the first unnatural base is [Chemical Formula 56] then the second unnatural base is [Chemical Formula 57] and the wavy line indicates the bond to the ribosyl moiety; or (f) (i) When the first unnatural base is [Chemical Formula 58] then the second unnatural base is 【Chemical Formula 59】 is, or (ii) the first unnatural base is 【Chemical Formula 60】 in the case of, the second unnatural base is 【Chemical Formula 61】 where the wavy line indicates the bond to the ribosyl moiety; or (g) the first unnatural base and the second unnatural base are each 【Chemical Formula 62】 where the wavy line indicates the bond to the ribosyl moiety; or (h) the first unnatural base is 【Chemical Formula 63】 selected from, the wavy line indicates the bond to the ribosyl moiety, the system according to claim 35.
37. The first unnatural base (X) is located at the central position (N-X-N) in the codon, and the first unnatural base is 【Chemical Formula 64】 selected from, the wavy line indicates the bond to the ribosyl moiety; or the first unnatural base (X) is located at the last position (N-N-X) in the codon, and the first unnatural base is 【Chemical Formula 65】 selected from, the wavy line indicates the bond to the ribosyl moiety; or the second unnatural base (Y) is located at the first position (Y-N-N) of the anticodon, and the second unnatural base is 【Chemical Formula 66】 selected from, the wavy line indicates the bond to the ribosyl moiety; or the second unnatural base (Y) is located at the central position (N-X-N) of the anticodon, and the second unnatural base is 【Chemical Formula 67】 selected from, the wavy line indicates the bond to the ribosyl moiety, the system according to claim 33. **Claim 38**: The first unnatural base (X) and the second unnatural base (Y) are identical and / or the first unnatural base (X) and the second unnatural base (Y) each **Figure 68** are selected from, where the wavy line indicates the bond to the ribosyl moiety, the system according to claim 33. **Claim 39**: The first unnatural base (X) and the second unnatural base (Y) each **Figure 69** are selected from, where the wavy line indicates the bond to the ribosyl moiety, the system according to claim 38. **Claim 40**: One or more of the first unnatural bases (X) is **Figure 70** selected from, one or more of the second unnatural bases (Y) is **Figure 71** and in each case the wavy line indicates the bond to the ribosyl moiety, the system according to claim 39. **Claim 41**: The first unnatural base (X) and the second unnatural base (Y) are different and / or the first unnatural base (X) and the second unnatural base (Y) each **Figure 72** are independently selected from, where the wavy line indicates the bond to the ribosyl moiety, the system according to claim 33. **Claim 42**: The first unnatural base (X) and the second unnatural base (Y) each **Figure 73** are independently selected from, where the wavy line indicates the bond to the ribosyl moiety, the system according to claim 41. **Claim 43**: One or more of the first unnatural bases (X) is **Figure 74** selected from, one or more of the second unnatural bases (Y) is **Figure 75** and in each case the wavy line indicates the bond to the ribosyl moiety, the system according to claim 42.
44. The three consecutive nucleobases in the codon are selected from AXC, GXC or GXU, where X is a first unnatural base, the system according to any one of claims 33 to 43.
45. The three consecutive nucleobases in the anticodon are selected from GYU, GYC and AYC, where Y is a second unnatural base, the system according to claim 44.
46. The codon is AXC and the anticodon is GYU, or the codon is GXC and the anticodon is GYC, or the codon is GXU and the anticodon is AYC, the system according to claim 45.
47. Pyrrolysyl or tyrosyl tRNA is derived from Methanococcus jannaschii, Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans; pyrrolysyl or tyrosyl tRNA synthetase is derived from Methanococcus jannaschii, Methanosarcina barkeri, Methanosarcina mazei or Methanosarcina acetivorans, the system according to any one of claims 33 to 46.
48. The system is in vitro or cell-free; the system contains a cell lysate; and / or the system is a reconstituted system of purified components, the system according to any one of claims 33 to 47.
49. The system is a system in a eukaryotic cell, the system according to any one of claims 33 to 47.
50. At least one unnatural amino acid is: a lysine analog; contains an aromatic side chain; contains an azide group; contains an alkyne group; or contains an aldehyde or ketone group, the system according to any one of claims 33 to 49.
51. The system of claim 50, wherein at least one unnatural amino acid comprises N6-((azidoethoxy)-carbonyl)-L-lysine (AzK). **Claim 52.** The system of any one of claims 33-51, wherein pyrrolyl or tyrosyl tRNA is charged with at least one unnatural amino acid.
Citation Information
Patent Citations
Means and methods for preparing genetically engineered proteins by genetic code expansion in insect cells
JP2018534943A
Methods for Incorporating Unnatural Amino Acids in Eukaryotic Cells
US20130183761A1
Methods of genetically encoding unnatural amino acids in eukaryotic cells using orthogonal trna / synthetase pairs
WO2008127900A1
Mutant pyrrolysyl-trna synthetase, and method for production of protein having non-natural amino acid integrated therein by using the same
WO2009038195A1
Improving unnatural amino acid incorporation in eukaryotic cells
WO2010141851A1