How to Optimize Protein Production

JP2024527392A5Pending Publication Date: 2025-09-02UNITED KINGDOM RESEARCH AND INNOVATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024501657
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-14
Filing Date
2022-07-14
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing methods for optimizing protein production using orthogonal ribosomes result in low and inconsistent yields, particularly when incorporating non-standard amino acids, due to inefficiencies in the translation initiation process.

Method used

A method for designing orthogonal messenger RNA (O-mRNA) sequences that optimize the 5' untranslated region (UTR) to enhance translation efficiency by orthogonal ribosomes, using thermodynamic modeling and simulated annealing to introduce modifications that improve the free energy difference, ensuring efficient and specific translation by O-ribosomes.

Benefits of technology

The method significantly enhances protein yield by orthogonal ribosomes, achieving up to 40 times more protein production compared to previous methods, with improved orthogonality and efficiency, comparable to wild-type mRNA translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to novel methods for optimizing protein production. These methods include methods for optimizing orthogonal mRNAs, methods for designing and producing optimal operons containing exogenous tRNAs, and methods for designing and producing optimal operons containing exogenous genes, such as exogenous genes encoding orthogonal aminoacyl-tRNA synthetases (O-aaRSs). The present invention also relates to the products of said methods. Also provided as part of the present invention are host cells containing the products of these innovations, methods of using said cells, and products thereof. The host cells of the present invention may be used for improved production of proteins and polypeptides containing genetically incorporated non-standard amino acids.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to novel methods for optimizing protein production. These methods include methods for optimizing orthogonal mRNAs, methods for designing and producing optimal operons containing exogenous tRNAs, and methods for designing and producing optimal operons containing exogenous genes, such as exogenous genes encoding orthogonal aminoacyl-tRNA synthetases (O-aaRSs). The present invention also relates to the products of said methods. Also provided as part of the present invention are host cells comprising the products of these innovations, methods of using said cells, and products thereof. The host cells of the present invention may be used for improved production of proteins and polypeptides that contain genetically incorporated non-standard amino acids. [Background technology]

[0002] The ability to genetically encode the incorporation of multiple distinct non-canonical amino acids (ncAAs) into proteins offers new opportunities for engineering and directed evolution of protein function, enables new strategies for biological discovery and understanding of biological processes, and provides the basis for encoded cellular synthesis of non-canonical biopolymers (Chin, JW Expanding and reprogramming the genetic code. Nature 550, 53-60 (2017); de la Torre, D. & Chin, JW Reprogramming the genetic code. Nat. Rev. Genet., 1-16 (2020)). Encoding multiple distinct ncAAs in a protein synthesized in a cell requires orthogonal codons in addition to the codons used to encode natural proteins synthesized in the same cell; these include quadruplet codons (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Wang, K. et al. Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393-403 (2014); Anderson, JC et al. An expanded genetic code with a functional quadruplet codon. Proc. Natl. Acad. Sci. USA 101, 7566-7571 (2004)), codons resulting from sense codon compression (Fredens, J. et al.Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514- 518 (2019), Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59-64 (2016)), and codons incorporating non-standard bases (Malyshev, D.A. et al. A semi-synthetic organism with an expanded genetic alphabet. Nature 509, 385-388 (2014), Zhang, Y. et al. A semi-synthetic organism that stores and retrieves increased genetic information. Nature 551, 644-647 (2017), Zhang, Y. et al. A semi-synthetic organism engineered for the stable expansion of the genetic alphabet. Proc. Natl. Acad. Sci. U.S.A. 114, 1317-1322 (2017), Fischer, E.C. et al. New codons for efficient production of unnatural proteins in a semi-synthetic organism. Nat. Chem. Biol.16, 570-576 (2020)). Orthogonal codons must be assigned to ncAAs using engineered mutually orthogonal aminoacyl-tRNA synthetase (aaRS) / tRNA pairs. These pairs should be orthogonal in their aminoacylation specificity with respect to the synthetases and tRNAs used by the host organism for natural translation, and with respect to other orthogonal aaRSs and tRNAs used to direct ncAAs in the same cell; furthermore, they should specifically recognize distinct ncAA monomers and decode distinct orthogonal codons (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Neumann, H., Slusarczyk, AL & Chin, JW De Novo Generation of Mutually Orthogonal Aminoacyl-tRNA Synthetase / tRNA Pairs. J. Am. Chem. Soc. 132, 2142-2144 (2010), Chatterjee, A., Sun, SB, Furman, JL, Xiao, H. & Schultz, PG A Versatile Platform for Single- and Multiple-Unnatural Amino Acid Mutagenesis in Escherichia coli. Biochemistry 52, 1828-1837 (2013), Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nat. Chem. 10, 831-837 (2018), Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JWEngineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non canonical amino acids. Nat. Chem. 12, 535-544 (2020)、Cervettini, D. et al. Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase-tRNA pairs. Nat. Biotechnol. 38, 989-999 (2020)、Zhang, M.S. et al. Biosynthesis and genetic encoding of phosphothreonine through parallel selection and deep sequencing. Nat. Methods 14, 729-736 (2017)、Italia, J. et al. Mutually Orthogonal Nonsense-Suppression Systems and Conjugation Chemistries for Precise Protein Labeling at up to Three Distinct Sites. J. Am. Chem. Soc. 141, 6204-6212 (2019))。.

[0003] Orthogonal ribosomes (O-ribosomes) are non-natural ribosomes directed against orthogonal mRNAs (O-mRNAs) that are not substrates for the wild-type (wt) ribosome in Escherichia coli (E. coli). These ribosomes operate in parallel with the natural ribosomes, but contain modifications in their ribosomal RNA that direct them to the O-ribosome binding site (O-RBS) in the 5' untranslated region (5'UTR) of the orthogonal message (Rackham, O. & Chin, JW A network of orthogonal ribosome·mRNA pairs. Nat. Chem. Biol. 1, 159-166 (2005)). Because O-ribosomes are not involved in the synthesis of the proteome, they can be engineered to perform new functions not accessed by native ribosomes, including de novo decoding and new intrinsic polymerization functions (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Wang, K., Neumann, H., Peak-Chew, SY & Chin, JW Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol. 25, 770-777 (2007); Schmied, WH et al. Controlling orthogonal ribosome subunit interactions enables evolution of new function. Nature 564, 444-448 (2018)).O-riboQ1 (evolved O-ribosome) efficiently decodes amber and quadruplet codons on the O-mRNA using cognate tRNAs, thus providing orthogonal codons that are selectively decoded on the orthogonal message (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Wang, K., Neumann, H., Peak-Chew, SY & Chin, JW Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol. 25, 770-777 (2007)).

[0004] Engineered mutually orthogonal aaRS / tRNA pairs that recognize distinct ncAAs and decode distinct codons have been used to incorporate two or three distinct ncAAs into proteins (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Wang, K. et al. Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393-403 (2014); Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nat. Chem. 10, 831-837). (2018), Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non canonical amino acids. Nat. Chem. 12, 535-544 (2020), Italia, J. et al. Mutually Orthogonal Nonsense-Suppression Systems and Conjugation Chemistries for Precise Protein Labeling at up to Three Distinct Sites. J. Am. Chem. Soc. 141, 6204-6212 (2019), Venkat, S. et al.Genetically Incorporating Two Distinct Post-translational Modifications into One Protein Simultaneously. ACS Synth. Biol. 7, 689-695 (2018)). The homologous Methanosarcina mazei (Mm) or Methanosarcina barkeri (Mb) pyrrolysyl-tRNA is a widely used orthogonal aaRS / tRNA pair for genetic code expansion (de la Torre, D. & Chin, JW Reprogramming the genetic code. Nat. Rev. Genet., 1-16 (2020); Chin, JW Expanding and Reprogramming the Genetic Code of Cells and Animals. Annu. Rev. Biochem. 83, 379-408 (2014)). We recently examined PylRS / tRNAPyl pairs from diverse organisms and found that natural PylRS and tRNAPyl sequences cluster into multiple subclasses with distinct specificities; this insight enabled us to engineer doubly and triply orthogonal PylRS / tRNAPyl pairs that recognize distinct ncAAs and decode distinct codons (Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nat. Chem. 10, 831-837 (2018); Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non canonical amino acids. Nat. Chem. 12, 535-544 (2020)). .

[0005] O(trans)-strepGFP (40TAG, 136AGGA or 150AGTA) His6 (It contains an O-ribosome binding site (O(trans)) and is translated from a previously described 5'UTR that contains two quadruplet codons (AGGA and AGTA) and an amber codon (TAG). Strep GFP His6 O-riboQ1-mediated translation of an open reading frame (ORF) (O-mRNA) was mediated by an engineered triply orthogonal PylRS / tRNA Pyl By combining with the pair, we have produced recombinant StrepGFP (40BocK, 136NmH, 150CbzK) His6demonstrated the incorporation of three ncAAs into tRNA (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non canonical amino acids. Nat. Chem. 12, 535-544 (2020)). However, as the inventors have noticed (Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nat. Chem. 10, 831-837 (2018); Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nat. Chem. 12, 535-544 (2020)), the protein yield from this expression system is low and not optimized. O(trans)- strep GFP His6 and has a 5'UTR containing a wt RBS strep GFP His6 Additional experiments with open reading frames demonstrated that O(trans)- strep GFP His6 Translation of is 31-fold lower than that produced by the wt ribosome. Strep GFP His6Furthermore, transferring the O(trans)5'UTR to another ORF also leads to a substantial reduction in the level of protein synthesis (Figure 1 and Supplementary Figure 1). The O(trans)5'UTR sequence was derived from a construct for producing a GST fusion protein in which the O(trans)5'UTR sequence directed O-ribosome-dependent translation at a level equivalent to O-ribosome-independent translation from a 5'UTR containing a wt RBS (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Wang, K., Neumann, H., Peak-Chew, SY & Chin, JW Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol. 25, 770-777 (2007)). These observations demonstrated that O(trans) sequences direct efficient orthogonal translation for some ORFs, but do not provide a general solution for the efficient translation of ORFs.

[0006] Therefore, there is a need for a general solution for the generation of O-mRNAs that maximize protein yields in orthogonal translation. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Chin, JW Expanding and reprogramming the genetic code. Nature 550, 53-60 (2017) [Non-Patent Document 2] de la Torre, D. & Chin, JW Reprogramming the genetic code. Nat. Rev. Genet., 1-16 (2020) [Non-Patent Document 3] Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010) [Non-Patent Document 4] Wang, K. et al. Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393-403 (2014) [Non-Patent Document 5] Anderson, JC et al. An expanded genetic code with a functional quadruplet codon. Proc. Natl. Acad. Sci. USA 101, 7566-7571 (2004) [Non-Patent Document 6] Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514- 518 (2019) [Non-Patent Document 7] Wang, K. et al. Defining synonymous codon compression schemes by genome recording. Nature 539, 59-64 (2016) [Non-Patent Document 8] Malyshev, DA et al. A semi-synthetic organism with an expanded genetic alphabet. Nature509, 385–388 (2014).

Outdoor Tools9

Outdoor Tools 10

Outdoor Content11

Outdoor Tools 12

Outdoor Content13

Outdoor Tools 14

Outdoor Tools 15

Outdoor Content 16

Outdoor Track 17

Outdoor Tools 18

[0008] We provide herein a highly efficient method for optimizing protein production. [Means for solving the problem]

[0009] In an embodiment of the invention, there is provided a method for designing a messenger RNA (mRNA), which is an orthogonal messenger RNA (O-mRNA) suitable for translation by an orthogonal ribosome (O-ribosome), the mRNA comprising a 5' untranslated region (5'UTR) and an open reading frame (ORF), the method comprising: (a) predicting the free energy difference (ΔGtot(O-ribo)) between the freely folded state of the mRNA and the initiation-competent state of the mRNA bound to the O-ribosome; (b) introducing an alteration into the 5'UTR; (c) predicting the new ΔGtot(O-ribo) (ΔGtotnew(O-ribo)) after modification; (d) tolerating the modification if said ΔGtotnew(O-ribo) is more negative than the preceding ΔGtot(O-ribo); accepting or rejecting the modification according to a probability distribution if the ΔGtotnew(O-ribo) is more positive than the preceding ΔGtot(O-ribo); and (e) generating an O-mRNA sequence comprising a 5'UTR containing the tolerated modification; A method is provided, comprising:

[0010] ΔG tot(O-ribo) is the free energy required to unfold mRNA (ΔG unfolding ) and the free energy released when mRNA binds to the O-ribosome to form an initiation-competent state bound to the O-ribosome (ΔG o-ribo binding ) may also be used.

[0011] The O-ribosome may contain an orthogonal 16S rRNA, the mRNA may contain a Shine-Dalgarno sequence, and ΔG tot (O-ribo) is as follows: ΔG tot (O-ribo) = (ΔG mRNA-O-rRNA +ΔG start +ΔG spacing -ΔG standby )+ΔG unfolding may be predicted according to; ΔG mRNA-O-rRNA is the free energy of the predicted cofolded secondary structure of the last 9 nucleotides of the orthogonal 16S rRNA and the mRNA; ΔG start is the energy released from the binding of the initiator tRNA to the start codon of the ORF; ΔG spacing is the energy penalty for a non-optimal spacing length between the Shine-Dalgarno sequence and the start codon; ΔG standby is the energy required to unfold the secondary structure separating the four nucleotides upstream of the Shine-Dalgarno sequence; ΔG unfolding is the energy required to unfold the secondary structure in the mRNA.

[0012] ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more positive than the ΔG tot new (O-ribo) and the above ΔG totThe magnitude of the difference between (O-ribo) may determine the probability of acceptance, with a smaller magnitude being associated with a higher likelihood of acceptance compared to a larger magnitude.

[0013] The probability distribution according to which a modification will be accepted or rejected is

number

[0014] T SA may be adjusted to maintain a tolerance rate of 5-20%.

[0015] In an embodiment, the method further comprises the step of: nd a method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains an O-ribosome, Step (a) is the synthesis of the freely folded state of mRNA and the 2 nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd -ribo)); Step (c) is the new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd -ribo); Step (d) is the above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo) to allow modification, Said ΔG tot new(O-ribo) precedes ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo), then accept or reject the modification according to a probability distribution.

[0016] In certain embodiments, the second ribosome (2 nd 1. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains an O-ribosome, the O-mRNA comprising a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd - predicting ribo); (b) introducing an alteration into the 5'UTR; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd - a step of predicting ribo; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo) to allow modification, Said ΔG totnew (O-ribo) precedes ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo), accepting or rejecting the modification according to a probability distribution; and (e) generating an O-mRNA sequence comprising a 5'UTR containing the tolerated modification; A method is provided, comprising:

[0017] ΔG tot (2 nd -ribo) is the free energy required to unfold mRNA (ΔG unfolding ) and mRNA is 2 nd - Binds to ribosomes, 2 nd - the free energy released upon forming an initiation-competent state bound to the ribosome (ΔG 2nd ribo binding ) may also be used.

[0018] 2 nd The ribosome may contain 16S rRNA, the mRNA may contain a Shine-Dalgarno sequence, and ΔG tot (2 nd -ribo) is as follows: ΔG tot (2 nd -ribo)=(ΔG mRNA-2nd-rRNA +ΔG start +ΔG spacing -ΔG standby )+ΔG unfolding may be predicted according to; ΔG mRNA-2nd-rRNA is the free energy of the predicted cofolded secondary structure of the last 9 nucleotides of 16S rRNA and the mRNA; ΔG start is the energy released from the binding of the initiator tRNA to the start codon of the ORF; ΔG spacingis the energy penalty for a non-optimal spacing length between the Shine-Dalgarno sequence and the start codon; ΔG standby is the energy required to unfold the secondary structure separating the four nucleotides upstream of the Shine-Dalgarno sequence; ΔG unfolding is the energy required to unfold the secondary structure in the mRNA.

[0019] In an embodiment, ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) or ΔG tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo), the ΔG tot new (O-ribo) and the above ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) and the ΔG tot (2 nd The magnitude of the difference between the ribo and the ribo determines the probability of acceptance, with smaller magnitudes being associated with a higher likelihood of acceptance compared to larger magnitudes.

[0020] In an embodiment, Step (a) is calculated using the formula: ΔG tot (opt)=ΔG tot (O-ribo)-X*ΔG tot (2 nd -ribo) according to ΔG tot Calculating (opt); Step (c) is calculated using the formula: ΔG tot new (opt)=ΔG tot new (O-ribo)-X*ΔG tot new (2 nd -ribo) according to ΔGtot new Calculating (opt); Step (d) is the above ΔG tot new ΔG preceded by (opt) tot Allow the modification if it is more negative than (opt), Said ΔG tot new ΔG preceded by (opt) tot accept or reject the modification according to a probability distribution if (opt) is more positive than (opt); X is 0.1 to 2, or X is 0.5.

[0021] In an embodiment, ΔG tot new ΔG preceded by (opt) tot If the ΔG tot new (opt) and the above ΔG tot The magnitude of the difference between (opt) determines the probability of acceptance, with smaller magnitudes being associated with a higher likelihood of acceptance compared to larger magnitudes.

[0022] The probability distribution according to which a modification will be accepted or rejected is

number

[0023] T SA may be adjusted to maintain a tolerance rate of 5-20%.

[0024] The modification may be or may involve a single nucleotide change, insertion or deletion.

[0025] In embodiments, step (b) comprises introducing an alteration in the 5'UTR or replacing any one of codons 2-20, 2-15, 2-12, 2-10, or 2-5 in the ORF with a synonymous codon; and step (e) comprises generating an O-mRNA sequence comprising the 5'UTR and the ORF containing the tolerated alteration.

[0026] In an embodiment, step (b) comprises introducing a modification into the 5'UTR, including a single nucleotide change, insertion, or deletion, or replacing any one of codons 2 to 12 in the ORF with a synonymous codon.

[0027] In embodiments, steps (b)-(d) are repeated at least 200, 300, 400, 500, 1000, 5000, or 10000 times; or steps (b)-(d) are repeated at least 10, 50, 100, 250, or 500 successive iterations resulting in a more negative ΔG tot new (O-ribo); or steps (b)-(d) are repeated until at least 10, 50, 100, 250, or 500 successive iterations result in a more negative ΔG tot new (O-ribo) or more positive ΔG tot new (2 nd or steps (b)-(d) are repeated until at least 10, 50, 100, 250, or 500 successive iterations lead to a more negative ΔG tot new This is repeated until it no longer leads to (opt).

[0028] In an embodiment, the 5'UTR in step (a) is 35 nucleotides in length; or the modification is in any of the 35 nucleotides of the 5'UTR closest to the start codon. The 5'UTR in step (a) may be from a randomly generated nucleic acid sequence. The 5'UTR in step (a) may comprise a wild-type Shine-Dalgarno sequence.

[0029] The O-ribosome may comprise an orthogonal anti-Shine-Dalgarno sequence, and the 5'UTR of step (a) may comprise an orthogonal Shine-Dalgarno sequence (O-SD) that is predicted to be perfectly complementary to the orthogonal anti-Shine-Dalgarno sequence.

[0030] In some embodiments, step (b) does not include introducing an alteration into the 5 nucleotide core of the O-SD.

[0031] The Shine-Dalgarno sequence may be 5 nucleotides from the start codon of the ORF.

[0032] In an embodiment, two nd The ribosome is a wild-type ribosome or nd The ribosome is an O-ribosome that is different from the first O-ribosome.

[0033] The method for designing an O-mRNA may be performed on a computer.

[0034] In an embodiment of the invention, a method is provided for producing a nucleic acid sequence encoding an exogenous protein for translation by the O-ribosome, in which the sequence of the O-mRNA is designed according to any of the methods for designing an O-mRNA disclosed herein, and then a nucleic acid molecule encoding said sequence is produced.

[0035] In an embodiment of the invention, there is provided a system for designing an orthogonal messenger RNA (O-mRNA) for translation by an orthogonal ribosome (O-ribosome), comprising: Processor; and One or more computer readable storage media having stored thereon instructions for execution on said processor for performing any of the methods of designing an O-mRNA disclosed herein. A system is provided that includes:

[0036] In an embodiment of the invention, a computer program product is provided that includes a non-transitory machine-readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to perform any of the methods for designing O-mRNA disclosed herein.

[0037] In another embodiment of the invention, there is provided a method for designing an operon encoding at least two exogenous tRNAs for expression in a host cell comprising an endogenous genome encoding an endogenous tRNA, comprising the steps of: (i) generating permutations of at least two exogenous tRNAs; (ii) identifying, within the endogenous genome, adjacent pairs of endogenous tRNAs that have the highest level of sequence identity to each adjacent pair of exogenous tRNAs within each permutation of the at least two exogenous tRNAs; (iii) identifying intergenic regions in the endogenous genome between each of the identified adjacent pairs of endogenous tRNAs; (iv) generating a plurality of sequences encoding each of the at least two exogenous tRNA permutations, the sequences including the identified intergenic regions located between each associated adjacent pair of the exogenous tRNAs; and (v) selecting sequences from said plurality of sequences for inclusion in an operon encoding at least two exogenous tRNAs. A method is provided, comprising:

[0038] The selection in step (v) may be made from a ranked list of a plurality of sequences, where the ranked list is generated by ranking each of the plurality of sequences based on the sum of sequence identities between at least two exogenous tRNAs and the corresponding endogenous tRNAs used to define the intergenic region.

[0039] The sequence identity in step (ii) may be calculated by comparing the acceptor stem sequence of the endogenous tRNA to the acceptor stem sequence of the exogenous tRNA. The first 7 nucleotides of the tRNA and the last 8 nucleotides not including the CCA tail may be compared.

[0040] The minimum intergenic region considered may be 5, 10, 15, 20, or 25 base pairs, and the maximum intergenic region may be 50, 75, 100, 125, or 150 base pairs. In an embodiment, the minimum intergenic region considered is 10 base pairs, and the maximum is 100 base pairs.

[0041] The method may be a method for designing an operon encoding at least three, at least four, at least five, or at least six exogenous tRNAs.

[0042] Any of the methods for designing an operon encoding at least two exogenous tRNAs may be performed on a computer.

[0043] In an embodiment of the invention, a method is provided for producing a nucleic acid sequence encoding an operon comprising at least two exogenous tRNAs, wherein the sequence of the nucleic acid is designed according to any of the methods for designing an operon encoding at least two exogenous tRNAs disclosed herein, and then the nucleic acid encoding said sequence is produced.

[0044] In an embodiment of the present invention, there is provided a system for designing an operon comprising at least two exogenous tRNAs, comprising: Processor; and One or more computer readable storage media having stored thereon instructions for execution on said processor for performing any of the methods for designing an operon encoding at least two exogenous tRNAs as disclosed herein. A system is provided that includes:

[0045] In an embodiment of the present invention, a computer program product is provided that includes a non-transitory machine-readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to perform any of the methods for designing an operon encoding at least two exogenous tRNAs disclosed herein.

[0046] In an embodiment of the present invention, a nucleic acid is provided that comprises an operon obtained or obtainable by any of the methods of designing an operon encoding at least two exogenous tRNAs disclosed herein.

[0047] In an embodiment of the invention, a host cell is provided that comprises an endogenous genome, wherein the host cell comprises a nucleic acid encoding an operon comprising at least two exogenous tRNAs, wherein the nucleic acid sequence between each pair of exogenous tRNAs is an intergenic sequence derived from the endogenous genome.

[0048] The host cell may comprise an operon obtained or obtainable by any of the methods disclosed herein for designing an operon encoding at least two exogenous tRNAs.

[0049] The host cell may be a prokaryotic cell, for example a bacterial cell. The bacterial cell may be E. coli and the endogenous genome may be the E. coli genome.

[0050] In another embodiment of the invention, there is provided a method for designing an operon comprising at least two exogenous ORFs for expression in a host cell, comprising the steps of: (i) generating a plurality of 5'UTR sequences for each of at least two exogenous ORFs, wherein each 5'UTR sequence is a sequence that is a negative predicted free energy difference (ΔG tot (ribo)) is optimized for the step; (ii) ΔG for each of the 5′ UTR sequences when the 5′ UTR is positioned 5′ to the optimized exogenous ORF and 3′ to each of the remaining exogenous ORFs of the at least two exogenous ORFs. tot predicting (ribo); and (iii) selecting a 5'UTR sequence and an arrangement of at least two exogenous ORFs; A method is provided, comprising:

[0051] Step (iii) may comprise selecting a 5' UTR sequence and an arrangement of the at least two exogenous ORFs; ΔG for all 5'UTR / exogenous ORF pairs tot The sum of (ribo) is most negative; and / or ΔG for all 5'UTR / exogenous ORF pairs tot The mean value of (ribo) is the most negative; and / or Each 5'UTR / exogenous ORF pair has a target ΔG tot More negative ΔG than (ribo) tot (ribo).

[0052] Step (i) may involve generating two, three, four, five, six or more 5'UTR sequences for each of the at least two exogenous ORFs.

[0053] In an embodiment, at least one or all of the at least two exogenous ORFs is an aminoacyl-tRNA synthetase.

[0054] The method may be a method for designing an operon encoding at least three, at least four, at least five, or at least six exogenous ORFs.

[0055] ΔG tot (ribo) is the free energy required to unfold mRNA (ΔG unfolding) and the free energy released when mRNA binds to the ribosome to form a ribosome-bound initiation-competent state (ΔG ribo binding ) may also be used.

[0056] ΔG tot (ribo) is as follows: ΔG tot (ribo)=(ΔG mRNA-rRNA +ΔG start +ΔG spacing -ΔG standby )+ΔG unfolding may be predicted according to; ΔG mRNA-rRNA is the free energy of the predicted cofolded secondary structure of the last 9 nucleotides of 16S rRNA and the mRNA; ΔG start is the energy released from the binding of the initiator tRNA to the start codon of the sequence encoding the exogenous ORF; ΔG spacing is the energy penalty for a non-optimal spacing length between the Shine-Dalgarno sequence and the start codon of the sequence encoding the exogenous ORF; ΔG standby is the energy required to unfold the secondary structure separating the four nucleotides upstream of the Shine-Dalgarno sequence; ΔG unfolding is the energy required to unfold the secondary structure in the mRNA.

[0057] In an embodiment, step (i) comprises: (a) introducing an alteration into the 5'UTR; (b) New ΔG after modification tot (ribo)(ΔG tot new predicting (ribo); (c) Said ΔG tot new ΔG preceded by (ribo) tot Allows modification if it is more negative than (ribo), Said ΔG tot new ΔG preceded by (ribo) tot accept or reject the modification according to a probability distribution if ribo is more positive than ribo; and (d) generating a 5'UTR sequence containing the tolerated modifications; Includes.

[0058] In an embodiment, ΔG tot new ΔG preceded by (ribo) tot If the ΔG tot new (ribo) and the above ΔG tot The magnitude of the difference between (ribo) determines the probability of acceptance, with smaller magnitudes being associated with a higher likelihood of acceptance compared to larger magnitudes.

[0059] The probability distribution according to which a modification will be accepted or rejected is

number

[0060] T SA may be adjusted to maintain a tolerance rate of 5-20%.

[0061] The modification may be or may include a single nucleotide change, insertion, or deletion. In embodiments, step (a) includes introducing a modification in the 5'UTR or replacing any one of codons 2-20, 2-15, 2-12, 2-10, or 2-5 with a synonymous codon in the sequence encoding the exogenous ORF; step (d) includes generating a sequence including the 5'UTR and ORF containing the tolerated modification. In certain embodiments, step (a) includes introducing a modification in the 5'UTR including a single nucleotide change, insertion, or deletion, or replacing any one of codons 2-12 in the ORF with a synonymous codon.

[0062] Steps (a)-(c) may be repeated at least 200, 300, 400, 500, 1000, 5000, or 10000 times. Alternatively, steps (a)-(c) may be repeated at least 10, 50, 100, 250, or 500 successive iterations resulting in a more negative ΔG tot new This may be repeated until it no longer leads to (ribo).

[0063] Any of the methods for designing an operon comprising at least two exogenous ORFs disclosed herein may be performed on a computer.

[0064] In an embodiment of the invention, a method is provided for producing a nucleic acid sequence encoding a polycistronic operon comprising at least two exogenous ORFs, wherein the sequence of the nucleic acid is designed according to any of the methods for designing an operon comprising at least two exogenous ORFs disclosed herein, and then a nucleic acid is produced according to said sequence.

[0065] In an embodiment of the present invention, there is provided a system for designing a polycistronic operon comprising at least two exogenous ORFs, comprising: Processor; and One or more computer readable storage media having stored thereon instructions for execution on said processor for performing any of the methods for designing an operon comprising at least two exogenous ORFs as disclosed herein. A system is provided that includes:

[0066] In an embodiment of the invention, a computer program product is provided that includes a non-transitory machine-readable medium that stores program code that, when executed by one or more processors of a computer system, causes the computer system to perform any of the methods for designing an operon comprising at least two exogenous ORFs disclosed herein.

[0067] In an embodiment of the invention, a nucleic acid is provided that comprises an operon obtained or obtainable by any of the methods disclosed herein for designing an operon comprising at least two exogenous ORFs.

[0068] In an embodiment of the invention, a host cell is provided that comprises a nucleic acid encoding an operon obtained or obtainable by any of the methods disclosed herein for designing an operon comprising at least two exogenous ORFs.

[0069] The host cell may be a prokaryotic cell, for example a bacterial cell. The bacterial cell may be E. coli and the endogenous genome may be the E. coli genome.

[0070] In an embodiment of the present invention, a nucleic acid sequence encoding an O-mRNA encoding an exogenous protein, the O-mRNA being obtained or obtainable by any of the methods for designing an O-mRNA disclosed herein, the O-mRNA comprising at least two types of orthogonal codons; a nucleic acid sequence comprising an O-tRNA operon encoding at least two orthogonal tRNAs, the at least two orthogonal tRNAs having the ability to decode the at least two types of orthogonal codons, the operon being obtained or obtainable by any of the methods for designing an O-tRNA operon disclosed herein; A nucleic acid sequence comprising an orthogonal aminoacyl-tRNA synthetase (O-aaRS) encoding at least two O-aaRS operons, wherein the at least two O-aaRSs form O-aaRS-O-tRNA pairs with at least two orthogonal tRNAs, the operon being obtained or obtainable by any of the methods for designing an operon encoding at least two exogenous genes disclosed herein; and Orthogonal Ribosomes A host cell is provided, comprising:

[0071] In an embodiment, The O-mRNA contains at least three types of orthogonal codons; the O-tRNA operon encodes at least three orthogonal tRNAs capable of decoding said at least three orthogonal codons; An O-aaRS operon encodes at least three O-aaRSs that form O-aaRS-O-tRNA pairs with at least three orthogonal tRNAs.

[0072] In an embodiment, O-mRNA contains at least four types of orthogonal codons; the O-tRNA operon encodes at least four orthogonal tRNAs capable of decoding said at least four orthogonal codons; The O-aaRS operon encodes at least four O-aaRSs that form O-aaRS-O-tRNA pairs with at least four orthogonal tRNAs.

[0073] The host cell may be a prokaryotic cell, for example a bacterial cell. The bacterial cell may be E. coli and the endogenous genome may be the E. coli genome.

[0074] In an embodiment of the invention there is provided a method for producing a polypeptide comprising the steps of: providing a host cell comprising an O-ribosome, an O-tRNA operon, and an O-aaRS operon as disclosed herein; incubating a host cell in the presence of a first non-standard amino acid, the first non-standard amino acid being a substrate for one of the O-aaRSs; and Incubating the host cell to allow incorporation of the first non-standard amino acid into the polypeptide via the O-aaRS - O-tRNA pair. A method is provided, comprising:

[0075] In an embodiment, the method comprises: incubating the host cell in the presence of a second non-standard amino acid, where the second non-standard amino acid is a substrate for one of the O-aaRSs; and Incubating the host cell to allow incorporation of the second non-standard amino acid into the polypeptide via the O-aaRS - O-tRNA pair. Includes.

[0076] In an embodiment, the method comprises: incubating the host cell in the presence of a third non-standard amino acid, where the third non-standard amino acid is a substrate for one of the O-aaRSs; and Incubating the host cell to allow incorporation of the third non-standard amino acid into the polypeptide via the O-aaRS - O-tRNA pair. Includes.

[0077] In an embodiment, the method comprises: incubating the host cell in the presence of a fourth non-standard amino acid, where the fourth non-standard amino acid is a substrate for one of the O-aaRSs; and Incubating the host cell to allow incorporation of the fourth non-standard amino acid into the polypeptide via the O-aaRS - O-tRNA pair. Includes.

[0078] In another aspect of the present invention there is provided a polypeptide obtained or obtainable by any method which produces a polypeptide as disclosed herein. [Brief description of the drawings]

[0079] [Figure 1-1] ~ [Figure 1-2]Automated discovery of O-mRNA sequences that are specifically and efficiently translated by O-ribosomes. a, Thermodynamic model for initiation of protein synthesis by wt and O-ribosomes on mRNA. The free energy for formation of the initiation complex (ΔGtot) is the sum of the free energy required to unfold the mRNA (ΔGunfolding) and the free energy released when the mRNA forms the initiation complex through binding to the ribosomal 30S subunit and tRNAfMet CAU (black trident and yellow star) (ΔGribo binding). The 30S subunit of the O-ribosome (light brown) contains an orthogonal anti-Shine-Dalgarno (O-aSD) at the 3' end of the O-16S rRNA, and the 30S subunit of the wt ribosome (dark brown) contains a wt anti-Shine-Dalgarno (wt aSD) at the 3' end of its 16S rRNA. The free energies released upon formation of the initiation complex from the unfolded mRNA with the wt and orthogonal 30S are ΔGwt ribo binding and ΔG0-ribo binding, respectively. Details on the calculations are provided in the Methods. ORF open reading frame (orange), start codon (purple), SD / O-SD Shine-Dalgarno sequence or orthogonal version (green), spacing between SD / O-SD and start codon (blue). The remaining part of the 5'UTR is shown in grey. b, Algorithm developed to predict O-mRNA sequences that are efficiently and specifically translated by the O-ribosome. The algorithm vol 1 generates a random 35 nucleotide 5'UTR containing the wt SD sequence and predicts its ΔGtot(O-ribo). In an iterative process, mutations are introduced into the 5'UTR (single nucleotide changes, insertions, or deletions). The algorithm then predicts the new orthogonal ΔGtot new(O-ribo). If ΔGtot new(O-ribo) is more negative than ΔGtot(O-ribo), the change is accepted; if the mutation leads to a more positive ΔGtot new(O-ribo), the change is rejected with some conditional probability (see Methods). The algorithm is terminated after 10,000 iterations.Algorithm vol 2 generates a random 35 nucleotide 5'UTR containing an O-SD sequence at an optimal 5 nucleotide spacing from the start codon and predicts its ΔGtot(wt ribo) and ΔGtot(O-ribo). In an iterative process, mutations are introduced into the 5'UTR (single nucleotide changes, insertions, or deletions). The algorithm then calculates new predicted values, ΔGtot new(wt ribo) and ΔGtot new(O-ribo). If ΔGtot new(wt ribo) is more positive than ΔGtot(wt ribo) and ΔGtot new(O-ribo) is more negative than ΔGtot(O-ribo), the change is accepted; otherwise, the mutation is rejected with some conditional probability (see Methods). If 500 consecutive iterations do not result in an improvement in the ΔGtot value (convergence criterion), the algorithm outputs the sequence and its predicted ΔGtot value. Algorithm vol 3 is based on vol 2, but with two notable differences: (1) Vol 3 also starts with an ORF in which codons 2-12 are randomly exchanged with synonymous codons, so that the encoded amino acid sequence is preserved. (2) In the iterative process, synonymous codon substitutions in the ORF, in addition to single nucleotide changes, insertions or deletions in the 5'UTR, are tolerated mutation mechanisms. c, The algorithm finds O-mRNA sequences that are specifically and efficiently translated by O-ribosomes. The y-axis shows the production of strepGFPHis6 from O-mRNA by O-ribosomes; data are shown as the percentage of strepGFPHis6 produced by wt ribosomes from wt message. The x-axis shows the orthogonality of O-mRNA; it is calculated by dividing strepGFPHis6 produced from O-mRNA in the presence of O-ribosomes by strepGFPHis6 produced from O-mRNA in the presence of wt ribosomes. Protein production levels were calculated from GFP absorption and fluorescence data; in our system, the wt system produces 30.6±1.6 mg / mL of strepGFPHis6. Each dot represents one O-mRNA. Trans (black dot) is O(trans)-strepGFPHis6.Colored dots represent sequences from the volume covered by the algorithm. d, e Same as c but for E2Crimson (d) and mCherry (e), respectively. [Figure 2-1] ~ [Figure 2-2]The new O-mRNA enables efficient production of proteins containing three distinct ncAAs. a, Structures of the amino acids used in this study. N6-(tert-butoxycarbonyl)-L-lysine (BocK) 1; Nπ-methyl-L-histidine (NmH) 2; N6-((benzyloxy)carbonyl)-L-lysine (CbzK) 3; N6-((allyloxy)carbonyl)-L-lysine (AllocK) 4; (S)-2-amino-3-(4-iodophenyl)propanoic acid (PheI) 5. b, Engineered triply orthogonal pyrrolysyl-tRNA synthetase tRNA pair for incorporation of three distinct ncAAs using two different orthogonal messages. One message contained the O1-strepGFPHis6 5'UTR generated by vol 1 of our algorithm, and the other message used the O-(trans) 5'UTR. c, Production of strepGFP(40BocK, 136NmH, 150CbzK)His6 from E. coli cells containing strepGFP(40TAG, 136AGGA and 150AGTA)His6 constructs with O(trans)- or O1-strepGFP-His6 5'UTR. Cells also contained O-riboQ1 and aaRS3 / tRNA3 operon (encoding MmPylRS / MspetRNAPyl CUA, MlumPylRS(NMH) / MinttRNAPyl-A17VC10 UCCU and M1r26PylRS(CbzK) / MalvtRNAPyl-8 UACU). ncAA BocK1, NmH2, CbzK3 were added to the cells. d, Positive electrospray TOF-MS results of nickel-NTA purified strepGFP(40BocK, 136NmH, 150CbzK)His6 purified from cells described in (b). StrepGFP(40BocK, 136NmH, 150Cbz)His6 predicted mass: 29314.5, observed mass: 29312.0. [Diagram 3]Four orthogonal aaRS / tRNA pairs, which decode four orthogonal quadruplet codons, are expressed from an aaRS operon and a computationally generated tRNA operon, are mutually orthogonal in their aminoacylation specificity, recognize distinct ncAAs, and decode distinct orthogonal codons. a-d, Fluorescence from cells containing O1-strepGFP(40XXXX)His6, where XXXX is the codon at position 40 in sfGFP: TAGA, CTAG, AGGA, or AGTA. E. coli also contained O-riboQ1 and an aaRS and tRNA operon (aaRS4_1-2 / tRNA4(quad)); these operons expressed MmPylRS / MspetRNAPyl-evol UCUA, MrumPylRS(NMH) / MinttRNAPyl-A17VC10 UCCU, AfTyrRS(PheI) / AftRNATyr-A01 CUAG, and Mg1PylRS(CbzK) / MalvtRNAPyl-8 UACU. The indicated ncAA: Nπ-methyl-L-histidine (NmH) 2, N6-((benzyloxy)carbonyl)-L-lysine (CbzK) 3, N6-((allyloxy)carbonyl)-L-lysine (AllocK) 4, (S)-2-amino-3-(4-iodophenyl)propanoic acid (PheI) 5 were added to cells or omitted (−). Each codon was efficiently decoded only in the presence of the cognate ncAA of the aaRS / tRNA pair assigned to the respective quadruplet codon: (a) O1-strepGFP(TAGA)His6 decoded by MmPylRS / MspetRNAPyl-evol UCUA, (b) O1-strepGFP(AGGA)His6 decoded by MrumPylRS(NMH) / MinttRNAPyl-A17VC10 UCCU, (c) O1-strepGFP(AGTA)His6 decoded by Mg1PylRS(CbzK) / MalvtRNAPyl-8 UACU, and (d) O1-strepGFP(CTAG)His6 decoded by AfTyrRS(PheI) / AftRNATyr-A01 CUAG.e-h, Positive electrospray TOF-MS of nickel-NTA purified strepGFPHis6 expressed from O1-strepGFP(40XXXX)His6 in the presence of NmH2, CbzK3, AllocK4, and PheI5, where XXXX is either TAGA (e), AGGA (f), AGTA (g), or CTAG (h). Cells also contained O-riboQ1 and operon aaRS4_2-1 / tRNA4(quad). strepGFP(40AllocK)His6 predicted mass 29113.2, observed mass 29114.8. strepGFP(40NmH)His6 predicted mass 29052.1, observed mass 29052.5. strepGFP(40CbzK)His6 predicted mass 29163.3, observed mass 29164.2. Predicted mass of strepGFP(40PheI)His6: 29174.03, observed mass: 29174.2. [Figure 4-1] ~ [Figure 4-2]It genetically encodes four distinct ncAAs into proteins using a 24 amino acid, 68 codon genetic code. a, Schematic representation of the four mutually orthogonal aaRS / tRNA pairs used for incorporation of four distinct ncAAs in response to four orthogonal quadruplet codons. b, Efficient production of full-length strepGFP (40PheI, 50AllocK, 136NmH, 150CbzK)His6 was dependent on the addition of all four ncAAs (Nπ-methyl-L-histidine (NmH) 2, N6-((benzyloxy)carbonyl)-L-lysine (CbzK) 3, N6-((allyloxy)carbonyl)-L-lysine (AllocK) 4, (S)-2-amino-3-(4-iodophenyl)propanoic acid (PheI) 5). Fluorescence from cells containing O1-strepGFP (40CTAG, 50TAGA, 136AGGA, 150AGTA)His6, O-riboQ1, and the operon aaRS4 / tRNA4(quad) (encoding MmPylRS / MspePyltRNAUCUA, MrumPylRS(NMH) / MintPyltRNA(A17,VC10)UCCU, AfTyrRS / AftRNACUAG, and Mg1PylRS(CbzK) / MalvPyltRNA(8)UACU) in the presence or absence of combinations of NmH (2), CbzK (3), AllocK (4), and PheI (5). c, Positive electrospray TOF-MS of nickel-NTA purified strepGFP(40PheI, 50AllocK, 136NmH, 150CbzK)His6 from cells containing O1-strepGFP(40CTAG, 50TAGA, 136AGGA, 150AGTA)His6, O-riboQ1 and aaRS4_1-2 / tRNA4(quad) in the presence of the indicated ncAA. Predicted mass 29470.4, observed mass 29468.2. [Figure 5-1] ~ [Figure 5-2](Supplementary Figure 1) Fluorescence measurements of reporter protein production from O-mRNA generated by the indicated algorithm. We cloned the sequences (O1-O12 strepGFPHis6 for strepGFPHis6 (a and b), O1-O8 E2Crimson for E2Crimson (c) as well as O1-O8 mCherry for mCherry (d)) into standardized p15A reporter constructs and produced proteins in the presence of plasmids encoding either the O-ribosome or an additional copy of the wt ribosome. Control experiments used a construct with a 5'UTR and RBS commonly used in our laboratory (wt) as well as a construct with an O(trans)5'UTR that was previously used for highly efficient O-GST-CaM production (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Wang, K., Neumann, H., Peak-Chew, SY & Chin, JW Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol. 25, 770-777 (2007)). Bars represent the mean ± standard deviation of three biological replicates. Dots represent individual experiments. [Figure 6-1] ~ [Figure 6-3](Supplementary Figure 2) MS / MS spectra of ncAA-containing peptides obtained after tryptic digestion of strepGFP (40BocK, 136NmH, 150CbzK)His6. Precursor ions confirm the incorporation of ncAAs. Fragmentation of each peptide is predicted to result in a series of b ions (blue) and y ions (red) as well as ions corresponding to the loss of lysine protecting groups during the fragmentation process (a and c). Ion peaks were manually assigned; together with the precursor ion masses, these confirmed the incorporation of each ncAA at its expected position. Mass spectrometry was performed in triplicate with similar results. a, MS / MS spectrum confirming BocK 1 incorporation at position 40. b, MS / MS spectrum confirming NmH 2 incorporation at position 136. c, MS / MS spectrum confirming CbzK 3 incorporation at position 150. [Figure 7](Supplementary Figure 3) Assembly pipeline for the generation of a polycistronic operon containing genes for four mutually orthogonal aaRSs (AfTyrRS (PheI), MrumPylRS (NmH), Mg1PylRS (CbzK) and MmPylRS).For each synthetase, the online RBS calculator was optimized for maximum ΔGtot(wt ribo) (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009); Salis, HM in Methods in Enzymology, Vol. 498 19-42 (Academic Press, Cambridge, MA, USA; 2011); Espah Borujeni, A., Channarasappa, AS & Salis, HM Translation rate is controlled by coupled trade-offs between site accessibility, selective RNA unfolding and sliding at upstream standby sites. Nucleic Acids Res. 42, 2646-2659 (2014); Espah Borujeni, A. & Salis, HM Translation Initiation is Controlled by RNA Folding Kinetics via a Ribosome Drafting Mechanism. J. Am. Chem. Soc. 138, 7016-7023 (2016); Espah Borujeni, A. et al. Precise quantification of translation inhibition by mRNA structures that overlap with the ribosomal footprint in N-terminal coding sequences. Nucleic Acids Res. 45, 5437-5448 (2017)) (herein incorporated by reference) were used to generate five 5'UTR sequences.Next, ΔGtot(wt ribo) for each alignment of the form aaRSX-5'_UTR(Y1-Y5)-aaRSY (X and Y refer to any combination of two of the four synthetases) was calculated using an online tool (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009); Salis, HM in Methods in Enzymology, Vol. 498 19-42 (Academic Press, Cambridge, MA, USA; 2011); Espah Borujeni, A., Channarasappa, AS & Salis, HM Translation rate is controlled by coupled trade-offs between site accessibility, selective RNA unfolding and sliding at upstream standby sites. Nucleic Acids Res. 42, 2646-2659 (2014), Espah Borujeni, A. & Salis, H. M. Translation Initiation is Controlled by RNA Folding Kinetics via a Ribosome Drafting Mechanism. J. Am. Chem. Soc. 138, 7016-7023 (2016), Espah Borujeni, A. et al. Precise quantification of translation inhibition by mRNA structures that overlap with the ribosomal footprint in N-terminal coding sequences. Nucleic Acids Res. 45, 5437-5448 (2017)).Finally, all four synthetases were manually aligned to ensure high ΔGtot(wt ribo) for each synthetase. Two independent solution methods yielded similar results. After experimental validation, the favorable sequence context of one synthetase was copied into the other operon to obtain the final construct (all 5'UTR sequences and ΔGtot(wt ribo) are given in Supplementary Table 3). [Figure 8] (Supplementary Figure 4) Fluorescence from cells containing O1-strepGFP(XXXX)His6, where XXXX is either TAG, CTAG, AGGA, or AGTA. E. coli also contained O-riboQ1 and MmPylRS / MspePyltRNACUAG, MrumPylRS(NMH) / MintPyltRNA(A17,VC10)UCCU, AfTyrRS / AfRNACUA and Mg1PylRS(CbzK) / MalvPyltRNA(8)UACU and one of the ncAA: NmH 2, CbzK 3, BocK 1, or PheI 5. Synthetases were initially placed in either operon RS4_1 / tRNA4 or RS4_2 / tRNA4 (see Supplementary Figure 3 and Supplementary Table 3). RS4_1 / tRNA4(a) gave better results for the repression of TAG, CTAG, and AGTA; however, AGGA was repressed only half as efficiently as in RS4_2 / tRNA4(b). Therefore, the 150 nt region upstream of MrumPylRS was copied into RS4_1 / tRNA4 to obtain operon 1 RS4_1-2 / tRNA4(c), which leads to an activity 2.6 higher than MrumPylRS. [Figure 9-1] ~ [Figure 9-2](Supplementary Figure 5) MS / MS spectra of ncAA-containing peptides obtained after tryptic digestion of strepGFP (40PheI, 50AllocK, 136NmH, 150CbzK)His6. Precursor ions confirm the incorporation of ncAAs. Fragmentation of each peptide is predicted to result in a series of b ions (blue) and y ions (red) as well as ions corresponding to the loss of lysine protecting groups during the fragmentation process (d). Ion peaks were manually assigned; together with the precursor ion masses, these confirmed the incorporation of each ncAA at its expected position. Mass spectrometry was performed in triplicate with similar results. a, MS / MS spectrum confirming 5 incorporations of PheI at position 40; b, MS / MS spectrum confirming 4 incorporations of AllocK at position 50; c, MS / MS spectrum confirming 2 incorporations of NmH at position 136. d, MS / MS spectrum confirming CbzK3 incorporation at position 150. [Figure 10](Supplementary Figure 6) Four orthogonal aaRS / tRNA pairs, decoding one amber codon and three orthogonal quadruplet codons, are expressed from an aaRS operon and a computationally generated tRNA operon, are mutually orthogonal in their aminoacylation specificity, recognize distinct ncAAs, and decode distinct orthogonal codons. a-d, Fluorescence from cells containing O1-strepGFP(40XXXX)His6, where XXXX is the codon at position 40 in sfGFP: TAG, CTAG, AGGA, or AGTA. The E. coli also contained ribo-Q1 and an aaRS and tRNA operon (aaRS4_1-2 / tRNA4); these operons expressed MmPylRS / MspetRNAPyl-evol CUAG, MrumPylRS(NMH) / MinttRNAPyl-A17VC10 UCCU, AfTyrRS(PheI) / AftRNATyr-A01 CUA, and Mg1PylRS(CbzK) / MalvtRNAPyl-8 UACU. The indicated ncAA: Nπ-methyl-L-histidine (NmH) 2, N6-((benzyloxy)carbonyl)-L-lysine (CbzK) 3, N6-(tertbutoxycarbonyl)-L-lysine (BocK) 1, (S)-2-amino-3-(4-iodophenyl)propanoic acid (PheI) 5 were added to cells or omitted (−). Each codon was efficiently decoded only in the presence of the cognate ncAA of the aaRS / tRNA pair assigned to the respective quadruplet codon: (a) O1-strepGFP(TAG)His6 decoded by AfTyrRS(PheI) / AftRNATyr-A01 CUA, (b) O1-strepGFP(AGGA)His6 decoded by MrumPylRS(NMH) / MinttRNAPyl-A17VC10 UCCU, (c) O1-strepGFP(AGTA)His6 decoded by Mg1PylRS(CbzK) / MalvtRNAPyl-8 UACU, and (d) O1-strepGFP(CTAG)His6 decoded by MmPylRS / MspetRNAPyl-evol CUAG.e-h, Positive electrospray TOF-MS of nickel-NTA purified strepGFPHis6 expressed from O1-strepGFP(40XXXX)His6 in the presence of 2 NmH, 3 CbzK, 1 BocK, and 5 PheI, where XXXX is either TAG (e), AGGA (f), AGTA (g), or CTAG (h). Cells also contained O-riboQ1 and operon aaRS4_2-1 / tRNA4. strepGFP(40PheI)His6 predicted mass 29174.03, observed mass 29174.2. strepGFP(40BocK)His6 predicted mass 29129.4, observed mass 29129.0. strepGFP(40NmH)His6 predicted mass 29052.1, observed mass 29052.5. strepGFP(40CbzK)His6 predicted mass 29163.3, observed mass 29164.2. strepGFP(40BocK)His6 predicted mass 29129.4, observed mass 29129.0. [Figure 11](Supplementary Figure 7) Genetically encoding four distinct ncAAs into proteins in response to an amber codon and three distinct quadruplet codons. a, Schematic representation of the four mutually orthogonal aaRS / tRNA pairs used for the incorporation of four distinct ncAAs in response to an amber codon and three distinct quadruplet codons. b, Efficient production of full-length strepGFP (40PheI, 50AllocK, 136NmH, 150CbzK)His6 was dependent on the addition of all four ncAAs (BocK 1, NmH 2, CbzK 3 and PheI 5). Fluorescence from cells containing O1-strepGFP (40TAG, 50CTAG, 136AGGA, 150AGTA)His6, O-riboQ1, operon aaRS4_1-2 / tRNA4 (encoding MmPylRS / MspetRNAPyl-evol CUAG, MrumPylRS(NMH) / MinttRNAPyl-A17VC10 UCCU, AfTyrRS(PheI) / AftRNATyr-A01 CUA and Mg1PylRS(CbzK) / MalvtRNAPyl-8 UACU) in the presence or absence of a combination of BocK1, NmH2, CbzK3, PheI5. c, TOF-MS ES+ of purified strepGFP(40PheI, 50BocK, 136NmH, 150CbzK)His6 purified from cells containing O1-strepGFP(40TAG, 50CTAG, 136AGGA, 150AGTA)His6, O-riboQ1 and operon RS4_1-2 / tRNA4 in the presence of 8mM BocK 1, 4mM NmH 2, 2mM PheI 5 and 2mM CbzK 3. Predicted mass 29482.0, observed mass 29483.0. [Figure 12-1] ~ [Figure 12-2](Supplementary Figure 8) MS / MS spectra of ncAA-containing peptides obtained after tryptic digestion of strepGFP (40PheI, 50BocK, 136NmH, 150CbzK)His6. Precursor ions confirm the incorporation of ncAAs. Fragmentation of each peptide is predicted to result in a series of b ions (blue) and y ions (red) as well as ions corresponding to the loss of lysine protecting groups during the fragmentation process (b and d). Ion peaks were manually assigned; together with the precursor ion masses, these confirmed the incorporation of each ncAA at its expected position. Mass spectrometry was performed in triplicate with similar results. a, MS / MS spectrum confirming 5 incorporations of PheI at position 40. b, MS / MS spectrum confirming 1 incorporation of BocK at position 50. c, MS / MS spectrum confirming 2 incorporations of NmH at position 136. d, MS / MS spectrum confirming CbzK3 incorporation at position 150. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0080] The use of cell-based protein expression systems to produce exogenous proteins, especially those containing unnatural amino acids, can be challenging for several reasons. One challenge is that the cell must also be able to produce endogenous proteins essential for survival. For example, if a protein expression system has been engineered to allow the incorporation of unnatural amino acids into an exogenous protein, it may be desirable to avoid the incorporation of unnatural amino acids into the endogenous protein. One approach to overcome this is to use a system that contains two ribosomes: a wild-type ribosome for the production of proteins endogenous to the host cell and an orthogonal ribosome that is capable of translating an orthogonal mRNA that encodes the exogenous protein.

[0081] Cell-based protein expression systems may also include O-ribosomes for other reasons. It has been found that mutations to endogenous ribosomes can be toxic, and some ribosome mutations can be lethal to cells even when present in only a few copies of endogenous ribosomes (i.e., the mutations can be dominant-lethal). However, O-ribosomes can tolerate these ribosome mutations because, as discussed herein, they are separated from other functions of the host cell. Therefore, O-ribosomes can be engineered for new desired functions. For example, O-ribosomes can be evolved to decode new orthogonal codons (quadruplet codons, (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010))) or new intrinsic polymerization functions (Schmied, WH et al. Controlling orthogonal ribosome subunit interactions enables evolution of new function. Nature 564, 444-448 (2018)).

[0082] However, as discussed in the Background section, protein yields from expression systems containing O-ribosomes can be low and suboptimal, and in particular, yields are not consistent when measured for different exogenous proteins.

[0083] The factors that determine protein yield for native translation are incompletely understood, and design of experimental studies suggests that only half of the variation in observed protein yield can be explained by known parameters (Cambray, G., Guimaraes, JC & Arkin, AP Evaluation of 244,000 synthetic sequences reveals design principles to optimize translation in Escherichia coli. Nat. Biotechnol. 36, 1005-1015 (2018)). Nevertheless, the inventors have noted that initiation of protein synthesis is generally a rate-limiting step of translation (Plotkin, JB & Kudla, G. Synonymous but not the same: the causes and consequences of codon bias. Nat. Rev. Genet. 12, 32-42 (2011)), and numerous studies suggest that RNA secondary structures in the 5'UTR and the first 30 nt of coding sequences are key determinants of translation initiation and protein yield (Cambray, G., Guimaraes, JC & Arkin, AP Evaluation of 244,000 synthetic sequences reveals design principles to optimize translation in Escherichia coli. Nat. Biotechnol. 36, 1005-1015 (2018); Tuller, T. & Zur, H. Multiple roles of the coding sequence 5' end in gene expression regulation. Nucleic Acids Res. 43, 13-28 (2015)). Indeed, the total free energy change (ΔG totThermodynamic models predicting the relative protein yield for native translation can be used to predict the relative protein yield for native translation (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009); Na, D., Lee, S. & Lee, D. Mathematical modeling of translation initiation for the estimation of its efficiency to computationally designed mRNA sequences with desired expression levels in prokaryotes. BMC Syst. Biol. 4, 1-16 (2010); Seo, SW et al. Predictive design of mRNA translation initiation region to control prokaryotic translation efficiency. Metab. Eng. 15, 67-74 (2013)) (herein incorporated by reference). Previous studies with different 35 nt in the 5′UTR immediately upstream of the start codon have demonstrated that the protein yield for a given ORF (interpreted as reflecting the rate of translation initiation) is proportional to the equilibrium constant for the formation of an initiation-competent state from folded mRNA (i.e., ΔG tot(Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009), Salis, HM in Methods in Enzymology, Vol. 498 19-42 (Academic Press, Cambridge, MA, USA; 2011), Espah Borujeni, A., Channarasappa, AS & Salis, HM Translation rate is controlled by coupled trade-offs between site accessibility, selective RNA unfolding and sliding at upstream standby sites. Nucleic Acids Res. 42, 2646-2659 (2014), Espah Borujeni, A. & Salis, HM Translation Initiation is Controlled by RNA Folding Kinetics via a Ribosome Drafting Mechanism. J. Am. Chem. Soc. 138, 7016-7023 (2016); Espah Borujeni, A. et al. Precise quantification of translation inhibition by mRNA structures that overlap with the ribosomal footprint in N-terminal coding sequences. Nucleic Acids Res. 45, 5437-5448 (2017) (incorporated herein by reference). ΔG tot (wt ribo) suppressed mRNA unfolding (ΔG unfolding), and the wt-ribosome and tRNA through base pairing at precise positions to the mRNA. fMet CAU Binding (ΔG wt ribo binding ) (Figure 1a).

[0084] Here, we use a thermodynamic model of initiation and a simulated annealing optimization algorithm (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009)) to automate the discovery of 5'UTR sequences for orthogonal translation of ORFs. We also develop an algorithm to explicitly select messages that bind the O-ribosome but not other ribosomes, and increase the freedom in the search by exploring variations in both the 5'UTR and synonymous codons that code for amino acids in the ORF, e.g., amino acids 2-12. Automating O-mRNA discovery has led to sequences that provide up to 40-fold more protein and are up to 50-fold more orthogonal than preceding O-mRNAs; protein yields from the new O-mRNAs are comparable or superior to those from WT mRNAs. These advances are the result of the development of an engineered triply orthogonal PylRS / tRNA. Pyl This directly translates into a 33-fold increase in yield for incorporating three distinct ncAAs in response to an amber codon and two quadruplet codons using the pair.

[0085] Thus, in an embodiment of the invention, there is provided a method of designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome, the mRNA comprising a 5'UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo) prediction step; (b) introducing an alteration into the 5'UTR; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo) prediction step; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than the modification, Said ΔG tot new (O-ribo) precedes ΔG tot accepting or rejecting the modification according to a probability distribution if the probability distribution is more positive than (O-ribo); and (e) generating an O-mRNA sequence comprising a 5'UTR containing the tolerated modification; A method is provided, comprising:

[0086] "O-ribosome," as used herein, is a ribosome that has reduced or no ability to translate mRNAs endogenous to a particular host cell compared to endogenous ribosomes; it is a ribosome that has the ability to translate an mRNA that is different from the endogenous mRNA (i.e., the O-mRNA).

[0087] "O-mRNA," as used herein, is a messenger RNA that is translated less efficiently by ribosomes endogenous to a particular host cell compared to the translation of endogenous mRNA; it is capable of being translated by ribosomes other than the endogenous ribosomes (i.e., O-ribosomes).

[0088] The adjective "orthogonal" as used herein describes a component or feature associated with the O-ribosome and O-mRNA, but not with endogenous ribosomes or mRNAs. For example, an orthogonal Shine-Dalgarno sequence is associated with the O-mRNA as having the ability to interact with the orthogonal anti-Shine-Dalgarno sequence of the O-ribosome. The orthogonal Shine-Dalgarno sequence allows only reduced binding to endogenous ribosomes, and the orthogonal anti-Shine-Dalgarno sequence allows only reduced binding to endogenous mRNAs.

[0089] As used herein, the O-ribosome and the O-mRNA function together such that the O-ribosome has the ability to translate the O-mRNA.

[0090] In embodiments featuring more than one O-ribosome, a first set of O-mRNAs may be applicable to only one of the O-ribosomes, and a second set of O-mRNAs may be applicable to the other O-ribosome.

[0091] In examples, the O-ribosome may be an artificially altered or modified ribosome that differs from a wild-type ribosome. The O-mRNA may be an mRNA that is not a substrate for the wild-type ribosome.

[0092] The O-ribosome may comprise an altered 16S rRNA. In particular, the 16 rRNA may be altered in a manner that affects binding to the ribosome binding site (RBS) of the mRNA. The O-ribosome may comprise an altered anti-Shine-Dalgarno sequence that has no or minimal ability to bind to the wild-type Shine-Dalgarno sequence. In such cases, the O-mRNA comprises an altered RBS, e.g., an altered Shine-Dalgarno sequence, that has the ability to bind to the O-ribosome.

[0093] In an embodiment, in the context of a host cell, the O-ribosome does not synthesize or synthesizes minimally the endogenous proteome. In such an embodiment, the O-mRNA is not translated or is minimally translated by the endogenous ribosome. The host cell is not particularly limited and may be any host cell, particularly any host cell suitable for heterologous protein production. In some cases, the host cell is a prokaryotic cell, such as a bacterial cell. In particular, the host cell may be an E. coli cell.

[0094] In some examples, the O-ribosome may be O-riboQ1. Additionally, the O-ribosome may be any O-ribosome disclosed in WO 2008 / 065398 A1 or obtainable by the method disclosed in WO 2008 / 065398 A1. Additionally, the O-ribosome may be any O-ribosome disclosed in WO 2011 / 077075 A1 or obtainable by the method disclosed in WO 2011 / 077075 A1. Both WO 2008 / 065398 A and WO 2011 / 077075 A1 are incorporated herein by reference. The O-ribosome may be any O-ribosome disclosed or obtainable by the disclosed methods in any of (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Wang, K., Neumann, H., Peak-Chew, SY & Chin, JW Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol. 25, 770-777 (2007); or Schmied, WH et al. Controlling orthogonal ribosome subunit interactions enables evolution of new function. Nature 564, 444-448 (2018)), each of which is incorporated herein by reference.

[0095] The term "5'UTR" is used herein according to its ordinary meaning in the art. Briefly, a 5'UTR is a region of an mRNA that is not translated into a polypeptide, is 5' to the ORF, and is involved in recognition by the ribosome.

[0096] The term "ORF" is used herein according to its ordinary meaning in the art. Briefly, an ORF is a portion of an mRNA that has the ability to be translated into an encoded protein.

[0097] The free folded state of mRNA is the state that exists when mRNA is not bound to ribosomes and is free to form secondary structures.

[0098] The ribosome-bound initiation-competent state of an mRNA is the state that exists when an mRNA is bound to a ribosome, an initiator tRNA is bound, and translation initiation can begin.

[0099] In an embodiment, the modification is or comprises a single nucleotide change, insertion or deletion introduced into the 5'UTR.

[0100] During the method of the present invention, the modification tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than ΔG. tot (O-ribo) is the predicted ΔG for the mRNA sequence before modification. tot (O-ribo). As discussed herein, the method of the present invention may be repeated, so that the preceding ΔG tot (O-ribo) is the ΔG calculated during the preceding iteration. tot new (O-ribo) may also be used.

[0101] Tolerating modifications during the methods of the present invention means that the sequence changes introduced by the modifications are maintained for the next iteration of the method, or, in the absence of further iterations of the method, are maintained in the sequence of the O-mRNA that is the output of the method.

[0102] During the method of the present invention, the modification tot new (O-ribo) precedes ΔG tot If it is more positive than (O-ribo), it is accepted or rejected according to the probability distribution. tot (O-ribo)" is as discussed above. The probability distribution is ΔG tot new (O-ribo) and ΔG tot The probability may be based on a conditional probability where the likelihood of acceptance decreases as the difference between (O-ribo) increases. The probability may be a Monte Carlo optimization.

[0103] In an embodiment, the probability distribution according to which a modification will be accepted or rejected is:

number

[0104] In certain embodiments, T SA is adjusted to maintain a tolerance of at least 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%. SA may be adjusted to maintain an acceptance rate of at least 5%.

[0105] In certain embodiments, T SA is adjusted to maintain a tolerance rate less than or equal to 75%, 50%, 40%, 30%, 25%, 20%, 15%, or 10%. SAmay be adjusted to maintain a tolerance rate less than or equal to 20%.

[0106] In certain embodiments, T SA is adjusted to maintain a tolerance of 0.1% to 75%, 1% to 50%, 2% to 40%, 3% to 30%, 4% to 25%, or, in particular, 5 to 20%.

[0107] T SA The adjustment of T is done when the tolerance falls outside the above values ​​for a certain number of iterations. SA may mean that T is increased or decreased to compensate. For example, if the acceptance rate falls below a lower threshold or exceeds an upper threshold for 5, 10, 20, 30, 40, 50, 60, 70, 100, 200, or 500 iterations, T may be increased or decreased so that the acceptance rate is corrected. SA may be lowered or raised. In certain embodiments, the tolerance rate is considered for 50 iterations. In certain embodiments, T SA is adjusted by doubling or halving the value.

[0108] Rejection of an alteration during the method of the invention means that the sequence change introduced by the alteration is reverted and therefore not maintained for the next iteration of the method or not maintained in the output sequence.

[0109] In some embodiments, modifications may be rejected if certain sequence constraints are violated. For example, modifications may be rejected if the modified sequence invalidates one of the assumptions of the underlying thermodynamic model. Optional step (d) of the method of designing an O-mRNA disclosed herein may include rejecting modifications based on these constraints, which may be included in addition to accepting or rejecting based on the probability distributions disclosed herein. The sequence constraints may be any as disclosed in (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009)) (incorporated by reference).

[0110] As an example of such a constraint, in an embodiment, a modification is rejected if the energy required to unfold a 16S rRNA binding site on the mRNA sequence exceeds a certain threshold, e.g., more than 6 kcal / mol. Alternatively or additionally, the presence of long-range nucleotide interactions may be quantified and modifications may be rejected if certain conditions are not met. For example, the equilibrium probability of nucleotides i and j forming a base pair in solution may be greater than P=|ij| -1.44 is considered to be proportional to, and when P is calculated for each base pair in sequence S, the minimum p is 6 × 10 -3 As another example of a constraint, which may be included in place of or in addition to any of the other constraints, the creation of a new AUG or GUG start codon within a ribosome binding sequence may be rejected, and thus any modification that introduces such a codon may be rejected.

[0111] In all embodiments disclosed herein where modifications are accepted or rejected according to a probability distribution, the modifications may alternatively simply be rejected. tot new(O-ribo) precedes ΔG tot If it is more negative than (O-ribo), the modification is rejected. This is also applicable to further embodiments disclosed herein.

[0112] The generation of an O-mRNA sequence means that the final sequence is the output that contains the cumulative effect of all of the allowed modifications.

[0113] In an embodiment, the first round of the method of the invention is performed on a potential mRNA sequence with a randomly generated 5'UTR. The length of the 5'UTR is not particularly limited. During the method, the length of the 5'UTR may be increased or decreased due to insertion or deletion modifications. In a particular embodiment, the initial 5'UTR is 30-40 nucleotides long, or in particular 35 nucleotides. Alternatively, the 5'UTR may be longer, but a 30-40, or in particular 35 nucleotide window is considered by the method of the invention for modification. The 35 nucleotide window may be 35 nucleotides of the 5'UTR closest to the start codon. In other embodiments, the initial 5'UTR may be a shorter 5'UTR, for example 15, 20, or 25 nucleotides, or longer, for example at least 40, 50, or 51 or more nucleotides. It may be desirable to generate a 5'UTR of a particular length, in which case a 15, 20, 25, 30, 35, 45, 50 nucleotide window may be considered so that a particular length of the output sequence can be achieved.

[0114] The 5'UTR to which step (a) is applied may comprise a wild-type Shine-Dalgarno sequence, or a 5 nucleotide core of the wild-type Shine-Dalgarno sequence. In other embodiments, the 5'UTR to which step (a) is applied may comprise an orthogonal Shine-Dalgarno sequence, as discussed herein. The 5'UTR may be a random sequence other than the Shine-Dalgarno sequence. The Shine-Dalgarno sequence may be 5 nucleotides from the start codon of the ORF that are predicted to be optimally spaced.

[0115] The method of the present invention is tot (O-ribo). The method for predicting this value is described in detail in the Examples section. tot (O-ribo) is the free energy required to unfold mRNA (ΔG unfolding ) and the free energy released when mRNA binds to the O-ribosome to form an initiation-competent state bound to the O-ribosome (ΔG o-ribo binding ) is the sum of

[0116] In an embodiment, ΔG tot (O-ribo) is as follows: ΔG tot (O-ribo) = (ΔG mRNA-O-rRNA +ΔG start +ΔG spacing -ΔG standby )+ΔG unfolding may be calculated as follows: ΔG mRNA-O-rRNA is the free energy of the predicted cofolded secondary structure of the last 9 nucleotides of the orthogonal 16S rRNA and the mRNA; ΔG start is the energy released from the binding of the initiator tRNA to the start codon of the ORF; ΔG spacing is the energy penalty for a non-optimal spacing length between the Shine-Dalgarno sequence and the start codon; ΔG standby is the energy required to unfold the secondary structure separating the four nucleotides upstream of the Shine-Dalgarno sequence; ΔG unfolding is the energy required to unfold the secondary structure in the mRNA.

[0117] In certain embodiments, the above values ​​are calculated as disclosed in (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009)) (incorporated by reference). For example, ΔG spacing may be calculated as disclosed in Section 3 of the Supplementary Methods of this publication.

[0118] The method of the invention may be iterative, such that multiple modifications are considered for acceptance or rejection, and the final output sequence includes the cumulative effect of all of said accepted modifications. In particular, steps (b)-(d) of the method of the invention may be iterative. In embodiments, the method is iterated at least 200, 300, 400, 500, 1000, 5000, or, in particular, 10000 times. In other embodiments, the method is iterative, such that successive iterations result in more negative ΔG tot new This may be repeated until it no longer leads to (O-ribo).

[0119] In certain embodiments, there is provided a method for designing an mRNA that is an O-mRNA suitable for translation by an O-ribosome, the mRNA comprising a 5'UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)), wherein the 5'UTR contains a wild-type Shine-Dalgarno sequence; (b) introducing an alteration into the 5'UTR that is, or includes, a single nucleotide change, insertion, or deletion; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo) prediction step; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than the modification, Said ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more positive than the modification according to a probability distribution; tot new (O-ribo) and the above ΔG tot (O-ribo) determines the probability of acceptance, with smaller magnitudes being associated with a higher likelihood of acceptance compared to larger magnitudes; and (e) repeating steps (b) to (d) at least 500, 1000, 5000, or, in particular, 10000 times; and then Generating an O-mRNA sequence that includes a 5'UTR that includes the tolerated modification. A method is provided, comprising:

[0120] In certain embodiments, there is provided a method for designing an mRNA that is an O-mRNA suitable for translation by an O-ribosome, the mRNA comprising a 5'UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)), wherein the 5'UTR contains a wild-type Shine-Dalgarno sequence; (b) introducing a modification comprising a single nucleotide change, insertion, or deletion in any one of the 35 nucleotides of the 5'UTR closest to the ORF; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo) prediction step; (d) The above ΔG tot new (O-ribo) precedes ΔG tot(O-ribo) is more negative than the modification, Said ΔG tot new (O-ribo) precedes ΔG tot If it is more positive than (O-ribo)

number

[0121] In some embodiments, the methods of the present invention provide a method for increasing the efficiency of translation by the O-ribosome and increasing the efficiency of translation by the second ribosome (2 nd The method may include optimizing the O-mRNA such that the efficiency of translation by the O-ribosome is reduced.

[0122] Therefore, in an additional embodiment, the method of the present invention comprises the step of: nd 1. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains an O-ribosome, the O-mRNA comprising a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd - predicting ribo); (b) introducing an alteration into the 5'UTR; (c) New ΔG after modification tot (O-ribo)(ΔG tot new(O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd - a step of predicting ribo; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo) to allow modification, Said ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo), accepting or rejecting the modification according to a probability distribution; and (e) generating an O-mRNA sequence comprising a 5'UTR containing the tolerated modification; The method includes:

[0123] The 5'UTR sequence of step (a) may comprise a Shine-Dalgarno sequence predicted to be fully complementary to the anti-Shine-Dalgarno sequence of the O-ribosome that is optimized for increased translation of the O-mRNA. This Shine-Dalgarno sequence is referred to as an orthogonal Shine-Dalgarno sequence (O-SD). The 5'UTR sequence of step (a) may comprise a 5 nucleotide core of the O-SD. In an embodiment, the O-SD is 5 nucleotides from the start codon of the ORF. In an embodiment, no modifications are introduced into the 5 nucleotide core of the O-SD. For example, the O-SD may be TAATCCCAT and no modifications are introduced into TCCCA. In some embodiments, no modifications are introduced into the O-SD. In other embodiments, the 5'UTR sequence may comprise a wild-type Shine-Dalgarno sequence.

[0124] The first round of the method of the present invention may be performed on a potential mRNA sequence with a randomly generated 5'UTR. The initial length, final length, or length of the window of nucleotides considered may be any disclosed herein. The 5'UTR may be a random sequence except for the Shine-Dalgarno sequence. The Shine-Dalgarno sequence may be 5 nucleotides from the start codon of the ORF that are predicted to be optimally spaced.

[0125] 2 nd The ribosome may be a wild-type ribosome ("WT-ribosome"). "WT-ribosome", as used herein, is a ribosome that has the ability to translate endogenous mRNAs in a host cell of interest and has a lower or no ability to translate O-mRNAs. For example, the WT-ribosome may include a wild-type region for interacting with the RBS of an mRNA. The 16S rRNA of the WT-ribosome (referred to as wild-type 16S rRNA) may include a wild-type sequence. In particular, the WT-ribosome may include a wild-type anti-Shine-Dalgarno sequence. In certain instances, all components of the WT-ribosome may be wild-type.

[0126] Alternatively, 2 nd The -ribosome may be another O-ribosome, e.g., 2 nd The ribosome may be an O-ribosome that includes a second orthogonal anti-Shine-Dalgarno sequence that is different from the orthogonal anti-Shine-Dalgarno sequence of the first ribosome (i.e., the ribosome that is optimized for increased translation of the mRNA). The second O-ribosome may efficiently translate a set of O-mRNAs that are different from the O-mRNAs efficiently translated by the first ribosome.

[0127] ΔG tot (2 nd A method for predicting the ΔG tot (2 nd -ribo) is the free energy required to unfold mRNA (ΔG unfolding ) and mRNA is 2 nd - Binds to ribosomes, 2 nd - the free energy released upon forming an initiation-competent state bound to the ribosome (ΔG 2nd-ribo binding ) is the sum of

[0128] In an embodiment, ΔG tot (2 nd -ribo) is as follows: ΔG tot (2 nd -ribo)=(ΔG mRNA-2nd-rRNA +ΔG start +ΔG spacing -ΔG standby )+ΔG unfolding may be calculated as follows: ΔG mRNA-2nd-rRNA is the free energy of the predicted cofolded secondary structure of the last 9 nucleotides of 16S rRNA and the mRNA; ΔG start is the energy released from the binding of the initiator tRNA to the start codon of the ORF; ΔG spacingis the energy penalty for a non-optimal spacing length between the Shine-Dalgarno sequence and the start codon; ΔG standby is the energy required to unfold the secondary structure separating the four nucleotides upstream of the Shine-Dalgarno sequence; ΔG unfolding is the energy required to unfold the secondary structure in the mRNA.

[0129] The above calculation is ΔG tot This may be done as discussed for (O-ribo).

[0130] In an embodiment, the modification is or comprises a single nucleotide change, insertion or deletion introduced into the 5'UTR.

[0131] The modification is the ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo) is accepted or rejected according to the probability distribution. tot (O-ribo)" is as discussed above, and "the preceding ΔG tot (2 nd -ribo) should be interpreted in the same way. The probability distribution is ΔG tot new (O-ribo) and ΔG tot (O-ribo) or ΔG tot new (2 nd -ribo) and ΔG tot (2 nd The probability may be based on a conditional probability where the likelihood of acceptance decreases as the difference between the sigma-ribo increases. The probability may be a Monte Carlo optimization.

[0132] The method of the invention may be iterative, such that multiple modifications are considered for acceptance or rejection, and the final output sequence includes the cumulative effect of all of the accepted modifications. In particular, steps (b)-(d) of the method of the invention may be iterative. The method may be repeated for at least 10, 50, 100, 250, 500, 1000, 2000, 3000, 5000, or 10000 successive iterations to produce a more negative ΔG tot new (O-ribo) or more positive ΔG tot new (2 nd -ribo)。 Alternatively, the method may be repeated a set number of times as disclosed herein.

[0133] In certain embodiments, the method of the present invention comprises the steps of: nd 5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd - predicting the 5'UTR containing an O-SD; (b) introducing a modification into the 5'UTR that is or includes a single nucleotide change, insertion, or deletion, where the modification is not introduced into the O-SD 5 nucleotide core; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new(2 nd - a step of predicting ribo; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo) to allow modification, Said ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo), accepting or rejecting the modification according to a probability distribution if ΔG tot new (O-ribo) and the above ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) and the ΔG tot (2 nd -ribo) determines the probability of acceptance, with smaller magnitudes being associated with a higher likelihood of acceptance compared to larger magnitudes; and (e) at least 10, 50, 100, 250, or, in particular, 500 consecutive iterations have a more negative ΔG tot new (O-ribo) or more positive ΔG tot new (2 nd Repeat steps (b) through (d) until no more ribos are connected; Generating an O-mRNA sequence that includes a 5'UTR that includes the tolerated modification. The method includes:

[0134] In certain embodiments, the method of the present invention comprises the steps of: nd5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd -ribo)) and the formula: ΔG tot (opt)=ΔG tot (O-ribo)-X*ΔG tot (2 nd -ribo) according to ΔG tot Calculating (opt); (b) introducing modifications into the 5'UTR, including single nucleotide changes, insertions, or deletions; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd -ribo) and the formula: ΔG tot new (opt)=ΔG tot new (O-ribo)-X*ΔG tot new (2 nd -ribo) according to ΔG tot new Calculating (opt); (d) The above ΔG tot new ΔG preceded by (opt) tot Allow the modification if it is more negative than (opt), Said ΔG tot new ΔG preceded by (opt) totaccepting or rejecting the modification according to a probability distribution if (opt) is more positive than (opt); (e) generating an O-mRNA sequence comprising a 5'UTR containing the tolerated modification; The method includes:

[0135] In some embodiments, X is a number between 0.1 and 2, particularly 0.5. In other examples, X may be a number between 0.1 and 2, between 0.15 and 1.5, between 0.2 and 1, between 0.25 and 0.9, between 0.3 and 0.8, between 0.35 and 0.7, between 0.4 and 0.6, between 0.45 and 0.55, or between 0.5. As one of ordinary skill in the art will appreciate, weighting may be used to determine the degree to which ΔG tot new (O-ribo), which is encompassed by the formula above. The weighting may be adjusted to favor certain properties, e.g., a higher X is 2 nd -ribosome, whereas a lower X favors maximizing translation by the first ribosome (i.e., the O-ribosome for which the O-mRNA is targeted).

[0136] The modification is the ΔG tot new ΔG preceded by (opt) tot It is permissible if it is more negative than (opt). tot (opt) is the predicted ΔG for the mRNA sequence before modification. tot (opt). As discussed herein, the method of the present invention may be iterative, so that the preceding ΔG tot (opt) is the ΔG calculated during the previous iteration tot new (opt) may also be used.

[0137] The modification is the ΔG tot new ΔG preceded by (opt) tot If the probability distribution is greater than ΔG tot (opt)" is as discussed above. The probability distribution is ΔGtot new (opt) and ΔG tot The probability may be based on a conditional probability where the likelihood of acceptance decreases as the difference between (opt) increases. The probability may be a Monte Carlo optimization.

[0138] In an embodiment, the probability distribution according to which a modification is accepted or rejected is:

number

[0139] T SA may be adjusted in any manner disclosed herein. In certain embodiments, T SA will be adjusted to maintain a tolerance rate of 5-20%.

[0140] In some embodiments, a modification may be rejected if certain sequence constraints are violated, as discussed herein. In addition to or as an alternative to the constraints already discussed, a modification may be rejected if a second O-SD or a second O-SD core is introduced into the sequence. This is to prevent initiation from the wrong site. For example, a modification may be rejected if the sequence "TCCCA" (an example of an O-SD core) is introduced.

[0141] The method of the invention may be iterative, such that multiple modifications are considered for acceptance or rejection, and the final output sequence includes the cumulative effect of all of the accepted modifications. In particular, steps (b)-(d) of the method of the invention may be iterative. The method may be repeated for at least 10, 50, 100, 250, 500, 1000, 2000, 3000, 5000, or 10000 successive iterations to produce a more negative ΔG tot new (opt) In other embodiments, the method may be repeated a set number of times, as discussed herein.

[0142] In certain embodiments, 2nd 5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd -ribo)) and the formula: ΔG tot (opt)=ΔG tot (O-ribo)-X*ΔG tot (2 nd -ribo) according to ΔG tot calculating (opt); wherein the 5'UTR comprises an O-SD; (b) introducing an alteration into the 5'UTR, wherein the alteration is not introduced into the O-SD 5 nucleotide core; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd -ribo) and the formula: ΔG tot new (opt)=ΔG tot new (O-ribo)-X*ΔG tot new (2 nd -ribo) according to ΔG tot new Calculating (opt); (d) The above ΔG tot new ΔG preceded by (opt) tot Allow the modification if it is more negative than (opt), Said ΔG tot newΔG preceded by (opt) tot (opt), accepting or rejecting the modification according to a probability distribution if ΔG tot new (opt) and the above ΔG tot The magnitude of the difference between (opt) determines the probability of acceptance, with a smaller magnitude being associated with a higher probability of acceptance compared to a larger magnitude; and (e) at least 10, 50, 100, 250, or, in particular, 500 consecutive iterations have a more negative ΔG tot new Repeat steps (b) through (d) until (opt) is no longer reached, then Generating an O-mRNA sequence that includes a 5'UTR that includes the tolerated modification. Includes; A method is provided wherein X is a number between 0.1 and 2, in particular 0.5.

[0143] In certain embodiments, 2 nd 5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd -ribo)) and the formula: ΔG tot (opt)=ΔG tot (O-ribo)-X*ΔG tot (2 nd --ribo) according to ΔG tot calculating (opt); wherein the 5'UTR comprises an O-SD; (b) introducing a modification into the 5'UTR that is a single nucleotide change, insertion, or deletion, where the modification is not introduced into the O-SD 5 nucleotide core; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd -ribo) and the formula: ΔG tot new (opt)=ΔG tot new (O-ribo)-X*ΔG tot new (2 nd -ribo) according to ΔG tot new Calculating (opt); (d) The above ΔG tot new ΔG preceded by (opt) tot Allow the modification if it is more negative than (opt), Said ΔG tot new ΔG preceded by (opt) tot If it is more positive than (opt)

number

[0144] The method of designing an O-mRNA can include optimizing the O-mRNA such that the efficiency of translation by a first O-ribosome is increased and the efficiency of translation by a second O-ribosome and a WT-ribosome is decreased.

[0145] Such methods are as disclosed above, in which the free energy difference between the freely folded state of the mRNA and the initiation-competent state bound to a ribosome is predicted for each of the ribosomes. As discussed, this prediction is made before and after the introduction of modifications to the mRNA sequence. In embodiments involving more than two ribosomes, the modifications are tot It may be acceptable if ΔG is more negative for the first ribosome (i.e., the ribosome for which translation efficiency is increased) and more positive for the other ribosomes. tot If the values ​​are not all favorably altered, the modification may be accepted or rejected according to a probability distribution as disclosed herein. tot Values ​​may be combined to form a single value that is considered for acceptance or rejection. For example, ΔG tot The value is calculated using the following formula ΔG tot (opt)=X*ΔG tot (1 st -O-ribo)-Y*ΔG tot (WT-ribo)-Z*ΔG tot (2 nd -O-ribo), where X, Y, and Z are weights. These weights may be adjusted to favor certain properties (e.g., optimization of translation by the first O-ribosome or reduction in translation by the second O-ribosome). ΔG tot (opt) may be considered for acceptance or rejection as disclosed herein.

[0146] As will be apparent to one of skill in the art, the above may be adapted such that the efficiency of translation by the first O-ribosome is increased and the efficiency of translation by two, three, four, five or more other ribosomes is decreased. The same or different weightings may be applied to the ΔG for each of the ribosomes for which the efficiency of translation is decreased. tot It may be associated with a value ΔG tot (1 st -O-ribo) may also be associated with a weighting.

[0147] The inventors further identified that replacing codons in the ORF with synonymous codons can lead to improved O-mRNA. Synonymous codons are codons that code for the same amino acid, and therefore, replacing a sense codon with a synonym does not change the sequence of the encoded protein. Thus, in embodiments, the modification in step (b) may include replacing any one of codons 2-20, 2-15, 2-12, 2-10, or 2-5 in the ORF with a synonymous codon. In certain embodiments, the modification in step (b) may include replacing any one of codons 2-12 in the ORF with a synonymous codon.

[0148] Exchanging codons for synonyms can be an alternative to introducing a single nucleotide change, insertion or deletion into the 5'UTR. Thus, step (b) may comprise introducing an alteration in the 5'UTR that is a single nucleotide change, insertion or deletion, or exchanging any one of codons 2-20, 2-15, 2-12, 2-10 or 2-5 (particularly 2-12) in the ORF with a synonymous codon.

[0149] In such embodiments, the O-mRNA sequence generated comprises a 5'UTR and an ORF that contain the tolerated modifications.

[0150] Thus, in an additional embodiment, a method of the invention provides a method for designing an mRNA that is an O-mRNA suitable for translation by an O-ribosome, the mRNA comprising a 5'UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo) prediction step; (b) introducing an alteration into the 5'UTR that is or includes a single nucleotide change, insertion, or deletion, or replacing any one of codons 2-20, 2-15, 2-12, 2-10, or 2-5 in the ORF with a synonymous codon; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo) prediction step; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than the modification, Said ΔG tot new (O-ribo) precedes ΔG tot accepting or rejecting the modification according to a probability distribution if the probability distribution is more positive than (O-ribo); and (e) generating an O-mRNA sequence that includes a 5'UTR and an ORF that includes the tolerated modifications. The method includes:

[0151] In an additional embodiment, the method of the present invention comprises the steps of: nd 5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd - predicting ribo); (b) introducing an alteration into the 5'UTR that is or includes a single nucleotide change, insertion, or deletion, or replacing any one of codons 2-20, 2-15, 2-12, 2-10, or 2-5 in the ORF with a synonymous codon; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd - a step of predicting ribo; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo) to allow modification, Said ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo), accepting or rejecting the modification according to a probability distribution; and (e) generating an O-mRNA sequence that includes a 5'UTR and an ORF that includes the tolerated modifications. The method includes:

[0152] In certain embodiments, the method of the present invention comprises the steps of: nd 5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd - predicting ribo); (b) introducing an alteration into the 5'UTR that is or includes a single nucleotide change, insertion, or deletion, or replacing any one of codons 2 to 12 in the ORF with a synonymous codon; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd - a step of predicting ribo; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo) to allow modification, Said ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) or the ΔG tot new (2 nd -ribo) precedes ΔG tot (2 nd -ribo), accepting or rejecting the modification according to a probability distribution if ΔG tot new (O-ribo) and the above ΔG tot (O-ribo) or the ΔG totnew (2 nd -ribo) and the ΔG tot (2 nd -ribo) determines the probability of acceptance, with smaller magnitudes being associated with a higher likelihood of acceptance compared to larger magnitudes; and (e) generating an O-mRNA sequence that includes a 5'UTR and an ORF that includes the tolerated modifications. The method includes:

[0153] In another embodiment, the method of the present invention comprises the steps of: nd 5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd -ribo)) and the formula: ΔG tot (opt)=ΔG tot (O-ribo)-X*ΔG tot (2 nd -ribo) according to ΔG tot Calculating (opt); (b) introducing an alteration into the 5'UTR that is or includes a single nucleotide change, insertion, or deletion, or replacing any one of codons 2-20, 2-15, 2-12, 2-10, or 2-5 in the ORF with a synonymous codon; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd-ribo)) and the formula: ΔG tot new (opt)=ΔG tot new (O-ribo)-X*ΔG tot new (2 nd -ribo) according to ΔG tot new Calculating (opt); (d) The above ΔG tot new ΔG preceded by (opt) tot Allow the modification if it is more negative than (opt), Said ΔG tot new ΔG preceded by (opt) tot accepting or rejecting the modification according to a probability distribution if (opt) is more positive than (opt); (e) generating an O-mRNA sequence that includes a 5'UTR and an ORF that includes the tolerated modifications. Optionally, X is a number as disclosed herein, such as 0.1 to 2, or, in particular, 0.5.

[0154] In another embodiment, the method of the present invention comprises the steps of: nd 5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd -ribo)) and the formula: ΔG tot (opt)=ΔG tot (O-ribo)-X*ΔG tot (2 nd -ribo) according to ΔG totcalculating (opt), wherein the 5'UTR comprises an O-SD; (b) introducing a modification into the 5'UTR that is a single nucleotide change, insertion, or deletion, where the modification is not into the O-SD 5 nucleotide core, or replacing any one of codons 2-12 in the ORF with a synonymous codon; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd -ribo) and the formula: ΔG tot new (opt)=ΔG tot new (O-ribo)-X*ΔG tot new (2 nd -ribo) according to ΔG tot new Calculating (opt); (d) The above ΔG tot new ΔG preceded by (opt) tot Allow the modification if it is more negative than (opt), Said ΔG tot new ΔG preceded by (opt) tot (opt), accepting or rejecting the modification according to a probability distribution if ΔG tot new (opt) and the above ΔG tot The magnitude of the difference between (opt) determines the probability of acceptance, with a smaller magnitude being associated with a higher probability of acceptance compared to a larger magnitude; and (e) at least 10, 50, 100, 250, or, in particular, 500 consecutive iterations have a more negative ΔG tot new Repeat steps (b) through (d) until (opt) is no longer reached, then Generating an O-mRNA sequence that includes a 5'UTR and an ORF that includes the tolerated modifications. Optionally, X is a number as disclosed herein, such as 0.1 to 2, or, in particular, 0.5.

[0155] In yet another embodiment, two nd 5. A method for designing an mRNA, which is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a 5′ UTR and an ORF, the method comprising: (a) The free energy difference (ΔG tot (O-ribo)) and the free folded state of mRNA and the 2-fold state of mRNA nd - the free energy difference between the initiation-competent state bound to the ribosome (ΔG tot (2 nd -ribo)) and the formula: ΔG tot (opt)=ΔG tot (O-ribo)-X*ΔG tot (2 nd -ribo) according to ΔG tot calculating (opt), wherein the 5'UTR comprises an O-SD; (b) introducing a modification into the 5'UTR that is a single nucleotide change, insertion, or deletion, where the modification is not introduced into the O-SD 5 nucleotide core, or replacing any one of codons 2-20, 2-15, 2-12, 2-10, or 2-5 in the ORF with a synonymous codon; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo)) and new ΔG tot (2 nd -ribo)(ΔG tot new (2 nd -ribo) and the formula: ΔG tot new (opt)=ΔG totnew (O-ribo)-X*ΔG tot new (2 nd -ribo) according to ΔG tot new Calculating (opt); (d) The above ΔG tot new ΔG preceded by (opt) tot Allow the modification if it is more negative than (opt), Said ΔG tot new ΔG preceded by (opt) tot If it is more positive than (opt)

number

[0156] Any of the methods for designing an O-mRNA may be used to optimize the O-mRNA to be translated by the O-ribosome at an enhanced rate and / or to optimize the O-mRNA to be more orthogonal. Optimized orthogonality refers to the efficiency of translation of the O-mRNA by the O-ribosome and / or the efficiency of translation of the O-mRNA by the O-ribosome. nd This may be such that the difference between the efficiency of translation of the O-mRNA by the O-ribosome (e.g., the WT-ribosome or a second O-ribosome) is increased by measuring the yield of protein produced from the O-mRNA in the presence of the O-ribosome and comparing it with the 2 nd- the yield of protein produced from the O-mRNA in the presence of ribosomes.

[0157] The yield obtained when the O-mRNA is in the presence of the O-ribosome may be increased by at least 2-fold, 5-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, or 40-fold compared to production from a non-optimized sequence.

[0158] The orthogonality of the O-mRNA may be increased by at least 2-fold, 5-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, 40-fold, 45-fold, or 50-fold compared to the orthogonality of a non-optimized sequence.

[0159] Any of the methods for designing an O-mRNA may further comprise the step of producing a nucleic acid molecule encoding said O-mRNA. The nucleic acid may be a DNA sequence and may be included in a vector suitable for delivery to a host cell of interest. Thus, a host cell comprising a nucleic acid molecule encoding said O-mRNA is also provided.

[0160] Any of the methods of designing an O-mRNA may further include a step of experimentally validating the O-mRNA. In such embodiments, the yield of the encoded protein from the O-mRNA may be compared to the yield of the protein from a non-optimized mRNA sequence, or to the yield of the protein when encoded by a WT-mRNA and translated by a WT-ribosome. Additionally, or alternatively, the experimental validation may include determining the orthogonality of the O-mRNA, and optionally comparing it to the orthogonality of a non-optimized mRNA, as discussed herein.

[0161] The method of designing an O-mRNA may be performed on a computer. Thus, a system is provided that includes a processor; and one or more computer readable storage media having stored thereon instructions for execution on said processor for performing the method of the invention. Additionally, a computer program product is also provided that includes a non-transitory machine readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to perform the method of the invention.

[0162] For example, in an embodiment, a computer-implemented method for designing an mRNA that is an O-mRNA suitable for translation by an O-ribosome, the mRNA comprising a 5′ UTR and an ORF, the method comprising executing program code on one or more processors to perform the following steps: (a) The free energy difference (ΔG tot (O-ribo) prediction step; (b) introducing an alteration into the 5'UTR; (c) New ΔG after modification tot (O-ribo)(ΔG tot new (O-ribo) prediction step; (d) The above ΔG tot new (O-ribo) precedes ΔG tot (O-ribo) is more negative than the modification, Said ΔG tot new (O-ribo) precedes ΔG tot accepting or rejecting the modification according to a probability distribution if the probability distribution is more positive than (O-ribo); and (e) generating an O-mRNA sequence comprising a 5'UTR containing the tolerated modification; The method may include any of the other features or limitations disclosed herein.

[0163] In addition to the above, the present inventors further provide a surprisingly efficient method for designing an operon comprising at least two exogenous tRNAs. The present inventors have automated the creation of an operon for compact, scalable expression of separate tRNAs, which may be orthogonal tRNAs. As an example, the present inventors have developed an engineered triply orthogonal PylRS / tRNA operon. Pyl Pair and Archaeoglobus fulgidus tyrosyl-tRNA synthetase (AfTyrRS) / tRNA Tyr We develop a compact operon expressing the derived pair and demonstrate that the operon is highly efficient.

[0164] Thus, in an embodiment of the invention, there is provided a method for designing an operon encoding at least two exogenous tRNAs for expression in a host cell comprising an endogenous genome encoding an endogenous tRNA, comprising: (i) generating permutations of at least two exogenous tRNAs; (ii) identifying, within the endogenous genome, adjacent pairs of endogenous tRNAs that have the highest level of sequence identity to each adjacent pair of exogenous tRNAs within each permutation of the at least two exogenous tRNAs; (iii) identifying intergenic regions in the endogenous genome between each of the identified adjacent pairs of endogenous tRNAs; (iv) generating a plurality of sequences encoding each of the at least two exogenous tRNA permutations, the sequences including the identified intergenic regions located between each associated adjacent pair of the exogenous tRNAs; and (v) selecting sequences from said plurality of sequences for inclusion in an operon encoding at least two exogenous tRNAs. A method is provided, comprising:

[0165] The method may be a method for designing an operon encoding at least three, four, five, six, or more than six exogenous tRNAs. However, the number of exogenous tRNAs does not have a specific upper limit for the applicability of the method of the present invention.

[0166] The resulting operon may include a first and a second exogenous tRNA, such that step (i) may include generating the arrangement: a) a first and then a second tRNA and b) a second and then a first tRNA. Other embodiments may include a first, second, and a third exogenous tRNA, such that step (i) may include generating the arrangement: a) a first, then a second, then a third tRNA, b) a first, then a third, then a second tRNA, c) a second, then a first, then a third tRNA, etc. In some embodiments, all possible permutations are generated.

[0167] For each of the above permutations, the method then includes associating each pair of exogenous tRNAs in the permutation with a pair of endogenous tRNAs in the endogenous genome of the host cell for which the operon is intended. The association is based on identifying adjacent tRNA pairs in the endogenous genome that have the highest level of sequence identity to the adjacent exogenous tRNA pairs. For example, if the permutation is "first, then third, then second tRNA", the endogenous adjacent tRNA pair that has the highest level of sequence identity to the first and third tRNA is identified, and the endogenous adjacent tRNA pair that has the highest level of sequence identity to the third and second tRNA is identified.

[0168] Sequence identity may be determined by comparing the acceptor stem sequence of the endogenous tRNA to the acceptor stem sequence of the exogenous tRNA, in particular the first 7 nucleotides of the tRNA and the last 8 nucleotides not including the CCA terminus.

[0169] If an intergenic region between an endogenous pair of tRNAs is identified, the method may optionally set limits on the minimum and / or maximum intergenic region to be considered. For example, the minimum intergenic region to be considered may be 5, 10, 15, 20, or 25 base pairs. In a particular embodiment, the minimum intergenic region to be considered is 10 base pairs. The maximum intergenic region to be considered may be 50, 75, 100, 125, or 150 base pairs. In a particular embodiment, the maximum intergenic region to be considered is 100 base pairs. In one embodiment, the minimum intergenic region to be considered is 10 base pairs, and the maximum is 100 base pairs.

[0170] A plurality of sequences encoding permutations of exogenous tRNAs and intergenic sequences may then be generated. For example, one of the sequences may encode the preceding example "first, then third, then second tRNA," where between the first tRNA and the third tRNA there is an intergenic sequence associated with the pair of endogenous tRNAs most similar to the first and third tRNA, and between the third tRNA and the second tRNA there is an intergenic sequence associated with the pair of endogenous tRNAs most similar to the third and second tRNA.

[0171] A sequence may then be selected from the plurality of sequences for inclusion in the operon encoding the exogenous tRNA. In some embodiments, the plurality of sequences is ranked based on the sum of sequence identities between at least two exogenous tRNAs and the corresponding endogenous tRNAs used to define the intergenic region. Selection may then be made from the ranked list, for example, the sequence with the highest identity may be selected.

[0172] The order of steps is not limited, except that the steps of the method for designing a tRNA operon are performed on the output of the preceding step. For example, adjacent pairs of endogenous tRNAs and intergenic regions in the endogenous genome may be identified before the method of the present invention is started or during the method. A list of adjacent pairs of endogenous tRNAs and intergenic regions in the endogenous genome may be prepared in advance before step (i) of the method of the present invention.

[0173] The method of designing a tRNA operon results in an operon that includes at least a first sequence encoding a first tRNA and a second sequence encoding a second tRNA, and an intergenic sequence derived from a host cell of interest.

[0174] In some embodiments, the operon may contain other ORFs. tRNAs may be used to place other ORFs so that multiple mRNAs can be produced from one promoter. In such embodiments, the method of designing tRNA operons may be used to optimize the flanking regions of these tRNAs.

[0175] Any of the methods for designing a tRNA operon may further comprise the step of producing a nucleic acid molecule encoding said tRNA operon. The nucleic acid may be a DNA sequence and may be included in a vector suitable for delivery to a host cell of interest. Thus, a host cell comprising a nucleic acid molecule encoding said tRNA operon is also provided.

[0176] Any of the methods of designing a tRNA operon may further include the step of experimentally validating the tRNA operon, in such embodiments, the yield of the encoded tRNA may be measured when the operon is inserted into a suitable host cell.

[0177] In an aspect of the invention, a host cell is provided that comprises an endogenous genome, the host cell comprising a nucleic acid encoding an operon comprising at least two exogenous tRNAs, the nucleic acid sequence between each pair of exogenous tRNAs being an intergenic sequence derived from the endogenous genome. The operon may be obtained or obtainable by the method of designing a tRNA operon of the invention. Thus, the intergenic sequence is the intergenic sequence from between the pair of endogenous tRNAs that has the highest identity to the exogenous tRNA. The host cell may also comprise the endogenous tRNA from which the intergenic sequence was derived. In another embodiment, one or more endogenous tRNAs are deleted from the host cell.

[0178] The method of designing a tRNA operon may be implemented on a computer. Thus, a system is provided that includes a processor; and one or more computer readable storage media having stored thereon instructions for execution on said processor to perform the method of the invention. The selection step may be performed manually or may be automated. Additionally, a computer program product is also provided that includes a non-transitory machine readable medium that stores program code that, when executed by one or more processors of a computer system, causes the computer system to perform the method of the invention.

[0179] Thus, a computer-implemented method for designing an operon encoding at least two exogenous tRNAs for expression in a host cell comprising an endogenous genome encoding an endogenous tRNA is provided, the method comprising executing program code on one or more processors to perform the following steps: (i) generating permutations of at least two exogenous tRNAs; (ii) identifying, within the endogenous genome, adjacent pairs of endogenous tRNAs that have the highest level of sequence identity to each adjacent pair of exogenous tRNAs within each permutation of the at least two exogenous tRNAs; (iii) identifying an intergenic region in the endogenous genome between each of the identified adjacent pairs of endogenous tRNAs; and (iv) generating a plurality of sequences encoding each of the at least two exogenous tRNA permutations and including the identified intergenic region located between each associated adjacent pair of the exogenous tRNAs; and optionally (v) selecting sequences from said plurality of sequences for inclusion in an operon encoding at least two exogenous tRNAs. The method may include any of the other features or limitations disclosed herein.

[0180] The present inventors further provide a surprisingly effective method for designing a polycistronic operon encoding at least two exogenous genes for expression in a host cell. The present inventors provide experimental data herein demonstrating that the method described herein can be used to achieve high expression of four exogenous aaRSs in a host cell.

[0181] Thus, in an embodiment of the invention, there is provided a method for designing an operon comprising at least two exogenous ORFs for expression in a host cell, comprising: (i) generating a plurality of 5'UTR sequences for each of at least two exogenous ORFs, wherein each 5'UTR sequence is a sequence that is a negative predicted free energy difference (ΔG tot (ribo)) is optimized for the step; (ii) ΔG for each of the 5′ UTR sequences when the 5′ UTR is positioned 5′ to the optimized exogenous ORF and 3′ to each of the remaining exogenous ORFs of the at least two exogenous ORFs. tot predicting (ribo); and (iii) selecting a 5'UTR sequence and an arrangement of at least two exogenous ORFs; A method is provided, comprising:

[0182] Step (i) may include generating 2, 3, 4, 5, or 6 or more 5'UTR sequences for each of at least two exogenous ORFs. In some examples, 6, 7, 8, 9, 10, 15, 20, or 21 or more 5'UTR sequences are generated. In certain embodiments, five 5'UTR sequences are generated for each exogenous ORF. For example, if an operon contains three exogenous ORFs, 15 5'UTR sequences may be generated, five sets for each exogenous ORF.

[0183] Each 5'UTR sequence reduces the predicted negative free energy difference (ΔG tot (ribo)), so that each 5'UTR is optimized for efficient translation by the ribosome.

[0184] ΔG tot The method for predicting (ribo) is described in detail in the Examples section. tot (ribo) is the free energy required to unfold mRNA (ΔG unfolding ) and the free energy released when mRNA binds to the ribosome to form a ribosome-bound initiation-competent state (ΔG ribo binding ) is the sum of

[0185] In an embodiment, ΔG tot (ribo) is as follows: ΔG tot (ribo)=(ΔG mRNA-rRNA +ΔG start +ΔG spacing -ΔG standby )+ΔG unfolding predicted according to; ΔG mRNA-rRNA is the free energy of the predicted cofolded secondary structure of the last 9 nucleotides of 16S rRNA and the mRNA; ΔG start is the energy released from the binding of the initiator tRNA to the start codon of the sequence encoding the exogenous ORF; ΔG spacing is the energy penalty for a non-optimal spacing length between the Shine-Dalgarno sequence and the start codon of the sequence encoding the exogenous ORF; ΔG standby is the energy required to unfold the secondary structure separating the four nucleotides upstream of the Shine-Dalgarno sequence; ΔG unfolding is the energy required to unfold the secondary structure in the mRNA.

[0186] Further information is provided regarding O-mRNA optimization.

[0187] Methods for optimizing 5'UTRs for efficient translation by ribosomes include: (a) introducing an alteration into the 5'UTR; (b) New ΔG after modification tot (ribo)(ΔG tot new predicting (ribo); (c) Said ΔG tot new ΔG preceded by (ribo) tot Allows modification if it is more negative than (ribo), Said ΔG tot new ΔG preceded by (ribo) tot accepting or rejecting the modification according to a probability distribution if ribo is more positive than ribo; and (d) generating a 5'UTR sequence containing the tolerated modifications. may include:

[0188] Methods may be as described for O-mRNA optimization.

[0189] During the method of the present invention, the ΔG tot new ΔG preceded by (ribo) tot A modification is allowed if it is more negative than (ribo). tot (ribo) is the predicted ΔG tot As discussed herein, the method of the present invention may be repeated, so that the preceding ΔG tot (ribo) is the ΔG calculated during the preceding iteration tot new (ribo) is also acceptable.

[0190] Tolerating modifications during the method of the present invention means that the sequence changes introduced by the modifications are maintained for the next iteration of the method, or are maintained in the sequence that is the output of the method if there are no further iterations of the method.

[0191] During the method of the present invention, the modification tot new ΔG preceded by (ribo) tot If the leading ΔG is more positive than (ribo), it is accepted or rejected according to the probability distribution. tot (ribo)" is as discussed above. The probability distribution is ΔG tot new (ribo) and ΔG tot The probability may be based on a conditional probability where the likelihood of acceptance decreases as the difference between (ribo) increases. The probability may be a Monte Carlo optimization.

[0192] In an embodiment, the probability distribution according to which a modification will be accepted or rejected is:

number

[0193] T SA may be adjusted in any manner disclosed herein. In certain embodiments, T SA will be adjusted to maintain a tolerance rate of 5-20%.

[0194] Rejection of an alteration during the method of the invention means that the sequence change introduced by the alteration is reverted and therefore not maintained for the next iteration of the method or not maintained in the output sequence.

[0195] In an embodiment, the modification is or comprises a single nucleotide change, insertion, or deletion. In another embodiment, the modification is introduced into the 5'UTR or replaces any one of codons 2-20, 2-15, 2-12, 2-10, or 2-5 with a synonymous codon within the sequence encoding the exogenous ORF. In a particular embodiment, the modification comprises a single nucleotide change, insertion, or deletion into the 5'UTR or replaces any one of codons 2-12 within the ORF with a synonymous codon.

[0196] The method of designing an operon comprising at least two exogenous ORFs may be iterative, such that multiple modifications are considered for acceptance or rejection, and the final output sequence includes the cumulative effect of all of the accepted modifications. The iterations may be any as disclosed herein. In particular, steps (a)-(c) of the method of the invention may be iterative. In an embodiment, the method is iterated at least 200, 300, 400, 500, 1000, 5000, or, in particular, 10000 times. In another embodiment, the method is iterative, such that successive iterations result in more negative ΔG tot new For example, steps (a)-(c) may be repeated until at least 10, 50, 100, 250, 500, 1000, 2000, 3000, 5000, or 10000 successive iterations result in a more negative ΔG tot new This may be repeated until it no longer leads to (ribo).

[0197] The initial 5'UTR considered for optimization may have the length and characteristics as described for O-mRNA optimization. In particular, the initial 5'UTR may be 30-40 nucleotides long, or in particular 35 nucleotides. Alternatively, the 5'UTR may be longer, but a 30-40, or in particular 35 nucleotide window is considered by the method of the invention for modification. The 35 nucleotide window may be 35 nucleotides of the 5'UTR closest to the start codon. In other embodiments, the initial 5'UTR may be shorter, for example a 15, 20, or 25 nucleotide 5'UTR, or longer, for example at least 40, 50, or 51 or more nucleotides. It may be desirable to generate a 5'UTR of a particular length, in which case a 15, 20, 25, 30, 35, 45, 50 nucleotide window may be considered so that a particular length of the output sequence can be achieved. The 5'UTR to which step (a) is applied may comprise a wild-type Shine-Dalgarno sequence, or a 5 nucleotide core of the wild-type Shine-Dalgarno sequence. The 5'UTR may be a random sequence other than the Shine-Dalgarno sequence. The Shine-Dalgarno sequence may be 5 nucleotides from the start codon of the ORF that are predicted to be optimally spaced.

[0198] The method for designing an operon containing at least two exogenous ORFs is not limited to a particular number of exogenous ORFs, for example, the method may be used to design an operon containing at least three, at least four, at least five, or at least six exogenous ORFs.

[0199] The method for designing an operon that comprises at least two exogenous ORFs is not limited to use with a particular type of exogenous ORF.The experimental data provided herein provides proof of principle for an operon that comprises multiple sequences that code for aaRS.Therefore, in an embodiment, at least one of the exogenous ORFs codes for aaRS.In another embodiment, the method may be a method for designing a polycistronic operon that codes for at least 2, 3, 4, 5, or 6 aaRSs.

[0200] Step (ii) of the method for designing a polycistronic operon comprises determining the ΔG for each of the 5′ UTR sequences when the 5′ UTR is positioned 5′ to the optimized exogenous ORF and 3′ to each remaining exogenous ORF of the at least two exogenous ORFs. tot This involves predicting the ribo (see FIG. 7, Supplementary FIG. 3). Thus, the 5'UTR optimized for translation of one of the exogenous ORFs is then considered in the context of being located 3' of one of the other exogenous ORFs, and the translation efficiency is measured again. This is done for each of the other exogenous ORFs. For example, in an embodiment where an operon has three exogenous ORFs, a particular 5'UTR optimized for a first exogenous ORF is considered when placed 3' of a second exogenous ORF, and separately, when placed 3' of a third exogenous ORF.

[0201] Step (iii) of the method for designing a polycistronic operon involves selecting an arrangement of the 5'UTR sequence and at least two exogenous ORFs. The arrangement may be selected such that each exogenous ORF is predicted to be translated at a high level. For example, the ΔG tot (ribo) may be predicted and added, the most negative cumulative ΔG tot In another embodiment, the most negative average ΔG for all 5'UTR / exogenous ORF pairs in the operon may be selected. totThe average may be a mean. Additionally, each 5'UTR / exogenous ORF pair may have a target ΔG tot More negative ΔG than (ribo) tot An arrangement with (ribo) may be selected. Targets may be selected to ensure a particular yield of the product of each exogenous ORF in the host cell. For example, targets may be at a level that ensures that the exogenous ORF is translated at a sufficient level for the protein product to achieve its function. For example, if the exogenous ORF encodes an aaRS, the target ΔG tot (ribo) may be such that sufficient aaRS protein is produced in a desired host cell to ensure that the aaRS functions with its cognate tRNA during protein synthesis.

[0202] In certain embodiments, step (iii) comprises determining the most negative average ΔG for all 5′ UTR / exogenous ORF pairs in the operon. tot Each 5'UTR / exogenous ORF pair contains a target ΔG tot More negative ΔG than (ribo) tot (ribo).

[0203] Any of the methods for designing an operon containing an exogenous ORF may further comprise the step of producing a nucleic acid molecule encoding said operon. The nucleic acid may be a DNA sequence and may be included in a vector suitable for delivery to the host cell of interest.

[0204] Any of the methods for designing an operon encoding an exogenous ORF may further comprise the step of experimentally validating the operon. In such embodiments, the yield of the encoded protein may be measured when the operon is inserted into a suitable host cell. The experimental validation may form part of selecting the 5'UTR sequence and the arrangement of the at least two exogenous ORFs.

[0205] In an embodiment of the present invention, a host cell is provided that comprises a nucleic acid encoding an operon comprising at least two exogenous ORFs, where the operon is obtained or obtainable by the method of designing an operon disclosed herein.

[0206] The method for designing a polycistronic operon comprising at least two exogenous ORFs may be performed on a computer. In some embodiments, step (iii) may be performed manually. Thus, a system is provided that includes a processor; and one or more computer-readable storage media having stored thereon instructions for execution on said processor to perform the method of the invention. Additionally, a computer program product is also provided that includes a non-transitory machine-readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to perform the method of the invention.

[0207] Thus, a computer-implemented method for designing an operon comprising at least two exogenous ORFs for expression in a host cell is provided, comprising running program code on one or more processors to perform the following steps: (i) generating a plurality of 5'UTR sequences for each of at least two exogenous ORFs, wherein each 5'UTR sequence is a sequence that is a negative predicted free energy difference (ΔG tot (ribo)) is optimized for the step; (ii) ΔG for each of the 5′ UTR sequences when the 5′ UTR is positioned 5′ to the optimized exogenous ORF and 3′ to each of the remaining exogenous ORFs of the at least two exogenous ORFs. tot predicting (ribo); and optionally (iii) selecting a 5'UTR sequence and an arrangement of at least two exogenous ORFs; The method may include any of the other features or limitations disclosed herein.

[0208] We have combined all of the above advances to create a 68-codon, 24-amino acid genetic code and have succeeded in efficiently incorporating four distinct ncAAs in response to four distinct orthogonal codons via O-ribosome-mediated translation of O-mRNA. As discussed in the Examples section, we use this system to generate, for the first time, a protein containing 20 standard and four non-standard amino acids.

[0209] Therefore, in an aspect of the invention there is provided a host cell comprising: a nucleic acid sequence encoding an O-mRNA encoding an exogenous protein, the O-mRNA being obtained or obtainable by any method of designing an O-mRNA of the invention, the O-mRNA comprising at least two types of orthogonal codons; A nucleic acid sequence comprising an O-tRNA operon encoding at least two orthogonal tRNAs, wherein the at least two orthogonal tRNAs have the ability to decode the at least two types of orthogonal codons, and the operon is obtained or obtainable by any method of designing a tRNA operon of the present invention; A nucleic acid sequence comprising an orthogonal aminoacyl-tRNA synthetase (O-aaRS) encoding at least two O-aaRS operons, wherein the at least two O-aaRSs form an O-aaRS-O-tRNA pair with at least two orthogonal tRNAs, the operon being obtained or obtainable by any method of designing a tRNA operon of the present invention; and Orthogonal Ribosomes A host cell is provided, comprising:

[0210] In embodiments, the O-tRNA and O-aaRS operons are present within the same nucleic acid sequence, e.g., the two operons can be introduced into the host cell via a single vector.

[0211] The exogenous protein encoded by the O-mRNA can be any protein whose production is desired. For example, the exogenous protein can be a therapeutic protein, such as an antibody or a cytokine.

[0212] The host cell comprises at least two O-aaRS and at least two O-tRNAs, which function in pairs, i.e., they form a first aaRS / tRNA pair and a second aaRS / tRNA pair. One pair has the ability to decode one of the types of orthogonal codons, and the other pair has the ability to decode the other type of orthogonal codons. Both pairs have the ability to function with the O-ribosome.

[0213] In other embodiments, the host cell of the invention comprises at least a third and optionally at least a fourth O-aaRS-O-tRNA pair. In such embodiments, the O-mRNA may comprise at least a third and optionally at least a fourth type of orthogonal codon. The third aaRS-tRNA pair has the ability to decode the third type of orthogonal codon, and the fourth aaRS-tRNA pair has the ability to decode the fourth type of orthogonal codon. Additional sets of O-aaRS, O-tRNA, and orthogonal codons may be included. All orthogonal components have the ability to function with the O-ribosome.

[0214] O-aaRSs do not recognize endogenous tRNAs and specifically aminoacylate orthogonal cognate tRNAs (which are not efficient substrates for endogenous synthetases) with non-standard amino acids provided to (or synthesized by) the cell (Chin, JW Expanding and reprogramming the genetic code. Nature 550, 53-60 (2017)).

[0215] The O-ribosome can be any disclosed herein. In particular, the O-ribosome may be O-riboQ1, any O-ribosome disclosed in WO 2008 / 065398 A1 or obtainable by the method disclosed therein, any O-ribosome disclosed in WO 2011 / 077075 A1 or obtainable by the method disclosed therein, or any O-ribosome disclosed in WO 2011 / 077075 A1 or obtainable by the method disclosed therein (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010);Wang, K., Neumann, H., Peak-Chew, SY & Chin, JW Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol. 25, 770-777 (2007); or Schmied, WH et al. Controlling orthogonal ribosome subunit interactions enables evolution of new function. Nature 564, 444-448 (2018)) or any O-ribosome obtainable by the methods disclosed therein.

[0216] The aminoacyl-tRNA synthetase used herein may be various. Although specific tRNA synthetase sequences may be used in the examples, the present invention is not intended to be limited to only those examples. In principle, any aminoacyl-tRNA synthetase that provides the function of charging tRNA (aminoacylation) and functions with O-ribosomes may be used. For example, the tRNA synthetase may be from any suitable species, such as an Archaea, for example a Methanosarcina, such as Alethanosarcina barkeri MS; Methanosarcina barkeri strain Fusaro; Methanosarcina mazei G01; Methanosarcina acetivorans C2A; Methanosarcina thermophila; or Methanococcoides, such as Methanococcoides burtonii. Alternatively, the tRNA synthetase may be from a bacterium, such as a Desulfitobacterium, e.g., Desulfitobacterium hafniense DCB-2; Desulfitobacterium hafniense Y51; Desulfitobacterium hafniense PCP1; or Desulfotomaculum acetoxidans DSM 771.

[0217] The aminoacyl-tRNA synthetase may be pyrrolysyl-tRNA synthetase (PylRS). The PylRS may be wild-type or engineered PylRS. Engineered PylRS is described, for example, in Neumann et al. (Nat Chem Biol 4:232, 2008) and Yanagisawa et al. (Chem Biol 2008, 15:1187), EP 2192185 A1, and WO 2016 / 066995, each of which is incorporated herein by reference. Preferably, an engineered tRNA synthetase gene is selected that increases the incorporation efficiency of non-standard amino acids. The PylRS may be Methanosarcina barkeri (MbPylRS) or Methanosarcina mazei (MmPylRS).

[0218] The tRNA used herein may vary. Although specific tRNAs may be used in the examples, the present invention is not intended to be limited to only those examples. In principle, any tRNA may be used as long as it is compatible with the selected tRNA synthetase and O-ribosome.

[0219] The tRNA may be from any suitable species, for example an archaea, for example Methanosarcina, for example Methanosarcina barkeri MS; Methanosarcina barkeri strain Fusaro; Methanosarcina mazei G01; Methanosarcina acetivorans C2A; Methanosarcina thermophila; or Methanococcoides, for example Methanococcoides brunthonii. Alternatively, the tRNA may be from a bacterium, for example Desulfitobacterium, for example Desulfitobacterium hafniense DCB-2; Desulfitobacterium hafniense Y51; Desulfitobacterium hafniense PCP1; or Desulfotomacrum acetoxydans DSM 771.

[0220] The tRNA gene can be a wild-type tRNA gene or a mutant tRNA gene. Preferably, a mutant tRNA gene is selected that increases the efficiency of incorporation of unnatural amino acids. In one embodiment, the mutant tRNA gene is the U25C variant of PylT described in Biochemistry (2013) 52, 10 (hereby incorporated by reference).

[0221] In one embodiment, the mutant tRNA gene is the Opt variant of PylT described in Fan et al. (Nucleic Acids Research doi:10.1093 / nar / gkv800), which is incorporated herein by reference.

[0222] In one embodiment, the mutant tRNA gene has both the U25C variant and the Opt variant of PylT, i.e., in this embodiment, the tRNA, e.g., PylT tRNA CUA The gene contains both a U25C mutation and an Opt mutation.

[0223] In one embodiment, the tRNA encoding sequence is the pyrrolysine tRNA (PylT) gene from Methanosarcina mazei pyrrolysine, which encodes a tRNAPyl.

[0224] The aminoacyl-tRNA synthetase and tRNA pairs may be disclosed or adapted from those disclosed in (Cervettini, D. et al. Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase-tRNA pairs. Nat. Biotechnol. 38, 989-999 (2020)) or (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nat. Chem. 12, 535-544 (2020)). Each of these documents is incorporated by reference.

[0225] The aaRS, tRNA, and codon set preferably function together and are orthogonal to each endogenous amino acid, group of the aaRS and isoacceptor tRNA, and their cognate groups of codons.

[0226] At least one of the orthogonal codons may be a quadruplet codon. At least one of the orthogonal codons may be a stop codon, such as an amber codon. At least one of the orthogonal codons may be a reassigned sense codon in a genome-recoded prokaryotic cell (see WO 2020 / 229592; or Robertson et al.; Sense codon reassignment enables viral resistance and encoded polymer synthesis; Science; 2021; Vol. 372, Issue 6546, pp. 1057-1062). In certain embodiments, all of the orthogonal codons may be quadruplet codons. In certain embodiments, the O-mRNA comprises a first, second, third, and fourth type of orthogonal codon, each of which is a quadruplet codon.

[0227] The host cell may be a prokaryotic cell. The host cell may be a bacterial cell, such as E. coli. The host cell may have the ability to produce proteins that contain all 20 standard amino acids and at least 4 non-standard amino acids.

[0228] The substrate for the orthogonal tRNA synthetase can be any non-standard amino acid. Thus, the cells of the invention can be used to produce polypeptides that include at least a first non-standard amino acid, at least a second non-standard amino acid, at least a third non-standard amino acid, and at least a fourth non-standard amino acid.

[0229] Therefore, in another aspect of the invention there is provided a method for producing a polypeptide comprising the steps of: Providing a host cell of the invention; incubating a host cell in the presence of a first non-standard amino acid, the first non-standard amino acid being a substrate for one of the O-aaRSs; and Incubating the host cell to allow incorporation of the first non-standard amino acid into the polypeptide via the O-aaRS - O-tRNA pair. A method is provided, comprising:

[0230] As discussed, the host cell may include a first, a second, a third, and a fourth orthogonal aaRS-tRNA pair, where the first pair is capable of decoding a first type of codon to incorporate a first non-standard amino acid, the second pair is capable of decoding a second type of codon to incorporate a second non-standard amino acid, the third pair is capable of decoding a third type of codon to incorporate a third non-standard amino acid, and the fourth pair is capable of decoding a fourth type of codon to incorporate a fourth non-standard amino acid.

[0231] As used herein, the term "non-standard amino acid" means any amino acid except L-alanine, L-cysteine, L-aspartic acid, L-glutamic acid, L-phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L-asparagine, L-proline, L-glutamine, L-arginine, L-serine, L-threonine, L-valine, L-tryptophan, and L-tyrosine.

[0232] A non-standard amino acid may be a non-natural amino acid. As used herein, a "non-natural amino acid" is any amino acid that is not naturally encoded or found in the genetic code. Such an amino acid may be a non-proteinogenic amino acid. Thus, a non-natural amino acid may be any amino acid except L-alanine, L-cysteine, L-aspartic acid, L-glutamic acid, L-phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L-asparagine, L-proline, L-glutamine, L-arginine, L-serine, L-threonine, L-valine, L-tryptophan, and L-tyrosine, L-pyrrolysine, and L-selenocysteine.

[0233] The non-standard amino acid suitable for use with the present invention is not particularly limited.Suitable non-standard amino acids are well known to those skilled in the art, for example, those disclosed in Neumann, H., 2012. FEBS letters, 586(15), pp.2057-2064; and Liu, CC and Schultz, PG, 2010. Annual review of biochemistry, 79, pp.413-444 (incorporated herein by reference). In some embodiments, the non-standard amino acid is p-acetylphenylalanine, m-acetylphenylalanine, O-allyltyrosine, phenylselenocysteine, selenocysteine, p-propargyloxyphenylalanine, p-azidophenylalanine, p-boronophenylalanine, O-methyltyrosine, p-aminophenylalanine, p-cyanophenylalanine, m-cyanophenylalanine, p-fluorophenylalanine, p-iodophenylalanine, p-bromophenylalanine, p-nitrophenylalanine, L-DOPA, 3-aminotyrosine, 3-iodotyrosine, p-isopropylphenylalanine, 3-(2-naphthyl)alanine. , Biphenylalanine, Homoglutamine, D-Tyrosine, p-Hydroxyphenyllactic acid, 2-Aminocaprylic acid, Bipyridylalanine, HQ-Alanine, p-Benzoylphenylalanine, o-Nitrobenzylcysteine, o-Nitrobenzylserine, 4,5-Dimethoxy-2-nitrobenzylserine, o-Nitrobenzyllysine, o-Nitrobenzyltyrosine, 2-Nitrophenylalanine, Dansylalanine, p-Carboxymethylphenylalanine, 3-Nitrotyrosine, Sulfotyrosine, Acetyllysine, Methylhistidine, 2-Aminonanoic acid, 2-Aminodecanoic acid, Pyrrolysine, Cbz-lysine, Boc-lysine, Allyloxycarbonyllysine, N ε -((tert-butoxy)carbonyl)-L-lysine (BocK), Nε-(carbobenzyloxy)-L-lysine (CbzK), N ε-allyloxycarbonyl-L-lysine (AllocK), (S)-2-amino-3-(4-iodophenyl)propanoic acid (pI-Phe), CypK, AlkK, 3-nitro-Tyr, and p-Az-Phe. The first, second, and third non-standard amino acids may be any combination of the non-standard amino acids listed above.

[0234] In certain embodiments, the non-standard amino acids may be any combination of BocK, CbzK, AllocK, pI-Phe, CypK, AlkK, 3-nitro-Tyr, and p-Az-Phe.

[0235] The host cells of the invention can be used to produce products not obtainable by any other method. Thus, in an embodiment of the invention, there is provided a polypeptide or protein containing at least four genetically incorporated non-standard amino acids obtained or obtainable by the methods disclosed herein.

[0236] Sequence comparisons can be performed with the aid of readily available sequence comparison programs. These publicly and commercially available computer programs can calculate the sequence identity between two or more sequences.

[0237] Those skilled in the art understand how to calculate the identity percentage between two nucleic acid sequences. To calculate the identity percentage between two nucleic acid sequences, the alignment of the two sequences must first be prepared, and then the sequence identity value is calculated. The identity percentage for two sequences may take different values ​​depending on (i) the method used to align the sequences, such as the Needleman-Wunsch algorithm (e.g., as applied by Needle (EMBOSS) or Stretcher (EMBOSS), the Smith-Waterman algorithm (e.g., as applied by Water (EMBOSS)), or the LALIGN application (e.g., as applied by Matcher (EMBOSS); and (ii) the parameters used by the alignment method, such as local vs. global alignment, the matrix used, and the parameters applied to gaps.

[0238] Once aligned, there are many different ways to calculate the identity percentage between two sequences. For example, the number of identities may be divided by (i) the length of the shortest sequence; (ii) the length of alignment; (iii) the average length of the sequences; (iv) the number of non-gap positions; or (iv) the number of equivalent positions excluding overhangs. Furthermore, it is also understood that the identity percentage is strongly length-dependent. Thus, the shorter the sequence pair, the higher the sequence identity that can be expected to occur by chance.

[0239] A calculation of the percentage identity between two nucleic acid sequences may then be calculated from such an alignment as (N / T)*100, where N is the number of positions where the sequences share identical residues and T is the total number of positions being compared, including gaps but excluding overhangs.

[0240] The sequence alignment may be a pairwise sequence alignment. Suitable services include Needle (EMBOSS), Stretcher (EMBOSS), Water (EMBOSS), Matcher (EMBOSS), LALIGN, or GeneWise. In an example, the identity between two amino acid sequences may be calculated using the service Needle (EMBOSS) set to default parameters, such as matrix (BLOSUM62), gap open (10), gap extension (0.5), end gap penalty (false), end gap open (10), and end gap extension (0.5). In another example, the identity between two amino acid sequences may be calculated using the service Matcher (EMBOSS) set to default parameters, such as matrix (BLOSUM62), gap open (14), gap extension (4), selective match (1). In an example, the identity between two nucleic acid sequences may be calculated using the service Needle (EMBOSS) set to default parameters, such as matrix (DNAfull), gap open (10), gap extension (0.5), end gap penalty (false), end gap open (10), and end gap extension (0.5). In another example, the identity between two nucleic acid sequences may be calculated using the service Matcher (EMBOSS) set to default parameters, such as matrix (DNAfull), gap open (16), gap extension (4), selective match (1).

[0241] All of the features described in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process disclosed therein, may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.

[0242] For a better understanding of the present invention and to show how embodiments thereof may be put into practice, reference will now be made to the examples which are not intended to limit the invention in any way. [Example]

[0243] We demonstrate a 68-codon genetic code for the incorporation of four distinct non-canonical amino acids, enabled by automated orthogonal mRNA discovery.

[0244] Orthogonal (O-) ribosome-mediated translation of O-mRNA allows the incorporation of up to three distinct non-canonical amino acids (ncAAs) into proteins in E. coli. However, general and efficient incorporation of multiple distinct ncAAs by the O-ribosome requires a scalable strategy for both the generation of efficiently and specifically translated O-mRNAs, as well as the compact expression of multiple O-aminoacyl-tRNA synthetase (O-aaRS) / O-tRNA pairs. We have automated the discovery of O-mRNAs that lead to up to 40-fold more proteins and are up to 50-fold more orthogonal than prior O-mRNAs; protein yields from our O-mRNAs are comparable or superior to those from wild-type mRNAs. These advances allow a 33-fold increase in yield for the incorporation of three distinct ncAAs. Additionally, we automate the generation of operons for O-tRNAs and develop an operon for the O-aaRS. Finally, we combine these advances to create a 68-codon, 24-amino acid genetic code that efficiently incorporates four distinct ncAAs in response to four distinct quadruplet codons. EXAMPLES

[0245] Automating discovery of 5'UTRs for efficient translation by O-ribosomes Our results on the 5′UTR containing the wt RBS strep GFP His6 Predicted ΔG for ORF tot(wt ribo) is -0.5 kcal / mol. In contrast, when the anti-Shine-Dalgarno sequence (aSD) used in the thermodynamic model is changed to that of the O-ribosome, O(trans)- Strep GFP His6 Calculated free energy change (ΔG tot (O-ribo)) was +3.5 kcal / mol. We found that the equilibrium model of initiation combined with a simulated annealing optimization algorithm developed for wt translation (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009)) yielded a 1.2 kcal / mol O(trans)- strep GFP His6 We decided to test whether the proposed method could be adapted to design an O-mRNA sequence that is more efficiently translated by the O-ribosome than the +1 transcription site (Fig. 1a, b). strep GFP His6 The 5'UTR sequence between the ORF was changed to create a very favorable ΔG tot Sequences with (O-ribo) were searched for via a simulated annealing optimization algorithm (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009)).

[0246] Using this algorithm (vol 1), we found that O-ribosomal strep GFP His6 Optimized 5'UTR region for protein production (O1- strep GFP His6 ~O4-strep GFP His6 ) with four new strep GFP His6 The constructs were identified. The ΔG tot (O-ribo) is O1- strep GFP His6 -5.8kcal / mol, O2- strep GFP His6 -4.9kcal / mol, O3- strep GFP His6 -5.1kcal / mol, O4- strep GFP His6 -6.6 kcal / mol. Therefore, these constructs are O(trans)- strep GFP His6 We predicted that this could lead to higher protein levels than the ribosome-containing constructs. strep GFP His6 The optimized sequence (O1- strep GFP His6 ~O4- strep GFP His6 ) is O(trans)- strep GFP His6 This led to a large (11- to 31-fold) increase in protein production with orthogonal translation compared to O1- (Fig. 1c and Supplementary Fig. 1). strep GFP His6 Produced from strep GFP His6 Protein levels were comparable to those from the original construct containing the wt RBS and translated by wt ribosomes. tot (wt-ribo) (Fig. 1a) was higher than +5 kcal / mol in all cases. Therefore, ΔG orthogonality (Figure 1a) predicts that these constructs will be preferentially translated by the O-ribosome. Consistent with this prediction, additional experiments demonstrated that the 5′ UTRs of each of the new constructs are preferentially translated by the O-ribosome. strep GFP His6 Translation of is O-ribosome dependent, and the orthogonality of the new sequence is O(trans)-strep GFP His6 We demonstrated that the α-methyltransferase activity was 12–19 times higher than that of the α-methyltransferase activity (Figure 1c and Supplementary Figure 1a). EXAMPLES

[0247] Automating 5'UTR and ORF discovery for scalable, efficient and selective orthogonal translation In an effort to fully automate the discovery of 5'UTRs that do not direct efficient translation by the wt ribosome but direct maximal protein production by the O-ribosome, we designed a new automated search (vol 2) (Fig. 1b). Our new search introduced an explicit penalty for 5'UTR sequences predicted to be substrates for the wt ribosome and biased towards sequences containing optimally spaced canonical O-RBS sequences.

[0248] The vol 2 search started with a 35 nt 5'UTR containing a 9 nt orthogonal SD (O-SD) sequence predicted to form a perfect Watson-Crick base pair with the orthogonal aSD sequence at the 3' end of the O-16S rRNA. The spacing between the O-SD sequence and the start codon was set to 5 nucleotides, and the sequence of the 5'UTR was randomized, except for the O-SD. We then calculated the ΔG tot (O-ribo) maximizes ΔG tot We searched for sequences that minimized the ribonucleotide sequence (wt ribo). We did not tolerate mutations in the five-nucleotide core (TGGGA) of the O-SD site, which is predicted to base-pair with O-16S rRNA but not with wt 16S rRNA, thus determining orthogonality. Using the vol 2 algorithm, we strep GFP His6 New 5'UTR for (O5-O8- strep GFP His6 These sequences have a higher average ΔG than those derived from vol 1 (−5.6 ± 0.8 kcal / mol). tot (O-ribo) (-7.7 ± 0.4 kcal / mol) (Supplementary Table 1). These sequences were O(trans)- strepGFP His6 Up to 18 times more strep GFP His6 To test the generality of the vol 2 algorithm for enhancing protein production, we examined the orthogonal translation of two additional ORFs, mCherry and E2Crimson. O(trans)-mCherry and O(trans)-E2Crimson, in which the O(trans) 5'UTR was placed between the +1 base of the transcript and the ATG start codon, led to low levels of orthogonal translation. Application of the vol 2 algorithm led to mCherry expression constructs that were up to 10-fold more active with the O-ribosome than O(trans)-mCherry, and up to 8-fold more orthogonal (Fig. 1d and Supplementary Fig. 1c). Similarly, application of the vol 2 algorithm led to orthogonal E2Crimson-producing constructs that were up to 14-fold more active with O-ribosomes than O(trans)-E2Crimson and up to 9-fold more active; E2Crimson was produced by the O-ribosome from O1-E2Crimson (discovered using the vol 2 algorithm) at levels comparable to those produced from the wt RBS using the wt ribosome (Fig. 1e and Supplementary Fig. 1d).

[0249] The first 35 nucleotides of an ORF sequence can contribute substantially to protein yield (Cambray, G., Guimaraes, JC & Arkin, AP Evaluation of 244,000 synthetic sequences reveals design principles to optimize translation in Escherichia coli. Nat. Biotechnol. 36, 1005-1015 (2018); Tuller, T. & Zur, H. Multiple roles of the coding sequence 5′ end in gene expression regulation. Nucleic Acids Res. 43, 13-28 (2015); Seo, SW et al. Predictive design of mRNA translation initiation region to control prokaryotic translation efficiency. Metab. Eng. 15, 67-74 (2013)). However, it remains controversial to what extent changing codons to their synonyms in this sequence affects translation through effects on mRNA secondary structure versus the effects resulting from decoding different synonyms with distinct isoacceptor tRNAs (Cambray, G., Guimaraes, JC & Arkin, AP Evaluation of 244,000 synthetic sequences reveals design principles to optimize translation in Escherichia coli. Nat. Biotechnol. 36, 1005-1015 (2018); Plotkin, JB & Kudla, G. Synonymous but not the same: the causes and consequences of codon bias. Nat. Rev. Genet. 12, 32-42 (2011); Tuller, T. & Zur, H.Multiple roles of the coding sequence 5′ end in gene expression regulation. Nucleic Acids Res. 43, 13-28 (2015), Kudla, G., Murray, AW, Tollervey, D. & Plotkin, JB Coding-Sequence Determinants of Gene Expression in Escherichia coli. Science 324, 255-258 (2009), Allert, M., Cox, JC & Hellinga, HW Multifactorial Determinants of Protein Expression in Prokaryotic Open Reading Frames. J. Mol. Biol. 402, 905-918 (2010), Goodman, DB, Church, GM & Kosuri, S. Causes and Effects of N-Terminal Codon Bias in Bacterial Genes. Science 342, 475-479 (2013)). Changing codons within the first 35 nucleotides of the ORF to synonymous codons results in ΔG. tot (O-ribo) maximizes ΔG tot We realized that this provides additional degrees of freedom in the computational search for mRNAs that minimize (wt ribo). We also hypothesized that in some cases this may allow us to find mRNAs that are more efficiently translated by the O-ribosome and more orthogonal with respect to translation by the wt ribosome. To test this hypothesis, we changed codons 2-12 of each ORF to their synonyms. We thereby created a third algorithm (vol 3) based on vol 2 to explore simultaneous variations in the ORF and 5'UTR (Figure 1b).

[0250] The vol 3 algorithm is the same as vol 2 ( strep GFP His6:-7.7±0.4; mCherry:-9.6±0.5kcal / mol; E2Crimson:-8.9±0.5kcal / mol) tot A significant increase in (O-ribo) strep GFP His6 mCherry: -12.6 ± 0.2 kcal / mol; mCherry: -13.5 ± 0.3 kcal / mol; E2Crimson: -13.2 ± 0.0 kcal / mol) and minimized ΔG tot We maintained that the ribonucleotide sequence (wt ribo) is more orthogonal than that from the vol 2 algorithm and produces proteins at higher levels than those produced by the wt ribosome from the wt message. strep GFP His6 and mCherry (Fig. 1c–e and Supplementary Fig. 1b–d). Overall, our vol 2 and vol 3 algorithms yielded 41, 31, and 14 times more accurate (respectively) results than when the O(trans)5'UTR was used with each ORF. strep GFP His6 , mCherry, and E2Crimson) and these yields are comparable to or exceed those from wt ribosomes on the wt message. The best sequences we found were 31-, 49-, and 9-fold more orthogonal than when the O(trans)5'UTR was used with each ORF (respectively). strep GFP His6 , mCherry, and E2Crimson) high. EXAMPLES

[0251] Optimized orthogonal mRNA allows for increased yield of proteins containing three distinct ncAAs Next, we demonstrated that the increase in protein expression yield from the optimized O-mRNA allows for an increase in the yield of a protein containing three distinct ncAAs via orthogonal translation. Because this work proceeded in parallel with the algorithm development described above, we used the best sequence available at the time, O1-, derived from the vol 1 algorithm. strep GFP His6 We carried out our experiments using O1- strep GFP(40TAG, 136AGGA, 150AGTA) His6 We created a triply orthogonal PylRS / tRNA Pyl Pair (MmPylRS / Methanosarcina spelaei (Mspe) tRNA Pyl CUA (N 6 -(tert-butoxycarbonyl)-L-lysine (BocK) 1), Methanomassiliicoccus luminyensis 1 (Mlum) PylRS (NmH) / Methanomassiliicoccus intestinalis (Mint) tRNA Pyl-A17VC10 UCCU (L121M, L125I, Y126F, M129A, V168V mutants) and Methanomethylophilus sp. 1R26 (M1r26) PylRS (CbzK) / Methanomethylophilus alvus (Malv) tRNA directing the incorporation of Nπ-methyl-L-histidine (NmH)2 Pyl-8 UACU (N 6 This was translated using O-riboQ1 in cells containing the riboQ1 construct (Y126G, M129L mutant, directing the incorporation of -((benzyloxy)carbonyl)-L-lysine (CbzK)3). strep GFP(40BocK, 136NmH, 150CbzK) His6was produced by addition of 1 BocK, 2 NmH, and 3 CbzK. Using this system, we obtained a 2.6±0.4 mg / L strep GFP(40BocK, 136NmH, 150CbzK) His6 The yield was O(trans)- strep GFP (40TAG, 136AGGA or 150AGTA) His6 The yield was 33-fold higher than that from O1- (Figure 2b and Supplementary Table 2). strep GFP(wt) His6 Produced from strep GFP(wt) His6 9% of the ribosomes and is translated from the wt RBS. strep GFP(wt) His6 Produced from strep GFP(wt) His6 This corresponds to 11% of the total yield. The observed yield suggests an average ncAA incorporation efficiency per step of 45%. Mass spectrometry confirmed the synthesis of the correct protein (Fig. 2c, Supplementary Fig. 2). EXAMPLES

[0252] Design of a functional operon for a quadruple orthogonal aaRS / tRNA pair Next, we attempted to build on the development of an efficient O-mRNA to enable the incorporation of four distinct ncAAs into a single protein, where each ncAA is encoded in response to a distinct quadruplet codon. This required four orthogonal aaRS / tRNA pairs that are (1) mutually orthogonal in aminoacylation specificity, (2) have four mutually orthogonal active sites, and (3) can be assigned to four mutually orthogonal quadruplet codons. We developed a PylRS / tRNA Pyl Triplet - Methanomassiliicoccales archaea RumEn M1(Mrum)Pyl(NmH)RS / Mint tRNA Pyl-A17VC10 UCCU(L121M, L125I, Y126F, M129A, V168V mutants directing the incorporation of NmH2), a methanogenic archaeal ISO4-G1(Mg1)Pyl(CbzK)RS / MalvtRNA Pyl-8 UACU (Y125G, M128L mutant, directing integration of CbzK3) and MmPylRS / MspetRNA Pyl-evol CUAG (BocK 1 or N 6 We chose AfTyrRS(PheI) / AftRNA -((allyloxy)carbonyl)-L-lysine (AllocK)4) -, which directs the incorporation of multiple ncAAs, as the starting point for our approach. Tyr-A01 CUA (Y36I, L69M, H74L, Q116E, D165T, I166G, F274V, L298G, D299R mutant, directing the incorporation of (S)-2-amino-3-(4-iodophenyl)propanoic acid (PheI) 5) was chosen as the starting point for the fourth aaRS / tRNA pair; we believe that this pair is capable of integrating multiple pyrrolysyl synthetases and tRNAs. Pyl We have previously shown that ncAAs are orthogonal to each other. Efforts to encode multiple ncAAs demand strategies for efficient and compact expression of the corresponding synthetases and tRNAs. We therefore established an operon-based system for co-expression of four exogenous tRNAs and their cognate synthetases.

[0253] In E. coli, many tRNAs are transcribed in polycistronic operons, and the 5' and 3' ends of mature tRNAs are generated by posttranscriptional RNase processing (Phizicky, EM & Hopper, AK tRNA biology charges to the front. Genes Dev. 24, 1832-1860 (2010); El Yacoubi, B., Bailly, M. & de Crecy-Lagard, V. Biosynthesis and Function of Posttranscriptional Modifications of Transfer RNAs. Annu. Rev. Genet. 46, 69-95 (2012)). We created a program to automatically design synthetic tRNA operons in which the intergenic sequences between exogenous tRNAs are derived from the sequences between the E. coli tRNAs that are most similar to the exogenous tRNAs. The program first generated all possible orderings of the exogenous tRNAs. For each pair of adjacent exogenous tRNAs in the ordering, it identifies the adjacent natural tRNA in the E. coli genome that has the highest sequence identity to the exogenous pair. It then inserts the sequence of the intergenic region found between these natural tRNAs between the exogenous tRNAs. This process generates a synthetic operon sequence for each ordering of the exogenous tRNAs. The program then compares the synthetic operons resulting from each tRNA order and ranks them based on the sum of the sequence identities between the exogenous tRNAs and the corresponding natural tRNAs used to define the intergenic regions in the operon.

[0254] We have Tyr-A01 , MspetRNA Pyl-evol , Mint tRNA Pyl-A17VC10 , and MalvtRNA Pyl-8 We used our program to generate tRNA operons with the top ranked operon, MinttRNA. Pyl-A17VC10 UCCU -inter(glyX, glyY)-MalvtRNA Pyl-8UACU -inter(glyW-cysT)-MspetRNA Pyl-evol CUAG -inter(argY,argZ)-AftRNA Tyr-A01 CUA (where inter(x,y) represents the intergenic spacer sequence between E. coli tRNAs x and y). To adapt this operon to express tRNAs that decode four separate quadruplet codons, we used MspetRNA Pyl-evol CUAG MspetRNA Pyl-evol UCUA (We previously reported that MbtRNA Pyl The anticodon stem evolved in MspetRNA Pyl (produced by transplantation into a mouse) and AftRNA Tyr-A01 CUA AftRNA Tyr-A01 CUAG (AftRNA Tyr-A01 CUA We replaced the tRNA operon with tRNA4(quad), which was generated by an anticodon mutation in tRNA4(quad). We named the resulting tRNA operon tRNA4(quad).

[0255] To identify an operon that allows high expression of the four exogenous aaRSs (MmPylRS, AfTyr(PheI)RS, Mg1(CbzK)PylRS and Mrum(NmH)PylRS), we first generated five optimized 5'UTR regions for each synthetase gene and then estimated the ΔG for progression from folded mRNA to initiation-competent translation complex for each 5'UTR using any of the other three aaRSs as the 5' sequence context. tot We predicted favorable ΔG for all four aaRSs. totWe selected two configurations, RS4_1 and RS4_2, that have the following structure: tRNA4-encoding 5′UTR, tRNA5-encoding 5′UTR, tRNA6-encoding 5′UTR, tRNA7-encoding 5′UTR, and tRNA8-encoding 5′UTR, tRNA9-encoding 5′UTR, tRNA10-encoding 5′UTR, tRNA11-encoding 5′UTR, tRNA12-encoding 5′UTR, and tRNA13-encoding 5′UTR, tRNA14-encoding 5′UTR, tRNA15-encoding 5′UTR, tRNA16-encoding 5′UTR, and tRNA17-encoding 5′UTR, tRNA18-encoding 5′UTR, tRNA19-encoding 5′UTR, and tRNA20-encoding 5′UTR. We cloned each of the aaRS operons into a plasmid encoding tRNA4 to generate compact synthetase and tRNA expression modules (RS4_1 / tRNA4 and RS4_2 / tRNA4) (Supplementary Fig. 3). We tested the activity of each aaRS in each operon (Supplementary Fig. 4, Supplementary Table 3). These experiments led us to design an optimized chimeric aaRS operon in which we transplanted 150 nt upstream of the optimized 5′UTR of Mrum(NmH)PylRS from RS4_2 to RS4_1 to generate RS4_1-2. This operon combined the best properties of RS4_2 and RS4_1 (Supplementary Fig. 4c).

[0256] We combined the RS4_1-2 and tRNA4(quad) operons in a single vector (279 RS4_1-2 / tRNA4(quad)) and strep GFP(40XXXX) His6 We systematically tested the activity and orthogonality of each aaRS / tRNA pair produced by measuring the GFP fluorescence produced from O-riboQ1 (where XXXX represents TAGA, AGGA, AGTA, or CTAG). Cells contained O-riboQ1, either without or with each individual ncAA (NmH2, CbzK3, AllocK4, PheI5), and RS4_1-2 / tRNA4(quad) (Figure 3a-d). O-riboQ1 was inhibited in the presence of RS4_1-2 / tRNA4(quad) and all four ncAAs (NmH2, CbzK3, PheI5, AllocK4). strep GFP(150XXXX) His6 (wherein XXXX represents TAGA, AGGA, AGTA, or TAGA) produced by O-riboQ1 strep GFP(40X) His6 ESI-MS of (where X represents NmH2, CbzK3, AllocK4, and PheI5) demonstrated that each aaRS, tRNA, and codon is functionally orthogonal to each other (Figure 3e-h). EXAMPLES

[0257] Genetically encoding four distinct ncAAs using four distinct quadruplet codons We combined our progress in generating aaRS / tRNA operons for orthogonal pairs to create optimized O-mRNAs that are efficiently read by O-riboQ1 to incorporate four distinct ncAAs into a single protein in response to four distinct quadruplet codons (Fig. 4a). We synthesized O1- in cells that contained RS4_1-2 / tRNA4(quad) and were provided with all four ncAA substrates (NmH2, CbzK3, PheI4, AllocK5). strep GFP(40CTAG, 50TAGA, 136AGGA, 50AGTA) His6 O-riboQ1-mediated translation of strep GFP (40PheI, 50AllocK, 136NmH, 150CbzK) His6 (Figure 4b). strep GFP (40PheI, 50AllocK, 136NmH, 150CbzK) His6 The production of ncAAs was dependent on the addition of all four ncAAs, with 0.41 ± 0.03 mg / mL of protein produced (Supplementary Table 2). The observed yield suggests an average ncAA incorporation efficiency per step of 38%. Mass spectrometry confirmed the incorporation of all four ncAAs in response to four distinct quadruplet codons (Figure 4c, Supplementary Figure 5). In additional experiments, we also demonstrated the incorporation of four distinct ncAAs in response to three quadruplet codons and an amber codon (Supplementary Figures 6-8, Supplementary Table 2). EXAMPLES

[0258] Considerations of Examples 1 to 5 We have developed a computational approach to design O-mRNA sequences that are efficiently and selectively translated by the O-ribosome. The new O-mRNAs lead to up to 40 times more protein and are up to 50 times more orthogonal than O-mRNAs created by grafting previously used 5'UTRs containing an O-RBS in front of the desired ORF. The O-mRNAs we created direct orthogonal protein production at levels equal to or higher than from wt mRNAs translated by wt ribosomes. Our automated, rapid and scalable method for O-mRNA discovery is a step forward in the design and directed evolution of orthogonal translation systems that incorporate multiple ncAAs and polymerize new monomers (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Rackham, O. & Chin, JW A network of orthogonal ribosome·mRNA pairs. Nat. Chem. Biol. 1, 159-166 (2005); Wang, K., Neumann, H., Peak-Chew, SY & Chin, JW Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol. 25, 770-777 (2007), Schmied, WH et al. Controlling orthogonal ribosome subunit interactions enables evolution of new function. Nature 564, 444-448 (2018)), as well as the creation and application of orthogonal gene expression systems (An, W. & Chin, JWSynthesis of orthogonal transcription-translation networks. Proc. Natl. Acad. Sci. USA 106, 8477-8482 (2009), Darlington, APS, Kim, J., Jimenez, JI & Bates, DG Dynamic allocation of orthogonal ribosomes facilitates uncoupling of co-expressed genes. Nat. Commun. 9, 1-12 (2018)).

[0259] Our O-mRNA optimization strategy involves explicit selection for orthogonality and co-optimization of 5'UTR and ORF sequences. We show that co-optimization of the 5'UTR and synonymous codon selection in the ORF results in a larger (more negative) ΔG tot These sequences also yielded the predicted values ​​for ΔG totWe found that the 5'UTR sequence and synonymous codons in the ORF had large (positive) predictive values ​​for protein yield (wt ribo). We found that co-optimization of 5'UTR sequences and synonymous codons within the ORF could improve protein yield, with testing of four clones leading to high levels of translation in each case tested. These observations are consistent with the notion that mRNA folding is the primary predictor of protein yield, among known parameters (Cambray, G., Guimaraes, JC & Arkin, AP Evaluation of 244,000 synthetic sequences reveals design principles to optimize translation in Escherichia coli. Nat. Biotechnol. 36, 1005-1015 (2018)). We note that other parameters, including codon compatibility, may affect protein yield, and it will be interesting to see whether including these considerations in future iterations of the algorithm will lead to even higher predictive power. Future studies will also explore co-optimization of 5'UTR sequences and coding sequences to improve production of difficult-to-express proteins from wt ribosomes.

[0260] Our automated O-mRNA design was performed using our previously developed triply orthogonal PylRS / tRNA PylBy combining with Pair, we increased the yield of a protein containing three distinct ncAAs by 33-fold. We established a pipeline for efficient and compact co-expression of many exogenous aaRSs and tRNAs. We developed a computer program to produce a polycistronic tRNA operon that mimics the endogenous transcription system in E. coli. Our algorithm provides a general solution for producing multiple distinct tRNAs in E. coli under the same promoter on one plasmid and can be easily adapted for other organisms. We also developed a polycistronic aaRS operon for efficient expression of four mutually orthogonal synthetases alongside a tRNA operon. We combined our progress to produce for the first time in vivo a protein consisting of 24 amino acids, i.e., the canonical 20 amino acids and four ncAAs. Each ncAA is alternatively translated on the O-mRNA and encoded using a quadruplet codon not used in natural translation, creating an organism with a 68-codon genetic code.

[0261] We anticipate that emerging developments in the creation of mutually orthogonal aaRS / tRNA pairs that recognize distinct ncAAs and decode distinct quadruplet codons may enable the expansion of the quadruplet code. The efficiency of quadruplet decoding can be further improved by selecting ribosomes that no longer read triplet codons or by developing quadruplet decoding in organisms with a compressed genetic code in which competing triplet-decoding tRNAs have been removed (Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514- 518 (2019); Chatterjee, A., Lajoie, MJ, Xiao, H., Church, GM & Schultz, PG A Bacterial Strain with a Unique Quadruplet Codon Specifying Non-native Amino Acids. ChemBioChem 15, 1782- 1786 (2014)).

[0262] References for Examples 1-6 and Figure Legends 5-12 (Supplementary Figures 1-8) 1. Chin, JW Expanding and reprogramming the genetic code. Nature 550, 53-60 (2017). 2. de la Torre, D. & Chin, JW Reprogramming the genetic code. Nat. Rev. Genet., 1-16 (2020). 3. Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, J.W. Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010). 4. Wang, K. et al. Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393-403 (2014). 5. Anderson, J.C. et al. An expanded genetic code with a functional quadruplet codon. Proc. Natl. Acad. Sci. U.S.A. 101, 7566-7571 (2004). 6. Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514-518 (2019). 7. Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59-64 (2016). 8. Malyshev, D.A. et al. A semi-synthetic organism with an expanded genetic alphabet. Nature509, 385-388 (2014). 9. Zhang, Y. et al. A semi-synthetic organism that stores and retrieves increased genetic information. Nature 551, 644-647 (2017). 10. Zhang, Y. et al. A semisynthetic organism engineered for the stable expansion of the genetic alphabet. Proc. Natl. Acad. Sci. U.S.A. 114, 1317-1322 (2017). 11. Fischer, E.C. et al. New codons for efficient production of unnatural proteins in a semisynthetic organism. Nat. Chem. Biol. 16, 570-576 (2020). 12. Neumann, H., Slusarczyk, A.L. & Chin, J.W. De Novo Generation of Mutually Orthogonal Aminoacyl-tRNA Synthetase / tRNA Pairs. J. Am. Chem. Soc. 132, 2142-2144 (2010). 13. Chatterjee, A., Sun, S.B., Furman, J.L., Xiao, H. & Schultz, P.G. A Versatile Platform for Single- and Multiple-Unnatural Amino Acid Mutagenesis in Escherichia coli. Biochemistry 52, 1828-1837 (2013). 14. Willis, J.C.W. & Chin, J.W. Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nat. Chem. 10, 831-837 (2018). 15. Dunkelmann, D.L., Willis, J.C.W., Beattie, A.T. & Chin, J.W. Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non canonical amino acids. Nat. Chem. 12, 535-544 (2020). 16. Cervettini, D. et al. Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase-tRNA pairs. Nat. Biotechnol. 38, 989-999 (2020). 17. Zhang, M.S. et al. Biosynthesis and genetic encoding of phosphothreonine through parallel selection and deep sequencing. Nat. Methods 14, 729-736 (2017). 18. Italia, J. et al. Mutually Orthogonal Nonsense-Suppression Systems and Conjugation Chemistries for Precise Protein Labeling at up to Three Distinct Sites. J. Am. Chem. Soc. 141, 6204-6212 (2019). 19. Rackham, O. & Chin, J.W. A network of orthogonal ribosome·mRNA pairs. Nat. Chem. Biol. 1, 159-166 (2005). 20. Wang, K., Neumann, H., Peak-Chew, S.Y. & Chin, J.W. Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol. 25, 770-777 (2007). 21. Schmied, W.H. et al. Controlling orthogonal ribosome subunit interactions enables evolution of new function. Nature 564, 444-448 (2018). 22. Venkat, S. et al. Genetically Incorporating Two Distinct Post-translational Modifications into One Protein Simultaneously. ACS Synth. Biol. 7, 689-695 (2018). 23. Chin, J.W. Expanding and Reprogramming the Genetic Code of Cells and Animals. Annu. Rev. Biochem. 83, 379-408 (2014). 24. Cambray, G., Guimaraes, J.C. & Arkin, A.P. Evaluation of 244,000 synthetic sequences reveals design principles to optimize translation in Escherichia coli. Nat. Biotechnol. 36, 1005-1015 (2018). 25. Plotkin, J.B. & Kudla, G. Synonymous but not the same: the causes and consequences of codon bias. Nat. Rev. Genet. 12, 32-42 (2011). 26. Tuller, T. & Zur, H. Multiple roles of the coding sequence 5’ end in gene expression regulation. Nucleic Acids Res. 43, 13-28 (2015). 27. Salis, H.M., Mirsky, E.A. & Voigt, C.A. Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009). 28. Na, D., Lee, S. & Lee, D. Mathematical modeling of translation initiation for the estimation of its efficiency to computationally design mRNA sequences with desired expression levels in prokaryotes. BMC Syst. Biol. 4, 1-16 (2010). 29. Seo, S.W. et al. Predictive design of mRNA translation initiation region to control prokaryotic translation efficiency. Metab. Eng. 15, 67-74 (2013). 30. Salis, H.M. in Methods in Enzymology, Vol. 498 19-42 (Academic Press, Cambridge, MA, USA; 2011). 31. Espah Borujeni, A., Channarasappa, A.S. & Salis, H.M. Translation rate is controlled by coupled trade-offs between site accessibility, selective RNA unfolding and sliding at upstream standby sites. Nucleic Acids Res. 42, 2646-2659 (2014). 32. Espah Borujeni, A. & Salis, H.M. Translation Initiation is Controlled by RNA Folding Kinetics via a Ribosome Drafting Mechanism. J. Am. Chem. Soc. 138, 7016-7023 (2016). 33. Espah Borujeni, A. et al. Precise quantification of translation inhibition by mRNA structures that overlap with the ribosomal footprint in N-terminal coding sequences. Nucleic Acids Res. 45, 5437-5448 (2017). 34. Kudla, G., Murray, A.W., Tollervey, D. & Plotkin, J.B. Coding-Sequence Determinants of Gene Expression in Escherichia coli. Science 324, 255-258 (2009). 35. Allert, M., Cox, J.C. & Hellinga, H.W. Multifactorial Determinants of Protein Expression in Prokaryotic Open Reading Frames. J. Mol. Biol. 402, 905-918 (2010). 36. Goodman, D.B., Church, G.M. & Kosuri, S. Causes and Effects of N-Terminal Codon Bias in Bacterial Genes. Science 342, 475-479 (2013). 37. Phizicky, E.M. & Hopper, A.K. tRNA biology charges to the front. Genes Dev. 24, 1832-1860 (2010). 38. El Yacoubi, B., Bailly, M. & de Crecy-Lagard, V. Biosynthesis and Function of Posttranscriptional Modifications of Transfer RNAs. Annu. Rev. Genet. 46, 69-95 (2012). 39. An, W. & Chin, J.W. Synthesis of orthogonal transcription-translation networks. Proc. Natl. Acad. Sci. U.S.A. 106, 8477-8482 (2009). 40. Darlington, A.P.S., Kim, J., Jimenez, J.I. & Bates, D.G. Dynamic allocation of orthogonal ribosomes facilitates uncoupling of co-expressed genes. Nat. Commun. 9, 1-12 (2018). 41. Chatterjee, A., Lajoie, M.J., Xiao, H., Church, G.M. & Schultz, P.G. A Bacterial Strain with a Unique Quadruplet Codon Specifying Non-native Amino Acids. ChemBioChem 15, 1782-1786 (2014).

[0263] method Thermodynamic model of translation initiation The thermodynamic model has been described previously (Chin, J. W. Expanding and reprogramming the genetic code. Nature 550, 53-60 (2017)). Briefly, the model calculates the predicted energy of a freely folded mRNA, ΔG unfolding , and the predicted energy of the initiation-competent ribosome-bound state, ΔG ribo_binding The free energy difference, ΔG tot Identify. ΔG tot = ΔG ribo_binding +ΔG unfolding

[0264] Here, ΔG unfolding is the energy required to unfold the mRNA secondary structure. The free energy released in the formation of the initiation-competent state, ΔG ribo_binding It consists of four components. ΔG ribo_binding = ΔG mRNA-rRNA +ΔG start +ΔG spacing -ΔG standby

[0265] ΔG mRNA-rRNA is the free energy of the predicted cofolded secondary structure of the last 9 nt of 16S rRNA and mRNA, where the major energetic contribution comes from the hybridization energy between the Shine-Dalgarno (SD) or orthogonal Shine-Dalgarno O-SD sequence of the mRNA and the 16S rRNA. Reflecting the ribosome footprint, mRNA folding downstream of the hybridization site is not allowed. ΔG start is the energy released from the binding of the initiator tRNA to the start codon. ΔG spacing is the energy penalty for non-optimal spacing length between the SD site and the start codon. ΔG standbyis the energy required to unfold the secondary structure separating the standby site, defined here as the four nucleotides upstream of the SD site.

[0266] A simulated annealing optimization algorithm for automated O-mRNA discovery RNA secondary structure prediction is performed in the NuPACK suite using the "mfe" algorithm. Calculations consider a window of up to 35 nt in the 5'UTR and ORF; if longer sequences are used, only the 35 nt closest to the start codon are considered.

[0267] The vol 1 algorithm is derived from a previously described simulated annealing optimization algorithm1 but with a ΔG mRNA-rRNA The final 9 nt of the orthogonal 16S rRNA (ATGGGATTA) is used instead of the standard sequence (ACCTCCTTA) for the calculation of ΔG tot (O-ribo) was evaluated using a thermodynamic model and the target function ΔG target is compared with ΔG target may be set arbitrarily or infinitely negative so that the target of the algorithm is as negative as possible. In an iterative procedure, mutations (either single nucleotide changes, insertions or deletions) are introduced into the 5'UTR, resulting in a new ΔG tot new (O-ribo) is calculated. If the mutated sequence violates the sequence constraints, the mutation is rejected. target ΔG closer to tot new The mutation is tolerated if it leads to (O-ribo). ΔG tot new (O-ribo) value is the original ΔG tot (O-ribo) target If it differs from, then the mutation is allowed with the following probability:

number

[0268] Here, T SA is the simulated annealing temperature, which is adjusted to maintain a tolerance rate of 5–20%. The algorithm terminates after 10,000 iterations, and the 5'UTR and predicted ΔG tot Output (O-ribo).

[0269] The vol 2 algorithm is based on the vol 1 algorithm. The random starting 5'UTR contains a 9 nucleotide O-SD site (TAATCCCAT) predicted to be perfectly complementary to O-16S rRNA (ATGGGATTA) at optimal spacing of 5 nucleotides from the ATG start codon. The ΔG tot (wt ribo) and ΔG tot (O-ribo) was evaluated using a thermodynamic model and a hypothetical ΔG tot (opt) is ΔG tot (opt)=ΔG tot (O-ribo)-0.5*ΔG tot In contrast to the vol 1 algorithm, ΔG target In an iterative procedure, mutations (either single nucleotide changes, insertions or deletions) are introduced into the 5'UTR to obtain new ΔG tot new A value is calculated. If the mutated sequence violates sequence constraints or removes the 5-nucleotide core (TCCCA) of the O-SD sequence, the mutation is rejected. The mutated sequence has an improved (more negative) ΔG tot new A mutation is tolerated if it leads to an (opt) value. ΔG tot new (opt) value is the original ΔG tot If it is greater (more positive) than (opt), the mutation is allowed with the following probability:

number

[0270] 500 consecutive iterations yield ΔG tot If it does not result in an improvement in (opt), the algorithm terminates and the 5'UTR and ΔG tot We typically run the algorithm multiple times and select the most favorable ΔG tot We found that this was computationally more efficient for identifying highly translated 5'UTRs than running the algorithm for many more iterations per starting sequence. In this study, we selected 4 sequences from 24 predicted 5'UTRs.

[0271] The vol 3 algorithm is based on the vol 2 algorithm. In addition to a random starting 5'UTR, amino acids at positions 2-12 are encoded by a randomly selected selection of synonymous codons. In addition to single nucleotide changes, insertions, or deletions in the 5'UTR, synonymous codon changes at positions 2-12 in the ORF are allowed as mutation mechanisms during simulated annealing optimization.

[0272] tRNA operon designer The program generates a list of all pairs of tRNAs in the host organism whose genes are adjacent to each other and on the same strand. It then extracts the gene sequences of these endogenous tRNA pairs as well as the corresponding intergenic sequences. Optionally, the user may specify the minimum and maximum length of the intergenic sequence to be considered by the program. For the tRNA operons used in this study, we used as the host genome the E. coli K-12 strain MG1655 substrain genome (version U00096.3, last updated: 24 September 2018), which has a minimum and maximum intergenic sequence length of 10 and 100 base pairs, respectively.

[0273] The program then generates all ordered pairs of exogenous tRNAs. For each ordered pair of exogenous tRNAs, the acceptor stem sequences of these tRNAs are compared to the acceptor stem sequences of the endogenous tRNA pairs. For consistency, we consider the first 7 nucleotides and the last 8 nucleotides of the tRNA (excluding the CCA end), which includes the standard E. coli tRNA acceptor stem and the discriminatory base region. Each endogenous tRNA pair is ranked by its similarity to the exogenous tRNA pair, calculated as the sequence identity of the acceptor stem. The exogenous tRNA pair is then assigned a score, defined as the sequence identity of the acceptor stem of the most similar endogenous tRNA pair.

[0274] Finally, the program generates all the orderings, or permutations, of the exogenous tRNAs. A synthetic tRNA operon corresponding to each permutation is created by inserting an endogenous tRNA intergenic region between each ordered pair of exogenous tRNA genes in the permutation. For each ordered exogenous pair, the intergenic region corresponding to the most similar endogenous tRNA pair is selected. Each operon is assigned a score calculated as the sum of the scores of all ordered pairs in the permutation.

[0275] The sequences and scores of the operons, together with information on the order of selected tRNAs and intergenic regions, are presented as a ranked list of entries in an Office Open XML spreadsheet.

[0276] Assembly of the aaRS operon Details on the assembly of the operon are given in Supplementary Figure 3. All predicted 5'UTRs along with ΔGtot(wt ribo) for the alignment are given in Supplementary Table 3.

[0277] DNA constructs Reporter gene ( strep GFP His6, mCherry and E2Crimson) were cloned by Gibson assembly into a p15A plasmid containing a tetracycline resistance cassette and expressed from the lac promoter. An optimized 5'UTR was inserted between the +1 transcription site and the ORF by quick-change PCR Gibson assembly. An optimized 5'UTR and ORF were inserted between the +1 transcription site and codon 13 by quick-change PCR Gibson assembly. O(trans)- strep GFP(40TAG, 136AGGA, 150AGTA) His6 was expressed from the p15A plasmid previously described (de la Torre, D. & Chin, JW Reprogramming the genetic code. Nat. Rev. Genet., 1-16 (2020)). strep GFP(40TAG, 136AGGA, 150AGTA) His6 , O1- strep GFP(40TAG, 50CTAG, 136AGGA, 150AGTA) His6 and O1- strep GFP(40CTAG, 50TAGA, 136AGGA, 150AGTA) His6were synthesized by IDT as gBlock double-stranded DNA fragments and cloned by Gibson assembly into a standard p15A reporter backbone. Ribosomes were encoded on a previously described pRSF plasmid containing a kanamycin resistance cassette and expressed from the trc promoter (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444 (2010); Wang, K. et al. Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393-403 (2014)).

[0278] The synthetase operon RS3 and tRNA operon tRNA3 were encoded on the previously described pMB1 plasmid (de la Torre, D. & Chin, JW Reprogramming the genetic code. Nat. Rev. Genet., 1-16 (2020)), which contains a spectinomycin resistance cassette. Synthetase operons RS4_1 and RS4_2 were synthesized as gBlocks by IDT and inserted after the +1 transcription site of the promoter of glnS' by Gibson assembly (de la Torre, D. & Chin, JW Reprogramming the genetic code. Nat. Rev. Genet., 1-16 (2020)). RS4_1-2 was assembled by Gibson cloning of fragments from RS4_1 and RS4_2. The tRNA operon tRNA4 was synthesized as gBlock by IDT and assembled in the same pMB1 plasmid as a synthetase operon under the control of the lpp promoter by Gibson cloning. tRNA4(quad) was assembled from tRNA4 by quick change PCR Gibson assembly.

[0279] Measurement of fluorescent reporter activity and orthogonality Each fluorescent reporter ( strep GFP His6To measure the activity and orthogonality of the ribosomes (E2Crimson, mCherry, and E2Crimson), we transformed 8 μL of chemically competent E. coli DH10B cells carrying pRSF plasmids encoding copies of the O-ribosome or wt ribosome with 0.5 μL of p15A plasmid encoding a fluorescent reporter. We recovered the transformed cells for 1 h at 37 °C and 750 rpm in 180 μL of SOC medium in a 96-well microtiter plate format. 30 μL of rescued cells were used to inoculate 500 μL of selective 2xYT-kt (2xYT medium containing 50 μg / mL kanamycin, 12.5 μg / mL tetracycline) medium in a 1.2 mL 96-well plate format, and the cultures were grown overnight at 37 °C and 750 rpm. 30 μL of the overnight culture was used to inoculate 500 μL of 2xYT-kt medium in a 1.2 mL 96-well plate format. Cells were grown for 2 h at 37° C. and 750 rpm, and production of the fluorescent reporter as well as ribosomes was induced by the addition of 10 μL of 0.1 M IPTG giving a final concentration of 2 mM IPTG. Cells were grown for 18 h at 37° C. and 750 rpm. 180 μL of each culture was transferred to a 96-well flat-bottom Costar plate, and fluorescence and optical density were measured using a PHERAstar FS plate reader.

[0280] O1- strep GFP(40TAG, 136AGGA, 150AGTA) His6 or O(trans)- strep GFP(40TAG, 136AGGA, 150AGTA) His6 Comparative analysis of the efficiency of triple integration from From reporters containing grafted or optimized orthogonal 5'UTRs strep GFP(40TAG, 136AGGA, 150AGTA) His6 To compare the efficiency of incorporation of three distinct ncAAs into O1- strep GFP His6 , O1- strep GFP(40TAG, 136AGGA, 150AGTA)His6 or O(trans)- strep GFP(40TAG, 136AGGA, 150AGTA) His6 We transformed 8 μL of chemically competent E. coli DH10B cells carrying a pRSF plasmid encoding a copy of O-riboQ1 with 0.4 μL of p15A plasmid encoding. We recovered the transformed cells for 1 h at 37°C and 750 rpm in 180 μL of SOC medium in a 96-well microtiter plate format. We used 30 μL of rescued cells to inoculate 500 μL of 2xYT-kts medium (2xYT containing 25 μg / mL kanamycin, 12.5 μg / mL tetracycline, and 37.5 μg / mL spectinomycin) in a 1.2 mL 96-well plate format, and the culture was grown overnight at 37°C and 750 rpm. 100 μL of the overnight culture was used to inoculate 4 mL of 2xYT-kts medium containing either 4 mM BocK1, 4 mM NmH2, and 2 mM CbzK3 or no ncAA in a 10 mL 24-well plate format. Cells were grown at 37°C and 220 rpm for 2 h and production of strepGFPHis6 and O-riboQ1 was induced by the addition of 8 μL of 1 M IPTG giving a final concentration of 2 mM IPTG. Cells were grown at 37°C and 750 rpm for 18 h. 180 μL of each culture was transferred to a 96-well flat-bottom Costar plate and fluorescence and optical density were measured using a PHERAstar FS. The remainder of the culture was centrifuged at 3200 rcf for 10 min and harvested in an OD600 adjusted amount of BugBuster containing Roche cOmplete proteinase inhibitor. Cells were lysed at room temperature under end-over-end rotation for 1 h. The lysates were transferred to 1.5 mL Eppendorf tubes and spun down at 15000 rcf for 20 minutes. 180 μL of the clarified cell lysates were transferred to a 96-well flat-bottom Costar plate and fluorescence was measured using a PHERAstar FS.

[0281] Assessment of activity and orthogonality of aaRS / tRNA operons To evaluate the activity and orthogonality of each aaRS / tRNA pair in our operon, we used 0.4 μL of pMB1 plasmid encoding the operon (aaRS4_1 / tRNA4, aaRS4_2 / tRNA4, aaRS4_1-2 / tRNA4, and aaRS4_1-2 / tRNA(quad)) in addition to the pRSF plasmid encoding a copy of O-riboQ1. strep GFP(40XXXX) His6 We transformed 8 μL of chemically competent E. coli DH10B cells harboring p15A plasmids encoding (where XXXX represents either TAG (all operons except aaRS4_1-2 / tRNA4(quad)), TAGA (only aaRS4_1-2 / tRNA4(quad)), AGGA, AGTA, or CTAG). We recovered the transformed cells for 1 h at 37° C. and 750 rpm in 180 μL of SOC medium in a 96-well microtiter plate format. 30 μL of rescued cells were used to inoculate 500 μL of selective 2xYT-kts medium in 1.2 mL of 96-well plate format, and the cultures were grown overnight at 37° C. and 750 rpm. 30 μL of the overnight culture was used to inoculate 500 μL of selective 2xYT-kts medium containing either 4 mM BocK 1, 4 mM NmH 2, 2 mM CbzK 3, 4 mM AllocK 4, 2 mM PheI 5 or no ncAA in a 1.2 mL 96-well plate format. Cells were grown at 37° C. and 750 rpm for 2 h. strep GFP(40XXXX) His6 In addition, expression of O-riboQ1 was induced by the addition of 10 μL of 0.1 M IPTG, giving a final concentration of 2 mM IPTG. Cells were grown for 18 h at 37° C. and 750 rpm. 180 μL of each culture was transferred to a 96-well flat-bottom Costar plate, and fluorescence and optical density were measured using a PHERAstar FS.

[0282] For MS analysis strep GFP(40X) His6 Production of To isolate proteins for MS analysis to assess the orthogonality of the aaRS / tRNA operons, 0.4 µL of pMB1 plasmid encoding the operon RS4_1-2 / tRNA4 or RS4_1-2 / tRNA4(quad) was added to O1- strep GFP(40XXXX) His6 (where XXXX stands for any of TAG (RS4_1-2 / tRNA4 only), TAGA (RS4_1-2 / tRNA4(quad) only), AGGA, AGTA, and CTAG) into 50 μL of chemically competent E. coli DH10B cells harboring a pRSF plasmid encoding a copy of O-riboQ1 together with 0.4 μL of p15A plasmid encoding either TAG (RS4_1-2 / tRNA4 only), TAGA (RS4_1-2 / tRNA4(quad) only), AGGA, AGTA, or CTAG, respectively. We recovered the transformed cells in 400 μL of SOC medium in a 1.5 mL Eppendorf tube at 37° C. and 750 rpm for 1 h. 100 μL of rescued cells were used to inoculate 50 mL of selective 2xYT-kts medium in a 250 mL Erlenmeyer flask and the culture was grown overnight at 37° C. and 220 rpm. Five mL of the overnight culture was used to inoculate 100 mL of selective 2xYTkts medium containing a combination of the ncAA BocK 1, NmH 2, CbzK 3, AllocK 4 and PheI 5 according to the construct used (RS4_1-2 / tRNA4 with 1, 2, 3, 5 - RS4_1-2 / tRNA4(quad) with 2, 3, 4, 5). The culture was grown at OD 600The cells were grown for 2-3 h at 37°C and 220 rpm until 0.5 and induced with 200 μL of 1M IPTG to a final concentration of 2 mM IPTG. Cells were grown for 18 h at 37°C and 220 rpm. Cells were centrifuged for 12 min at 3200 rcf, resuspended in 10 mL BugBuster containing Roche cOmplete proteinase inhibitor, sonicated for 1.5 min (2 sec on, 2 sec off at 40% amplitude), and the lysate was centrifuged for 20 min at 4°C, 15000 rcf. The lysate was bound to 40 μL of nickel-NTA beads overnight. Beads were washed six times with 240 μL of 20 mM imidazole in PBS. Proteins were eluted nine times in 20 μL of 250 mM imidazole. Buffer was exchanged into water using a 3 kDa Amicon ultra column for MS and MS / MS analysis.

[0283] O1- strep GFP(40CTAG, 50TAGA, 136AGGA, 150AGTA) His6 Assessment of orthogonality and efficiency of incorporation of four distinct ncAAs in response to four distinct quadruplet codons from To evaluate the efficiency and orthogonality of the incorporation of four distinct ncAAs into four distinct quadruplet codons, we transformed O1- strep GFP His6 Or O1- strep GFP(40CTAG, 50TAGA, 136AGGA, 150AGTA) His6We transformed 8 μL of chemically competent E. coli DH10B cells carrying a pRSF plasmid encoding a copy of O-riboQ1 with 0.4 μL of p15A plasmid encoding either riboQ1 or riboQ2. We recovered the transformed cells for 1 h at 37 °C and 750 rpm in 180 μL of SOC medium in a 96-well microtiter plate format. We used 30 μL of rescued cells to inoculate 500 μL of selective 2xYT-kts medium in a 1.2 mL 96-well plate format and grew the culture overnight at 37 °C and 750 rpm. 100 μL of the overnight culture was used to inoculate 4 mL of selective 2xYT-kts medium containing none or all of the four ncAAs: 4 mM NmH2, 2 mM CbzK3, 4 mM AllocK4, and 2 mM PheI5 in a 24-well plate format (O1- strep GFP His6 (All ncAAs were grown only in the presence of ncAAs). Cells were grown at 37°C and 220 rpm for 2 h. strep GFP His6 and O-riboQ1 production was induced by the addition of 8 μL of 1 M IPTG to give a final concentration of 2 mM IPTG. Cells were grown for 18 h at 37° C. and 750 rpm. 180 μL of each culture was transferred to a 96-well flat-bottom costar plate, and fluorescence and optical density were measured using a PHERAstar FS.

[0284] O1- strep GFP(40TAG, 50CTAG, 136AGGA, 150AGTA) His6 The same procedure was used for orthogonality and efficiency evaluation of the incorporation of four distinct ncAAs in response to one amber codon and three distinct quadruplet codons into tRNA4. However, RS4_1-2 / tRNA4 was used as an operon, and O1- strep GFP(40TAG, 50CTAG, 136AGGA, 150AGTA) His6was used as a reporter for quadruplet integration. 4 mM BocK1 was used instead of 4 mM AllocK4.

[0285] For MS analysis of the incorporation of three and four distinct ncAAs and determination of isolated yields strep GFP(XXXX) His6 Production of strep GFP(XXXX) His6 The same procedure for protein production as for mass spectrometry was used with the following combination of reporter, operon and ncAA: O1- with RS3 / tRNA3 and 4 mM BocK1, 4 mM NmH2, 2 mM CbzK3. strep GFP(40TAG, 136AGGA, 150AGTA) His6 , or RS4_1-2 / tRNA4 and O1- with 4 mM BocK 1, 4 mM NmH 2, 2 mM CbzK 3, and 2 mM PheI 5 strep GFP(40TAG, 50CTAG, 136AGGA, 150AGTA) His6 , or RS4_1-2 / tRNA4(quad) and O1- with 4 mM NmH2, 2 mM CbzK3, 4 mM AllocK4, and 2 mM PheI5. strep GFP(40CTAG, 50TAGA, 136AGGA, 150AGTA) His6 .

[0286] To determine the isolated yield, the fluorescence of 180 μL of isolated protein was measured using a PHERAstar FS. strep GFP His6 Protein concentrations were calculated based on a standard curve generated using standards. The buffer was exchanged into water using a 3 kDa Amicon ultra column for MS and MS / MS analysis.

[0287] Electrospray ionization mass spectrometry Denatured protein samples (~10 μM) were subjected to LC-MS analysis. Briefly, proteins were separated on a C4 BEH 1.7 μm, 1.0 × 100 mm UPLC column (Waters, UK) using a nanoAcquity (Waters, UK) modified to give a flow rate of ~50 μl / min. The column was developed with a gradient of acetonitrile (2% v / v → 80% v / v) in 0.1% v / v formic acid over 20 min. The outlet of the analytical column was directly connected to a hybrid quadrupole time-of-flight mass spectrometer (Xevo G2, Waters, UK) via an electrospray ionization source. Data were acquired over an m / z range of 300-2000 in positive ion mode with a cone voltage of 30 V. Scans were manually summed and deconvoluted using MaxEnt1 (Masslynx, Waters, UK). The theoretical molecular weight of the wild-type protein was first calculated using an online tool (http: / / web.expasy.org / protparam / ), and then the theoretical molecular weight of the protein with ncAA was calculated by manually performing a correction for the theoretical molecular weight of the ncAA.

[0288] Tandem MS / MS analysis Proteins were run on a 4-12% NuPAGE Bis-Tris gel (Invitrogen) with MES buffer and briefly stained with InstantBlue (Expedeon). Bands were excised and stored in water. Trypsin digestion and tandem MS / MS analysis were performed by Mark Skehel (Biological Mass Spectrometry and Proteomics Laboratory, MRC Laboratory of Molecular BIology).

[0289] References for the Methods Section 1. Salis, H. M., Mirsky, E. A. & Voigt, C. A. Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950, doi:10.1038 / nbt.1568 (2009). 2. Dunkelmann, D. L., Willis, J. C. W., Beattie, A. T. & Chin, J. W. Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nat. Chem. 12, 535-544, doi:10.1038 / s41557-020-0472-x (2020). 3. Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, J. W. Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441-444, doi:10.1038 / nature08817 (2010). 4. Wang, K. et al. Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393-403, doi:10.1038 / nchem.1919 (2014).

[0290]

Table 1

[0291] [Table 2]

[0292] Supplementary Tables 3 and 4 are found in the publication: Daniel L. Dunkelmann1, Sebastian B. Oehm, Adam T. Beattie, Jason W. Chin; "A 68-codon genetic code to incorporate four distinct non-canonical amino acids enabled by automated orthogonal mRNA discovery," incorporated by reference in its entirety.

Claims

1. 1. A method for designing a messenger RNA (mRNA), which is an orthogonal messenger RNA (O-mRNA) suitable for translation by an orthogonal ribosome (O-ribosome), wherein the mRNA comprises a 5′ untranslated region (5′UTR) and an open reading frame (ORF), the method comprising: (a) the free energy difference (ΔG tot predicting (O-ribo); (b) introducing an alteration into the 5'UTR; (c) New ΔG after modification tot (O-ribo) (ΔG tot new predicting (O-ribo); (d) the ΔG tot new (O-ribo) is the preceding ΔG tot (O-ribo) is more negative than (O-ribo), and the modification is allowed; Said ΔG tot new (O-ribo) is the preceding ΔG tot accepting or rejecting the modification according to a probability distribution if (O-ribo) is more positive than (O-ribo); and (e) generating an O-mRNA sequence comprising the 5'UTR containing the tolerated modification; The method comprising:

2. (i) ΔG tot (O-ribo) is the sum of the free energy required to unfold mRNA (ΔG unfolding ) and the free energy released when the mRNA binds to the O-ribosome to form an O-ribosome-bound initiation-competent state (ΔG o-ribo binding ); (ii) if ΔG tot new (O-ribo) is more positive than the preceding ΔG tot (O-ribo), the magnitude of the difference between said ΔG tot new (O-ribo) and said ΔG tot (O-ribo) determines the probability of acceptance, with smaller magnitudes being associated with a higher likelihood of acceptance compared to larger magnitudes; (iii) the probability distribution according to which modifications will be accepted or rejected, [Equation 1] where T SA is the simulated annealing temperature; (iv) the modification is or includes a single nucleotide change, insertion, or deletion; (v) steps (b) through (d) are repeated at least 200, 300, 400, 500, 1000, 5000, or 10,000 times; Steps (b) through (d) are repeated until at least 10, 50, 100, 250, or 500 successive iterations do not lead to a more negative ΔG tot new (O-ribo); (vi) the 5'UTR of step (a) is 35 nucleotides in length; or the modification is in any of the 35 nucleotides of the 5'UTR closest to the start codon; (vii) the 5′UTR of step (a) is a randomly generated nucleic acid sequence; (viii) the 5′ UTR of step (a) comprises a wild-type Shine-Dalgarno sequence; (ix) the second ribosome is a wild-type ribosome; or the second ribosome is an O-ribosome different from the first O-ribosome; and / or (x) is performed on a computer; The method of claim 1.

3. (i) The O-ribosome comprises an orthogonal 16S rRNA, the mRNA comprises a Shine-Dalgarno sequence, and ΔG tot (O-ribo) is as follows: ΔG tot (O-ribo) = (ΔG mRNA-O-rRNA + ΔG start + ΔG spacing - ΔG standby) + ΔG unfolding predicted according to ΔG mRNA-O-rRNA is the free energy of the predicted co-folded secondary structure of the last 9 nucleotides of the orthogonal 16S rRNA and the mRNA; ΔG start is the energy released from the binding of the initiator tRNA to the start codon of the ORF; ΔG spacing is the energy penalty for a non-optimal spacing length between the Shine-Dalgarno sequence and the start codon; ΔG standby is the energy required to unfold the secondary structure separating the four nucleotides upstream of the Shine-Dalgarno sequence; ΔG unfolding is the energy required to unfold the secondary structure in the mRNA; (ii) the T SA is adjusted to maintain an acceptance rate of 5-20%; and / or (iii) The Shine-Dalgarno sequence is 5 nucleotides from the start codon of the ORF. The method of claim 2.

4. A method according to claim 1 for designing an mRNA that is an O-mRNA suitable for translation by an O-ribosome in a cell that also contains a second ribosome (2nd-ribosome), comprising: step (a) comprises predicting the free energy difference (ΔG tot (2 nd -ribo)) between the freely folded state of said mRNA and the initiation-competent state of said mRNA bound to said 2 nd -ribosome; Step (c) comprises predicting the new modified ΔG tot (2 nd -ribo) (ΔG tot new (2 nd -ribo); Step (d) allows the modification if ΔG tot new (O-ribo) is more negative than the preceding ΔG tot (O-ribo) and if the ΔG tot new (2 nd -ribo) is more positive than the preceding ΔG tot (2 nd -ribo); accepting or rejecting the modification according to a probability distribution when the ΔG tot new (O-ribo) is more positive than the preceding ΔG tot (O-ribo) or when the ΔG tot new (2 nd -ribo) is more negative than the preceding ΔG tot (2 nd -ribo); The method.

5. The method according to claim 4, wherein ΔG tot (2nd -ribo) is the sum of the free energy required to unfold mRNA (ΔG unfolding ) and the free energy released when the mRNA binds to the 2nd -ribosome to form an initiation-competent state bound to the 2nd -ribosome (ΔG 2nd ribo binding ).

6. A method for producing a nucleic acid sequence comprising the steps of: a 2nd ribosome containing 16S rRNA, an mRNA containing a Shine-Dalgarno sequence, and ΔG tot (2nd-ribo) being: ΔG tot (2nd - ribo) = (ΔG mRNA - 2nd - rRNA + ΔG start + ΔG spacing - ΔG standby ) + ΔG unfolding predicted according to ΔG mRNA-2nd-rRNA is the free energy of the predicted cofolded secondary structure of the last 9 nucleotides of the 16S rRNA and the mRNA; ΔG start is the energy released from the binding of the initiator tRNA to the start codon of the ORF; ΔG spacing is the energy penalty for a non-optimal spacing length between the Shine-Dalgarno sequence and the start codon; ΔG standby is the energy required to unfold the secondary structure separating the four nucleotides upstream of the Shine-Dalgarno sequence; ΔG unfolding is the energy required to unfold the secondary structure in the mRNA; The method of claim 5.

7. The method of claim 6, wherein step (a) comprises reacting a compound of formula: ΔG tot (opt)=ΔG tot (O-ribo) -X*ΔG tot (2nd -ribo) calculating ΔG tot (opt) according to: Step (c) comprises reacting a compound of formula: ΔG tot new (opt)=ΔG tot new (O-ribo)-X*ΔG tot new (2nd -ribo) calculating ΔG tot new (opt) according to Step (d) allows modification if said ΔG tot new (opt) is more negative than said previous ΔG tot (opt); Accepting or rejecting the modification according to a probability distribution if the ΔG tot new (opt) is more positive than the previous ΔG tot (opt). Including; X is 0.1 to 2, or X is 0.5; The method of claim 4.

8. (i) if ΔG tot new (opt) is more positive than the preceding ΔG tot (opt), the magnitude of the difference between said ΔG tot new (opt) and said ΔG tot (opt) determines the probability of acceptance, with smaller magnitudes being associated with a higher likelihood of acceptance compared to larger magnitudes; (ii) the probability distribution according to which modifications will be accepted or rejected; [Equation 2] and T SA is the simulated annealing temperature; and / or (iii) steps (b) through (d) are repeated until at least 10, 50, 100, 250, or 500 successive iterations do not lead to a more negative ΔG tot new (opt); The method of claim 7.

9. The method of claim 8, wherein T SA is adjusted to maintain an acceptance rate of 5-20%.

10. The method of claim 1, wherein step (b) comprises introducing a modification into the 5'UTR or replacing any one of codons 2 to 20, 2 to 15, 2 to 12, 2 to 10, or 2 to 5 in the ORF with a synonymous codon; step (e) comprising generating an O-mRNA sequence comprising said 5'UTR and said ORF containing the tolerated modifications; The method of claim 1.

11. The method described in claim 10, wherein step (b) comprises introducing a modification into the 5'UTR comprising a single nucleotide change, insertion, or deletion, or replacing any one of codons 2 to 12 in the ORF with a synonymous codon.

12. The method described in claim 1, wherein the O-ribosome comprises an orthogonal anti-Shine-Dalgarno sequence and the 5'UTR of step (a) comprises an orthogonal Shine-Dalgarno sequence (O-SD) that is predicted to be fully complementary to the orthogonal anti-Shine-Dalgarno sequence.

13. The method of claim 12, wherein step (b) does not include introducing a modification into the five-nucleotide core of the O-SD.

14. A method for producing a nucleic acid sequence encoding an exogenous protein for translation by the O-ribosome, wherein the sequence of an O-mRNA is designed according to the method of claim 1, and then a nucleic acid molecule encoding said sequence is produced.

15. A nucleic acid sequence encoding an O-mRNA that encodes an exogenous protein, the O-mRNA being obtained or obtainable by the method of claim 14, the O-mRNA comprising at least two types of orthogonal codons; and Orthogonal ribosomes; A host cell comprising: