Compositions and methods relating to RNA polymerases engineered with capping enzymes
Patent Information
- Application Number
- JP2024563078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-23
- Filing Date
- 2023-04-25
- Publication Date
- 2026-03-18
AI Technical Summary
The prior art lacks 5' modifications when using T7 RNA polymerase for eukaryotic mRNA expression, limiting the efficiency of protein expression, especially in eukaryotic chassis.
An engineered enzyme with a connector was designed that combines T7 RNA polymerase and unity capping enzyme (NP868R) to achieve co-transcription and capping, improving the functionality of mRNA and protein expression efficiency.
Through this method, protein expression levels are significantly improved, enzyme polymerization and capping activity are enhanced, and are suitable for a variety of eukaryotic chassis, including yeast and human cells.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 334,406, filed April 25, 2022, and U.S. Provisional Application No. 63 / 409,353, filed September 23, 2022, both of which are incorporated by reference in their entireties herein. [Background technology]
[0002] The role of T7 RNA polymerase (RNAP)-based transcription has been central to recombinant protein expression in prokaryotic chassis. Beyond such systems, the simplicity of T7 RNAP-catalyzed transcription forms a cornerstone for the in vitro production of therapeutic RNAs and other biotechnological applications. However, the lack of 5' modifications in T7 RNAP-derived transcripts has limited its use in protein expression in eukaryotic chassis organisms as well as in the generation of functional eukaryotic mRNAs in vitro. Specifically, in the latter case, viral capping enzymes can be used to modify transcripts separately, suggesting that when both enzymes are used together (fused or separately), the cooperative activity of the two enzymes may result in the production of functional mRNAs in eukaryotes independent of the host transcription and capping machinery. The well-characterized and commonly used capping enzyme obtained from vaccinia virus consists of two subunits. Recently, it was reported that a single-subunit capping enzyme (NP868R) from African swine fever virus can catalyze all three reactions involved in the generation of capped RNA. Therefore, the use of this capping enzyme compared to vaccinia can greatly simplify its implementation for T7 RNAP-coupled mRNA / protein expression.
[0003] Previously, the ability of a wild-type version of the fusion enzyme (NP868R fused to T7 RNAP via a flexible glycine-serine linker) to generate capped transcripts was determined in mammalian cells (Jais 2019, Eaton 2017). In particular, the fusion enzyme was used specifically for the cytoplasmic expression of a target gene under the control of the T7 RNAP promoter, and the level of protein produced was used as a proxy for the efficiency of production of capped transcripts. Although the fusion of the capping enzyme increased the expressed protein compared to T7 RNAP alone, the efficiency of capping was lower than that observed for Pol II-derived transcripts (Jais 2019, Eaton 2017). Furthermore, this system has not yet been characterized in other eukaryotic chassis, and characterization of the enzyme for nuclear expression of a target gene in any chassis has not been reported.
[0004] What is needed in the art are engineered enzymes that combine RNA polymerase and capping capabilities, which engineered enzymes have significantly increased polymerase and / or capping activity. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Jais PH, Decroly E, Jacquet E, Le Boulch M, Jais A, Jean-Jean O, et al.C3P3-G1:first generation of a eukaryotic artificial cytoplasmic expression system.Nucleic Acids Research.2019;47(5):2681-98. [Non-Patent Document 2] Eaton HE, Kobayashi T, Dermody TS, Johnston RN, Jais PH, Shmulevitz M.African Swine Fever Virus NP868R Capping Enzyme Promotes Reovirus Rescue during Reverse Genetics by Promoting Reovirus Protein Expression, Virion Assembly, and RNA Incorporation into Infectious Virions.J Virol.2017;91(11). Summary of the Invention
[0006] Disclosed herein is an engineered enzyme comprising a T7 RNA polymerase component and a capping enzyme component separated by a linker, wherein the T7 RNA polymerase component comprises 90% or greater identity to SEQ ID NO:3, and the capping enzyme component comprises 90% or greater identity to SEQ ID NO:5.
[0007] Also disclosed herein is an engineered enzyme comprising SEQ ID NO:1 having at least one substitution that confers at least one improved property compared to SEQ ID NO:1 without the substitution, and further comprising a linker at positions 881-896 of the engineered enzyme, which may vary in length or amino acid composition.
[0008] Also disclosed herein is an engineered enzyme comprising an amino acid sequence having at least 90% identity to any one of SEQ ID NOs:6-24.
[0009] Further disclosed are nucleic acids encoding the engineered enzymes, expression vectors containing the nucleic acids, and host cells containing the expression vectors.
[0010] Disclosed are methods of producing and using the engineered enzymes disclosed herein.
[0011] A method for selecting one or more engineered enzymes comprising a non-eukaryotic polymerase component and a capping enzyme component is disclosed herein, wherein the engineered enzyme comprises enhanced activity compared to a control, and the method includes: (a) making a nucleic acid encoding one or more engineered enzyme variants, wherein the variant comprises a variant of a naturally occurring non-eukaryotic polymerase and a variant of a naturally occurring capping enzyme component; (b) incorporating the nucleic acid encoding one or more engineered enzyme variants into one or more eukaryotic cells, wherein the eukaryotic cells comprise a reporter, wherein the reporter is under the control of a polymerase promoter specific to the polymerase of the engineered enzyme, and further wherein the reporter is expressed only when the eukaryotic cell is capped by the capping enzyme; (c) expressing the nucleic acid encoding one or more engineered enzyme variants; and (d) determining which of the one or more variants confers enhanced activity compared to a control, and selecting the engineered enzyme variants. An example of such a method can be seen in Example 1.
[0012] Also disclosed herein is a system that utilizes the above-mentioned method for directed evolution.Therefore, described herein is a system for selecting one or more engineered enzymes that comprise non-eukaryotic polymerase components and capping enzyme components, wherein the engineered enzyme comprises enhanced activity, the system comprises a transformed eukaryotic cell, the eukaryotic cell comprises a reporter plasmid, the reporter plasmid is under the control of a polymerase promoter specific for the engineered enzyme polymerase, and the reporter is expressed only when capped by the capping enzyme.Eukaryotic cells can be designed to incorporate one or more variant nucleic acids.
[0013] Further disclosed is a method for selecting one or more engineered enzymes comprising a non-eukaryotic polymerase component and a capping enzyme component, comprising: a) providing a nucleic acid encoding the engineered enzyme, wherein expression of the engineered enzyme is under the control of a promoter, the promoter being recognized by the non-eukaryotic polymerase of the engineered enzyme; b) placing the nucleic acid encoding the engineered enzyme under conditions suitable for its expression; and c) detecting mRNA produced by the engineered enzyme and selecting the enzyme for further analysis.
[0014] Additional aspects and advantages of the present disclosure will be set forth in part in the following detailed description and any claims, and in part will be derived from the detailed description, or may be learned by practice of various aspects of the present disclosure. The advantages described below will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure.
[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain examples of the present disclosure and, together with the description, serve to explain the principles of the disclosure, without limitation. Like numbers refer to the same elements throughout the drawings. [Brief description of the drawings]
[0016] [Figure 1] We present the design of a selection scheme for the evolution of NP868R:T7 RNAP for cotranscriptional capping and subsequent protein expression in Saccharomyces cerevisiae. [Diagram 2] We present characterization of the evolved capping:T7 enzyme compared to the wild-type enzyme (245) and a catalytically dead enzyme (246). [Diagram 3]The 3D structure of T7 RNAP is shown. T7 is capable of orthogonal transcription, has heterologous function, programmable promoter strength and high levels of activity. However, it does not generate capped transcripts for eukaryotic expression. [Figure 4] Figure 1 is a bar graph showing bulk fluorescence measurements of capping-T7 polymerase variants expressing ZsGreen in yeast. The "disrupted" enzyme negative control consists of a K282A mutation that renders the capping domain inactive. "WT" indicates the capping-T7 fusion enzyme with the SV40 NLS but no additional mutations. V1, V2, V3 correspond to ES-230, ES-368, ES-443, respectively. ES-443 is 76-fold more fluorescent than the disrupted control. ES-230 (V1) is from round 17 of selection, whereas ES-368 (V2) and ES-443 (V3) are both from round 20. Fluorescence was measured on a Tecan M200 plate reader. [Diagram 5] Single cell fluorescence of yeast populations containing WT, V1 or V3 enzymes expressing the ZsGreen reporter is shown. Fluorescence is measured as fluorescence intensity (height) on a Sony SA3800 spectrum analyzer. [Figure 6A-B] All fusion variants of NPT7 were placed under the control of a galactose-responsive promoter followed by the tENO2 terminator and integrated into the HO locus of the Saccharomyces cerevisiae BY4741 genome. The target gene (ZsGreen) was placed under the control of the T7 promoter followed by the SV40 polyadenylation signal and the T7 terminator (A) and cloned into the 2 micron plasmid. Fold induction was calculated after addition of 5% galactose relative to no induction (B). [Figure 7A-B] For strains containing fusion enzyme variants-WT, 433 and 443 and the target plasmid (A), reporter expression was determined by adding different levels of galactose to obtain a dose response (B). [Figure 8A-C]To compare the activity of the variants, we show the same reporter gene (ZsGreen) under the control of the pGal promoter in two different contexts: first, it was integrated into the genome at the HO locus (IV target) (A), and then it was cloned into the same 2 micron plasmid (B). Galactose dose response for all constructs, and fold induction were calculated based on the levels of ZsGreen observed compared to the uninduced control (C). [Figure 9A-D] We show that the strength of T7-based expression can be controlled using mutant T7 promoters. Three different mutant T7 promoters were chosen that were predicted to provide an expression panel that controlled the expression of ZsGreen (Panel B, wild type (WT) is SEQ ID NO: 25, variant 2 (V2) is SEQ ID NO: 26, variant 3 (V3) is SEQ ID NO: 27, and variant 4 (V4) is SEQ ID NO: 28). All mutant promoters were cloned into the same plasmid backbone (C, D). [Figure 10A-C] It is shown that promoter specificity of T7 RNAP can be obtained by introducing specific mutations into the DNA binding region of the gene. Specifically disclosed herein is that the specificity of fusion proteins for panel orthogonal promoters can be similarly obtained by grafting mutations into v443 (A, B). Each variant showed highly specific activity for its own promoter, and minimal crosstalk between variants was observed (C). Panel C shows SEQ ID NO:25 (PT7), SEQ ID NO:29 (Portho1), SEQ ID NO:30 (Portho2), SEQ ID NO:31 (Portho3), SEQ ID NO:32 (Portho4) and SEQ ID NO:33 (Portho5). [Figure 11] Other reporter genes-BFP and mScarlet-I were cloned under the T7 promoter and transformed into strains containing v433 and v443. Expression of each gene was determined upon induction with galactose compared to the uninduced control. [Figure 12A-C]Shown are the levels of cargo-BFP and mScarlet-I (A) controlled using the set of mutant T7 promoters previously described. The relative order of strength of each promoter was conserved across the three reporter genes (B,C), thus conclusively demonstrating that control of gene expression is exclusively controlled by the interaction of the fusion protein with its promoter. [Figure 13A-D] Plasmids encoding all three reporter genes-ZsGreen, mScarlet-I and BFP were cloned (A). Each gene was placed under the control of a T7 promoter (B). Wild type (WT) is SEQ ID NO: 25. PT7 V2 (variant 2) is SEQ ID NO: 26. PT7 V3 (variant 3) is SEQ ID NO: 27. Versions of the same plasmid were constructed by placing ZsGreen under the control of mutant promoters (v2 and v3). These reporter plasmids were transformed into strains containing fusion proteins-v433 and v443. The relative expression of each gene relative to BFP upon induction of the fusion proteins with galactose was determined. As shown, the levels of Zsgreen can be predictably controlled while the expression of the other two genes remains consistent, demonstrating multiplexed control of expression using fusion proteins (C, D). [Figure 14A-B] A two-plasmid system for assaying the cytoplasmic activity of fusion enzymes in mammalian cells is shown. First, the NLS was removed from WT NPT7 and v443 and placed under the control of a strong constitutive promoter - CMV. The reporter plasmid consisted of ZsGreen under the control of the T7 promoter followed by a Kozak sequence. A synthetic sequence of a series of 120 A's was added followed by the T7 terminator (A). Upon transfection of both plasmids in HEK293T, the levels of ZsGreen were determined 48 h after transfection. Variant 443 showed approximately 1.8-fold higher expression compared to WT NPT7 (B). [Figure 15A-B]Expression and purification of the fusion protein are shown. First, an affinity tag (Twin Strep) was cloned into the T7 RNAP coding region of the fusion protein (A). After expression in BL21 cells and subsequent StrepTactin-based purification, a pure fraction of T7 RNAP was obtained and the yield was compared to commercial in vitro transcription mixtures (Thermo and Promega) (B). [Figure 16A-C] To assess the activity of purified T7 RNAP variants, including the WT, a reporter plasmid was linearized and used as a transcription template (A). In vitro transcription was performed using 200 ng of template and 250 ug of purified T7 RNAP. In vitro transcription reactions were performed at two different temperatures (37 and 30) and RNA yields were analyzed using an Agilent TapeStation 4200 (B). Considering the presence of a larger amount of enzyme, a higher yield was obtained with the Promega mix (C). [Figure 17A-C] For expression and purification of the full-length fusion protein, the affinity tag was cloned under the control of the strong inducible E. coli promoter T5-lac (A). After expression in BL21 cells and subsequent StrepTactin-based purification, the elution fractions were analyzed using SDS-PAGE gel electrophoresis (B). The band was excised from the gel and the full-length sequence was confirmed using mass spectrometry (C is the full-length sequence, SEQ ID NO: 34). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] definition Unless otherwise defined, all technical and scientific terms used herein generally have the same meaning as commonly understood by those skilled in the art to which this invention belongs. In general, the nomenclature used herein and the laboratory procedures of cell culture, molecular genetics, microbiology, biochemistry, organic chemistry, analytical chemistry and nucleic acid chemistry described below are well known and commonly used in the art. Such techniques are well known and described in many texts and references well known to those skilled in the art. Standard techniques or modifications thereof are used for chemical synthesis and chemical analysis. All patents, patent applications, papers and publications mentioned herein, both above and below, are expressly incorporated herein by reference.
[0018] Any suitable method and material similar or equivalent to those described herein can be used to carry out the present invention, and some methods and materials are described herein.It should be understood that the present invention is not limited to the specific methodology, protocols, and reagents described, which may vary according to the context in which they are used by those skilled in the art.Therefore, the terms defined immediately below are more fully explained by referring to this application as a whole.All patents, patent applications, papers and publications mentioned in this specification, both above and below, are expressly incorporated herein by reference.
[0019] Also, as used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0020] Numeric ranges are inclusive of the numbers that define the range. Accordingly, every numerical range disclosed herein is intended to include every narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were all expressly written herein. Every maximum (or minimum) numerical limitation disclosed herein is also intended to include every lower (or higher) numerical limitation, as if such lower (or higher) numerical limitations were all expressly written herein.
[0021] The term "about" refers to an acceptable error for a particular value. In some cases, "about" means within 0.05%, 0.5%, 1.0%, or 2.0% of the range of a given value. In some cases, "about" means within 1, 2, 3, or 4 standard deviations of a given value.
[0022] Moreover, the headings provided herein are not limitations of the various aspects or embodiments of the invention that one might have by reference to the application as a whole.
[0023] Accordingly, the terms defined immediately below are more fully defined by reference to the application as a whole. Nonetheless, to facilitate the understanding of the present invention, certain terms are defined below.
[0024] Unless otherwise indicated, nucleic acids are written left to right in 5' to 3' orientation, respectively; amino acid sequences are written left to right in amino to carboxy orientation, respectively.
[0025] As used herein, the term "comprising" and its cognates are used in an inclusive sense (i.e., equivalent to the term "including" and its corresponding cognates).
[0026] The "EC" numbers refer to the International Union of Biochemistry and Molecular Biology Commission's (NC-IUBMB) Enzyme Nomenclature. The IUBMB biochemical classification is a numerical classification system for enzymes based on the chemical reaction they catalyze.
[0027] "ATCC" refers to the American Type Culture Collection, whose biorepository collection includes genes and strains.
[0028] "NCBI" refers to the National Center for Biological Information and the sequence databases provided therein.
[0029] As used herein, "T7 RNA polymerase" refers to the DNA-directed RNA polymerase encoded by the T7 bacteriophage that catalyzes the formation of RNA in a 5' to 3' direction.
[0030] As used herein, the term "cap" refers to a guanine nucleoside linked through its 5" carbon to a triphosphate group, which is attached to the 5' carbon of the 5'-most nucleotide of an mRNA transcript. In some embodiments, the nitrogen at position 7 of the guanine in the cap is methylated.
[0031] As used herein, the terms "capped RNA," "5' capped RNA," and "capped mRNA" refer to RNA and mRNA, respectively, that contain a cap.
[0032] As used herein, "polynucleotide" and "nucleic acid" refer to two or more nucleosides covalently linked to each other. A polynucleotide may be composed entirely of ribonucleotides (i.e., NA), entirely of deoxyribonucleotides (i.e., DNA), or may be composed of a mixture of ribonucleotides and deoxyribonucleotides. Nucleosides are typically linked together via standard phosphodiester bonds, but a polynucleotide may contain one or more non-standard bonds. A polynucleotide may be single-stranded or double-stranded, or may contain both single-stranded and double-stranded regions. Additionally, a polynucleotide is typically composed of naturally occurring coding nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), but may contain one or more modified and / or synthetic nucleobases, such as, for example, inosine, xanthine, hypoxanthine, etc. In some embodiments, such modified or synthetic nucleobases are nucleobases that code for an amino acid sequence.
[0033] "Protein," "polypeptide," and "peptide" are used interchangeably herein to refer to a polymer of at least two amino acids covalently joined by amide bonds, regardless of length or post-translational modification (e.g., glycosylation or phosphorylation).
[0034] "Amino acids" are referred to herein by their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Similarly, nucleotides may be referred to by their commonly accepted one-letter codes. The abbreviations used for the genetically encoded amino acids are conventional and are as follows: alanine (Ala or A), arginine (Are or R), asparagine (Asn or N), aspartic acid (Asp or D), cysteine (Cys or C), glutamic acid (Glu or E), glutamine (Gin or Q), histidine (His or H), isoleucine (Leu or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Tip or W), tyrosine (Tyr or Y), and valine (Val or V). Θ054]When three-letter abbreviations are used, the amino acids may be in either the L- or D-configuration about the a-carbon (C<<), unless specifically preceded by "L" or "D" or otherwise clear from the context in which the abbreviation is used. For example, "Ala" indicates alanine without specifying the configuration about the α-carbon, whereas "D-Ala" and "L-A3a" indicate D-alanine and L-alanine, respectively. When one-letter abbreviations are used, capital letters indicate amino acids in the L-configuration about the a-carbon, and lower case letters indicate amino acids in the D-configuration about the a-carbon. For example, "A" indicates L-alanine and "a" indicates D-alanine. When a polypeptide sequence is presented as a series of one-letter or three-letter abbreviations (or mixtures thereof), the sequence is presented in the amino (N) to carboxy (C) direction according to common convention.
[0035] The abbreviations used for the genetically encoded nucleosides are conventional and are as follows: adenosine (A); guanosine (G); cytidine (C); thymidine (T); and uridine (U). Unless specifically indicated, the abbreviated nucleoside may be either a ribonucleoside or a deoxyribonucleoside. Nucleosides, either individually or as aggregates, may be specified as either a ribonucleoside or a deoxyribonucleoside. When a nucleic acid sequence is presented as a series of one-letter abbreviations, the sequence is presented in the 5' to 3' direction according to common convention, with no phosphate indicated.
[0036] The terms "engineered," "recombinant," "non-naturally occurring," and "variant," when used with respect to a cell, polynucleotide, or polypeptide, refer to material that does not occur in nature, or that is identical to but modified in a manner produced or derived from synthetic material and / or by manipulation using recombinant techniques, or material that corresponds to the natural or native form of the material.
[0037] As used herein, "wild type" and "naturally occurring" refer to the form found in nature. For example, a wild type polypeptide or polynucleotide sequence is a sequence that can be isolated from a source in nature and is present in an organism that has not been intentionally modified by human manipulation. In this disclosure, "wild type" also refers to a fusion of a wild type RNAP with a wild type capping enzyme from another organism. A fusion protein means that it is not further mutated but does not exist in nature because it is a fusion of enzymes from two different organisms. For example, a "wild type" fusion protein is found in SEQ ID NO:1.
[0038] "Coding sequence" refers to that portion of a nucleic acid (eg, a gene) that codes for the amino acid sequence of a protein.
[0039] The term "percent (%) sequence identity" is used herein to refer to a comparison between polynucleotides and polypeptides, and is determined by comparing two optimally aligned sequences over a comparison window, where the portion of the polynucleotide or polypeptide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence due to optimal alignment of the two sequences. The percentage may be calculated by determining the number of positions where the identical nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Alternatively, the percentage may be calculated by determining the number of positions where the identical nucleic acid base or amino acid residue occurs in both sequences, or by aligning the nucleic acid base or amino acid residue with gaps, obtaining the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Those skilled in the art will appreciate that there are many established algorithms available for aligning two sequences. Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith and Waterman (Smith and Waterman, Adv. Appl. Math., 2:482
[1981] ), by the homology alignment algorithm of Needleman and Wunsch (Needleman and Wunsch, J. Mol. Biol, 48:443
[1970] ), by the search for similarity method of Pearson and Lipman (Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444
[1988] ), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin software package), or by visual inspection as known in the art.Examples of algorithms suitable for determining percent sequence identity and percent sequence similarity include, but are not limited to, the BLAST and BLAST 2.0 algorithms described by Altschul et al. (See, respectively, Altschul et al., J. Mol. Biol,, 215:403-410
[1990] ; and Altschul et al., Nucleic Acids Res,, 25:3389-3402
[1977] ). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website. This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words W in the query sequence that match or meet some positive threshold score T when aligned with words of the same length in database sequences. T is referred to as the neighborhood word score threshold (see Altschul et al., supra). These initial neighborhood word hits serve as seeds for initiating searches to find longer HSPs that contain them. The word hits are then extended in both directions along each sequence as far as the cumulative alignment score can be increased. The cumulative score is calculated using the parameters M (reward score for a pair of matching residues; always greater than 0) and N (penalty score for mismatching residues; always less than 0) for nucleotide sequences. For amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of the word hits in each direction is stopped when the cumulative alignment score decreases by an amount X from its maximum performance value, when the cumulative score becomes zero or less due to the accumulation of one or more negative-scoring residue alignments, or when the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLAST program (for nucleotide sequences) uses by default a word length (W) of 11, an expectation (E) of 10, M=5, N=-4, and a comparison of both strands.For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915
[1989] . Exemplary determinations of sequence alignments and percent sequence identity can use the BESTFIT or GAP programs in the GCG Wisconsin Software package (Accelrys, Madison WI) using the default parameters provided.
[0040] "Reference sequence" refers to a defined sequence used as a basis for sequence comparison. A reference sequence can be a subset of a larger sequence, such as a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is a nucleic acid or polypeptide that is at least 20 nucleotides or amino acid residues long, at least 25 residues long, at least 50 residues long, at least 100 residues long, or full-length. Since two polynucleotides or polypeptides each may (1) contain a sequence that is similar between the two sequences (i.e., a portion of the complete sequence) and (2) further contain a sequence that differs between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides is typically performed by comparing the sequences of the two polynucleotides or polypeptides over a "comparison window" to identify and compare local regions of sequence similarity. In some embodiments, a "reference sequence" can be based on a primary amino acid sequence, and the reference sequence is a sequence that may have one or more changes in the primary sequence.
[0041] A "comparison window" refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acid residues, where a sequence can be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids, and the portion of the sequence within the comparison window can include no more than 20% additions or deletions (i.e., gaps) compared to the reference sequence (no additions or deletions) for optimal alignment of the two sequences. The comparison window can be longer than 20 contiguous residues, and can optionally include a window of 30, 40, 50, 100, or more.
[0042] "Corresponding," "referring to," or "with respect to," when used in the context of the numbering of a given amino acid or polynucleotide sequence, refers to the numbering of residues in a particular reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue numbers or residue positions of a given polymer are specified with respect to the reference sequence, not by the actual numerical position of the residues in the given amino acid or polynucleotide sequence. For example, a given amino acid sequence, such as the amino acid sequence of an engineered T7 RNA polymerase, can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, although gaps are present, the numbering of residues in a given amino acid or polynucleotide sequence is done with respect to the reference sequence to which it is aligned.
[0043] An "amino acid difference" or "residue difference" refers to the difference in an amino acid residue at a position of a polypeptide sequence relative to an amino acid residue at the corresponding position in a reference sequence. The position of an amino acid difference is generally referred to herein as "Xn", where n refers to the corresponding position in the reference sequence on which the residue difference is based. For example, a "residue difference at position K9 compared to SEQ ID NO:1" refers to the difference in the amino acid residue at the polypeptide position corresponding to position 9 of SEQ ID NO:1. Thus, if a reference polypeptide of SEQ ID NO:1 has a lysine at position 9, then a "residue difference at position K9 compared to SEQ ID NO:1" refers to an amino acid substitution of any residue other than lysine at the polypeptide position corresponding to position 9 of SEQ ID NO:1. In most cases herein, a specific amino acid residue difference at a position is designated as "XnY", where "Xn" identifies the corresponding position as above, and "Y" is the one-letter identifier of the amino acid found in the engineered polypeptide (i.e., the residue that differs from the reference polypeptide). In some cases (e.g., in the tables provided in the Examples herein), the disclosure also provides specific amino acid differences, indicated by the conventional designation "AnB," where A is a one-letter identifier of the residue in the reference sequence, "n" is the number of the residue position in the reference sequence, and B is a one-letter identifier of the residue substitution in the sequence of the engineered polypeptide. Referring again to the above example, a substitution of asparagine for lysine at position K would read "K9N." In some cases, the polypeptides of the disclosure may include one or more amino acid residue differences relative to the reference sequence, indicated by a list of the specific positions where the residue difference exists relative to the reference sequence. In some embodiments, the enzyme variants include more than one substitution. These substitutions are separated by slashes for ease of reading (e.g., R10K / R10I). The application includes engineered polypeptide sequences that include one or more amino acid differences, including either or both conservative and non-conservative amino acid substitutions.
[0044] "Conservative amino acid substitution" refers to the replacement of a residue with a different residue having a similar side chain, and thus typically includes the replacement of an amino acid in a polypeptide with an amino acid within the same or similar defined amino acid class.By way of example and not limitation, an amino acid having an aliphatic side chain may be replaced with another aliphatic amino acid (e.g., alanine, valine, leucine and isoleucine); an amino acid having a hydroxyl side chain is replaced with another amino acid having a hydroxyl side chain (e.g., serine and threonine); an amino acid having an aromatic side chain is replaced with another amino acid having an aromatic side chain (e.g., phenylalanine, tyrosine, tryptophan and histidine); an amino acid having a basic side chain is replaced with another amino acid having a basic side chain (e.g., lysine and arginine); an amino acid having an acidic side chain is replaced with another amino acid having an acidic side chain (e.g., aspartic acid or glutamic acid); and / or a hydrophobic or hydrophilic amino acid is replaced with another hydrophobic or hydrophilic amino acid, respectively.
[0045] "Non-conservative substitution" refers to the substitution of an amino acid in a polypeptide with an amino acid that has significantly different side chain properties. Non-conservative substitutions may use amino acids between defined groups rather than within defined groups, and affect (a) the structure of the peptide backbone in the area of substitution (e.g., proline for glycine), (b) the charge or hydrophobicity, or (c) the bulk of the side chain. By way of example and not limitation, exemplary non-conservative substitutions may be an acidic amino acid substituted with a basic or aliphatic amino acid; an aromatic amino acid substituted with a small amino acid; and a hydrophilic amino acid substituted with a hydrophobic amino acid.
[0046] "Deletion" refers to a modification to a polypeptide by removal of one or more amino acids from a reference polypeptide. A deletion may include removal of one or more amino acids, two or more amino acids, five or more amino acids, ten or more amino acids, fifteen or more amino acids, or twenty or more amino acids, up to 10% of the total number of amino acids, or up to 20% of the total number of amino acids that make up the reference enzyme, while retaining the enzyme activity and / or retaining the improved properties of the engineered enzyme. Deletions may be directed to internal and / or terminal portions of the polypeptide. In various embodiments, deletions may include continuous segments or may be discontinuous.
[0047] "Insertion" refers to a modification to a polypeptide by the addition of one or more amino acids from a reference polypeptide, and the insertion can be at the inside of the polypeptide or at the carboxy or amino terminus.As used herein, an insertion includes fusion proteins known in the art.Insertion can be a continuous segment of amino acids in a naturally occurring polypeptide, or can be separated by one or more of the amino acids.
[0048] "Isolated polypeptide" refers to a polypeptide that is substantially separated from other contaminants (e.g., proteins, lipids, and polynucleotides) that naturally accompany it. This term encompasses polypeptides that have been removed or purified from the naturally occurring environment or expression system (e.g., host cells or in vitro synthesis). Recombinant T7 RNA polymerase polypeptides may be present in cells, in cell culture media, or prepared in various forms, such as lysates or isolated preparations. Thus, in some embodiments, recombinant T7 RNA polymerase polypeptides may be isolated polypeptides.
[0049] "Substantially pure polypeptide" refers to a composition in which the polypeptide species is the predominant species present (i.e., more abundant than any other individual macromolecular species in the composition, on a molar or weight basis), and is generally a substantially purified composition when the species of interest constitutes at least about 50% of the macromolecular species present, on a molar or weight percent basis. Generally, a substantially pure T7 RNA polymerase composition comprises about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more of all macromolecular species present in the composition, on a molar or weight percent basis, and in some embodiments, the species of interest is purified to essential homogeneity (i.e., contaminating species cannot be detected in the composition by conventional detection methods), and the composition consists essentially of a single macromolecular species. Solvent species, small molecules (<500 Daltons), and elemental ion species are not considered macromolecular species, and in some embodiments, an isolated recombinant T7 RNA polymerase polypeptide is a substantially pure polypeptide composition.
[0050] "Improved enzymatic properties" of a T7 RNA polymerase and / or capping enzyme refers to an engineered T7 RNA polymerase polypeptide and / or capping enzyme that exhibits an improvement in any enzymatic property compared to a reference T7 RNA polymerase polypeptide and / or a wild-type T7 RNA polymerase polypeptide and / or another engineered T7 RNA polymerase polypeptide, or compared to a reference capping enzyme and / or a wild-type capping enzyme and / or another engineered capping enzyme. Improved properties include, but are not limited to, improved capping enzyme-specific properties, improved T7 RNA polymerase properties, and such properties with respect to improved properties resulting from both enzymes together. These include increased selectivity of cap analogs relative to GTP, increased replication fidelity, increased RNA yield, increased protein expression, increased thermal activity, increased pseudouridine incorporation or other RNA base analogs, increased thermostability, increased pH activity, increased stability, increased enzymatic activity, increased substrate specificity or affinity, increased specific activity, increased resistance to substrate or end product inhibition (including pyrophosphate), increased chemical stability, improved solvent stability, increased resistance to acidic or basic pH, increased resistance to proteolytic activity (i.e., reduced susceptibility to proteolysis), reduced aggregation, increased solubility, and altered temperature profile. Improved capping properties include, for example, increased RNA triphosphatase activity, guanylyltransferase activity, and methyltransferase activity. Increased chromatin remodeling or epigenetic modifications in eukaryotic cells are also included. Another improvement is fewer abortive transcripts using T7 variants compared to wild type. Improved mating properties may also include altered T7 kinetics for initiation and elongation that can enhance capping efficiency.
[0051] "Increased enzymatic activity" or "enhanced catalytic activity" refers to an improved property of an engineered T7 RNA polymerase, which can be expressed by an increase in specific activity (e.g., product produced / time / weight protein) or an increase in the percent conversion of substrate to product (e.g., the percent conversion of a starting amount of substrate to product in a specified period of time using a specified amount of variant T7 RNA polymerase compared to a reference T7 RNA polymerase). Exemplary methods for determining enzymatic activity are provided in the Examples. Any property related to enzymatic activity can be affected.
[0052] "Hybridization stringency" refers to hybridization conditions, such as washing conditions, in the hybridization of nucleic acids. Generally, hybridization reactions are performed under conditions of lower stringency, followed by washing of various but higher stringency. The term "moderately stringent hybridization" refers to conditions that allow a target DNA having more than about 90% identity to a target polynucleotide to bind to a complementary nucleic acid having about 60% identity to the target DNA, preferably about 75% identity, about 85% identity. Exemplary moderately stringent conditions are conditions equivalent to hybridization in 50% formamide, 5x Denhardt's solution, 5xSSPE, 0.2% SDS at 42°C, followed by washing in 0,2,xSSPE, 0.2% SDS at 42°C. "High stringency hybridization" generally refers to conditions that are about 10°C or less from the thermal melting temperature Tm determined under solution conditions for a defined polynucleotide sequence. In some embodiments, high stringency conditions refer to conditions that allow hybridization of only those nucleic acid sequences that form stable hybrids in 0.018M NaCl at 65° C. (i.e., if a hybrid is not stable in 0.018M NaCl at 65° C., it is not stable under high stringency conditions as contemplated herein). High stringency conditions can be provided, for example, by conditions equivalent to hybridization in 50% formamide, 5*Denhardt's solution, 5*SSPE, 0.2% SDS at 42° C., followed by washing in 0.1xSSPE and 0.1% SDS at 65° C. Another high stringency condition is hybridization in 5xSSC containing 0.1% (w:v) SDS at 65° C. and washing in 0.1xSSC containing 0.1% SDS at 65° C. Other highly stringent hybridization conditions, as well as moderately stringent conditions, are described in the references cited above.
[0053] "Optimized codons" refers to changes in the codons of a polynucleotide encoding a protein to those preferentially used in a particular organism so that the encoded protein is more efficiently expressed in the organism of interest. Although the genetic code is degenerate in that most amino acids are represented by a few codons, called "synonymous" or "synonymous" codons, it is well known that the frequency of codon usage by a particular organism is non-random and biased toward certain codon triplets. This bias in codon usage can be higher for a given gene, genes of common function or ancestral origin, highly expressed proteins versus low copy number proteins, and aggregated protein-coding regions of an organism's genome. In some embodiments, a polynucleotide encoding a T7 RNA polymerase enzyme can be codon-optimized for optimal production from the host organism selected for expression.
[0054] "Control sequences" as used herein refers to include all components necessary or advantageous for the expression of the polynucleotides and / or polypeptides of the present application. Each control sequence may be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control sequences include, but are not limited to, a leader, polyadenylation sequence, propeptide sequence, promoter sequence, signal peptide sequence, initiation sequence, and transcription terminator. At a minimum, control sequences include a promoter, and transcriptional and translational stop signals. The control sequences may be provided with linkers for the purpose of introducing specific restriction sites that facilitate ligation of the control sequences with the coding region of the nucleic acid sequence encoding the polypeptide.
[0055] "Operably linked" is defined herein as a configuration in which a control sequence is appropriately positioned (i.e., in a functional relationship) with a polynucleotide of interest such that the control sequence directs or regulates expression of the polynucleotide and / or polypeptide of interest.
[0056] "Promoter sequence" refers to a nucleic acid sequence recognized by a host cell for expression of a polynucleotide of interest, such as a coding sequence. The promoter sequence contains a transcriptional control sequence that mediates the expression of the polynucleotide of interest. The promoter may be any nucleic acid sequence that exhibits transcriptional activity in a selected host cell, including mutant, truncated and hybrid promoters, and may be obtained from a gene encoding an extracellular or intracellular polypeptide that is homologous or heterologous to the host cell.
[0057] "Suitable reaction conditions" refers to conditions in an enzyme conversion reaction solution (e.g., ranges of enzyme loading, substrate loading, temperature, pH, buffers, co-solvents, etc.) that allow the 17 RNA polymerase polypeptide of the present application to convert a substrate into a desired product compound.
[0058] "Substrate," in the context of an enzymatic conversion reaction process, refers to a compound or molecule that is acted upon by a T7 RNA polymerase polypeptide.
[0059] "Product," in the context of an enzymatic conversion process, refers to a compound or molecule that results from the action of a T7 RNA polymerase polypeptide on a substrate.
[0060] As used herein, the term "culture" refers to the growth of a population of microbial cells under any suitable conditions (eg, using a liquid, gel, or solid medium).
[0061] Recombinant polypeptides can be produced using any suitable method known in the art. A gene encoding a wild-type polypeptide of interest can be cloned into a vector, such as a plasmid, and expressed in a desired host, such as E. coli, S. cerevisiae, etc. Variants of recombinant polypeptides can be generated by a variety of methods known in the art. Indeed, there are a wide variety of different mutagenesis techniques well known to those skilled in the art. In addition, mutagenesis kits are also available from many commercial molecular biology suppliers. Methods are available for making specific substitutions at defined amino acids (site-directed), specific or random mutations in localized regions of a gene (region-specific), or random mutagenesis throughout a gene (e.g., saturation mutagenesis). Numerous suitable methods for generating enzyme variants are known to those skilled in the art, including, but not limited to, site-directed mutagenesis of single-stranded or double-stranded DNA using PCR, cassette mutagenesis, gene synthesis, error-prone PCR, shuffling, and chemical saturation mutagenesis, or any other suitable method known in the art. Non-limiting examples of methods used in DNA and protein manipulation are provided in the following patents: U.S. Patent No. 6,117,679; U.S. Patent No. 6,420,175; U.S. Patent No. 6,376,246; U.S. Patent No. 6,586,182; U.S. Patent No. 7,747,391; U.S. Patent No. 7,747,393; U.S. Patent No. 7,783,428; and U.S. Patent No. 8,383,346. After variants are produced, they can be screened for any desired property (e.g., high or increased activity, or low or reduced activity, increased thermal activity, increased thermal stability, and / or acidic pH stability, etc.).
[0062] In some embodiments, "recombinant T7 RNA polymerase polypeptides" (also referred to herein as "engineered T7 RNA polymerase polypeptides," "variant T7 RNA polymerase enzymes," and "T7 RNA polymerase variants") are used.
[0063] As used herein, a "vector" is a DNA construct for introducing a DNA sequence into a cell. In some embodiments, the vector is an expression vector that is operably linked to a suitable control sequence that can affect the expression of the polypeptide encoded by the DNA sequence in a suitable host. In some embodiments, an "expression vector" has a promoter sequence operably linked to a DNA sequence (e.g., a transgene) to drive expression in a host cell, and in some embodiments, also includes a transcription terminator sequence.
[0064] As used herein, the term "expression" includes any step involved in the production of a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of the polypeptide from the cell.
[0065] As used herein, the term "produce" refers to the production of a protein and / or other compound by a cell. The term is intended to encompass any step involved in the production of a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses the secretion of a polypeptide from a cell.
[0066] As used herein, an amino acid or nucleotide sequence (e.g., a promoter sequence, signal peptide, terminator sequence, etc.) is "heterologous" to another sequence to which it is operably linked if the two sequences are not related in nature.
[0067] As used herein, the terms "host cell" and "host strain" refer to a suitable host for an expression vector comprising a DNA provided herein (e.g., a polynucleotide encoding a T7 RNA polymerase variant). In some embodiments, a host cell is a prokaryotic or eukaryotic cell that is transformed or transfected with a vector constructed using recombinant DNA techniques known in the art.
[0068] The term "analog", when used in reference to a polypeptide, refers to a polypeptide having greater than 70% sequence identity but less than 100% sequence identity (e.g., greater than 75%, 78%, 80%, 83%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity) to a reference polypeptide. In some embodiments, an analog refers to a polypeptide that contains one or more non-naturally occurring amino acid residues, including but not limited to homoarginine, ornithine, and norvaline, as well as naturally occurring amino acids. In some embodiments, an analog also includes one or more D-amino acid residues and non-peptide bonds between two or more amino acid residues.
[0069] The term "effective amount" means an amount sufficient to effect a desired result. One of ordinary skill in the art can determine the effective amount using routine experimentation.
[0070] The terms "isolated" and "purified" are used to refer to a molecule (e.g., an isolated nucleic acid, polypeptide, etc.) or other component that has been removed from at least one other component with which it is naturally associated. The term "purified" does not require absolute purity, but rather is intended as a relative definition.
[0071] As used herein, "composition" and "formulation" encompass products comprising at least one engineered T7 RNA polymerase of the invention intended for any suitable use (e.g., research, diagnostics, etc.).
[0072] The term "transcription" is used to refer to the process by which a portion of a DNA template is copied into RNA by the action of an RNA polymerase enzyme.
[0073] The term "DNA template" is used to refer to a double- or single-stranded DNA molecule that contains a promoter sequence and a sequence that encodes the RNA product of transcription.
[0074] The term "promoter" is used to refer to a DNA sequence that is recognized by RNA polymerase as the start site of transcription. The promoter recruits RNA polymerase and, in the case of T7 RNA polymerase, determines the start site of transcription.
[0075] The term "RNA polymerase" is used to refer to a DNA-directed RNA polymerase that copies a DNA template into an RNA polynucleotide by stepwise incorporation of nucleotide triphosphates into the growing RNA polymer.
[0076] The terms "messenger RNA" and "mRNA" are used to refer to RNA molecules that code for proteins, which are decoded by the act of translation.
[0077] The terms "7-methylguanosine cap", "7meG", "5-prime cap" and "5'' cap" are used in reference to a specific modified nucleotide structure present at the 5' end of eukaryotic mRNA. The 7-methylguanosine cap structure is attached to the first nucleotide in the mRNA via a 5' to 5' triphosphate linkage. In vivo, this cap structure is added to the 5' end of the nascent mRNA by the sequential activity of multiple enzymes. In vitro, the cap can be incorporated directly at the initiation of transcription by RNA polymerase through the use of cap analogs.
[0078] The term "cap analog" refers to a dinucleotide that contains a 5'-5' di-, tri-, or tetra-phosphate linkage. One end of the dinucleotide terminates in either a guanosine or substituted guanosine residue; it is at this end that RNA polymerase initiates transcription by extending from the 3' hydroxyl. The other end of the dinucleotide is a guanosine that mimics the eukaryotic cap structure, typically with a 7-methyl-, 7-benzyl-, or 7-ethyl-substitution and / or a 7-aminomethyl or 7-aminoethyl substitution. In some cases, the nucleotide is also substituted with a 3' hydroxyl group to prevent initiation of transcription from the capped end of the molecule.
[0079] The terms "ARCA" and "anti-reverse cap analog" refer to chemically modified forms of cap analogs designed to maximize the efficiency of in vitro translation by ensuring that the cap analog is properly incorporated into the transcript in the correct orientation. These analogs are used to enhance translation. In some embodiments, ARCAs known in the art are used (e.g., Peng et al., Org. Lett., 4:161-164
[2002] ).
[0080] As used herein, the term "endogenous DNA-dependent RNA polymerase" refers to the endogenous DNA-dependent RNA polymerase of the host cell. When the host cell is a eukaryotic cell, the endogenous DNA-dependent RNA polymerase is RNA polymerase II.
[0081] As used herein, the term "endogenous capping enzyme" refers to a capping enzyme endogenous to the host cell.
[0082] As used herein, the term "inhibits expression of a protein" refers to a decrease of at least 20%, particularly at least 35%, at least 50%, more particularly at least 65%, at least 80%, at least 90% of the expression of said protein. Inhibition of protein expression can be determined by techniques well known to those skilled in the art, including but not limited to Northern blot, Western blot, RT-PCR.
[0083] The term "riboswitch" is used to refer to an autocatalytic RNA enzyme that cleaves itself or another RNA in the presence of a ligand.
[0084] The term "fidelity" is used to refer to the accuracy of an RNA polymerase in transcribing or replicating a DNA template into an RNA polynucleotide. Inaccurate transcription can result in single nucleotide polymorphisms (SNPs) or indels.
[0085] The term "single nucleotide polymorphism" or "SNP" refers to a nucleotide change occurring at a single position in a polynucleotide. In the context of transcription, a SNP can result from the misincorporation of a non-complementary ribonucleotide (A, C, G, or U) by an RNA polymerase at a position on a DNA template.
[0086] The term "indel" is used to refer to the insertion or deletion of one or more polynucleotides. In the context of transcription by RNA polymerase, an indel error can result from the addition of one or more extra ribonucleotides or the failure to incorporate one or more nucleotides at a position on a DNA template.
[0087] The term "selectivity" is used to refer to the quality of an enzyme having higher activity for one substrate compared to another during a catalysis reaction. In the context of co-transcriptional capping, an RNA polymerase may have higher or lower selectivity for a cap analog over GTP.
[0088] The term "inorganic pyrophosphatase" is used to refer to the enzyme that breaks down inorganic pyrophosphate to orthophosphate.
[0089] General Description Prior to the present invention, state-of-the-art techniques for orthogonal protein expression in eukaryotes relied on the use of synthetic transcription factors in combination with characterized endogenous or viral promoters. However, the entire process still relied on host RNA polymerase II and its regulatory-associated limitations. The use of T7 RNAP for transcription and a viral capping enzyme (NP868R) for expression of the target gene completely decouples the process from Pol II, thus providing an orthogonal mode of control, potentially allowing overexpression of the target gene. The use of a fully orthogonal RNA polymerase for expression of functional mRNA greatly expands expression and control capabilities and expands the repertoire of synthetic circuit elements. This engineered enzyme has been evolved to function optimally in a yeast background as well as in other eukaryotic hosts.
[0090] This orthogonal mode of gene expression is highly desirable for protein overexpression because it proceeds without the inhibitory feedback characteristic of cellular stress. Orthogonal gene expression is also used in cellular circuits used outside the laboratory, where environmental stresses can affect gene expression.
[0091] Perhaps more importantly, beyond in vivo applications, this fusion enzyme also streamlines the process of generating capped mRNA in vitro in a single reaction. Typically, this workflow proceeds with two distinct steps: T7-driven RNA transcription followed by capping with a viral capping enzyme (typically Vaccinia Virus Capping Enzyme). The present invention allows fewer reagents to be required and simplifies the reaction to a single step. Furthermore, mutations observed only in the T7 RNAP domain of the evolved variants may result in higher processivity and less sterile products compared to the wild-type enzyme. In the case of NP868R, the evolved version may result in improved capping efficiency compared to the industrial workhorse enzyme - Vaccinia Capping Enzyme.
[0092] Furthermore, the present invention illustrates a method for engineering capping enzyme for mRNA vaccine production.Current bottleneck in mRNA vaccine is the production and activity of capping enzyme.The method described herein greatly aids in the generation of capping enzyme variants for scale-up production of mRNA for vaccines and therapeutics.
[0093] Disclosed herein is an engineered enzyme comprising T7 RNA polymerase (SEQ ID NO:3) and a single subunit capping enzyme from African swine fever virus (NP8968R) (SEQ ID NO:5). These two components may be linked by a linker. Such linkers are known to those skilled in the art and may be 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 amino acid residues long (longer or shorter as can be determined by one skilled in the art). The linker may vary in content and length. An example of a linker can be found in SEQ ID NO:4. The engineered enzyme may also comprise a signal peptide, such as a nuclear localization signal (NLS). An example of a nuclear localization sequence (NLS) can be found in SEQ ID NO:2, as well as positions 1-13 of SEQ ID NO:1. Mutation of the NLS can result in increased activity of the engineered enzyme.
[0094] Variants of any of SEQ ID NOs: 2-6 are also disclosed, which include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more amino acid variations, including deletions, insertions, or substitutions. These variations can confer improved properties, which are described in detail below. In other words, variants of any of SEQ ID NOs: 2-6 that have 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% identity with SEQ ID NOs: 2-6, respectively, are disclosed herein.
[0095] When the above components are combined, the result is SEQ ID NO:1, which includes the NLS, T7 RNAP, linker and capping enzyme. Variants of SEQ ID NO:1 can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 or more amino acid composition differences compared to SEQ ID NO:1. Some of these mutations are represented by SEQ ID NOs:6-24. Positions 881-896 of SEQ ID NO:1 can include a linker. As mentioned above, the linker can be changed and the enzyme can still retain its function. When referring to percent identity to SEQ ID NO:1, the linker may or may not be taken into account. For example, a variant of SEQ ID NO: 1 can have 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% relative to SEQ ID NO: 1. This percentage may or may not include a linker.
[0096] Mutations to the T7 RNAP, capping enzyme, and / or NLS have been found to confer surprising and unexpected benefits to the activity of the engineered enzyme. For example, the enzyme can have higher protein expression compared to the wild-type enzyme. In another example, the improved properties can be selected from improved selectivity for capping, improved processability of capping, improved protein expression, improved RNA yield, improved stability in storage buffer, improved stability under reaction conditions, improved processability of translation, improved thermostability, and improved transcription fidelity. Improved properties can also include improved capping enzyme activity, such as improved activity of RNA triphosphatase guanyltransferase and / or methyltransferase.
[0097] "Improved" means that the above properties are improved by 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 102%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 109%, 1 This means 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, or 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 times or more improvement.
[0098] As mentioned above, the sequence may contain mutations, including substitutions, deletions, or insertions. Substitutions include, but are not limited to, those found in Tables 1 and 2. One or more of these substitutions may occur in the same engineered enzyme. Examples of engineered enzymes containing these substitutions can be found in SEQ ID NOs: 6-24.
[0099] Further disclosed herein is a capping enzyme (NP868R) derived from African swine fever virus, comprising one or more mutations that confer improved properties to the enzyme. These improved properties are disclosed elsewhere herein. The capping enzyme may comprise 90, 91, 92, 93, 94, 95, 96, 79, 98, or 99% or more identity to SEQ ID NO:5. In other words, the capping enzyme may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more amino acid deletions, insertions, or substitutions that confer improved properties.
[0100] Also disclosed are T7 RNA polymerases that contain one or more mutations that confer improved properties of the enzyme. These improved properties are disclosed elsewhere herein. The T7 RNA polymerase can contain 90, 91, 92, 93, 94, 95, 96, 79, 98, or 99% or more identity to SEQ ID NO:3. In other words, the T7 RNA polymerase can contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more amino acid deletions, insertions, or substitutions that confer improved properties.
[0101] Nucleic acids encoding the engineered enzymes of the invention are also disclosed herein. The group of isolated nucleic acid molecules encoding the engineered enzymes of the invention can include all of the nucleic acid molecules necessary and sufficient to obtain the engineered enzymes of the invention by their expression. The nucleic acid encoding the engineered enzyme can be operably linked to a control sequence.
[0102] In particular, the nucleic acid molecule according to the present invention can be operably linked to a promoter. Linking the nucleic acid to the promoter of eukaryotic DNA-dependent RNA polymerase, preferably RNA polymerase II, has the significant advantage that when the chimeric enzyme of the present invention is expressed in eukaryotic host cells, the expression of the chimeric enzyme is driven by eukaryotic RNA polymerase, preferably RNA polymerase II. These chimeric enzymes can then initiate the transcription of transgenes. When a tissue-specific RNA polymerase II promoter is used, the chimeric enzyme of the present invention can be selectively expressed in target tissues / cells. The promoter can be a constitutive promoter or an inducible promoter well known to those skilled in the art. The promoter can be developmentally regulated, inducible, or tissue-specific.
[0103] The present invention also relates to a vector comprising the nucleic acid molecule according to the present invention. The vector can be used for semi-stable or stable expression. The present invention also relates to a group of vectors comprising the group of isolated nucleic acid molecules according to the present invention. In particular, the vector according to the present invention is a cloning or expression vector.
[0104] The present invention also relates to a host cell comprising a nucleic acid molecule according to the invention or a vector according to the invention or a group of vectors according to the invention. The host cell according to the invention may be useful for large-scale protein production.
[0105] The present invention also relates to genetically engineered eukaryotic organisms expressing engineered enzymes encoded by the nucleic acid molecules or groups of isolated nucleic acid molecules according to the present invention, in particular engineered enzymes according to the present invention. The eukaryotic organism may be any unicellular eukaryotic organism, such as yeast. Organisms such as mammals, or any other animal or plant organisms are also contemplated. Examples of yeast include, but are not limited to, Saccharomyces cerevisiae and Pichia pastoris. Examples of mammalian cells that can be used with the present invention include, but are not limited to, HEK 293, Jurkat, CHO, COS, and primary human cells, including immune cells and stem cells. The present invention can be used in vivo or in vitro.
[0106] The present invention also relates to the use of an engineered enzyme according to the invention for the production of an RNA molecule with a 5'-end cap, in particular an RNA molecule that can be synthesized by a bacteriophage DNA-dependent RNA polymerase, such as T7 RNAP.
[0107] The invention also relates to the use of an engineered enzyme according to the invention, or an isolated nucleic acid molecule or a group of isolated nucleic acid molecules according to the invention, for the production of proteins, especially proteins of therapeutic interest such as vaccines or antibodies, especially in eukaryotic systems such as in vitro synthetic protein assays or cultured cells, which can be used in conjunction with purified protein therapeutics or as cell factories in which proteins are continuously produced within or outside a host organism.
[0108] The present invention also relates to a method for producing an RNA molecule with a 5'-end cap, comprising expressing in a host cell a nucleic acid molecule or a group of isolated nucleic acid molecules according to the present invention, said DNA sequence being covalently linked to at least one sequence encoding an RNA element of said protein-RNA tethering system that specifically binds to said RNA binding domain. As used herein, the term "RNA element of a protein-RNA tethering system that specifically binds to said RNA binding domain" generally relates to an RNA sequence that forms a stem-loop that can bind with high affinity to the corresponding RNA binding domain of the protein-RNA tethering system.
[0109] In particular, the DNA sequence is operably linked to a promoter for a bacteriophage DNA-dependent RNA polymerase or a promoter for the DNA-dependent RNA polymerase of the chimera of the present invention. In particular, when the RNA binding domain of the protein-RNA tethering system is the RNA binding domain of the lambdoid N antitermination protein-RNA tethering system, the element that specifically binds to the RNA binding domain can be a boxBL and / or boxBR stem-loop RNA structure (Das 1993, Greenblatt, Nodwell et al. 1993, Friedman and Court 1995).
[0110] In particular, the DNA sequence is operably linked to a promoter for a bacteriophage DNA-dependent RNA polymerase or a promoter for the DNA-dependent RNA polymerase of a chimera of the invention and is covalently linked at its 3' end to at least one, preferably at least two, at least three, more preferably at least four sequences encoding elements that specifically bind to the RNA-binding domain.
[0111] In particular, the method according to the present invention further comprises contacting the DNA sequence encoding the RNA molecule with the enzyme of the present invention.For example, the DNA sequence can be operably linked to the promoter for bacteriophage DNA-dependent RNA polymerase or the promoter for the DNA-dependent RNA polymerase of the chimera of the present invention, and at its 3' end can be covalently linked to at least one sequence encoding an element that specifically binds to the RNA binding domain covalently linked to a poly(A) track sequence consisting of at least 10, particularly at least 20, 30, more particularly at least 40 deoxyadenosine residues.Also, a PolyA signal sequence can be used to recruit polyadenylation enzyme.
[0112] In particular, the poly(A) track sequence can be covalently linked at its 3' end to a self-cleaving RNA sequence and optionally to a transcription termination sequence.The self-cleaving RNA sequence can be from the group comprising the genome false-negative ribozyme of hepatitis D virus (Genbank accession number AJ000558.1), the antigenomic false-negative ribozyme of hepatitis D virus (Genbank accession number AJ000558.1), the tobacco ringspot virus satellite hairpin ribozyme (Genbank accession number NC_003889.1) or artificial short hairpin RNA (shRNA).
[0113] In particular, the method according to the invention may further comprise the step of introducing said DNA sequence and / or nucleic acid according to the invention into a host cell using methods well known to those skilled in the art, such as by transfection using calcium phosphate, electroporation or by mixing cationic lipids with DNA to produce liposomes.
[0114] In one embodiment, the method according to the invention further comprises the step of inhibiting, in particular silencing, the cellular transcriptional and post-transcriptional machinery of said host cell, preferably by means of siRNA (small interfering RNA), miRNA (microRNA) or shRNA.
[0115] In one embodiment, the method according to the invention further comprises the step of inhibiting expression of an endogenous DNA-dependent RNA polymerase and / or an endogenous capping enzyme in the host cell.
[0116] The step of inhibiting the expression of endogenous DNA-dependent RNA polymerase and / or endogenous capping enzyme in the host cell can be carried out by any technique known to those skilled in the art, including, but not limited to, siRNA technique targeting the endogenous DNA-dependent RNA polymerase and / or endogenous capping enzyme, antisense RNA technique targeting the endogenous DNA-dependent RNA polymerase and / or endogenous capping enzyme, shRNA technique targeting the endogenous DNA-dependent RNA polymerase and / or endogenous capping enzyme.
[0117] In addition to siRNA (or shRNA), other inhibitory sequences, including DNA or RNA antisense (Liu and Carmichael 1994, Dias and Stein 2002), hammerhead ribozyme (Salehi-Ashtiani and Szostak 2001), hairpin ribozyme (Lian, De Young et al. 1999) or chimeric snRNA U1 antisense targeting sequence (Fortes, Cuevas et al. 2003), can be considered for the same purpose. In addition, other cellular target genes can be considered for inhibition, including other genes involved in cellular transcription (e.g., RNA polymerase II or other subunits of transcription factors), post-transcriptional processing (e.g., other subunits of capping enzyme, as well as polyadenylation or spliceosome factors), and mRNA nuclear export pathways.
[0118] In one embodiment of the method according to the invention, the RNA molecule is capable of encoding a therapeutic polypeptide.
[0119] In another embodiment, said RNA molecule can be a non-coding RNA molecule selected from the group comprising siRNA, ribozyme, shRNA and antisense RNA.In particular, said DNA sequence can code the RNA molecule selected from the group consisting of mRNA, non-coding RNA, particularly siRNA, ribozyme, shRNA and antisense RNA.
[0120] The present invention also relates to the use of the engineered enzymes according to the invention as capping enzymes, preferably pol(A) polymerases and DNA-dependent RNA polymerases.
[0121] The present invention also relates to a kit for the production of RNA molecules having a 5'-terminal cap, in particular a 5'-terminal m7GpppN cap, comprising at least one engineered enzyme according to the invention as defined above, and / or an isolated nucleic acid molecule and / or a group of nucleic acid molecules according to the invention as defined above, and / or a vector according to the invention as defined above, or a protein comprising an engineered enzyme as disclosed herein.
[0122] Advantageously, the kit or composition of the present invention can be used as an orthogonal gene expression system. As used herein, the term "orthogonal" refers to biological systems whose basic structures are independent and generally originate from different species.
[0123] The present invention also relates to an engineered enzyme according to the invention, an isolated nucleic acid molecule according to the invention, a group of nucleic acid molecules according to the invention or a vector according to the invention for use in the prevention and / or treatment of a human or animal pathology, preferably by gene therapy.
[0124] The present invention also relates to a pharmaceutical composition comprising a chimeric enzyme according to the invention, and / or an isolated nucleic acid molecule according to the invention, and / or a group of nucleic acid molecules according to the invention, and / or a vector according to the invention. Preferably, said pharmaceutical composition according to the invention is formulated in a pharma- ceutically acceptable carrier.
[0125] Pharmaceutically acceptable carriers are well known to those skilled in the art.
[0126] The pharmaceutical composition according to the invention may further comprise at least one DNA sequence of interest, which is operably linked to a promoter for said catalytic domain of a DNA-dependent RNA polymerase and covalently linked to at least one sequence encoding an element that specifically binds to said RNA-binding domain.
[0127] Such components (in particular selected in the group consisting of the chimeric enzyme according to the invention, the isolated nucleic acid molecule according to the invention, the vector according to the invention and at least one DNA sequence of interest) can be present in the pharmaceutical composition or medicament according to the invention in therapeutic amounts (active and non-toxic amounts).
[0128] Such therapeutic amount can be determined by one skilled in the art by routine trials, including evaluation of the effect of administration of the components on the pathology and / or disorder sought to be prevented and / or treated by administration of the pharmaceutical composition or medicament according to the invention.
[0129] For example, such testing can be carried out by analyzing both the quantitative and qualitative effects of administration of different amounts of the above components (in particular selected in the group consisting of the chimeric enzyme according to the invention, the isolated nucleic acid molecule according to the invention, the vector according to the invention and at least one DNA sequence of interest), in particular from a biological sample of a subject, on a set of marker (biological and / or clinical) characteristics of the pathology and / or disorder.
[0130] The present invention also relates to a method of treatment comprising administering to a subject in need thereof an engineered enzyme according to the invention, and / or an isolated nucleic acid molecule according to the invention, and / or a group of nucleic acid molecules according to the invention and / or a vector according to the invention in a therapeutic amount. The method of treatment according to the present invention may further comprise administering to a subject in need thereof a therapeutic amount of at least one DNA sequence of interest, said DNA sequence being operably linked to a promoter for said catalytic domain of a DNA-dependent RNA polymerase and covalently linked to at least one sequence encoding an element that specifically binds to said RNA-binding domain.
[0131] Said engineered enzymes, nucleic acid molecules and / or said vectors according to the invention may be administered simultaneously with said DNA sequence of interest, separately or sequentially, in particular prior to said DNA sequence of interest.
[0132] The present invention also relates to a pharmaceutical composition according to the invention for its use for the prevention and / or treatment of human or animal pathologies, in particular by gene therapy.
[0133] The pathology may be selected from the group consisting of pathologies that can be ameliorated by administration of the at least one DNA sequence of interest.
[0134] The present invention also relates to the use of an engineered enzyme according to the invention and / or an isolated nucleic acid molecule according to the invention and / or a group of nucleic acid molecules according to the invention and / or a vector according to the invention for the preparation of a medicament for the prevention and / or treatment of a human or animal pathology, in particular by gene therapy.
[0135] The invention may further comprise at least one vector comprising and / or expressing an engineered enzyme according to the invention and / or at least one nucleic acid molecule according to the invention and / or a group of nucleic acid molecules according to the invention and / or a nucleic acid molecule according to the invention; and at least one DNA sequence of interest, said DNA sequence being operably linked to a promoter for said catalytic domain of a DNA-dependent RNA polymerase and covalently linked to at least one sequence encoding an element that specifically binds to said RNA-binding domain, said active ingredients being formulated for separate, simultaneous or sequential administration.
[0136] Said DNA sequence of interest can be an anti-cancer gene (tumor suppressor gene). Said DNA sequence of interest can code for a therapeutic polypeptide or a non-coding RNA selected in the group comprising siRNA, ribozyme, shRNA and antisense RNA. Said therapeutic polypeptide can be selected from monoclonal antibodies or fragments thereof, growth factors, cytokines, cell or nuclear receptors, ligands, coagulation factors, CFTR proteins, insulin, dystrophin, hormones, enzymes, enzyme inhibitors, polypeptides with antineoplastic effect, polypeptides capable of inhibiting bacterial, parasitic or viral infections, in particular HIV, antibodies, toxins, immunotoxins. Preferably, the combination product according to the invention can be formulated in a pharma-ceutically acceptable carrier. In one embodiment of the combination product according to the invention, said vector is administered before said DNA sequence of interest.
[0137] The present invention also relates to a combination product according to the invention for use as a medicament in the prevention and / or treatment, in particular by gene therapy, of a human or animal pathology, which may be selected from the group consisting of pathologies that can be ameliorated by the administration of at least one DNA sequence of interest, as described above.
[0138] For example, the pathologies, as well as their clinical, biological or genetic subtypes, include liver disorders (e.g., acute liver failure due to acetaminophen poisoning or other causes, prevention of liver failure after hepatectomy, primary cancers of the liver including hepatocellular carcinoma or cholangiocarcinoma, nonalcoholic steatohepatitis, and hepatic monogenic disorders such as hemochromatosis, ornithine transcarbamylase deficiency, argininosuccinate deficiency, argininosuccinate synthetase 1, hemochromatosis or Wilson's disease), disorders resulting from or associated with a deficiency of a secretory protein (e.g., lysosomal storage, such as Gaucher's disease, Niemann-Pick disease, Tay-Sachs disease or Sandhoff disease, Hunter's syndrome or Hurler's disease; factor VIIIc, factor IX, Von deficiencies of clotting factors including Willebrand factor, fibrinogen factor or other clotting proteins, and colony-stimulating factors including erythropoietin, granulocyte colony-stimulating factor and thrombopoietin), cancers and predispositions thereto (e.g., breast, colorectal, pancreatic, gastric, esophageal and lung cancer, and melanoma), malignant hemopathies (e.g., leukemia, Hodgkin's and non-Hodgkin's lymphoma, myeloma), hemoglobinopathies (e.g., sickle cell anemia, glucose-6-phosphate dehydrogenase deficiency) and and thalassemia, autoimmune disorders (e.g., systemic lupus erythematosus, scleroderma, autoimmune hepatitis), cardiovascular disorders (e.g., cardiac rhythm and conduction disorders, hypertrophic cardiomyopathy, cardiovascular disease, or chronic heart failure), metabolic disorders (e.g., diabetes mellitus type I and type II and their complications, dyslipidemia, atherosclerosis and their complications), infectious disorders (e.g., AIDS, viral hepatitis B, viral hepatitis C, influenza flu, Zika, Ebola and other viral diseases; botulism, tetanus and other bacterial disorders;malaria and other parasitic disorders), muscle disorders (e.g. Duchenne muscular dystrophy and Steinert myotonic muscular dystrophy), respiratory diseases (e.g. cystic fibrosis, alpha-1 antitrypsin deficiency, acute respiratory distress syndrome, pulmonary arterial hypertension, pulmonary veno-occlusive disease), renal diseases (e.g. polycystic kidney disease, glomerulopathies), colorectal disorders (e.g. Crohn's disease and ulcerative colitis), eye disorders, especially retinal diseases (e.g. Leber's black eye). cataracts, retinitis pigmentosa, age-related macular degeneration), central nervous system disorders (e.g., Alzheimer's disease, Parkinson's disease, amyotrophic lateral sclerosis, multiple sclerosis, Huntington's disease, neurofibromatosis, adrenoleukodystrophy, bipolar disorder, schizophrenia and autism), bone and joint disorders (e.g., rheumatoid arthritis, ankylosing spondylitis, osteoarthritis) and skin and connective tissue disorders (e.g., neurofibromatosis and psoriasis);
[0139] The present invention also relates to a method for producing a chimeric enzyme according to the invention, comprising the step of expressing in at least one host cell said nucleic acid molecule or a group of said nucleic acid molecules encoding the chimeric enzyme of the invention under conditions allowing the expression of said nucleic acid molecule(s) in said host cell.
[0140] A method for selecting one or more engineered enzymes comprising a non-eukaryotic polymerase component and a capping enzyme component is disclosed herein, wherein the engineered enzyme comprises enhanced activity compared to a control, and the method includes: (a) making a nucleic acid encoding one or more engineered enzyme variants, wherein the variant comprises a variant of a naturally occurring non-eukaryotic polymerase and a variant of a naturally occurring capping enzyme component; (b) incorporating the nucleic acid encoding one or more engineered enzyme variants into one or more eukaryotic cells, wherein the eukaryotic cells comprise a reporter, wherein the reporter is under the control of a polymerase promoter specific to the polymerase of the engineered enzyme, and further wherein the reporter is expressed only when the eukaryotic cell is capped by the capping enzyme; (c) expressing the nucleic acid encoding one or more engineered enzyme variants; and (d) determining which of the one or more variants confers enhanced activity compared to a control, and selecting the engineered enzyme variants. An example of such a method can be seen in Example 1.
[0141] In one embodiment, the naturally occurring non-eukaryotic polymerase and the naturally occurring capping enzyme components do not naturally occur in the same organism. For example, the polymerase can be T7 RNA polymerase and the capping enzyme can be NP868R. The polymerase and the capping enzyme can be separated by a linker. Examples of these enzymes and their linkers are described herein. The variant encoding the fusion protein can also encode a nuclear localization signal (NLS) that can be at the N-terminus of the fusion protein. Again, such NLSs are described elsewhere herein. The eukaryotic cell can be a yeast cell, such as, for example, Saccharomyces cerevisiae.
[0142] The nucleic acid encoding one or more engineered enzyme variants may be under the control of a promoter, which allows the practitioner to initiate expression of the fusion protein as desired. Such promoters are known to those skilled in the art, and one example includes the galactose-inducible promoter.
[0143] The reporter used to detect the expression of the fusion protein can be found in the plasmid. Examples of such reporters are known to those skilled in the art. Having the reporter in a separate plasmid allows customization of the system. The reporter plasmid can, for example, contain a fluorescent molecule that can be easily detected upon expression of the desired product.
[0144] The desired fusion protein product can then be identified and isolated. Optionally, the selection can then be further repeated. For example, the desired fusion protein product can be further mutated to determine additional mutations that confer the desired benefit. These additional mutants can then be subjected to the above selection method. This method can be performed 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90 or more times. As discussed in Example 1, the top 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2.0% or more of the desired fluorescent clones can be gated and selected for further directed evolution. "Sexual PCR" can then be used to minimize deleterious mutations along the pathway to select for fusion proteins, for example, this can be done in a high-throughput manner.
[0145] After the desired engineered enzyme (also referred to herein as a "fusion protein") has been identified and isolated, it can be sequenced. Sequencing methods are known to those skilled in the art. This sequencing can be performed in a high-throughput manner or by fluorescent (Sanger) sequencing.
[0146] Disclosed herein are engineered enzyme variants discovered as a result of the methods of selecting fusion proteins described herein. Also disclosed are nucleic acid molecules encoding the fusion proteins.
[0147] The controls used in the above methods can be naturally occurring non-eukaryotic polymerases and / or naturally occurring capping enzyme components corresponding to the variant(s). Other controls include, but are not limited to, non-functional proteins, proteins from other organisms, or mutant proteins from other rounds of selection.
[0148] Also disclosed herein is a system that utilizes the above-mentioned method for directed evolution.Therefore, described herein is a system for selecting one or more engineered enzymes that comprise non-eukaryotic polymerase components and capping enzyme components, wherein the engineered enzyme comprises enhanced activity, the system comprises a transformed eukaryotic cell, the eukaryotic cell comprises a reporter plasmid, the reporter plasmid is under the control of a polymerase promoter specific for the polymerase of the engineered enzyme, and the reporter is expressed only when capped by the capping enzyme.The eukaryotic cell can be designed to incorporate one or more variant nucleic acids.
[0149] Further disclosed is a method for selecting one or more engineered enzymes that include a non-eukaryotic polymerase component and a capping enzyme component, comprising: a) providing a nucleic acid encoding the engineered enzyme, where expression of the engineered enzyme is under the control of a promoter, the promoter being recognized by the non-eukaryotic polymerase of the engineered enzyme; b) placing the nucleic acid encoding the engineered enzyme under conditions suitable for its expression; and c) detecting the mRNA produced by the engineered enzyme and selecting the enzyme for further analysis. In this method, the polymerase of the engineered enzyme controls the promoter for its own expression. This allows for a "feedback loop" that can result in an mRNA product, which can then be detected and / or quantified, for example, to determine the efficiency of transcription or the total amount present. This feedback loop can be used to determine whether the engineered enzyme designed is actually functional. Further analysis can include sequencing or amplification of the produced mRNA. Methods for sequencing and amplification are described elsewhere herein. Separate reporters can also be included. The reporters are known to those skilled in the art. EXAMPLES
[0150] To further illustrate the principles of the present disclosure, the following examples are set forth to provide those skilled in the art with a complete disclosure and description of how the compositions, articles, and methods claimed herein are made and evaluated. They are intended to be purely illustrative of the present invention and are not intended to limit the scope of what the inventors regard as their disclosure. Efforts have been made to ensure accuracy with respect to numbers (e.g., amounts, temperatures, etc.). However, some errors and deviations should be accounted for. Unless otherwise indicated, temperatures are in °C or at ambient temperature, and pressures are at or near atmospheric pressure. There are many variations and combinations of process conditions that can be used to optimize product quality and performance. Only reasonable and routine experimentation is required to optimize such process conditions.
[0151] Example 1 For further characterization of the fusion enzyme disclosed herein and its ability to generate capped RNA in the nucleus, the model eukaryote Saccharomyces cerevisiae was used. First, to adapt the system for nuclear expression, a nuclear localization signal (SV40 NLS) was fused to the N-terminus of the fusion protein. The gene was placed under the control of a tightly regulated Gal promoter, and the entire cassette was integrated into the HO locus of S. cerevisiae. Next, a nuclear episomal reporter plasmid (2 micron) was constructed containing a fluorescent protein (ZsGreen1) under the control of an insulated T7 RNAP promoter followed by a polyadenylation signal (SV40). Given the critical role of 5' methylguanylate in the efficient translation of any mRNA in a eukaryotic host, the level of expressed reporter protein was used as a direct proxy to determine the transcriptional capping ability of the enzyme (Figure 1). As a negative control, a fusion protein containing the K282A mutation in NP868R was designed to abrogate the capping activity of the enzyme (ES246). When initially characterized in yeast to drive reporter expression (in the nucleus), the activity of the wild-type enzyme (ES245) was minimal compared to the negative control (Figure 2).
[0152] It was hypothesized that this fusion enzyme could be engineered to produce capped transcripts more efficiently, potentially resulting in higher protein expression. To engineer this enzyme, the entire gene (approximately 5.5 kbp) was mutagenized using error-prone PCR. The library of variants was subcloned into Escherichia coli (E. coli) and then integrated into a yeast strain containing a reporter plasmid. After induction with galactose, the top 0.5-1% of fluorescent clones were gated and selected for the next round of directed evolution. After multiple rounds of selection, a specific variant containing multiple mutations in both proteins showed highly enhanced activity (approximately 75-fold higher protein expression compared to the wild-type enzyme) Figure 2. "Sexual PCR" was used to minimize deleterious mutations along the pathway to select for the fusion protein. A complete list of mutations obtained from these variants is listed in Table 1. Additionally, machine learning tools such as convolutional neural networks were used to identify more beneficial mutations in addition to those obtained from our selection (Table 2).
[0153] The capabilities of the enzyme variants will be characterized for improved protein production across human cell lines and other industrially relevant eukaryotic chassis organisms such as plants. Additionally, the in vitro activity of the variants to generate 5'-capped RNA will be evaluated, both as fusion and separate enzymes, and their performance will be compared to wild-type T7 RNAP and Vaccinia capping enzyme for improved production of mRNA therapeutics and vaccines.
[0154] table [Table 1] Table 1: Summary of all mutations from active variants (SEQ ID NOs: 6-24) [Table 2] Table 2: 51 positions for improving stability of engineered fusion proteins
[0155] Finally, while the present disclosure has been provided in detail with respect to certain illustrative and particular embodiments thereof, it should be understood that it should not be considered as limited thereto, since many modifications are possible without departing from the broader spirit and scope of the present disclosure as defined in the appended claims.
[0156] It will be apparent to those skilled in the art that various modifications and variations can be made in the present disclosure without departing from the scope or spirit of the invention. Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the methods disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims.
[0157] References 1.Jais PH, Decroly E, Jacquet E, Le Boulch M, Jais A, Jean-Jean O, et al.C3P3-G1:first generation of a eukaryotic artificial cytoplasmic expression system.Nucleic Acids Research.2019;47(5):2681-98. 2.Eaton HE, Kobayashi T, Dermody TS, Johnston RN, Jais PH, Shmulevitz M.African Swine Fever Virus NP868R Capping Enzyme Promotes Reovirus Rescue during Reverse Protein Genetics by Promoting Reovirus Expression, Virion Assembly, and RNA Incorporation into Infectious Virions.J Virol.2017;91(11). Sequence Listing
[0158] array SEQ ID NO:1 Wild-type NP868R:T7 RNAP Sequence number 2 Nuclear localization signal MFLEPPKKKRKVV Sequence number 3 T7 RNAP NTINIAKNDFSDIELAAIPFNTLADHYGERLAREQLALEHESYEMGEARFRKMFERQLKAGEVADNAAAKPLITTLLPKMIARINDWFEEVKAKRGKRPTAFQFLQEIKPEAVAYITIKTTLACLTSADNTTVQAVASAIGRAIEDEARFGRIRDLEAKHFKKNVEEQLNKRVGHVYKKAFMQVVEADMLSKGLLGGEAWSSWHKEDSIHVGVRCIEMLIESTGMVSLHRQNAGVVGQDSETIELAPEYAEAIATRAGALAGISPMFQPCVVPPKPWTGITGGGYWANGRRPLALVRTHSKKALMRYEDVYMPEVYKAINIAQNTAWKINKKVLAVANVITKWKHCPVEDIPAIEREELPMKPEDIDMNPEALTAWKRAAAAVYRKDKARKSRRISLEFMLEQANKFANHKAIWFPYNMDWRGRVYAVSMFNPQGNDMTKGLLTLAKGKPIGKEGYYWLKIHGANCAGVDKVPFPERIKFIEENHENIMACAKSPLENTWWAEQDSPFCFLAFCFEYAGVQHHGLSYNCSLPLAFDGSCSGIQHFSAMLRDEVGGRAVNLLPSETVQDIYGIVAKKVNEILQADAINGTDNEVVTVTDENTGEISEKVKLGTKALAGQWLAYGVTRSVTKRSVMTLAYGSKEFGFRQQVLEDTIQPAIDSGKGLMFTQPNQAAGYMAKLIWESVSVTVVAAVEAMNWLKSAAKLLAAEVKDKKTGEILRKRCAVHWVTPDGFPVWQEYKKPIQTRLNLMFLGQFRLQPTINTNKDSEIDAHKQESGIAPNFVHSQDGSHLRKTVVWAHEKYGIESFALIHDSFGTIPADAANLFKAVRETMVDTYESCDVLADFYDQFADQLHESQLDKMPALPAKGNLNLRDILESDFAFA SEQ ID NO:4 Linker GGGGSGGGGSGGGGSL SEQ ID NO:5 NP868R ASLDNLVARYQRCFNDQSLKNSTIELEIRFQQINFLLFKTVYEALVAQEIPSTISHSIRCIKKVHHENHCREKILPSENLYFKKQPLMFFKFSEPASLGCKVSLAIEQ PIRKFILDSSVLVRLKNRTTFRVSELWKIELTIVKQLMGSEVSAKLAAFKTLLFDTPEQQTTKNMMTLINPDGEYLYEIEIEYTGKPESLTAADVIKIKNTVLTLISP NHLMLTAYHQAIEFIASHILSSEILLARIKSGKWGLKRLLPRVKSMTKADYMKFYPPVGYYVTDKADGIRGIAVIQDTQIYVVADQLYSLGTTGIEPLKPTILDGEFM PEKKEFYGFDVIMYEGNLLTQQGFETRIESLSKGIKVLQAFNIKAEMKPFISLTSADPNVLLKNFESIFKKKTRPYSIDGIILVEPGNSYLNTNTFKWKPTWDNTLDFL VRKCPESLNVPEYAPKKGFSLHLLFVGISGELFKKLALNWCPGYTKLFPVTQRNQNYFPVQFQPSDFPLAFLYYHPDTSSFSNIDGKVLEMRCLKREINYVRWEIVKI REDRQQDLKTGGYFGNDFKTAELTWLNYMDPFSFEELAKGPSGMYFAGAKTGIYRAQTALISFIKQEIIQKISHQSWVIDLGIGKGQDLGRYLDAGVRHLVGIDKDQTA LAELVYRKFSHATTRQHKHATNIYVLHQDLAEPAKEISEKVHQIYGFPKEGASSIVSNLFIHYLMKNTQQVENLAVLCHKLLQPGGMVWFTTMLGEQVLELLHENRIE LNEVWEARENEVVKFAIKRLFKEDILQETGQEIGVLLPFSNGDFYNEYLVNTAFLIKIFKHHGFSLVQKQSFKDWIPEFQNFSKSLYKILTEADKTWTSLFGFICLRKN SEQ ID NO:6 Variant ES 230 SEQ ID NO:7 Variant ES 368 SEQ ID NO:8 Variant ES 369 SEQ ID NO:9 Variant ES-430 SEQ ID NO:10 Variant ES-431 SEQ ID NO:11 Variant ES-432 SEQ ID NO:12 Variant ES-433 SEQ ID NO:13 Variant ES-434 SEQ ID NO:14 Variant ES-436 SEQ ID NO:15 Variant ES-440 SEQ ID NO:16 Variant ES-441 SEQ ID NO:17 Variant ES-442 SEQ ID NO:18 Variant ES-443 SEQ ID NO:19 Variant ES-444 SEQ ID NO:20 Variant ES-446 SEQ ID NO:21 Variant ES-447 SEQ ID NO:22 Variant ES-448 SEQ ID NO:23 Variant ES-45: SEQ ID NO:24 Variant ES-451 SEQ ID NO:25 TAATACGACTCACTATA SEQ ID NO:26 TAATACGACTCACTAAA SEQ ID NO:27 TAATACGACTCACTGTA SEQ ID NO:28 TAATACGACTCACTCTA SEQ ID NO:29 TAATACCGGTCACTATA SEQ ID NO:30 TAATACCTGACACTATA SEQ ID NO:31 TAATAACCCTCACTATA SEQ ID NO:32 TAATAACTATCACTATA SEQ ID NO:33 TAATAACCCACACTATA SEQ ID NO:34 WT-NPT7
Claims
1. An engineered enzyme comprising a T7 RNA polymerase component and a capping enzyme component separated by a linker, wherein the T7 RNA polymerase component has 90% or more identity with SEQ ID NO: 3, and the capping enzyme component has 90% or more identity with SEQ ID NO:
5.
2. The manipulated enzyme according to claim 1, wherein the manipulated enzyme comprises at least one substitution that confers at least one improved property compared to the manipulated enzyme without substitution.
3. The manipulated enzyme according to claim 1, wherein the linker may vary in length or amino acid composition.
4. An engineered enzyme comprising SEQ ID NO: 1 having at least one substitution that confers at least one improved property compared to SEQ ID NO: 1 without substitution, further comprising a linker at positions 881-896 of the engineered enzyme, the linker of which may differ in length or amino acid composition.
5. The manipulated enzyme according to claim 4, wherein the substitution is located at at least one of the following positions in Sequence ID No. 1: D279, H1667, K9, H1195, N379, Q624, L740, K1058, L1156, R1202, Q1543, S1581, D1769.
6. The manipulated enzyme according to claim 4, wherein the substitution is located at one or more of the following positions in SEQ ID NO: R10, G160, K553, H690, F753, K831, D921, W983, I1012, A1024, N1026, G1133 and / or S1390.
7. The manipulated enzyme according to claim 4, wherein the substitution is located at one or more of the following positions in Sequence ID No. 1: K184, Q624, D982, N1060 and / or Q1551.
8. The substitution occurs at the following positions in sequence number 1: R22, Q30, E38, Q45, Q45, Q98, L100, R124, R143, M159, I324, N396, T497, Q498, N502, M598, R678, Q679, Q760, H832, L911, I914, H922, R952, R994, T1022, D1025, T1027, Y1073, L The manipulated enzyme according to claim 4, which is present in one or more of 1090, L1091, W1096, H1100, V1131, F1163, W1239, M1257, M1264, M1296, I1476, N1487, I1500, F1539, M1561, C1618, L1647, I1705, V1723, M1727, and / or C1734.
9. The engineered enzyme according to claim 8, comprising at least one improved property compared to SEQ ID NO:
1.
10. The engineered enzyme according to claim 4, wherein the at least one improved property is selected from improved selectivity for capping, improved processability of capping, improved capping enzyme activity, improved protein expression, improved RNA yield, improved stability in storage buffer, improved stability under reaction conditions, improved processability of translation, improved thermal stability, and improved transcription fidelity.
11. The engineered enzyme according to claim 10, wherein the improved enzyme activity includes an improvement in the activity of RNA triphothphatase guanyltransferase and / or methyltransferase.
12. A nucleic acid encoding the manipulated enzyme described in claim 1.
13. The nucleic acid according to claim 12, wherein the polynucleotide sequence is operably linked to a control sequence.
14. An expression vector comprising the nucleic acid sequence described in claim 12.
15. A host cell comprising the expression vector described in claim 14.
16. A method for selecting one or more engineered enzymes comprising a non-eukaryotic polymerase component and a capping enzyme component, wherein the engineered enzymes have enhanced activity compared to a control, and the method is a. To produce nucleic acids encoding one or more manipulated enzyme variants, wherein the variants include variants of naturally occurring non-eukaryotic polymerases and / or variants of naturally occurring capping enzyme components; b. Incorporating the nucleic acids encoding one or more manipulated enzyme variants into one or more eukaryotic cells, wherein the eukaryotic cells include a reporter, the reporter is under the control of a polymerase promoter specific to the polymerase of the manipulated enzyme, and the reporter is expressed only when capped by the capping enzyme; c. Expressing the nucleic acid encoding one or more manipulated enzyme variants; and d. Determining which of one or more variants confers enhanced activity compared to the control, and selecting the manipulated enzyme variant. Methods that include...
17. The method according to claim 16, wherein the naturally occurring non-eukaryotic polymerase and the naturally occurring capping enzyme component are not naturally present in the same organism.
18. The method according to claim 16, wherein the polymerase is T7 RNA polymerase.
19. The method according to claim 16, wherein the capping enzyme is NP868R.
20. The method according to claim 16, wherein the polymerase and capping enzyme are separated by a linker.