Compositions and methods relating to engineered RNA polymerases with capping enzymes

WO2026206837A1PCT designated stage Publication Date: 2026-10-01BOARD OF RGT THE UNIV OF TEXAS SYST +8
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020360
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure IMGF000041_0001
    Figure IMGF000041_0001
  • Figure IMGF000041_0002
    Figure IMGF000041_0002
  • Figure IMGF000042_0001
    Figure IMGF000042_0001
Patent Text Reader

Abstract

The present disclosure relates to engineered enzymes, compositions, and systems thereof comprising T7 polymerases and mRNA capping enzymes, wherein at least one substitution mutation has been introduced to said engineered enzyme to improve it functionality. The present disclosure also relates to methods of using said engineered enzymes, compositions, and systems thereof including generating a capped mRNA transcript, generating a polypeptide, and / or generating a bulk amount of RNA.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Docket No. 10046-673W01

[0002] COMPOSITIONS AND METHODS RELATING TO ENGINEERED RNA POLYMERASES WITH CAPPING ENZYMES

[0003] RELATED APPLICATION

[0004] This PCT application claims priority to, and the benefit of, U. S. Provisional Patent Application No. 63 / 776,669, filed March 24, 2025, entitled “COMPOSITIONS AND METHODS RELATING TO ENGINEERED RNA POLYMERASES WITH CAPPING ENZYMES,” which is incorporated by reference herein in its entirety.

[0005] REFERENCE TO SEQUENCE LISTING

[0006] The sequence listing submitted on March 23, 2026, as an. XML file entitled “10046-673WO1_ST26” created on February 10, 2026, and having a file size of 38,246 bytes is hereby incorporated by reference pursuant to 37 C. F. R. § 1.52(e)(5).

[0007] FIELD

[0008] The present disclosure relates to compositions of engineered T7 polymerases, mRNA capping enzymes, and methods of use thereof.

[0009] BACKGROUND

[0010] Over the past decade, the fields of synthetic biology and protein engineering have made remarkable strides, driving progress across therapeutic, molecular, and industrial biotechnology domains. At the core of many of these advancements are RNA polymerases, which play a fundamental role in transcription — the first and indispensable step of gene expression. Among these enzymes, bacteriophage T7 RNA polymerase (T7 RNAP) stands out as a versatile tool due to its exceptional efficiency and specificity in RNA synthesis. This has established T7 RNAP as a cornerstone in synthetic biology, particularly in applications involving in vitro transcription and mRNA production.

[0011] Presently T7 RNAP has shown substantial challenges in eukaryotic systems mainly because eukaryotic systems require complex post-transcriptional modifications, such as the addition of 5'-m7G caps and poly(A) tails, to stabilize mRNA and promote efficient translation. These additional requirements, coupled with processes like splicing and nuclear export, necessitate innovative approaches to expand T7 RNAP’s functionality in eukaryotic contexts.

[0012] The engineered enzymes and methods disclosed herein address these and other needs.Docket No. 10046-673W01

[0013] SUMMARY

[0014] The present disclosure provides engineered polymerases, engineered capping enzymes, and / or enzyme complexes thereof, for enhanced transcription and translation. The present disclosure also provides compositions comprising engineered polymerases, engineered capping enzymes, and / or enzyme complexes thereof. The present disclosure also provides methods using engineered polymerases, engineered capping enzymes, and / or enzyme complexes thereof, to efficiently and / or rapidly producing ribonucleic acids and / or polypeptides.

[0015] In some aspects, disclosed herein is an engineered T7 RNA polymerase (T7 RNAP), wherein said T7 RNAP comprises SEQ ID NO: 15, and further wherein SEQ ID NO: 15 comprises at least one substitution mutation selected from Q58S, Q404L, A584K, V609S, W698F, and / or Q786L, or any combination thereof.

[0016] In some embodiments, SEQ ID NO: 1 comprises at least one additional mutation other than Q58S, Q404L, A584K, V609S, W698F, and / or Q786L. In some embodiments, the T7 RNAP is operably linked to a capping enzyme comprising SEQ ID NO: 11, or a variant thereof. In some embodiments, the variant of SEQ ID NO: 11 comprises at least one substitution mutation of S199A, N267W, and / or Q426V.

[0017] In some aspects, disclosed herein is an capping enzyme, wherein the engineered capping enzyme comprises SEQ ID NO: 11, and wherein SEQ ID NO: 11 comprises at least one substitution mutation comprising S199A, N267W, and / or Q426V.

[0018] In some embodiments, the engineered capping enzyme is operably linked to a T7 RNA polymerase comprising SEQ ID NO: 15, or a variant thereof. In some embodiments, the variant of SEQ ID NO: 15 comprises at least one substitution comprising Q58S, Q404L, A584K, V609S, W698F, and / or Q786L.

[0019] In some aspects, disclosed herein is an enzyme complex comprising a T7 RNA polymerase (T7 RNAP) component and a capping enzyme component, wherein the T7 RNAP component comprises SEQ ID NO: 15, or a variant thereof, and further wherein the capping component comprises SEQ ID NO: 11, or a variant thereof, wherein a linker comprising at least 90% sequence identity to SEQ ID NO: 2 separates the T7 RNAP component and the capping enzyme component, wherein the variant of SEQ ID NO: 15 comprises at least one mutation selected from Q58S, Q404L, A584K, V609S, W698F, and / or Q786L, and wherein the variant of SEQ ID NO: 11 comprises at least one mutation selected from S199A, N267W, and / or Q426V.Docket No. 10046-673W01

[0020] In some embodiments, the enzyme complex is a component in a transcription system, a post-transcription system, or a combination thereof. In some embodiments, the post¬ transcription system comprises protein translation.

[0021] In some aspects, disclosed herein is a polynucleotide sequence encoding the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any preceding aspect.

[0022] In some aspects, disclosed herein is a cell comprising the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any preceding aspect.

[0023] In some embodiments, the cell is a eukaryotic cell or a prokaryotic cell. In some embodiments, the cell proliferates into a stable cell line or a transient cell line.

[0024] In some aspects, disclosed herein is a cell comprising a polynucleotide sequence encoding the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any- preceding aspect.

[0025] In some aspects, disclosed herein is a cell culture composition comprising the cell of any preceding aspect, wherein said cell includes but is not limited to a tissue culture cell.

[0026] In some aspects, disclosed herein is an in vitro system comprising one or more polynucleotide sequences encoding the enzyme complex of any preceding aspect and a reporter gene, wherein the polynucleotide sequence is under the control of a T7 RN AP promoter.

[0027] In some aspects, disclosed herein is a method of producing a capped RNA transcript, the method comprising performing a transcription reaction utilizing the enzyme complex of any preceding aspect, wherein the method further comprises a composition comprising a) one or more components needed to produce a methylguanylate cap; and b) a DNA template, wherein said DNA template is exposed to said composition under conditions such that the enzyme complex produces a capped RNA transcript.

[0028] In some aspects, disclosed herein is a method of producing RNA, the method comprising a) expressing the enzyme complex of any preceding aspect in a cell, and b) isolating an amount of RNA from the cell that is increased relative to an otherwise identical control cell expressing the T7 RNA polymerase lacking the capping enzyme.

[0029] BRIEF DESCRIPTION OF FIGURES

[0030] The accompanying figures, which are incorporated in and constitute a part of this specification, illustrate several aspects described below.

[0031] Figures 1A, IB, 1C, and ID show the engineering of T7 RNA polymerase using structure-aware machine learning predictions for efficient target gene expression. Figure 1A shows the two-plasmid system for screening T7 RNAP variants. The expression plasmid,Docket No. 10046-673W01

[0032] containing the ASFVCE: T7 RNAP fusion protein, was selected using the HIS3 marker, while the reporter plasmid, encoding the ZsGreen reporter gene, was selected using the URA3 marker. Both plasmids were transformed into a AGal2 derivative of 5. cerevisiae for screening gene expression. Figure IB show's the schematic representation of the ASFVCE: T7 RNAP expression cassette and reporter plasmid. T7 RNAP variants fused with ASFVCE via a GS linker were expressed under the control of the PGalpromoter. An SV40 nuclear localization signal (NLS) was appended to the N-terminal of ASFVCE to facilitate nuclear transport. The ASFVCE: T7 RNAP cassette was integrated into the HO locus of a AGal2 derivative of S. cerevisiae. The ZsGreen reporter gene, controlled by the PT7promoter, was expressed from a plasmid maintained using the 2-micron system. An HDV ribozyme was positioned upstream of the PT7promoter to ensure proper transcriptional regulation. Figure 1C shows the screening of single-point mutations in T7 RNAP. A total of 60 single-point mutations were introduced into T7 RNAP fused with ASFVCE to evaluate their impact on reporter gene expression. Figure ID shows the screening of combinatorial mutations in T7 RNAP. A total of 26 combinatorial mutations, including double, triple, quadruple, and quintuple mutations, were introduced into T7 RNAP fused with ASFVCE to assess their effect on reporter gene expression. For each mutation, fold change was calculated as the ratio of ZsGreen expression, measured by flow cytometry, relative to WT T7 RNAP (set as 1) under identical galactose induction conditions. The " Null" construct, containing the Y629A mutation, was used as a negative control since this mutation impairs T7 RNAP activity. Each dot represents the average value obtained from a transformation experiment, with three biological replicates per transformation. Bar graphs indicate the mean and standard deviation (SD).

[0033] Figures 2A and 2B show the engineering of quadruple and quintuple T7 RNAP variants using evolution-aw'are machine learning. Figure 2 A shows the single and combinatorial mutations in the quadruple T7 RNAP variant (T7 RNAPQ58S / A584K / V609S / W698F). A total of 24 mutations, including single, double, and triple mutations, were introduced into the quadruple T7 RNAP variant fused with ASFVCE to assess their effect on reporter gene expression. Figure 2B shows the screening of single and combinatorial mutations in the quintuple T7 RNAP variant (T7 RNAPQ58S / S128V / A584K / V609S / W698F). A total of 22 mutations, including single and double mutations, were introduced into the quintuple T7 RNAP variant fused with ASFVCE to evaluate their impact on reporter gene expression. Fold change was determined by comparing ZsGreen fluorescence intensity, as measured via flow cytometry, to that of the respective quadruple or quintuple T7 RNAP variant (set as 1) under identical galactose induction conditions. The fold change of the quadruple and quintuple T7 RNAP variants isDocket No. 10046-673W01

[0034] shown as a horizontal dashed line in (A) and (B), respectively. Each dot represents the average value from a transformation experiment, with three biological replicates per transformation. Bar graphs indicate the mean and standard deviation (SD).

[0035] Figures 3 A, 3B, 3C, and 3D show the screening and characterization of single-subunit CEs. Figure 3A shows the phylogenetic tree of single-subunit CE candidates. Phylogenetic relationships among 10 viral-derived single-subunit CEs, including ASFVCE and other enzymes, are shown. Figure 3B shows the schematic representation of the CE: T7 RNAP expression cassette and reporter plasmid. Single-subunit CE variants were fused to WT T7 RNAP via a GS linker and expressed under the control of the PGalpromoter. An SV40 nuclear localization signal (NLS) was attached to the N-terminal of each CE: T7 RNAP to facilitate nuclear transport. The CE: T7 RNAP cassette was integrated into the HO locus of a AGal2 derivative of S. cerevisiae. The ZsGreen reporter gene, under the control of the PT7promoter, was expressed from a plasmid maintained using the 2-micron system. An HDV ribozyme was positioned upstream of the PT7promoter to ensure proper transcription regulation. Figure 3C shows the screening of single-subunit CEs. A total of 10 viral -derived single-subunit CEs, including WT ASFVCE and the ASFVCE (K 82N) null mutant, were tested for their impact on reporter gene expression. Fold change in ZsGreen fluorescence was determined by normalizing each CE variant to the WT ASF VCE (set as 1 ), with measurements obtained under identical galactose induction conditions via flow cytometry. Each dot represents the average value obtained from three colonies per transformation, and the bar graphs indicate the mean and standard deviation (SD). Figure 3D shows the evaluation of ZsGreen reporter gene fluorescence. Flow cytometry histograms illustrate the fluorescence intensity distribution, where a rightward shift represents increased CE activity relative to the negative control. ASFVCE(K282N) served as the null mutant control, displaying baseline fluorescence levels, while the negative control (NC) represented WT T7 RNAP expressed without a capping enzyme.

[0036] Figures 4A, 4B, and 4C show the engineering of BMCE for enhanced target gene expression. Figure 4A shows the schematic representation of the BMCE: T7 RNAP expression cassette and reporter plasmid. BMCE variants fused to T7 RNAP via a GS linker were expressed under the control of the PGalpromoter. An SV40 nuclear localization signal (NLS) was appended to the N-terminal of BMCE to facilitate nuclear transport. The BMCE: T7 RNAP cassette was integrated into the HO locus of a ΔGal2 derivative of S. cerevisiae. The ZsGreen reporter gene, controlled by the P77 promoter, was expressed from a plasmid maintained using the 2-micron system. An HDV ribozyme was positioned upstream of the P77 promoter to ensureDocket No. 10046-673W01

[0037] proper transcription regulation. Figure 4B shows the screening of single-point mutations in BMCE. A total of 58 single-point mutations were introduced into BMCE fused with T7 RNAP to evaluate their impact on reporter gene expression. Figure 4C shows the screening of combinatorial mutations in BMCE. A total of 36 combinatorial mutations, including double, triple, quadruple, and quintuple mutations, were introduced into BMCE fused with T7 RNAP to assess their effect on reporter gene expression. For each mutation, fold change was calculated as the ratio of ZsGreen expression, measured by flow cytometry, relative to WT BMCE (set as 1) under identical galactose induction conditions. Each dot represents the average value obtained from a transformation experiment, with three biological replicates per transformation. Bar graphs indicate the mean and standard deviation (SD).

[0038] Figures 5A and 5B show the comparative activity of combinatorial capping enzyme (CE) and T7 RNAP variants on reporter gene expression. Figure 5A shows the schematic representation of combinatorial mutations. Four capping enzyme (CE) variants (ASFVCE(WT), BMCE(WT), EvoBMCE, and ASFVCE(443)) and four T7 RNAP variants (T7 RNAP(WT), T7 RNAP(443), EvoT7, and EvoT7(443)) were combined to generate multiple CE-T7 RNAP fusion constructs for activity assessment. EvoBMCE corresponds to the triple variant BMCES199A / N267W / Q426V, and EvoT7(443) comprises 12 mutations (six from T7 RNAP(443) and six from EvoT7), integrating both directed evolution and machine learning-derived substitutions. Figure 5B shows the characterization of combinatorial variants for reporter gene expression. Reporter gene expression was quantified by measuring ZsGreen fluorescence normalized to OD600nm. ZsGreen / OD600nm values were assessed for each combinatorial CE-T7 RNAP construct and compared to the baseline ASFVCE(WT)-T7 RNAP(WT), which was set to 1. Each dot represents the average value obtained from three colonies per transformation, with bar graphs indicating the mean and standard deviation (SD).

[0039] Figures 6A and 6B show a comparison of single mutants tested in EvoT7, epT7, and G47A / 884G, alongside the mutations identified in v443. Figure 6A shows a Venn diagram comparing single mutants based on exact residue matches. Each circle represents the number of single mutants tested in each study, except for v443, where the six mutations present in the final variant (obtained through random mutagenesis) were considered instead of individually tested single mutations. Overlapping regions indicate identical mutations at both the same position and with the same amino acid substitution. EvoT7 tested 72 single mutants, among which only one (H300R) was found in v443, while no exact matches were observed with epT7 and G47A / 884G. Figure 6B shows a Venn diagram comparing mutation site overlap regardless of residue substitution. EvoT7 contained 68 unique mutation sites, among which three (V134,Docket No. 10046-673W01

[0040] V177, and V273) were also present in epT7. No overlapping mutation sites were found between EvoT7 and G47A / 884G, while v443 shared only one site (H300) with EvoT7.

[0041] Figures 7A, 7B, 7C, and 7D show a structural comparison of mutation sites and local microenvironments in EvoT7 and EvoBMCE. Figure 7 A shows the locations of mutation sites mapped onto the crystal structure of T7 RNAP (PDB 1H38). Figure 7B compares the local microenvironment of EvoT7 mutation sites between the AlphaFold-predicted structure of EvoT7 (bottom panels) and the crystal structure of WT T7 RNAP (top panels, PDB 1H38). Figure 7C shows the mapped locations of mutation sites onto the AlphaFold-predicted structure of WT BMCE. Figure 7D compares the local microenvironments of EvoBMCE mutation sites between the AlphaFold-predicted structure of EvoBMCE (bottom panels) and the AlphaFold-predicted structure of WT BMCE (top panels).

[0042] Figure 8 shows the activity of T7 RNA polymerase (T7 RNAP) variants in a bacterial cell-free assay system. Each assay was performed in triplicate. The cell-free reactions were incubated at 37°C for 4 hours, and GFP fluorescence was determined using a plate reader.

[0043] Figures 9 A and 9B show the evaluation of the capping enzyme -T7 RNAP fusion system in tissue culture. Figure 9A shows the schematic of reporter constructs used to test the transcriptional and translational activity of the various capping enzyme-T7 RNAP fusion protein in tissue culture (HEK293T). Figure 9B shows the GFP fluorescence signals from tissue culture cells transfected with the ASFVCE(443)-T7RNAP(443) expression vector and various reporter plasmids. The result confirms that the fusion protein successfully produces functional, 5’ capped mRNA capable of efficient translation within a tissue culture environment.

[0044] Figure 10 shows the performance of optimization of evolved capping enzyme-T7 RNAP fusion variants. GFP fluorescence signals from HEK293T cells demonstrate the superior efficiency of engineered fusion constructs relative to the wild-type ASFVCE-T7 RNAP. The data indicates that utilizing specifically evolved components leads to a significant increase in expression levels, providing highly effective tools for protein production in tissue culture.

[0045] Figure 11 shows that multiple, integrated variants support expression from a T7 RNAP promoter. Percentage of GFP-positive cells is shown for HEK293T cell lines with the following various constructs stably integrated into the AAVS1 locus.

[0046] • ASFV 1: HEK293T A AVS 1:: PCAG- ASFVCE(443)-T7 RNAP(443)

[0047] • ASFV2: HEK293T AAVS1:: PTRE3GV-ASFVCE(443)-T7 RNAP(443)

[0048] • BM1: HEK293T AAVS1:: PCAG-EVOBMCE-EVOT7

[0049] • BM2: HEK293T AAVS1:: PTRE3GV-EVOBMCE-EVOT7Docket No. 10046-673W01

[0050] Figure 12 shows the impact of 3' optimization of the reporter constructs for maximizing T7-mediated gene expression in stable cell lines. GFP fluorescence signals are shown for the HEK293T cell line stably expressing the ASFVCE(443)-T7RNAP(443), with GFP reporter constructs.

[0051] Figure 13 shows the optimization of 5' reporter architecture for maximized T7-mediated gene expression. GFP fluorescence signals were measured in HEK293T cell lines stably expressing the ASFVCE(443)-T7RNAP(443) fusion protein. Lanes 1 through 4 represent reporter constructs featuring various 5’ UTR variants. These results demonstrate that the system maintains high activity across various 5’ UTRs and ensures robust and consistent protein production in a tissue culture environment.

[0052] DETAILED DESCRIPTION

[0053] The following description of the disclosure is provided as an enabling teaching of the disclosure in its best, currently known embodiments). To this end, those skilled in the relevant art will recognize and appreciate that many changes can be made to the various embodiments of the invention described herein, while still obtaining the beneficial results of the present disclosure. It will also be apparent that some of the desired benefits of the present disclosure can be obtained by selecting some of the features of the present disclosure without utilizing other features. Accordingly, those who work in the art will recognize that many modifications and adaptations to the present disclosure are possible and can even be desirable in certain circumstances and are a part of the present disclosure. Thus, the following description is provided as illustrative of the principles of the present disclosure and not in limitation thereof.

[0054] Reference will now be made in detail to the embodiments of the invention, examples of which are illustrated in the drawings and the examples. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein.

[0055] Terminology

[0056] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which this disclosure belongs. The term “comprising” and variations thereof as used herein is used synonymously with the term “including” and variations thereof and are open, non-limiting terms. Although the terms “comprising” and “including” have been used herein to describe various embodiments, the terms “consisting essentially of’ and “consisting of’ can be used in place ofDocket No. 10046-673W01

[0057] “comprising” and “including” to provide for more specific embodiments and are also disclosed. As used in this disclosure and in the appended claims, the singular forms “a”, “an”, “the”, include plural referents unless the context clearly dictates otherwise.

[0058] The following definitions are provided for the full understanding of terms used in this specification.

[0059] Ranges can be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. It is also understood that when a value is disclosed that “less than or equal to” the value, “greater than or equal to the value” and possible ranges between values are also disclosed, as appropriately understood by the skilled artisan. For example, if the value “10” is disclosed the “less than or equal to 10” as well as “greater than or equal to 10” is also disclosed. It is also understood that the throughout the application, data is provided in a number of different formats, and that this data, represents endpoints and starting points, and ranges for any combination of the data points. For example, if a particular data point “10” and a particular data point 15 are disclosed, it is understood that greater than, greater than or equal to, less than, less than or equal to, and equal to 10 and 15 are considered disclosed as well as between 10 and 15. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.

[0060] “Optional” or “optionally” means that the subsequently described event or circumstance may or may not occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.

[0061] An "increase" can refer to any change that results in a greater amount of a symptom, disease, composition, condition, or activity. An increase can be any individual, median, or average increase in a condition, symptom, activity, composition in a statistically significant amount. Thus, the increase can be a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100% or more increase so long as the increase is statistically significant.Docket No. 10046-673W01

[0062] A "decrease" can refer to any change that results in a smaller amount of a symptom, disease, composition, condition, or activity. A substance is also understood to decrease the genetic output of a gene when the genetic output of the gene product with the substance is less relative to the output of tire gene product without the substance. Also, for example, a decrease can be a change in the symptoms of a disorder such that the symptoms are less than previously observed. A decrease can be any individual, median, or average decrease in a condition, symptom, activity, composition in a statistically significant amount. Thus, the decrease can be a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100%, or more decrease so long as the decrease is statistically significant.

[0063] “Composition” refers to any agent that has a beneficial biological effect. Beneficial biological effects include both therapeutic effects, e.g., treatment of a disorder or other undesirable physiological condition, and prophylactic effects, e.g., prevention of a disorder or other undesirable physiological condition. The terms also encompass pharmaceutically acceptable, pharmacologically active derivatives of beneficial agents specifically mentioned herein, including, but not limited to, a vector, polynucleotide, cells, salts, esters, amides, proagents, active metabolites, isomers, fragments, analogs, and the like. When the term “composition” is used, then, or when a particular composition is specifically identified, it is to be understood that the term includes the composition per se as well as pharmaceutically acceptable, pharmacologically active vector, polynucleotide, salts, esters, amides, proagents, conjugates, active metabolites, isomers, fragments, analogs, etc.

[0064] " Comprising" is intended to mean that the compositions, methods, etc. include the recited elements, but do not exclude others. " Consisting essentially of when used to define compositions and methods, shall mean including the recited elements, but excluding other elements of any essential significance to the combination. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants from the isolation and purification method and pharmaceutically acceptable carriers, such as phosphate buffered saline, preservatives, and the like. " Consisting of' shall mean excluding more than trace elements of other ingredients and substantial method steps for administering the compositions provided and / or claimed in this disclosure. Embodiments defined by each of these transition terms are within the scope of this disclosure.

[0065] " Inhibit," "inhibiting," and "inhibition" mean to decrease an activity, response, condition, disease, or other biological parameter. This can include but is not limited to the complete ablation of the activity, response, condition, or disease. This may also include, for example, a 10% reduction in the activity, response, condition, or disease as compared to theDocket No. 10046-673W01

[0066] native or control level. Thus, the reduction can be a 10, 20, 30, 40, 50, 60, 70, 80, 90, 100%’, or any amount of reduction below, above, or in between the given ranges as compared to native or control levels.

[0067] By “reduce” or other forms of the word, such as “reducing” or “reduction,” means lowering of an event or characteristic (e.g., tumor growth). It is understood that this is typically in relation to some standard or expected value, in other words it is relative, but that it is not always necessary for the standard or relative value to be referred to. For example, “reduces tumor growth” means reducing the rate of growth of a tumor relative to a standard or a control.

[0068] By “prevent” or other forms of the word, such as “preventing” or “prevention,” is meant to stop a particular event or characteristic, to stabilize or delay the development or progression of a particular event or characteristic, or to minimize the chances that a particular event or characteristic will occur. Prevent does not require comparison to a control as it is typically more absolute than, for example, reduce. As used herein, something could be reduced but not prevented, but something that is reduced could also be prevented. Likewise, something could be prevented but not reduced, but something that is prevented could also be reduced. It is understood that where reduce or prevent are used, unless specifically indicated otherwise, the use of the other word is also expressly disclosed.

[0069] The terms “treat,” “treating,” and grammatical variations thereof as used herein, include partially or completely delaying, alleviating, mitigating or reducing the intensity of one or more attendant symptoms of a disorder or condition and / or alleviating, mitigating or impeding one or more causes of a disorder or condition. Treatments according to the disclosure may be applied preventively, prophylactically, palliatively or remedialiy. Treatments are administered to a subject prior to onset (e.g., before obvious signs of a disease), during early onset (e.g., upon initial signs and symptoms of a disease), or after an established development of a disease.

[0070] The term “subject” refers to any individual who is the target of administration or treatment. The subject can be a vertebrate, for example, a mammal. In one aspect, the subject can be human, non-human primate, bovine, equine, porcine, canine, or feline. The subject can also be a guinea pig, rat, hamster, rabbit, mouse, or mole. Thus, the subject can be a human or veterinary patient. The term “patient” refers to a subject under the treatment of a clinician, e.g., physician.

[0071] The term “treatment” refers to the medical management of a patient with the intent to cure, ameliorate, stabilize, or prevent a disease, pathological condition, or disorder. This term includes active treatment, that is, treatment directed specifically toward the improvement of a disease, pathological condition, or disorder, and also includes causal treatment, that is,Docket No. 10046-673W01

[0072] treatment directed toward removal of the cause of the associated disease, pathological condition, or disorder. In addition, this term includes palliative treatment, that is, treatment designed for the relief of symptoms rather than the curing of the disease, pathological condition, or disorder; preventative treatment, that is, treatment directed to minimizing or partially or completely inhibiting the development of the associated disease, pathological condition, or disorder; and supportive treatment, that is, treatment employed to supplement another specific therapy directed toward the improvement of the associated disease, pathological condition, or disorder.

[0073] The term “amino acid,” includes but is not limited to amino acids contained in the group consisting of alanine (Ala or A), cysteine (Cys or C), aspartic acid (Asp or D), glutamic acid (Glu or E), phenylalanine (Phe or F), glycine (Gly or G), histidine (His or H), isoleucine (Ile or I), lysine (Lys or K), leucine (Leu or L), methionine (Met or M), asparagine (Asn or N), proline (Pro or P), glutamine (Gin or Q), arginine (Arg or R), serine (Ser or S), threonine (Thr or T), valine (Val or V), tryptophan (Trp or W), and tyrosine (Tyr or Y) residues. The term “amino acid residue” also may include amino acid residues contained in the group consisting of homocysteine, 2-Aminoadipic acid, N-Ethylasparagine, 3-Aminoadipic acid, Hydroxylysine, P-alanine, P- Amino-propionic acid, allo-Hydroxylysine acid, 2- Aminobutyric acid, 3-Hydroxyproline, 4- Aminobutyric acid, 4-Hydroxyproline, piperidinic acid, 6-Aminocaproic acid, Isodesmosine, 2-Aminoheptanoic acid, allo-Isoleucine, 2-Aminoisobutyric acid, N-Methylglycine, sarcosine, 3-Aminoisobutyric acid, N- Methylisoleucine, 2-Aminopimelic acid, 6-N-Methyllysine, 2,4-Diaminobutyric acid, N-Methylvaline, Desmosine, Norvaline, 2,2'-Diaminopimelic acid, Norleucine, 2,3-Diaminopropionic acid, Ornithine, and N-Ethylglycine. Typically, the amide linkages of the peptides are formed from an amino group of the backbone of one amino acid and a carboxyl group of the backbone of another amino acid.

[0074] Reference also is made herein to peptides, polypeptides, proteins, and compositions comprising peptides, polypeptides, and proteins. As used herein, a polypeptide and / or protein is defined as a polymer of amino acids, typically of length>100 amino acids (Garrett & Grisham, Biochemistry, 2nd edition, 1999, Brooks / Cole, 110). A peptide is defined as a short polymer of amino acids, of a length typically of 20 or less amino acids, and more typically of a length of 12 or less amino acids (Garrett & Grisham, Biochemistry, 2nd edition, 1999, Brooks / Cole, 110).

[0075] The peptides, polypeptides, and proteins disclosed herein may be modified to include non-amino acid moieties. Modifications may include but are not limited to carboxylation (e.g.,Docket No. 10046-673W01

[0076] N-terminal carboxylation via addition of a di-carboxylic acid having 4-7 straight-chain or branched carbon atoms, such as glutaric acid, succinic acid, adipic acid, and 4,4-dimethylglutaric acid), amidation (e.g., C-terminal amidation via addition of an amide or substituted amide such as alkylamide or dialkylamide), PEGylation (e.g., N-terminal or C-terminal PEGylation via additional of polyethylene glycol), acylation (e.g., O-acylation (esters), N-acylation (amides), S-acylation (thioesters)), acetylation (e.g., the addition of an acetyl group, either at the N-terminus of the protein or at lysine residues), formylation lipoylation (e.g., attachment of a lipoate, a C8 functional group), myristoylation (e.g., attachment of myristate, a C14 saturated acid), palmitoylation (e.g., attachment of palmitate, a C16 saturated acid), alkylation (e.g., the addition of an alkyl group, such as an methyl at a lysine or arginine residue), isoprenylation or prenylation (e.g., the addition of an isoprenoid group such as farnesol or geranylgeraniol), amidation at C-terminus, glycosylation (e.g., the addition of a glycosyl group to either asparagine, hydroxylysine, serine, or threonine, resulting in a glycoprotein). Distinct from glycation, which is regarded as a nonenzymatic attachment of sugars, polysialylation (e.g., the addition of polysialic acid), glypiation (e.g., glycosylphosphatidylinositol (GPI) anchor formation, hydroxylation, iodination (e.g., of thyroid hormones), and phosphorylation (e.g., the addition of a phosphate group, usually to serine, tyrosine, threonine, or histidine).

[0077] The phrases “percent identity” and “% identity,” as applied to polypeptide sequences, refer to the percentage of residue matches between at least two polypeptide sequences aligned using a standardized algorithm. Methods of polypeptide sequence alignment are well-known. Some alignment methods consider conservative amino acid substitutions. Such conservative substitutions, explained in more detail above, generally preserve the charge and hydrophobicity at the site of substitution, thus preserving the structure (and therefore function) of the polypeptide. Percent identity for amino acid sequences may be determined as understood in the art. (See, e.g., U. S. Pat. No. 7,396,664, which is incorporated herein by reference in its entirety). A suite of commonly used and freely available sequence comparison algorithms is provided by the National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST) (Altschul, S. F. et al. (1990) J. Mol. Biol. 215:403 410), which is available from several sources, including the NCBI, Bethesda, Md., at its website. The BLAST software suite includes various sequence analysis programs including “blastp,” that is used to align a known amino acid sequence with other amino acids sequences from a variety of databases.

[0078] Percent identity may be measured over the length of an entire defined polypeptide sequence or may be measured over a shorter length, for example, over the length of a fragmentDocket No. 10046-673W01

[0079] taken from a larger, defined polypeptide sequence, for instance, a fragment of at least 15, at least 20, at least 30, at least 40, at least 50, at least 70 or at least 150 contiguous residues. Such lengths are exemplary only, and it is understood that any fragment length may be used to describe a length over which percentage identity may be measured.

[0080] The term “variant” means a polypeptide derived from a parent polypeptide by one or more (several) alteration(s), i.e., a substitution, insertion, and / or deletion, at one or more (several) positions. A substitution means a replacement of an amino acid occupying a position with a different amino acid; a deletion means removal of an amino acid occupying a position; and an insertion means adding 1 or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, preferably 1-3 amino acids immediately adjacent an amino acid occupying a position. In relation to substitutions, ‘immediately adjacent’ may be to the N-side (‘upstream’) or C-side (‘downstream’) of the amino acid occupying a position (‘the named amino acid’). Therefore, for an amino acid named / numbered ‘X,’ the insertion may be at position ‘X+l’ (‘downstream’) or at position ‘X-l’ (‘upstream’).

[0081] A “variant” of a particular polypeptide sequence may be defined as a polypeptide sequence having at least 50% sequence identity to the particular polypeptide sequence over a certain length of one of the polypeptide sequences using blastp with the “BLAST 2 Sequences” tool available at the National Center for Biotechnology Information's website. (See Tatiana A. Tatusova, Thomas L. Madden (1999), “Blast 2 sequences — a new tool for comparing protein and nucleotide sequences”, FEMS Microbiol Lett. 174:247-250). In some embodiments a variant polypeptide may show, for example, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or greater sequence identity over a certain defined length relative to a reference polypeptide.

[0082] A variant polypeptide may have substantially the same functional activity as a reference polypeptide. For example, a variant polypeptide may exhibit or more biological activities associated with binding a ligand and / or binding DNA at a specific binding site.

[0083] Variants comprising a fragment of a reference amino acid sequence or nucleotide sequence are contemplated herein. A “fragment” is a portion of an amino acid sequence or a nucleotide sequence which is identical in sequence to but shorter in length than the reference sequence. A fragment may comprise up to the entire length of the reference sequence, minus at least one nucleotide / amino acid residue. For example, a fragment may comprise from 5 to 1000 contiguous nucleotides or contiguous amino acid residues of a reference polynucleotide or reference polypeptide, respectively. In some embodiments, a fragment may comprise at leastDocket No. 10046-673W01

[0084] 5, 10, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 150, 250, or 500 contiguous nucleotides or contiguous amino acid residues of a reference polynucleotide or reference polypeptide, respectively. Fragments may be preferentially selected from certain regions of a molecule, for example the N-terminal region and / or the C -terminal region of a polypeptide or the '-terminal region and / or the 3' terminal region of a polynucleotide. The term “at least a fragment” encompasses the full length polynucleotide or full length polypeptide.

[0085] Variants comprising insertions or additions relative to a reference sequence are contemplated herein. The words “insertion” and “addition” refer to changes in an amino acid or nucleotide sequence resulting in the addition of one or more amino acid residues or nucleotides. An insertion or addition may refer to 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid residues or nucleotides.

[0086] An “enzyme” is a biological molecule, usually a protein or peptide, that acts under certain conditions, such as pH, temperature, and / or salt concentration, to accelerate biochemical reactions, either inside or outside of a tissue or cell. The molecules upon which enzymes initiate a reaction are called “substrates”, and the enzyme converts the substrates into different molecules called “products”. Enzyme functions are usually measured based the enzyme “activity” which refers to the amount of substrate converted into a product or products by the enzyme within a given amount of time.

[0087] As used herein, the term “polymerase” refers to an enzyme that synthesizes long chains of polymers or nucleic acids. DNA polymerase and RNA polymerase are used to assemble DNA and RNA molecules, respectively, by copying a DNA template strand using base-pairing interactions.

[0088] A “nucleic acid” is a chemical compound that serves as the primary informationcarrying molecules in cells and make up the cellular genetic material. Nucleic acids comprise nucleotides, which are the monomers made of a 5-carbon sugar (usually ribose or deoxyribose), a phosphate group, and a nitrogenous base. A nucleic acid can also be a deoxyribonucleic acid (DNA) or a ribonucleic acid (RNA). A chimeric nucleic acid comprises two or more of the same kind of nucleic acid fused together to form one compound comprising genetic material.

[0089] The terms “percent identity” and “% identity,” as applied to polynucleotide sequences, refer to the percentage of residue matches between at least two polynucleotide sequences aligned using a standardized algorithm. Such an algorithm may insert, in a standardized and reproducible way, gaps in the sequences being compared in order to optimize alignment between two sequences, and therefore achieve a more meaningful comparison of the twoDocket No. 10046-673W01

[0090] sequences. Percent identity for a nucleic acid sequence may be determined as understood in the art. (See, e.g., U. S. Pat. No. 7,396,664, which is incorporated herein by reference in its entirety). A suite of commonly used and freely available sequence comparison algorithms is provided by the National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST) (Altschul, S. F. et al. (1990) J. Mol. Biol. 215:403 410), which is available from several sources, including the NCBI, Bethesda, Md., at its website. The BLAST software suite includes various sequence analysis programs including “blastn,” that is used to align a known polynucleotide sequence with other polynucleotide sequences from a variety of databases. Also available is a tool called “BLAST 2 Sequences” that is used for direct pairwise comparison of two nucleotide sequences. “BLAST 2 Sequences” can be accessed and used interactively at the NCBI website. The “BLAST 2 Sequences” tool can be used for both blastn and blastp (discussed above).

[0091] Percent identity may be measured over the length of an entire defined polynucleotide sequence or may be measured over a shorter length, for example, over the length of a fragment taken from a larger, defined sequence, for instance, a fragment of at least 20, at least 30, at least 40, at least 50, at least 70, at least 100, or at least 200 contiguous nucleotides. Such lengths are exemplary only, and it is understood that any fragment length may be used to describe a length over which percentage identity may be measured.

[0092] A “full length” polynucleotide sequence is one containing at least a translation initiation codon (e.g., methionine) followed by an open reading frame and a translation termination codon. A “full length” polynucleotide sequence encodes a “full length” polypeptide sequence.

[0093] A “variant,” “mutant,” or “derivative” of a particular nucleic acid sequence may be defined as a nucleic acid sequence having at least 50% sequence identity to the particular nucleic acid sequence over a certain length of one of the nucleic acid sequences using blastn with the “BLAST 2 Sequences” tool available at the National Center for Biotechnology Information's website. (See Tatiana A. Tatusova, Thomas L. Madden (1999), “Blast 2 sequences — a new tool for comparing protein and nucleotide sequences”, FEMS Microbiol Lett. 174:247-250). In some embodiments a variant polynucleotide may show, for example, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or greater sequence identity over a certain defined length relative to a reference polynucleotide.

[0094] The term “mRNA” refers to messenger ribonucleic acid, or single stranded molecule of RNA that corresponds to the genetic sequence of a gene, and is translated by a ribosome in the process of synthesizing a protein. mRNA is created during the process of transcription, whereDocket No. 10046-673W01

[0095] a gene is converted into a primary transcript mRNA (or pre-mRNA). The primary transcript is further processed through RNA splicing to only contain regions that will encode protein. mRNA can also be targeted for epigenetic modifications, such as methylation, to impact mRNA translation, nuclear retention, nuclear export, processing, and splicing.

[0096] As used herein, “wild-type” refers to the genetic and physical characteristics of the typical form of a species as it occurs in nature. A wild-type or wild type characteristic is conceptualized as a product of the standard “normal” allele at a gene locus, in contrast to that produced by a non-standard “mutant” allele.

[0097] Enzymes Complexes

[0098] The present disclosure provides enzyme complexes and components thereof for enhanced transcription and translation. The present disclosure also provides engineered T7 RNA Polymerase (T7 RNAP) and mRNA capping enzyme (CE) to enhance transcriptional and post-transcriptional processes in eukaryotic systems, including but not limited to yeast cells. Unlike traditional methods relying on random mutagenesis and extensive screening, this approach uses ML models such as MutCompute to predict beneficial mutations across enzyme structures, enabling efficient identification of critical residues, including those outside the active site, that enhance enzyme functionality.

[0099] A) Engineered T7 RNAP

[0100] T7 RNA Polymerase (T7 RNAP) is an RNA polymerase derived from the T7 bacteriophage that catalyzes the formation of RNA from DNA in the 5’ to 3’ direction. The T7 RNAP is promoter-specific, wherein it only transcribes double stranded DNA downstream of a T7 promoter. As used herein, “downstream” refers to a direction of transcription, the direction of transcription being from a promoter sequence to a RNA-encoding sequence. For a template strand of a double-stranded DNA molecule, the direction of transcription is 3’ to 5’. For a non¬ template strand of the double-stranded DNA molecule, the direction of transcription is 5’ to 3’.

[0101] In some aspects, disclosed herein is an engineered T7 RNA polymerase (T7 RNAP), wherein said 1’7 RNAP comprises SEQ ID NO: 15, and further wherein SEQ ID NO: 15 comprises at least one substitution mutation selected from Q58S, Q404L, A584K, V609S, W698F, and / or Q786L, or any combination thereof. It should be noted that the substitution mutations Q58S, Q404L, A584K, V609S, W698F, and Q786L refer to the amino acids found in the T7 RNAP component of the fusion polypeptide.

[0102] In some embodiments, the T7 RNAP comprises at least 70% sequence identity to SEQ ID NO: 15. In some embodiments, the T7 RNAP comprises 70%, 75%, 80%, 85%, 90%, 95%,Docket No. 10046-673W01

[0103] 98%, 99%, 100%, or any percentage in-between the mentioned percentages sequence identity to SEQ ID NO: 15.

[0104] In some embodiments, SEQ ID NO: 15 comprises at least one additional mutation other than Q58S, Q404L, A584K, V609S, W698F, and / or Q786L. In some embodiments, the T7 RNAP is operably linked to a capping enzyme comprising SEQ ID NO: 11, or a variant thereof. In some embodiments, the variant of SEQ ID NO: 11 comprises at least one substitution mutation of S199A, N267W, and / or Q426V.

[0105] In some embodiments, the T7 RNAP comprises a nuclear localization sequence (NLS ). As used herein, a “NLS” refers to a short amino acid sequence (roughly 4-15 amino acids in length), generally comprising positively charged amino acids (i.e., lysine (K) and arginine (R)), that labels proteins for import into the nucleus of a cell. In some embodiments, the NLS comprises at least 90% sequence identity to SEQ ID NO: 1. In some embodiments, the NLS comprises 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 1.

[0106] B) Engineered Capase

[0107] An mRNA capping enzyme (CE) is an enzyme that catalyzes the attachment of a 5’ cap to the messenger RNA (mRNA) after said mRNA is synthesized / transcribed in the nucleus. The addition of the 5’ cap occurs co-transcriptionally after the mRNA has grown at least 20 nucleotides long. The capping of mRNA is one of three modifications (5’ capping, splicing, and 3’ polyadenylation) that occur prior to the mRNA becomes a mature molecule that exits the nucleus to be translated into a peptide, polypeptide, or protein. In some embodiments, the capping enzymes adds a 7-methylguanosine cap to the 5’ end of the mRNA transcript.

[0108] In some aspects, disclosed herein is a capping enzyme, wherein the capping enzyme comprises SEQ ID NO: 11, and wherein SEQ ID NO: 11 comprises at least one substitution mutation comprising S199A, N267W, and / or Q426V. It should be noted that the substitution mutations S199A, N267W, and Q426V refer to the amino acids found in the capping enzyme component of the fusion polypeptide.

[0109] It should be understood that the terms “capase” and “capping enzyme” can be used interchangeably throughout the present disclosure.

[0110] In some embodiments, the capping enzyme comprises at least 70% sequence identity to SEQ ID NO: 11. In some embodiments, the capping enzyme comprises 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any percentage in-between the mentioned percentages sequence identity to SEQ ID NO: 11.Docket No. 10046-673W01

[0111] In some embodiments, the engineered capping enzyme is operably linked to a T7 RNA polymerase comprising SEQ ID NO: 15, or a variant thereof. In some embodiments, the variant of SEQ ID NO: 15 comprises at least one substitution comprising Q58S, Q404L, A584K, V609S, W698F, and / or Q786L.

[0112] C) Enzyme Complexes and Systems Thereof

[0113] “Enzyme complex” as used herein refers to any system, arrangement, assembly, combination, construct, composition, or functional association comprising (a) a T7 RNA polymerase, or a variant, fragment, mutant, homolog, derivative, or functional equivalent thereof, and (b) a capping enzyme, or a variant, fragment, mutant, homolog, derivative, or functional equivalent thereof, wherein the T7 RNA polymerase and the capping enzyme are configured such that RNA synthesized by the polymerase is capable of being capped by the capping enzyme during, immediately following, or in connection with transcription. The term “enzyme complex” is intended to be construed broadly and is not limited to a particular structural format, physical linkage, mode of expression, mode of assembly, or spatial arrangement, so long as the T7 RNA polymerase and the capping enzyme are present in a manner that permits coordinated or functionally coupled transcription and capping activity.

[0114] It should be understood that term “enzyme complex” may be used interchangeably with “fusion peptide”, “fusion polypeptide”, “chimeric peptide”, or “chimeric polypeptide” throughout the present disclosure. For instance, the enzyme complex could also be considered a fusion polypeptide or a chimeric peptide comprising the T7 RNAP and capping enzyme operably linked together using a linker peptide.

[0115] In some aspects, disclosed herein is an enzyme complex comprising a T7 RNA polymerase (T7 RNAP) component and a capping enzyme component, wherein the T7 RNAP component comprises SEQ ID NO: 15, or a variant thereof, and further wherein the capping component comprises SEQ ID NO: 11, or a variant thereof, wherein a linker comprising at least 70% sequence identity to SEQ ID NO: 2 separates the T7 RNAP component and the capping enzyme component, wherein the variant of SEQ ID NO: 15 comprises at least one mutation selected from Q58S, Q404L, A584K, V609S, W698F, and / or Q786L, and wherein the variant of SEQ ID NO: 11 comprises at least one mutation selected from S199A, N267W, and / or Q426V.

[0116] Non-limiting examples of enzyme complexes used herein include a complex comprising a T7 RNAP component with a Q58S substitution mutation and a capping enzyme component with a S199A substitution mutation; a T7 RNAP component with a Q404L substitution mutation and a capping enzyme with a N267W substitution mutation; and a T7Docket No. 10046-673W01

[0117] RNAP component with a A584K substitution mutation and a capping enzyme with a Q426V substitution mutation. Thus, in some embodiments the enzyme complex can contain at least one substitution mutation disclosed herein. In some embodiments, the enzyme complex can contain any combination of substitution mutations disclosed herein. In some embodiments, the enzyme complex can also contain all of the substitution mutations disclosed herein. In some embodiments, the enzyme complex can also not contain any mutations.

[0118] In some embodiments, the at least one substitution mutation of any preceding aspect confers one or more improved properties in the T7 RNA polymerase compared to a T7 RNA polymerase without a substitution mutation. In some embodiments, the one or more improved properties comprise increased fidelity of replication, increased RNA yield, increased protein expression, increased thermoactivity, increased pseudouridine incorporation, increased incorporation of other RNA base analogs, increased thermostability, increased pH activity, increased stability, increased enzymatic activity, increased substrate specificity or affinity, increased specific activity, increased resistance to substrate or end-product inhibition, increased chemical stability, improved solvent stability, increased tolerance to acidic or basic pH, increased tolerance to proteolytic activity, reduced aggregation, increased solubility, altered temperature profile, and / or less abortive transcripts.

[0119] In some embodiments, the enzyme complex may comprise a directly linked complex, including but not limited to a fusion protein, chimeric polypeptide, multidomain protein, covalently linked construct, chemically conjugated construct, or recombinant polypeptide in which the T7 RNA polymerase and the capping enzyme are joined directly or indirectly, optionally through a linker, spacer, peptide sequence, affinity interaction domain, protease-cleavable region, flexible linker, rigid linker, or other connecting moiety.

[0120] In some embodiments, the enzyme complex may comprise an indirectly associated complex in which the T7 RNA polymerase and the capping enzyme are separate molecular entities but are associated through non-covalent interaction, affinity pairing, scaffold-mediated interaction, adaptor-mediated interaction, co-localization tag, binding domain interaction, encapsulation, immobilization on a common support, or co-compartmentalization within a shared reaction environment.

[0121] In certain embodiments, the enzyme complex may comprise a cell -free complex or reaction complex in which the T7 RNA polymerase and the capping enzyme are separately produced, purified, synthesized, or obtained, and then combined in a common reaction mixture under conditions that permit transcription and capping to occur in a coordinated manner. InDocket No. 10046-673W01

[0122] such embodiments, the two enzymes need not be physically bound to one another, provided they function together in the same system to generate capped RNA.

[0123] Accordingly, the term “enzyme complex” includes, without limitation: (1) a fusion protein comprising T7 RNA polymerase and a capping enzyme; (2) separate T7 RNA polymerase and capping enzyme polypeptides that physically associate; (3) separate T7 RNA polymerase and capping enzyme polypeptides that do not physically associate in a stable manner but function together in the same reaction or cellular environment; (4) expression systems encoding both enzymes on a single nucleic acid construct or on multiple nucleic acid constructs;

[0124] (5) enzymes tethered through scaffolds, tags, adaptors, nanoparticles, beads, solid supports, membranes, vesicles, droplets, condensates, or compartments; and (6) any other arrangement that enables transcription by the T7 RNA polymerase and capping by the capping enzyme in a coordinated or functionally coupled manner.

[0125] The enzyme complex may be pre-assembled, self-assembled, co-translated, coexpressed, post-translationally assembled, chemically assembled, reconstituted, or formed in situ.

[0126] The term also encompasses complexes in which the T7 RNA polymerase and the capping enzyme are present in stoichiometric or non-stoichiometric amounts, in fixed or variable ratios, and in transient, stable, reversible, irreversible, constitutive, inducible, or condition-dependent association.

[0127] In some embodiments, the enzyme complex is a component in a transcription system, a post-transcription system, or a combination thereof. In some embodiments, the posttranscription system comprises protein translation. In some embodiments, the enzyme complex of any preceding aspect is active in the nucleus of a cell (i.e. transcription). In some embodiments, the enzyme complex of any preceding aspect contributes to protein translation in the cytoplasm.

[0128] In some embodiments, the linker of any preceding aspect comprises 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any percentage value in-between thereof sequence identity to SEQ ID NO: 2.

[0129] In some aspects, disclosed herein is a polynucleotide sequence encoding the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any preceding aspect.

[0130] In some embodiments, the polynucleotide sequence of any preceding aspect further encodes a linker having at least 70% sequence identity (include but not limited to 70%, 75%,Docket No. 10046-673W01

[0131] 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any percentage value in-between thereof) to SEQ ID NO: 2.

[0132] In some embodiments, the polynucleotide sequence comprises deoxyribonucleic acid (DNA). In some embodiments, the polynucleotide sequence comprises ribonucleic acid (RNA). In some embodiments, the polynucleotide sequence is inserted into a vector or expression vector.

[0133] In some embodiments, the enzyme complex may comprise a co-expressed system in which the T7 RNA polymerase and the capping enzyme are encoded by one or more nucleic acid constructs. Such constructs may include a single vector, multiple vectors, a polycistronic construct, separate expression cassetes on the same vector, separate vectors introduced into the same host cell, or any other genetic arrangement that results in expression of both activities in a shared system. Thus, the enzyme complex may include embodiments in which the polymerase and capping enzyme are encoded on the same vector, on different vectors, under the same promoter, under different promoters, or otherwise arranged to permit concurrent or sequential production and functional cooperation.

[0134] As used herein, a “vector” or an “expression vector” refers to any vehicle that carries a polynucleotide into a cell for the expression of the polynucleotide in the cell. Non-limiting examples of an expression vector includes a plasmid, a virus, a phage particle, or a nanoparticle. Said expression vectors are capable of extrachromosomal replication or, optionally, can integrate into tire host genome. As used herein, the term "integrated" used in reference to an expression vector (e.g., a plasmid or viral vector) means the expression vector, or a portion thereof, is incorporated (physically inserted or ligated) into the chromosomal DNA of a host cell. As used herein, a “viral vector” refers to a virus-like particle containing genetic material which can be introduced into a eukaryotic cell without causing substantial pathogenic effects to the eukaryotic cell. A wide range of viruses or viral vectors can be used for transduction but should be compatible with the cell type the virus or viral vector are transduced into (e.g., low toxicity, capability to enter cells). Suitable viruses and viral vectors include adenovirus, lentivirus, retrovirus, among others. In some embodiments, the expression vector encoding a chimeric polypeptide is a naked DNA or is comprised in a nanoparticle (e.g., liposomal vesicle, porous silicon nanoparticle, gold-DNA conjugate particle, polyethyleneimine polymer particle, cationic peptides, etc.).

[0135] In some embodiments, the expression vector of any preceding aspect further comprises one or more regulatory genes to optimize expression of the enzyme complex. Non-limiting example of the one or more regulatory genes include promoters (e.g. T7 promoters),Docket No. 10046-673W01

[0136] operators / repressors, ribosome binding sites, enhancers / silencers, and transcription terminators.

[0137] In some aspects, disclosed herein is an in vitro system comprising one or more polynucleotide sequences encoding the enzyme complex of any preceding aspect and a reporter gene, wherein the polynucleotide sequence is under the control of a T7 RNAP promoter.

[0138] In some aspects, disclosed herein is an in vitro system comprising one or more expression vectors comprising the polynucleotide sequence of any preceding aspect.

[0139] In some embodiments, tire in vitro system comprises one, two, three, or more polynucleotide sequences. In some embodiments, each component of the enzyme complex (i.e. the T7 RNAP, the capping enzyme, or the linker) are encoded on a separate polynucleotide sequence. In some embodiments, the whole enzyme complex (i.e. the T7 RNAP, the capping enzyme, and the linker) are encoded on the same polynucleotide sequence. In some embodiments, any combination of enzyme complex components (i.e. the T7 RNAP, the capping enzyme, or the linker) are encoded on the same polynucleotide sequence.

[0140] In some embodiments, each component of the enzyme complex (i.e. the T7 RNAP, the capping enzyme, or the linker) is encoded with at least one reporter protein on a polynucleotide sequence. In some embodiments, the whole enzyme complex (i.e. the T7 RNAP, the capping enzyme, or the linker) and at least one reporter protein are encoded on the same polynucleotide sequence. In some embodiments, any combination of enzyme complex components (i.e. the T7 RNAP, the capping enzyme, or the linker) and reporter proteins are encoded on the same polynucleotide sequence.

[0141] In some embodiments, the reporter gene is a polynucleotide encoding at least one reporter protein including, but not limited to luciferase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), blue fluorescent protein (BFP), cyan fluorescent protein (CFP), monomeric red fluorescent protein (mRFP), Discosoma striata (DsRed), mCherry, mOrange, tdTomato, mSTrawberry, mPlum, photoactivatable GFP (PA-GFP), Venus, Kaede, monomeric kusabira orange (mKO), Dronpa, enhanced CFP (ECFP), Emerald, Cyan fluorescent protein for energy transfer (CyPet), super CFP (SCFP), Cerulean, photoswitchable CFP (PS-CFP2), photoactivatable RFP1 (PA-RFP1), photoactivatable mCherry (PA-mCherry), monomeric teal fluorescent protein (mTFPl), Eos fluorescent protein (EosFP), Dendra, TagBFP, TagRFP, enhanced YFP (EYFP), Topaz, Citrine, yellow fluorescent protein for energy transfer (YPet), super YFP (SYFP), enhanced GFP (EGFP), Superfolder GFP, T-Sapphire, Fucci, mK02, mOrange2, mApple, Sirius, Azurite, EBFP, and / or EBFP2.Docket No. 10046-673W01

[0142] A non-limiting example of the functionality of the system includes synthesis of a GFP, wherein the system of any aspect disclosed herein expresses, produces, and / or makes the T7 RNAP, then the T7 RNAP synthesizes the GFP mRN A, and then said system further expresses, produces, and / or makes the GFP. In some embodiments, the system can operate within a host cell including but not limited to bacterial cells and / or yeast cells.

[0143] In some embodiments, said system is an in vitro transcription system. In some embodiments, said system is an in vitro post-transcription system. In some embodiments, said system is an in vitro transcription and post-transcription system. In some embodiments, the post-transcription system is a translation system.

[0144] D) Cell Compositions

[0145] In some aspects, disclosed herein is a cell comprising the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any preceding aspect.

[0146] In some aspects, disclosed herein is a cell comprising a polynucleotide sequence encoding the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any preceding aspect.

[0147] In some embodiments, the cell is a eukaryotic cell or a prokaryotic cell. Non-limiting examples of a eukaryotic cell includes an animal cell (such as for example neurons, myocytes, erythrocytes, leukocytes, hepatocytes, adipocytes, gametes, and osteocytes), plant cells (such as for example parenchyma cells, collenchyma cells, xylem cells, phloem cells, epidermal cells, and sclerenchyma cells), unicellular or multicellular fungal cells (such as for example yeast and hyphae, and protists (Protozoa / algae) (such as for example Amoeba, Paramecium, and Euglena. Non-limiting examples of a prokaryotic cell includes bacteria (such as for example gram-positive bacteria or gram-negative bacteria) and archaea.

[0148] In some embodiments, the cell proliferates into a stable cell line or a transient cell line. As used herein, a “stable cell line” refers to a population of cultured cells that have been genetically modified with one or more transgenes integrated into the cell’s genome and is expressed reliably over several cell generations. Said transgene is heritable and maintained during cell division, therefore these cells show consistent, long-term expression of the desired protein (i.e. a T7:capping enzyme complex). In some embodiments, the stable cell line expresses the enzyme complex of any preceding aspect for more than seven days. In some embodiments, the stable cell line expresses the enzyme complex for 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 days, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 months, or 1, 2, 3, 4, 5, or more years.Docket No. 10046-673W01

[0149] As used herein, a “transient cell line” refers to a population of cells that has been temporarily transfected with one or more transgenes so that said transgene(s) is only expressed for a short time and is not integrated into the cell’s genome. In some embodiments, the transient cell line expresses the enzyme complex of any preceding aspect for about 6 weeks or less. In some embodiments, the transient cell line expresses the enzyme complex of any preceding aspect for about 3 weeks to about 6 weeks. In some embodiments, the transient cell line expresses the enzyme complex of any preceding aspect for about 1, 2, 3, 4, 5, 6, 7 days, or 1, 2, 3, 4, 5, or 6 weeks.

[0150] In some aspects, disclosed herein is a cell culture composition comprising the cell of any preceding aspect, wherein said cell includes but is not limited to a tissue culture cell. Nonlimiting examples of tissue culture cells includes human cervical epithelial cells (such as for example HeLa cells), human kidney cells (such as for example HEK293T cells, or variants thereof), Chinese hamster ovarian cells (CHO cells), and human T-cell leukemia cells (such as for example Jurkat cells).

[0151] In some embodiments, the cell culture composition further comprises a buffer or media composition. In some embodiments, the buffer includes, but is not limited to a HEPES buffer and a phosphate buffered saline (PBS). In some embodiments, the medium composition includes, but is not limited to a basal media (such as for example MEM, DMEM, RPMI), a reduced-serum media, and a serum-free / chemically defined media. In some embodiments, the cell culture composition can also contain glycerol or dimethyl sulfoxide (DMSO) as a cryoprotective agent, wherein the glycerol or DMSO is present at a final concentration of about 5%, 7.5%, 10%, 15%, or 20%, depending on the cell type.

[0152] Methods of Use

[0153] The biological process of transcription involves copying segment of a DNA molecule into an RNA molecule for the purpose of gene expression. Some segments of the DNA molecule are transcribed into RNA molecules that can encode proteins, referred to as mRNA. Other segments of DNA are transcribed into RNA molecules called non-coding RNAs (ncRNAs). Transcription is divided into 1) initiation, 2) promoter escape, 3) elongation, and 4) termination. During initiation, transcription begins with the RNA polymerase, such as the T7 RNAP, and one or more transcription factors (TFs) binding to a DNA promoter sequence. During elongation, in which nucleotides (adenine, guanosine, cytosine, and uracil) are added to make an RNA molecule, a CE adds a 7-methylguanosine cap to the 5’ end of the mRNA transcript.Docket No. 10046-673W01

[0154] Current challenges in synthetic biology and gene expression technologies include, but are not limited to inefficient transcriptional and post-transcriptional processes for eukaryotic systems. Naturally-occurring T7 RNAP are highly efficient in prokaryotic systems, however it struggles to meet the demands of eukaryotic gene expression, mRNA stability, and efficient translation. Similarly, naturally occurring CEs often fail to achieve optimal capping efficiency leading to mRNA instability and inefficient translation. Thus, the present disclosure enhances the T7 RNAP and CE, which leads to improved gene expression and / or other mechanisms involved in gene expression. The present disclosure also provides methods of improving transcriptional and post-transcriptional processes in eukaryotic systems using the 1'7 RNAP, the CE, or enzyme complexes of any aspect disclosed herein.

[0155] In some aspects, disclosed herein is a method of producing a capped RNA transcript, the method comprising performing a transcription reaction utilizing the enzyme complex of any preceding aspect, wherein the method further comprises providing a composition comprising a) one or more components needed to produce a methylguanyl ate cap; and b) a DNA template, wherein said DNA template is exposed to said composition under conditions such that the enzyme complex produces a capped RNA transcript.

[0156] In some embodiments, the capped RNA transcript includes but is not limited to a capped mRNA transcript.

[0157] In some embodiments, the enzyme complex of any preceding aspect converts the DNA template into a RNA transcript. In some embodiments, the one or more components needed to produce the methylguanylate cap (or 5" RNA cap) includes but is not limited to guanosine triphosphate (GTP), S-adenosyl-L-methionine (SAM), a salt (such as for example potassium chloride (KC1) or sodium chloride (NaCl)), a divalent cation (such as for example magnesium (Mg2+)), and a reducing agent (such as for example ditliiotlireitol (DTT)).

[0158] In some aspects, disclosed herein is a method of producing RNA, the method comprising a) expressing the enzyme complex of any preceding aspect in a cell, and b) isolating an amount of RNA from the cell that is increased relative to an otherwise identical control cell expressing the T7 RNA polymerase lacking the capping enzyme.

[0159] Generating RNA within a cell by using the enzyme complex described herein results in a measurable elevation in the amount of RNA obtained from the cell. This increase is relative to an appropriate control cell (e.g., an otherwise identical cell expressing T7 RNA polymerase in the absence of the capping enzyme). In certain embodiments, the increase is at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 75%, 100%, 150%, 200%, or any value in¬ between the percentages mentioned herein, relative to the control. Such increase may ariseDocket No. 10046-673W01

[0160] from one or more contributing factors, including but not limited to enhanced accumulation, improved stability, reduced degradation, and / or improved recovery of RNA. For example, enhanced accumulation may be at least about 10%, 20%, 50%, or 100%; improved stability may be at least about 10%, 20%, 50%, or 100%; reduced degradation may be at least about 10%, 20%, 50%, or 100% relative to the control; and improved recovery may be at least about 5%, 10%, 20%, 50%, or 100%. In certain embodiments, the RNA produced in the presence of the enzyme complex exhibits increased resistance to cellular nucleases, resulting in higher steady-state levels of RNA within the cell. In other embodiments, the RNA, including capped RNA species, demonstrates improved integrity (e.g., increased proportion of full-length transcripts) and / or enhanced recoverability during isolation procedures, thereby yielding a greater amount of RNA upon extraction. The increase may be assessed using any suitable quantitative or semi -quantitative method known in the art, including but not limited to spectrophotometric measurement, electrophoretic analysis, quantitative PCR, or sequencing¬ based approaches. Accordingly, “increased” does not require that transcriptional activity per se is elevated, but rather encompasses any condition in which a greater amount of RNA is present in or obtainable from the cell relative to the control.

[0161] In some embodiments, the method of any preceding aspect involves mass producing RNA and / or an amino acid sequence of any preceding aspect. In some embodiments, the method of any preceding aspect involves mass producing a therapeutic agent comprising either the capped mRNA, the capped RNA transcript, the RNA, or the amino acid sequence of any- preceding aspect. In some embodiments, the capped RNA transcript, the RNA, or the amino acid of any preceding aspect is a therapeutic agent. In some embodiments, the therapeutic agent of any preceding aspect includes, but is not limited to an immunotherapeutic agent, a vaccine (such as, for example an mRNA vaccine), an antibiotic, an antiviral, or any other therapeutic agents thereof.

[0162] In some embodiments, the capped RNA transcript or the RNA is further translated into a peptide. In some embodiments, capped RNA transcript or RNA is further translated into an amino acid sequence. In some embodiments, the amino acid sequence is further processed via protein folding and / or post-translation modifications to produce a polypeptide or protein.

[0163] In some embodiments, the method of any preceding aspect confers one or more improved properties in the engineered capping enzyme compared to a capping enzyme without a substitution mutation. In some embodiments, the one or more improved properties comprise improved selectivity for capping, improved processivity for capping, improved capping enzymatic activity, improved protein expression, improved RNA yield, improved stability in aDocket No. 10046-673W01

[0164] storage buffer, improved stability under reaction conditions, improved processivity of translation, improved thermostability, and / or improved transcription fidelity. In some embodiments, the improved capping enzymatic activity comprises improved RNA triphosphatase activity and / or improved RNA methyltransferase activity.

[0165] In some embodiments, the method of any preceding aspect comprises optimizing gene expression. In some embodiments, the method of any preceding aspect comprises producing stable and efficient mRNA. In some embodiments, the method of any preceding aspect is used for mRNA-based therapeutics and / or vaccine development.

[0166] In some embodiments, the transcription reaction is performed using an in vitro transcription system. In some embodiments, the translation reaction is performed using an in vitro translation system.

[0167] The present disclosure overcomes the current challenges in the field by employing machine learning-guided approaches to engineer T7 RNAP and mRNA capping enzymes. The T7 RNAP, CE, and variants thereof, were developed to enhance transcriptional activity and mRNA capping efficiency, addressing major bottlenecks in the gene expression pathway. Their combinatorial use resulted in a 12.68-fold increase in ZsGreen fluorescence in a yeast-based system, demonstrating a synergistic improvement in overall gene expression efficiency. However, ZsGreen fluorescence serves as a proxy, reflecting combined transcriptional and translational processes.

[0168] Furthermore, the present disclosure overcomes the limitations of traditional directed evolution methods by leveraging machine learning to efficiently identify beneficial mutations. This approach reduces the time and resources required for optimization while enabling exploration of mutations across enzyme structures. The yeast-based validation system provides a scalable and cost-effective alternative to resource-intensive mammalian systems, broadening the applicability to synthetic biology research, industrial RNA production, and potential therapeutic applications.

[0169] The present disclosure possesses several advantages over current technologies in the engineering of T7 RNA polymerase (T7 RNAP) and mRNA capping enzymes (CEs). By employing machine learning tools such as MutCompute, it enables precise prediction of beneficial mutations, facilitating efficient exploration of sequence spaces and faster optimization compared to traditional directed evolution, which relies on random mutagenesis and extensive screening. This approach reduces time and resource demands while uncovering mutations that enhance enzyme performance.Docket No. 10046-673W01

[0170] The T7 RNAP, CE, and variants thereof show substantial improvements in gene expression efficiency as measured by ZsGreen fluorescence, with a 6.2-fold increase for EvoT7 and a 3-fold enhancement for EvoBMCE. These improvements show enhanced transcriptional activity and mRNA capping efficiency, though ZsGreen fluorescence reflects the combined outcomes of transcription and translation processes. Wien used together, the T7 RNAP and CE of any preceding aspect, produce a synergistic effect, resulting in a 12.7-fold improvement in ZsGreen fluorescence in a yeast-based system. This dual optimization addresses multiple bottlenecks in transcriptional and post-transcriptional processes, representing a significant advancement over existing technologies that typically target individual steps in the gene expression pathway.

[0171] The present disclosure further benefits from a yeast-based validation system, which provides a cost-effective, scalable, and high-throughput platform for testing. Unlike resource¬ intensive mammalian models commonly used in current technologies, this system offers flexibility for application in non-model organisms and experimental environments. Additionally, the modular and robust design of the invention ensures broad applicability across diverse biological contexts, including synthetic biology research, industrial-scale RNA production, and therapeutic applications.

[0172] A number of embodiments of the disclosure have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the invention. Accordingly, other embodiments are within the scope of the following claims.

[0173] By way of non-limiting illustration, examples of certain embodiments of the present disclosure are given below.

[0174] EXAMPLES

[0175] The following examples are set forth below to illustrate the compositions, devices, methods, and results according to the disclosed subject matter. These examples are not intended to be inclusive of all aspects of the subject matter disclosed herein, but rather to illustrate representative methods and results. These examples are not intended to exclude equivalents and variations of the present invention which are apparent to one skilled in the ait.

[0176] Example 1: Machine Learning-Guided Engineering of T7 RNA Polymerase and mRNA Capping Enzymes for Enhanced Gene Expression in Eukaryotic SystemsDocket No. 10046-673W01

[0177] The integration of synthetic biology tools into eukaryotic systems offers both significant opportunities and challenges, particularly in optimizing transcriptional and post- transcriptional processes. T7 RNA polymerase (T7 RNAP) and mRN A capping enzymes (CEs) have been fused to enable eukaryotic mRNA production within a single construct. However, the activity of the fusion construct between the African Swine Fever Virus capping enzyme (ASFVCE) and T7 RNAP was relatively low. To address this, the Brazilian Marseillevirus capping enzyme (BMCE) was fused to T'7 RNAP, and a machine learning pipeline was developed to engineer greatly improved fusion variants. This approach enabled the additive integration of nine predicted single substitutions that improved gene expression in yeast, thereby generating fusion polymerases that exhibited over 10-fold improvements in gene expression efficiency relative to the original ASFVCE: T7 RNAP fusion enzyme. Not only were machine learning substitutions additive for gene expression, but they could also be further combined with variants identified via directed evolution for even higher activities. By allowing ML predictions to guide experimental validations the sequence landscape for enzyme optimization could rapidly be explored, achieving superior results even when compared to directed evolution. The improved enzymes herein can impact numerous synthetic biology applications, including metabolic engineering, mRNA therapeutics, and cell free systems.

[0178] While T7 RNAP has proven highly effective in prokaryotic systems, its adaptation for eukaryotic expression presents substantial challenges. Unlike prokaryotic systems, eukaryotic cells require complex post-transcriptional modifications, such as the addition of 5'-m7G caps and poly(A) tails, to stabilize mRNA and promote efficient translation. These additional requirements, coupled with processes like splicing and nuclear export, necessitate innovative approaches to expand T7 RNAP's functionality in eukaryotic contexts. A promising solution has been the fusion of T7 RNAP with viral capping enzymes, such as the African Swine Fever Virus capping enzyme (ASFVCE), to facilitate co-transcriptional capping. This approach has also been extended to the Faustovirus capping enzyme (FCE), where its fusion with T7 RNAP demonstrated up to 90% Cap-1 incorporation, significantly simplifying mRNA synthesis and enabling efficient production for therapeutic applications, including mRNA vaccines and protein therapeutics.

[0179] Despite producing abundant mRNA, these fusion constructs often fail to achieve the anticipated levels of protein expression, underscoring the need for further optimization. Efforts to overcome these limitations have included engineering poly(A) polymerases to extend mRNA tails, resulting in moderate increases in protein yield. Additionally, directed evolutionDocket No. 10046-673W01

[0180] of single-chain T7 RNAP variants fused with ASFV capping enzymes has enhanced activity by up to four-fold in mammalian cells.

[0181] The advent of machine learning (ML)-based protein engineering provides opportunities for the rapid exploration of very large sequence spaces, even relative to methods like directed evolution, which is sometimes constrained by the limitations of local optima in complex fitness landscapes. For example, EVOLVEpro has utilized Al-driven in silico evolution to engineer T7 RNAP variants optimized for in vitro applications, achieving enhanced RNA yield, improved mRNA translation efficiency, and reduced immunogenicity. In parallel, Rosetta¬ based structural modeling was employed to assess the impact of amino acid substitutions on T7 RNAP stability and function. This approach identified the G47A + 884G variant, which significantly reduced immunostimulatory double-stranded RNA (dsRNA) formation by altering RNAP-RNA interactions, thereby lowering cytokine responses in mammalian systems and streamlining mRNA purification.

[0182] Directed evolution has previously been used to identify improvements to a fusion between T7 RNAP and a capase enzyme that improved gene expression in yeast, ultimately identifying an ASFVCE(443): T7 RNAP(443) variant that shows a 5-fold improvement in activity relative to the original wild-type fusion combination. Now, using a NIL framework that integrated structure-based (MutCompute), stability-based (Stability Oracle), and evolutionbased (MutRank) predictions, a unique set of sequence substitutions relative to other algorithms have been generated. Experimental validation of improvements in gene expression in yeast led to additive combinations that improved overall activity, allowing efficient and rapid enzyme optimization. In consequence, a variant, EvoBMCE: EvoT7, has now been identified that contains nine targeted substitutions and exhibits over a 10-fold improvement in activity in yeast. These mutations can be combined with those previously identified through directed evolution, yielding a polymerase fusion (EvoBMCE: EvoT7(443)) that demonstrated an almost 13-fold improvement in activity.

[0183] Results

[0184] Engineering of T7 RNA polymerase: capase based on structurally-aware machine learning predictions

[0185] Machine learning methodologies offer transformative opportunities to explore protein sequence space more deeply, enabling the identification of key residues critical for protein function, even when located outside the traditional active site. Many of these approaches model the relationship between protein structure and function, predicting the effects of mutations onDocket No. 10046-673W01

[0186] protein activity or stability. Herein, MutConipute, a structure-informed machine learning algorithm for protein engineering was employed to optimize T7 RNAP for improved function in yeast. Both MutCompuie, a convolutional neural network, and MutComputeX, a residual neural network, were applied to analyze multiple structural conformations from the Protein Data Bank (PDB), incorporating both initiation complex (IC) and elongation complex (EC) states (PDB IDs: 1QLN, 1CEZ, 2PI4, 2PI5 for IC; 1S0V, 1S76, 1S77, 1H38, and 1MSW for EC). To mitigate biases introduced by multiple structural representations within a single PDB file, predictions were averaged across all chains. There were attempts to identify mutations that were predicted across multiple different structure files, including both the IC and the EC, and point mutants were selected based on high log ratio probability comparisons relative to the wild-type residue. To further improve the robustness of the predictions, it was evaluated whether excluding charge and solvent-accessible surface area (SASA) as input features would impact the stability-focused predictions of 1'7 RNAP. This led to the development of MutComputeX.2, which generated an additional set of 20 predicted variants without these input features.

[0187] These predictions ultimately led to 60 single variants spanning 56 positions (20 from MutConipute, 20 from MutComputeX, and 20 from MutComputeX.2), with four positions (L534, V710, L853, and H854) having two, different predicted substitutions. These variants were then incorporated into a previously studied fusion protein consisting of T7 RNAP and the African Swine Fever Virus (ASFV) NP868R mRNA capping enzyme (Table 2). Interestingly, MutCompute primarily predicted mutations in the IC (14 out of 20), whereas MutComputeX predictions were more balanced between the IC (9 out of 20) and EC (11 out of 20). It is possible that MutComputeX, which employs a residual neural network, may have captured higher-order structural dependencies that are relevant to both IC and EC states.

[0188] The ASFV NP868R capase was initially chosen based on its demonstrated ability to enhance protein expression in both mammalian systems and in yeast. The fusion domains were linked by a GS linker, containing an SV40 nuclear' localization signal (N S), and were expressed under the control of the yeast GAL promoter (Figure 1A). Fusion proteins carrying the T7 RNAP variants were integrated into the HO locus of the yeast chromosome. To determine the activity of the fusion proteins in the AGal2 derivative yeast strain, induction was carried out with 5% D-Galactose. Although 2% D-Galactose is commonly used for induction, 5% is used with the Gal2 transporter knockout to better control fusion protein expression. Following induction, the variants transcribed the ZsGreen reporter gene under the control ofDocket No. 10046-673W01

[0189] the T7 RNAP promoter on an episomal reporter plasmid (pT7P-ZsGreen; Figures 1 A and IB). A null variant, the substitution Y639A, was used to validate that the system (Figure 1C).

[0190] Out of the 60 single variants tested, five (Q58S, S128V, A584K, V609S, and W698F) led to up to 1.5-fold increases in expression of the ZsGreen reporter gene compared to wild¬ type (WT) 1'7 RNAP (Figure 1C). These beneficial mutations all originated from different sources: V609S from MutCompute; Q58S, S128V, and W698F from MuiComputeX; and A584K from MutComputeX.2. Then, the five beneficial single variants (Q58S, S128V, A584K, V609S, and W698F) were systematically combined to generate all possible double, triple, quadruple, and quintuple combinations, resulting in 26 unique multi-mutant variants for further analysis (Figure ID). In general, the combinations showed increasing activity, with the quadruple variant T7 RNAPQ58S / A584K,^609S / W698Fand the quintuple variant (T7 RNAPQ58S / S128V / A:’84K / V6U9S / W698F) having 3.9- and 3.5-fold increases in activity, respectively, relative to WT ASFVCE: T7 RNAP (Figure ID).

[0191] Further engineering of T7 RNAP based on evolutionarily aware machine learning predictions

[0192] As has been seen with numerous other proteins, MutCompute and other structure-aware machine learning algorithms were able to identify substitutions beneficial for activity, and these substitutions could generally be stacked in a roughly additive way. However, the exhaustive procedure used for stacking (all combinations of the best variants) becomes less feasible as the number of mutations (and thus the number of paths for stacking) increases. Therefore, MutRank, a self-supervised learning-based ranking framework that prioritizes beneficial mutations based on evolutionary likelihood based on multiple sequence alignment (MSA) data, was employed. It was contemplated that MutRank could see additional mutations that MutCompute and other structurally aware algorithms could not.

[0193] To test this, twelve MutRank-predicted single variants were selected (Table 2) and they were individually introduced into the previous quadruple and quintuple backgrounds, which served as scaffolds for the next round of engineering. Of the twelve single substitutions added to the quadruple background (T7 RN3Q58S / A584K'V6WS''W698F), five (C125I, C347P, Q404L, N419T, and Q786L) exhibited increased activity (Figure 2A). An additional 12 combinations were evaluated, and three additional variants demonstrated significant improvements in activity, with the best variant containing both Q404L and Q786L, and having a 2.3-fold increase in activity compared to the original quadruple variant, corresponding to a 6.2-fold increase relative to ASFVCE(WT): T7 RNAP(WT) (Figure 2A). This variant was designatedDocket No. 10046-673W01

[0194] EvoT7, and it was carried forward into further experiments, as described below'. A similar approach was applied using the quintuple variant as the parent enzyme, and quintuple047^4191exhibited the highest activity, with a 5.2-fold increase relative to ASFVCE(WT): T7 RNAP(WT) (Figure 2B).

[0195] Identification of a more highly active virus-derived mRNA capping enzyme

[0196] The addition of a 5'-m7G cap to eukaryotic mRNA is crucial, serving to prevent degradation and facilitate translation. In mammalian cells, several mRNA capping enzymes, including NP868R from African Swine Fever Virus (hereafter referred to as ASFVCE), have been identified and their activities compared through biochemical and functional assays. Among these, ASFVCE exhibited the highest activity and was subsequently engineered to enhance its performance in yeast-based expression systems. While these modifications improved their functionality, they were still based on a single viral source, prompting an investigation into whether alternative wild-type viral capping enzymes exhibit even greater activity in yeast. To identify a single-subunit RNA capping enzyme with higher activity than the ASFVCE, nine capping enzyme (CE) fusion candidates were screened, again using the fluorescent reporter system (Figures 3A and 3B). Among the nine tested enzymes, MACE and APMCE showed minimal or undetectable reporter gene expression, comparable to the null mutant of ASFVCE (ASFVCEK282N) and the negative control (NC), which consisted of WT T7 RNAP without a capping enzyme (Figures 3C and 3D). The enzymes FCE, PCE, BSVCE, GMCE, and NCE displayed activity levels comparable to or slightly lower than ASFVCE, while BMCE (from Brazilian Marseillevirus) exhibited an activity level around 2-fold higher than the previously used ASFVCE.

[0197] As with T7 RNA polymerase, further enhancement of BMCE activity was pursued through ML-guided predictions. Starting from the AlphaFol d -prediced structure of BMCE, three ML models - MutCompute, MutRank, and Stability Oracle - were applied to predict beneficial mutations. Each model identified 20 mutations, and after removing duplicates (S515G, H786F), a total of 58 unique single variants were generated (Table 3). To evaluate the activity of these variants, the same two-plasmid system was employed as in Figure 1 (Figure 4a). Among the 58 single variants, seven (D59E, M165L, S199A, N267W, Q426V, Q426Y, and D582F) exhibited up to 1.4-fold increased ZsGreen expression compared to WT BMCE (Figure 4B). As before, the top-performing single mutants ware combined to generate all possible double mutants, resulting in 20 variants (Figure 4C). Two of the double mutants (M165L / Q426Y and S199A / Q426Y) yielded over a 2-fold increase in activity compared toDocket No. 10046-673W01

[0198] WT BMCE. Additionally, four double mutants (M165L / Q426V, M165L / D582F, S199A / D582F, and N267W / Q426V) exhibited moderate improvements of over 1.5-fold (Figure 4C). Based on the six selected double mutants that exhibited enhanced activity, 13 triple mutants were systematically generated. Three of these variants had additive improvements, up to a 3-fold final increase in activity. Further stacking of quadruple and quintuple mutants did not yield further improvements in activity.

[0199] Combining the best engineered capase and T7 RNAP enzymes

[0200] The triple variant BMCES!99A / N267W / Q426V(hereafter referred to as EvoBMCE) was chosen for further testing with improved 1'7 RNAP variants. Similarly, in order to best combine previous directed evolution efforts with substitutions garnered from machine learning, EvoT7(443) was generated, which contained 12 mutations (6 from T7 RNAP(443) and 6 from EvoT7). Fusion proteins were constructed using the most effective capping enzymes (ASFVCE(443) and EvoBMCE) and T7 RNAP variants (T7 RNAP(443), Evol'7, and EvoT7(443)), derived either from directed evolution, and their activities were assessed (Figure 5A) relative to suitable controls. Gene expression was evaluated by measuring ZsGreen fluorescence normalized to ODgoonm, using ASFVCE(WT)-T7 RNAP(WT) as the baseline (set to 1).

[0201] Significant further enhancement in gene expression efficiency relative to previous capase-polymerase combinations were found, and these were due to both improvements in the capase and improvements in the polymerase (Figure 5B). A roughly 2-fold improvement of activity relative to wild-type was already observed when BMCE capase was used in place of the ASFVCE enzyme (Figure 2C and Figure 5 A). The introduction of the improved EvoBMCE further improved transcriptional efficiency, up to 5.6-fold above the previously developed wild-type enzyme combination of ASFVCE: T7 RNAP; these improvements were similar to those seen in fusions between the evolved ASFVCE(443) capase and the wild-type enzyme. The use of improved T7 RNAP variants also proved to be roughly additive, with the directed evolution (443) and machine learning (EvoT7) variants both leading to 2- to 3-fold improvements over parental enzyme combinations. Remarkably, the directed evolution and machine learning variants could be stacked together (EvoBMCE-EvoT7(443)) to attain even higher activity, upwards of an overall 7-fold improvement over the wild-type enzyme combination (BMCE-T7 RNAP(WT)) and greater than 12-fold improvement over the previously studied ASFVCE(WT)-T7 RNAP(WT) fusion.Docket No. 10046-673W01

[0202] Discussion

[0203] In the present disclosure, a machine learning (ML)-driven approach was employed to optimize fusion constructs between capase enzymes (either ASFVCE or BMCE) and T7 RNA polymerase (T7 RNAP) for improved transcription and gene expression in yeast. MutCompute, a structure-based ML model, was initially used to predict beneficial mutations by evaluating multiple T7 RNAP conformations derived from Protein Data Bank (PDB) structures. Specifically, MutCompute, a convolutional neural network-based model, and MutComputeX, a residual neural network-based approach, were applied to predict mutations that could enhance polymerase function by analyzing structural features of both the initiation complex (IC) and elongation complex (EC) states. The model’s predictions were subsequently filtered to select the most promising variants based on a log-ratio scoring system, prioritizing mutations predicted to be more favorable than the wild-type residue. Additionally, MutComputeX.2 was developed by modifying input features--- excluding charge and solvent-accessible surface area (SASA) — to assess whether this adjustment could improve stability-focused predictions.

[0204] Each ML algorithm contributed distinct yet complementary insights and overall described a broad fitness landscape for experimental validation. MutCompute and its variants focused on structural constraints, MutRank prioritized evolutionarily favorable changes, and Stability Oracle identified stability-enhancing substitutions. Each program eventually identified useful, non-overlapping substitutions. For T7 RNAP, one improved substitution originated from MutCompute, three were from MutComputeX, and one was from MutComputeX.2. The relative underperformance of MutComputeX.2 compared to MutComputeX shows that removing charge and SASA features, both of which play critical roles in protein stability and function, may have impeded the selection of beneficial substitutions. Similarly, for the BMCE capase variants, seven beneficial mutations were found: one from MutCompute, one from MutRank, and five from Stability Oracle.

[0205] Interestingly, machine-learning-based optimizations of T'7 RNAP led to significantly different results compared to previous studies (Table 4). In particular', EVOLVEpro-assisted directed evolution led to epT7 (T3M / G47A / E643G); structure-based rational engineering with Rosetta resulted in G47A / 884G; and high-throughput mutagenesis coupled with FACS-based selection led to the previous best capase-polymerase combination, v443 (N131K7L261M / H300R / 307H / Q648R / / H772R). Along the paths to each of these polymerases, numerous single mutations were tested and either incorporated or discarded (Figure 6): epT7 tested 42 single substitutions, while G47A / 884G tested 21 single substitutions; v443 accumulated mutations iteratively over the course of directed evolution. Only oneDocket No. 10046-673W01

[0206] mutation (H300R) overlapped between EvoT7 and v443, while no exact residue matches were observed between EvoT7 and epl’7 or G47A / 884G (Figure 6A). When considering only potential sites for substitution, as opposed to exact substitutions (Figure 6B), Evo'1'7 contained 68 unique mutation sites, epT7 had 33, G47A / 884G had 8, and v443 had (by the conclusion of directed evolution). EvoT7 shared three mutation sites (V134, V177, and V273) with epT7, one with v443 (H300), and none with G47A / 884G.

[0207] To better understand the structural basis for the improved performance of EvoT7 and EvoBMCE, an in-silico enzyme structure was generated with AlphaFold 3 and the atomic interactions gained or lost were examined compared to either the crystal structure of WI’ T7 RNAP (PDB: 1H38) or the AF3 wild-type BMCE structure (Figure 7). The six mutation sites (Q58S / Q404E / A584K / V609SAV698F / Q786E) in EvoT7 were distributed throughout T7 RNAP, and possible structural improvements span a gamut of chemistries (Figure 7A). It seems that the Q58S substitution may create a stable hydrogen bond with Ser58 (Figure 7B), and V609S may form hydrogen bonds with Asn592 and Gln669 (Figure 7B). The Q404L substitution enhances the hydrophobic core formed by Phe400, Phe408, and Phe432 (Figure 7B), while Q786L places a hydrophobic residue within the greasy pocket formed by Met549, Thr729, Pro730, Phe782, and Val841 (Figure 7B). In contrast, A584K mutation installs a positively charged lysine within the solvent-exposed domain and allows polar interactions with nearby polar residues (Glu580, Asp585, and Asn588) or water molecules (Figure 7B).

[0208] It is more difficult to interpret improvements to EvoBMCE, as there is no known structural determination of this enzyme. In the two predicted BMCE structures (EvoBMCE and WT BMCE), S199A, N267W, and Q426V appear to enhance the hydrophobic core of the enzyme (Figures 7C and 7D). It is interesting to note that the mutations S199A, N267W, and Q426V were exclusively predicted by the Stability Oracle model, which leverages a structure¬ based deep learning approach to identify thermodynamically favorable substitutions. Notably, Stability Oracle predicted the majority (5 out of 7) of the beneficial mutations for BMCE capase variants. The improvements in enzyme activity for EvoBMCE are particularly impressive when recalling that predictions were done entirely computationally, without an experimentally determined structure.

[0209] It is instructive that the different algorithms and models predicted diverse, mostly nonoverlapping variants. The fact that different methods all work roughly equally well in improving enzymes is probably a function of the fact that these programs were all trained on roughly the same data: the preponderance of information about wild-type proteins from nature. In essence, the different machine learning approaches to gleaning a ‘Platonic ideal’ of a wild-Docket No. 10046-673W01

[0210] type protein, the best protein that accords with what nature has already done, will provide variable but similarly useful insights into how to engineer a protein. There is no one best way to do machine learning for protein improvement. That said, the methods herein (especially the ability to test in a yeast background, where basal transcription was very low, and improved variants could be more easily identified than in bacteria) allowed for exploration of much larger landscapes than previous studies that focused on a limited number of beneficial mutations, and the catalytic superiority of EvoT7 is likely due in pail to this broadened exploration.

[0211] Fortunately, the improved single substitutions identified by the structure-based machine learning algorithms could generally be combined to generate variants with increasingly higher activities. It was also observed that machine learning builds on the results of previous directed evolution experiments, as evidenced by the combination of the original T7 RNAP(443) with predicted substitutions.

[0212] Overall, diverse, ML-driven strategy was used to engineer highly active variants of the fusion protein between 1'7 RNA polymerase (T7 RNAP) and the Brazilian Marseillevirus capping enzyme (BMCE), achieving significant improvements in gene expression in yeast over previous directed evolution efforts. By integrating ML predictions with high-throughput experimental validation, the 9-tuple substitution EvoBMCE (BMCEsl99A'N267W / Q426V): EvoT7 (T7 RNAPQ58S / Q404L / A584K / V609S, AV69SF / Q786I0 was developed, which have greater than 10-fold enhancement in gene expression efficiency compared to wild-type capase-polymerase combinations, and which could be further combined with substitutions identified by directed evolution to yield even higher activity. The improved fusion protein can be assayed and improved in important non-model yeast strains with unique metabolic traits, such as Yarrowia lipolytica, Pichia pastoris, and Kluyveromyces marxianus, and in mammalian cells, leading to high-efficiency gene expression platforms for synthetic biology, mRNA therapeutics, and even cell free systems.

[0213] Materials and Methods

[0214] Strains and culture media

[0215] For plasmid construction and maintenance, Escherichia coll NEB 5-alpha was cultured in lysogeny broth (LB) medium, which contained 10 g / L tryptone, 5 g / L yeast extract, and 10 g / L NaCl, at 37 °C with agitation at 225 rpm. When necessary, cultures were supplemented with 100 pg / mL ampicillin or 50 pg / mL kanamycin for plasmid selection. For yeast-based experiments, Saccharomyces cerevisiae BY4741 (MATa; his3Al; leu2 0; metl5A0; ura3A0) and its AGal2 derivative were utilized. Yeast strains were cultivated in yeast extract peptoneDocket No. 10046-673W01

[0216] dextrose (YPD) medium, consisting of 10 g / L yeast extract, 20 g / L peptone, and 20 g / L glucose. For selective growth, cultures were maintained in synthetic defined (SD) medium (Takara Bio), supplemented with drop-out amino acids (BUFFERAD) at 30 °C with shaking at 225 rpm. Induction of gene expression was performed using yeast nitrogen base (YNB) medium lacking amino acids and carbon sources, supplemented with ammonium sulfate, raffinose, and galactose (Sigma-Aldrich). To ensure reproducibility, all chemicals and media were of the highest available purity.

[0217] Plasmid construction

[0218] Yeast-codon-optimized parts encoding all target proteins were obtained from Twist Bioscience. Assembly of DNA fragments were performed using the NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs) in accordance with the manufacturer’s protocol. PCR amplification was conducted with the Q5 High-Fidelity 2X Master Mix (New England Biolabs), and the resulting products were purified using the Monarch® PCR & DNA Cleanup Kit (New England Biolabs). Wien necessary, DNA fragments were excised and extracted using the Monarch® DNA Gel Extraction Kit (New England Biolabs). Primers were designed via SnapGene 8.0.1 and synthesized by IDT (USA). Plasmid preparations were carried out using the Monarch® Plasmid Miniprep Kit (New England Biolabs). Site-directed mutagenesis was introduced using the Q5 Site-Directed Mutagenesis Kit (New England Biolabs). To verify plasmid construction, sequencing was performed by Plasmidsaurus (USA).

[0219] Yeast transformation

[0220] Transformation of S. cerevisiae strains was performed using the Yeast Transformation Kit (YEAST 1, Sigma- Aldrich) according to the manufacturer’s protocol. For genomic integration, 2 pg of linearized plasmid (digested with Notl-HF, New England Biolabs) was introduced into 50 pL of competent yeast cells. In the case of plasmid-based transformation, 1 pg of circular plasmid DNA was used per 50 L of yeast cells. Following transformation, the cells were spread on SD-URA-HIS selection plates and incubated at 30°C for up to 72 hours to enable colony formation. Successful genomic integration was verified through colony PCR, followed by sequencing analysis to confirm correct insertion.

[0221] Tissue cell cultureDocket No. 10046-673W01

[0222] HEK293T cells were cultured in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% (v / v) fetal bovine serum (FBS) and 1% (v / v) penicillin-streptomycin. The cells were maintained at 37°C in a humidified incubator containing 5% CO2.

[0223] Transient transfections were performed using the Lipofectamine 3000 (Thermo Fisher Scientific) according to the manufacturer’s instructions. One day prior to transfection, HEK293T cells were seeded into multi-well tissue culture plates at an appropriate density to reach 70-80% confluence on the day of transfection. For experiments utilizing a doxycycline -inducible promoter (e.g., TRE3G promoter), cells were treated with 2 pg / mL doxycycline simultaneously with transfection. Cells were incubated for 24 to 72 hours post-transfection before being subjected to subsequent assays.

[0224] Generation of stable ceil lines

[0225] Stable integration of the engineered capping enzyme -T7 RNAP fusion constructs into the AAVS1 safe harbor locus was achieved utilizing a Bxbl integrase-mediated site-specific recombination system. An engineered HEK293T master cell line stably expressing the reverse tetracycline-controlled transactivator and harboring a pre-integrated landing pad cassette comprising a Bxbl attachment site at the AAVS1 locus was utilized.. To generate stable cell lines, cells were co-transfected with a donor plasmid and a Bxbl integrase expression vector using the transient transfection method described above. The donor plasmid contained a matching attB site alongside the expression cassette of interest. Following the transient expression of the Bxbl integrase, a recombination event occurred between the attP and attB sites, precisely integrating the donor construct into the AAVS1 landing pad. To identify and isolate successfully recombined clones, a reporter plasmid encoding EGFP under the T7 promoter was subsequently transfected into the cell population. The EGFP-positive cells were then isolated using fluorescence-activated cell sorting and subsequently expanded to establish functional stable cell lines for downstream analyses.

[0226] Flow cytometry characterization

[0227] To evaluate fluorescence levels, engineered yeast strains were cultured in SD-URA- HIS medium at 30°C for 48 hours, allowing them to reach stationary phase. Cultures were subsequently diluted 1:10 in yeast nitrogen base (YNB) medium supplemented with 5% D-galactose for GALI promoter induction and 2% D-raffinose as an additional carbon source. The cultures were incubated overnight at 30°C with agitation in a 96-well plate shaker. Following incubation, cells were harvested by centrifugation, washed twice with phosphate-Docket No. 10046-673W01

[0228] buffered saline (PBS), and resuspended in fresh PBS. For flow cytometric analysis, samples were further diluted in PBS and transferred into a 96-weli plate for automated analysis using an SA38OO Spectral Cell Analyser (Sony Biotechnology). Single-cell populations were gated based on forward scatter (FSC) and side scatter (SSC) parameters in a logarithmic scale, with 10,000 events recorded per sample. ZsGreen fluorescence was detected using a 488 nm excitation laser and a 495-510 nm emission filter. 'The geometric mean fluorescence from three independent replicates was used for comparative analysis, with fold changes calculated as the ratio of mean fluorescence intensity of each T7 RNAP or BMCE variant relative to the wild¬ type control. Data analysis and processing were performed using FlowJo (version 10.10.0).

[0229] For the quantitative analysis of reporter gene expression in tissue culture cells, flow cytometry was performed. At 48 to 72 hours post-transfection or post-induction, HEK293T cells were gently washed with PBS, dissociated using 0.05% Trypsin-EDTA, and resuspended in complete culture medium to neutralize the trypsin. The cells were then pelleted by centrifugation and resuspended in an appropriate buffer. Cell suspensions were analyzed using a flow cytometer. FSC and SSC were used to gate viable single-cell populations, and EGFP fluorescence was measured using a standard FITC channel. For each sample, at least 10,000 events were recorded, and the median fluorescence intensity and the percentage of EGFP- positive cells were calculated to evaluate the transcriptional activity of the engineered T7 RNAP variants. Data analysis and processing were performed using FlowJo (version 10.10.0).

[0230] ML-predictions for T7 RNAP engineering

[0231] T7 RNA polymerase (T7 RNAP) transitions between two key conformational states during its catalytic cycle: the initiation complex (IC) and the elongation complex (EC). The IC is unstable and produces short RNA fragments, known as abortive transcripts, through a process called abortive cycling (PDB: 1QLN, 1CEZ, 2PI4, 2PI5), whereas the EC is stable and processive (PDB: 1S0V, 1S76, 1S77, 1H38, 1MSW). MutCompute models were used to generate predictions by analyzing the nine crystal structures of T7 RNAP. Both MutCompute (a convolutional neural network) and MutComputeX (a residual neural network) were used to analyze each of the nine PDB structures. Point mutants were selected based on high probability ratios, calculated as either the log? ratio (for MutCompute) or the natural log ratio (for MutComputeX) of the probability score of the highest-ranked amino acid to that of the wild-type residue at the same position (log2(-— ■" ~ for MutCompute,

[0232]

[0233] LAAyprobabihty forMutComp1uteX). The choice of log base reflects differences in model

[0234]

[0235] Docket No. 10046-673W01

[0236] design: MutCompute, optimized for discrete probability comparisons, applies logz for intuitive interpretation in terms of fold changes, while MutComputeX, trained to predict continuous stability changes (e.g., ATm, AAG), uses the natural log to better correlate with thermodynamic parameters. These values were averaged across multiple structures, irrespective of whether the structure was an IC or an EC (Table 2). Initially, 20 mutants were selected from the MutCompute predictions and 20 mutants from the MutComputeX predictions. In general, predictions by MutComputeX gave more consistent results across multiple input files. As the results were analyzed, it was noted that excluding charge and SASA led to differences in zero shot performance on thermostability benchmarks, and thus an additional 20 predictions were made without these input features, a process referred to as MutComputeX.2. Variants were tested and mutations stacked based on their relative activities in yeast. From the top five single substitutions, all possible double, triple, quadruple, and quintuple substitutions were assayed, and the best quadruple and quintuple substitutions were carried forward.

[0237] To further evaluate predictive models, the program MutRank was also employed to identify mutations with functional benefits. MutRank leverages EvoRank, a self-supervised learning-based ranking framework, which incorporates evolutionary information from multiple sequence alignments (MSAs) to prioritize beneficial mutations. Unlike traditional wild-type accuracy-based models, which focus on recovering known amino acids, EvoRank ranks amino acid substitutions based on their evolutionary likelihood, improving its ability to predict functionally advantageous mutations. MutRank was run on ten PDB structures, including the nine used in MutCompute predictions (1QLN, 1CEZ, 2PI4, 2PI5, 1S0V, 1S76, 1S77, 1H38, and 1MSW), with the addition of 1ARO, which represents 1'7 RNAP complexed with 1'7 lysozyme. Variants were ranked based on their predjprob values. MutRank employs a rankingbased learning framework, referred to as EvoRank loss, to identify amino acid substitutions that align more closely with evolutionary constraints and functional requirements. Unlike MutCompute and MutComputeX, which assess mutations through absolute probability ratios, MutRank instead estimates the relative ranking of amino acids at a given site. This is formulated in Equation 1:

[0238] P}MSA(aa+) 1

[0239] 7 [ (dQ, 0.0. ) r< M SA ( “7 tT

[0240]

[0241] P-!^\acr) + Pj \aa ) 2

[0242] Equation 1

[0243] where P

[0244]

[0245] ftSA(aaA)' and PJMSA(a.a") denote the probability values derived from multiple sequence alignments (MSAs) for two competing amino acids at position j. Unlike previousDocket No. 10046-673W01

[0246] models that focused on predicting individual amino acid probabilities, MutRank learns the evolutionary hierarchy of amino acids, enabling the model to capture functional fitness landscapes beyond conventional probability-based approaches. In this context, pred_prob values generated by MutRank do not represent absolute confidence scores but rather the relative probability that a given mutation would be observed in an evolutionary setting. A higher pred_prob shows that the model assigns greater likelihood to the mutated residue over alternative substitutions, reflecting both evolutionary constraints inferred from sequence conservation patterns and functional fitness predicted from structural stability data. This transition from probability-based assessment to ranking-based evaluation facilitates a more biologically meaningful approach to prioritizing mutations, particularly in scenarios where evolutionary selection pressures extend beyond thermodynamic stability alone. Those variants with pred...prob scores higher than 2 were selected for manual examination via visualization software; twelve mutations were ultimately chosen based on their apparent ability to better fit into the chemical microenvironment (based on evaluation of hydrophobicity, solventaccessibility, flexibility, and steric hindrance) compared to the wild-type amino acid. The MutRank predictions were further introduced into the best quadruple and quintuple mutant backgrounds, previously identified through MutCompute predictions, and subsequently assayed. The complete list of selected single-substitution predictions (72 mutations = 20 MutCompute, 20 MutComputeX, 20 MutComputeX.2, 12 MutRank), along with their predicted scores, is provided in Table 2.

[0247] ML-predictions for BMCE engineering

[0248] To improve the activity of BMCE, a systematic engineering approach was implemented incorporating MutCompute, MutRank, and Stability Oracle predictions. This method facilitated the discovery of mutations with functional benefits by integrating machine learning¬ based structure-function analysis. For the initial stage, MutCompute was applied to analyze BMCE, using its AlphaFold2 -predicted structure as the input model. The predictions generated by MutCompute were ranked based on the log ratio comparing wild-type residues to substitutions. In addition, MutRank was employed to independently identify 20 single mutations from its prediction dataset, selecting variants with a high probability of contributing to BMCE activity. Finally, a distinct set of 20 mutations was selected using Stability Oracle, a graph-transformer-based model designed to assess the thermodynamic stability (AAG) of protein variants. Unlike previous methods like Rosetta and FoldX, which require explicit modelling of both wild-type and mutant structures, Stability Oracle leverages a structure-Docket No. 10046-673W01

[0249] informed approach by integrating local structural data with graph-based attention mechanisms to predict stability changes. This method uses a graph-transformer architecture, where atoms are represented as nodes and atomic distances are treated as edges, guiding the attention mechanism to focus on relevant structural features. The core prediction is based on thermodynamic permutations (TP), a technique that enhances AAG predictions by leveraging the Gibbs free energy state-function property. AAG predictions indicate stabilizing or destabilizing mutations, where negative values represent stabilizing mutations and positive values indicate destabilizing mutations. The prediction model is based on Equation 2:

[0250] AA G = W ■ (emut- ewt)

[0251] Equation 2

[0252] In Equation 2, emutand ewtrepresent the embedding vectors of the mutated and wild-type amino acids, respectively, capturing their structural and chemical properties. The difference between these embeddings is scaled by a learned weight matrix W, transforming the difference into a prediction of the thermodynamic stability change (AAG). AAG quantifies the relative impact of a mutation on stability, considering the local structural context. A higher AAG indicates a destabilizing effect, while a lower AAG indicates stabilization. Tills transition from physics-based methods to graph-based learning provides a more biologically relevant approach for predicting mutations, particularly in cases where evolutionary pressures go beyond simple thermodynamic calculations. In total, the final BMCE dataset comprised 60 mutations, with 20 mutations selected from each of MutCompute, MutRank, and Stability Oracle. Among them, 58 unique mutations were chosen (with two overlapping across methods) to maximize BMCE activity and efficiency. A detailed list of selected mutations and their predicted scores can be found in Table 3.

[0253] Example 2: T7 RNA Polymerase Cell-free Assays

[0254] The cell-free assay was designed to compare the activities of T7 RNA polymerase (T7 RNAP) variants between yeast and bacterial systems. Five single-mutant variants that exhibited the highest increase in activity (Q58S, S128V, A584K, V609S, and W698F) and five single¬ mutant variants that displayed the greatest decrease in activity (V1 4G, Y418R, R423L, A638T, and Y639A) compared to the wild-type in the yeast system were selected. These variants were optimized for expression in E. coll, with a sigma 70 promoter added upstream of the genes and a T500 terminator added downstream. The genes were ordered as gBlocks through Twist Bioscience and amplified using PCR, incorporating 210 base pairs upstream of the promoter and 123 base pairs downstream of terminator. The PCR products containing theDocket No. 10046-673W01

[0255] T7 RNAP genes were added to 15 ul of the cell-free mixture at 9, 19 and 29 pM, along with 9 nM of the reporter GFP gene under the 1'7 promoter.

[0256] Example 3: Capase: T7 RNAP in Tissue Culture

[0257] The capase: T7 RNAP variants also drive functional gene expression in higher eukaryotic cells, including tissue culture cells. To evaluate and optimize T7 transcription and translation in tissue culture cells, various reporter constructs were generated in which a reporter gene, EGFP, was placed downstream of a T7 RNAP promoter element (Figure 9A). The EGFP gene further had sequences (Kozak, poly(A)) to help ensure its expression. Upon transient co¬ transfection of these reporter vectors and a vector containing the capase: T7 RNAP variant ASFVCE(443): T7RNAP(443), robust transcription and translation of EGFP was observed (Figure 9B). This contrasted with negative controls in which there was either no promoter, or a series of stop codons at the beginning of the gene, which showed no expression. Capase: T7 RNAP-mediated expression of a reporter gene could also be observed across a similar series of experiments that contained any of the evolved and / or machine learning-derived variants reported (BMCE: T7RNAP; EvoBMCE: T7RNAP(443); EvoBMCE: EvoT7; EvoBMCE: EvoT7(443) (Figure 10).

[0258] The capase: T7 RNAP variants work when integrated into the HEK293T genome, as well as when introduced transiently. A genomic construct was generated at the AAVS1 site in which ASFVCE(443): T7RNAP(443) was driven by the PolII promoter pCAG and terminated by poly(A)bGH). The nature of the integrated construct made only modest differences: irrespective of whether the capase: T7 RNAP fusion was driven by either pCAG or pTRE3GV, different capase: T7 RNAP fusion variants that substitutions derived from directed evolution and / or machine learning approaches showed high activity relative to negative controls (Figure 11). Genome -integrated stable cell lines were obtained by sorting EGFP-positive cells via FACS. In subsequent experiments utilizing a cell line with integrated ASFVCE(443)-T7RNAP(443) under the C AG promoter, the presence of an integrated, stable capase: T7 RNAP construct allowed for screening many possible combinations of gene expression architectures via transient transfection. Upon transient transfection of reporter plasmids, expression was observed relative to negative controls (Figure 12). As a result, a distinct EGFP signal was observed compared to the negative control. And notably, the greatest expression was observed with an encoded poiy(A) tail (poly(A)syn), as opposed to a signal for polyad enylation (poly(A)bGH). This was contemplated to be due to the ability of T7 RNAP transcripts with anDocket No. 10046-673W01

[0259] encoded poly(A) sequence that can undergo direct translation in the cytosol, bypassing the requirement for nuclear processing and export processing.

[0260] Next, it was investigated whether the 5’ UTR elements previously reported to enhance T7 promoter strength could yield similar effects within mammalian cells. When reporter plasmids containing the 1'7 RNAP promoter and an EGFP gene were introduced with additional 5’ UTR elements that served as enhancers (Figure 9 A and Figure 13), the enhancers provided little additional activity in a system that was orthogonal and insulated in large measure from PolII-mediated transcription. The capase: T7 RNAP variants can function on their own to drive transcription from a T7 RNAP promoter, as long as they are paired with other elements that allow translation initiation (Kozak) and termination (poly(A)).

[0261] It will be apparent to those skilled in the art that various modifications and variations can be made in the present disclosure without departing from the scope or spirit of the invention. Other embodiments of the disclosure will be apparent to those skilled in the an from consideration of the specification and practice of the methods disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the invention being indicated by the following claims.Docket No. 10046-673W01

[0262] TABLES

[0263] Table 1. Summary of additional 1'7 RNAP variants.

[0264] Position WT Mutation

[0265] T7 RNAP

[0266] 905 D N

[0267] 921 D G

[0268] 963 A T

[0269] 982 D N

[0270] 983 W R

[0271] 1012 I T

[0272] 1024 A T

[0273] 1026 N K

[0274] 1058 K R

[0275] 1076 A V

[0276] 1078 M V

[0277] 1133 G C

[0278] 1156 L M

[0279] 1172 P T

[0280] 1178 G C

[0281] 1195 H R

[0282] 1199 A V

[0283] 1202 R H

[0284] 1227 K E

[0285] 1301 N K

[0286] 1390 S F

[0287] 1401 D G

[0288] 1407 L I

[0289] 1415 Ci E

[0290]

[0291] Table 2. ML -predictions for T7 RNAP engineering. IC and EC represent the Initiation and Elongation Complex states of T7 RNAP, respectively. CountMut indicates the number of occurrences of a specific position / amino acid combination in the analyzed PDB structures.

[0292] Prediction PDB ID

[0293] Position Wt_AA Pred_AA avgjog_ratio CountMut Model (T7 RNAP State)

[0294] ICEZ(IC) 710 VAL GLY 19.075 1 1QLN{IC) 853 LEU GLY 16.805 2 MutCompute

[0295] 1QLN(IC) 596 THR GLY 14.801 1 1QLN(IC) 710 VAL SER 13.257 1

[0296]

[0297] Docket No. 10046-673W01

[0298] 2PI4(IC) 185 VAL ALA 12.878 1 1QLN{IC) 384 VAL SER 12.646 1 ISOV(EC) 16 LEU SER 11.776 2 1S76(EC) 585 ASP GLN 11.757 1 2PI4(IC) 763 THR LEU 11.689 1 2PI5(IC) 854 HIS ASP 11.630 1 ICEZ(IC) 642 LYS TRP 11.513 1 2PI4(IC) 349 VAL LEU 11.331 1 ICEZ(IC) 637 LEU SER 11.110 1 1H38(EC) 134 VAL GLY 10.890 1 2PI4(IC) 778 ILE LEU 9.775 2 1CEZ(IC) 854 HIS GLY 9.626 2 1QLN(IC) 609 VAL SER 9.280 2 1S76(EC) 724 ALA GLY 8.422 2 ISOV(EC) 534 LEU VAL 6.281 4 ISOV(EC) 418 TYR ARG 4.489 5 ICEZ(IC) 853 LEU ALA 11.204 1 1H38(EC) 19 ILE TRP 10.053 3 2PI4(IC) 273 VAL LEU 8.502 4 IMSW(EC) 128 SER VAL 8.459 2 1H38(EC) 6 ILE ARG 8.450 4 2PI4(IC) 423 ARG LEU 7.994 4 IQLN(iC) 834 ASP ARG 7.923 8 1H38(EC) 253 ALA ASN 7.379 4 1H38(EC) 15 GLU SER 7.036 4 1H38(EC) 58 GLN SER 6.477 2 MutComputeX

[0299] 1H38(EC) 169 GLN LYS 5.710 4 1H38(EC) 567 VAL PRO 5.336 2 1H38(EC) 181 ALA LYS 5.190 6 1CEZ(IC) 207 GLU ARG 5.156 5 1H38(EC) 491 ALA ARG 5.022 3 1CEZ(IC) 698 TRP PHE 4.869 4 2PI4(IC) 841 VAL PRO 4.680 5 1H38(EC) 771 ALA PRO 4.379 3 IQLN(IC) 638 ALA THR 4.374 5 2PI4(IC) 300 HIS ARG 4.155 5 MutComputeX.2 2PI5(IC) 223 SER ALA 4.37 4

[0300]

[0301] Docket No. 10046-673W01

[0302] 2PI4(IC) 584 ALA LYS 4.055 2 2PI4(IC) 772 HIS LYS 4.029 4 2PI5(IC) 208 ASP HIS 3.866 4 1H38(EC) 534 LEU ILE 3.588 5 ICEZ(IC) 234 ALA PRO 3.381 4 IMSW(EC) 568 GLN HIS 3.224 4 IMSW(EC) 350 GLU LYS 3.219 4 2PI4(IC) 177 VAL LYS 3.194 4 2PI4(IC) 587 ILE LYS 3.066 4 1QLN{IC) 242 GLU ASN 2.835 4 1H38(EC) 879 ASP PRO 2.709 4 1CEZ(IC) 593 GLU LYS 2.335 1 1QLN(IC) 218 GLU GLN 2.31 4 2PI4(IC) 654 THR VAL 2.165 4 ICEZ(IC) 306 MET LEU 1.863 4 1H38(EC) 401 MET TYR 1.857 4 2PI4(IC) 403 GLU ASN 1.854 4 1H38(EC) 461 LYS MET 1.847 4 IMSW(EC) 684 SER ALA 1.678 4 PDB ID

[0303] Prediction model Position Wt_AA Pred_AA Pred_prob CountMut (T7 RNAP State)

[0304] 1ARO 873 ARG LEU 3.385 1 1H38(EC) 216 CYS LEU 3.073 1 ISOV(EC) 530 CYS ILE 2.725 1 2PI4(IC) 492 CYS VAL 2.673 1 ISOV(EC) 774 GLN ILE 2.662 1 IMSW(EC) 540 CYS LEU 2.572 1 MutRank

[0305] ICEZ(IC) 347 CYS PRO 2.435 1 ISOV(EC) 419 ASN THR 2.258 1 2PI4(IC) 291 ARG VAL 2.155 1 1S77(EC) 125 CYS ILE 2.122 1 1ARO 404 GLN LEU 2.113 1 1ARO 786 GLN LEU 2.051 1

[0306]

[0307] Docket No. 10046-673W01

[0308] Table 3. ML-predictions for BMCE engineering.

[0309] Prediction Model Position Wt_AA Pred__AA avgjog_ratio 7 3 GLU LEU 13.280 274 ARG LEU 12.877 216 ASN THR 10.611 25 GLU LEU 9.100 230 MET GLN 8.885 165 MET LEU 8.711 311 LEU ASP 8.327 585 SER GLY 7.978 530 ASN ASP 7.680 176 GLY ALA 7.677 MutCompute

[0310] 337 LYS ASN 7.534 113 TRP TYR 7.321 690 MET LEU 7.292 538 VAL THR 7.275 184 GLU LEU 7.201 218 iLE LEU 7.140 725 ASN ASP 7.024 214 GLN SER 6.603 290 PHE ASN 6.489 520 HIS LYS 6.131 Prediction model Position Wt_AA Pred__AA Pred_prob 582 ASP VAL 3.915 267 ASN PRO 3.375 142 ARG PHE 2.227 92 ARG VAL 2.196 251 GLN GLY 2.135 321 CYS VAL 2.091 90 ASN VAL 1.639 MutRank

[0311] 191 LYS ILE 1.588 182 GLU VAL 1.582 107 LYS VAL 1.446 426 GLN TYR 1.425 524 THR PRO 1.403 786 HIS PHE 1.383 95 LYS VAL 1.305Docket No. 10046-673W01

[0312] 268 LYS TRP 1.295 109 ARG VAL 1.277 563 ASN LEU 1.276 123 ALA PHE 1.274 515 SER GLY 1.206 341 THR GLU 1.193 Prediction model Position Wt_AA Pred__AA Pred_G 458 ASN TRP -2.777 593 LYS LEU -2.559 480 LYS PHE -2.504 515 SER GLY -2.484 109 ARG TYR -2.429 426 GLN VAL -2.158 182 GLU LYS -2.134 786 HIS PHE -2.125 530 ASN PHE -2.028 142 ARG VAL -1.993 Stability Oracle

[0313] 251 GLN TRP -1.928 439 LYS ILE -1.901 690 MET PHE -1.892 582 ASP PHE -1.891 267 ASN TRP -1.873 542! LE PHE -1.828 679 SER TRP -1.799 199 SER ALA -1.797 563 ASN TRP -1.771 59 ASP GLU -1.712

[0314] Table 4. List of T7 RNA polymerase mutants identified in this study and reported in selected previous studies. _

[0315] Name Key mutations

[0316] EvoT7 Q58S, Q404L, A584K, V609S, W698F, Q786L

[0317] v443 N131K, L261M, H300R, R307H, Q648R, H772R

[0318] G47A / 884G G47A, 884G

[0319] epT7 T3M, G47A, E643GDocket No. 10046-673W01

[0320] SEQUENCES

[0321] 1. SEQ ID NO: 1 - Nuclear localization signal of SV40 (simian virus 40) large T antigen MFI. EPPKKKRK VV

[0322] 2. SEQ ID NO: 2 - Glycine-serine linker

[0323] GGGGSGGGGSGGGGS

[0324] 3. SEQ ID NO: 3 - ASFVCE ASLDNLVARYQRCFNDQSLKNSTIELEIRFQQINFLLFKTVYEALVAQEIPSTISIISIRCI KKVHHENHCREKILPSENLYFKKQPLMFFKFSEPASLGCKVSLAIEQPIRKEILDSSVL VRLKNRTTFRVSELWKffiLTIVKQLMGSEVSAKLAAFKTLLFDTPEQQTTKNMMTLI NPDGEYI. YEIEIEYTGKPESLTAADVIKIKNTVLTUSPNIILMI. TAYIIQAIEFIASIin. SS EILLARIKSGKWGLKRLLPRVKSMTKADYMKFYPPVGYYVTDKADGIRGIAVIQDTQ IYVVADQLYSLGTTGIEPLKPTILDGEFMPEKKEFYGFDVIMYEGNLLTQQGFETRIES LSKGIKVLQAFNIKAEMKPFISLTSADPNVLLKNFESIFKKKTRPYSIDG1ILVEPGNSY I. NTNTFKWKPTWDNTLDH. VRKCPESLNVPEYAPKKGFSU LFVGISGEI, FKKLAL NWCPGYTKLFPVTQRNQNYFPVQFQPSDFPLAFLYYHPDTSSFSNIDGKVLEMRCLK REINYVRWEIVKIREDRQQDLKTGGYFGNDFKTAELTWLNYMDPFSFEELAKGPSG MYFAGAKTGIYRAQTALISFIKQEIIQKISHQSWVIDLGIGKGQDLGRYLDAGVRHLV GIDKDQTALAELVYRKFSHATTRQUKHATNIYVLHQDLAEPAKEISEKVHQIYGFPK EGASSIVSNLFIHYLMKNTQQVENLAVLCHKLLQPGGMVWFTTMLGEQVLELLHEN RIELNEVWEARENEVVKFA1KRLFKEDILQETGQEIGVLLPFSNGDFYNEYLVNTAFLI KIFKHHGFSLVQKQSFKDWIPEFQNFSKSLYKIIMEADKTWTSIJFGFICLRKN

[0325] 4. SEQ ID NO: 4 - FCE AKRLQRCQDVNQVCEIYNSKGGIGELELRFDKLPQNLFAGVFDKLKPDGEIQTTMRV SNRDGVAREITFGGGVKTNEIFVKKQNICVFDVVDIFSYKVAVSTEETVVEKPTMETT AGVRFKIRLSVEDVVKDWRIDLTAVKTAELGKIAQHTASIVQRTFPDNLLKLTGAEV AKLAADSYELELEYTGKSPATNEKVNVAAKYAVELLSSVRNANSTAAASFGESVSD LCRVAKIIHTHEYANWCRTPSFKMLLPQWSLTKSSYYGGLYPPENLWLAGKTDGV RALVVCEDGVAKVITAESVDITIIGVCSATTILDCELNVDAKILYVFDVIISNNTQVYT QPFSTRITTDISDIKIDGYKIEMKPFVKVVKADEATFKSAYKAPUNEGLIMIEDGAAY AATKTYKWKPLSHNTIDFI JKACPKQLINVDPYKPRAGYKLV’LLFTriSLDQQRELGIDocket No. 10046-673W01

[0326] EFIPAWKILFTDINMFGSRVPIQFQPAINPLAYVCYLPEDVNVNDGDIVEMRAVDGYD TIPKWELVRSRNDRKNEPGFYGNNYKIASDIYLNYIDVFHFEDLYKYNPGYFEKNKS DIYVAPNKYRRYLIKSLFGRYLRDAKWVIDAAAGRGADLHLYKAECVEHLLAIDIDP TAISELVRRRNEITGYNKSHRGGRNMHSHRGQSHCAKSTSLHALVADLRENPDVLIP KIIQSRPHERCYDAIVINFAIHYLCD'IDEHIRDFEITVSRLLAPNGVFIFTTMDGESIVKL LADHKVRPGEAWTIHTGDVNSPDSTVPKYSIRRLYDSDKLTKTGQQIEVI PMSGEM KAEPLCNIKNIISMARKMGLDLVESANFSVLYEAYARDYPDIYARMTPDDKLYNDLH TYAVFKRKK

[0327] 5. SEQ ID NO: 5 - PCE SSNFTDKMLNEMINKYAQETANLSAGENffiLEIRFKDTTRDAFEAVYNAINGNPEFA NPVLECSVNVISENVYERSVGGKTDETQYIRKMTFDKGSIISDEYMQKIRLMRPIQMN DYIKYSVGLSRESKSKKFSTSNNALVRFKVRVSFDYIGSPSHPARWRFDLTAVKFGVL SEIGASLKMIKNT FrEALSPSNFIRELNFNMIDSYETEIEYIDLPANFrVDDLHIAKKVF SLVNPQYISEIAYQEEIYHIAEHIILNLNVLHMYKNPTHRLKQLGNQAIALSKNTYVN DVFPPDQYFATDKADGKRAIISVNGNRCRVLMSDGMMEFMQGEQFTPGEITIADAEL IYDSQLVKSKKTDKFTLYVFDVMVYKNDNVSKSGFAVRATHLVD ASKLID ALVNLE GNHAQSKNYVRLEKDNLENGLREVWENDFPYIVDGLIITEPDEPYMTTKNYKWKPY ERNTIDFLAVKCPQKLLGIKPYDVIKGKDLYILFVGINHQMREKLGLGFIPQYIILLFPK SDSGYYPIQFTPSADPLAYIYYHDSKFGDIDRQIVELARDKDNTTWLFHQVRTDRKLE KNYFGNDFRIAELrFLNYIDPFNFEDLWKPSGSYFTKTAGDIYMASNKFKRFIISILLK DNLSGrKVA'IDEAAGRGGDI. HRYQEIGVENAl. FIDIDPTAIAElJRRKFAFFAVKKRH VRNWMQNGKQGGDTTGKTTKVVTAYDRVHDIEYDKLIVKDMKSLTVHTLVADLK APASDLIAHTFQFGLNVGIVDGVVCNFALHYMCDTIEHIRNLLIFNAKMLKIGGLFIFT VMDGKAVFELLKPLSRGQQWEVRENGVVKYAIKKDYAGDKLSQTGQMISVLLPFS DEMYQEPI. CNVDAVISEASKLGEEVEI. NSSMGTMFDRFAKADRALSDRLT ADDKEY ISLHRFVSLRKIKDIKKLQDE

[0328] 6. SEQ ID NO: 6 -- TVCE DIPKRIFDILVDRTENAMKDSQLELEARIVPGHNKEKALNRVGFDRVLRYLSKTKSV YEPEHTVI. DVTRGMFRVSLTDENRIAEQCSMWNKRSLIQMYESSQDSIIAVRKEPVQ KPIDVSEYNILISLKRETNQQDPTEALKYLSSPNRPMTFRLKKRYSFVSSDGSYRIDCT AVKSASSTKIDDLASADEMFEIEVELLDGETVASKNPLSRANAETIAKGMLNTLSEIL VVLRDTHSMTLITSTERSKVINEYSKIYSWRSTQGPKPVTLERGNLI ETVDNFSIRDDocket No. 10046-673W01

[0329] DGPVGLYTATDKADGNR 'LMFINSQGKGYLIDDSRNVIPTDISVPGAANSLLDGEYV RKSRLETSINLFLVFDVYFSRGKDLRSLPLVDFDNPDTETRLRSMRKDGDIGAPLSNL RNPE1RLKEFVPLSKAEQLYKRYDGTTGPEYN1DGLILTPAKLSVFQERPDEPPTKKSG TWTRVYKWKPPKDNSVDFEVQFDKTSQGDMAVYSGKMKARLKVGMQVAEGNDP YEILSGALLEQDYRRRTFEAPDYKMRVMSNPDSAGSGMYIHNGRELPRCLQPPNDTI YDGSIVEFTWDETNGVWSPLRVRHDKDSPNDIDTALSVWRSIAFPVQIQDIVDPDRV TDAGAGSKDMIYFDEVTAGDGAATDAMRRFHRTWIIDRMMFRMGAAFVRSMRPPS STSDDLRIIDMACGRGADIRSWMRNGYDVVVGLDLLEDNLMGTAKTAAYARLAKIR DKLKDQGMRYSFVPMDVSKPVGPEAVEEISNPSLKKIGQTLWRTKKGKSDPRLIPYD GLALKPFDVVSCQFALHYFTESRLSLETFTDNVASNLDSGGIFIGTCMDGMAVDEEFr KRESELQSDNVTMEAFSESGSLVWRIEKRYNGRYDPVEGDEDEEIFGKPIGVYMETIN KVNQEYLVHFPTLVTALEQRGMRLLNVKELDQYGIDKSSGLFGSVYNSTSWEEVER DVENPYEASIAKMIKEMSPDLKTFSFMNRYFVF KI

[0330] 7. SEQ ID NO: 7 - BSVCE QKSNIQSKIGNDNVKQIEDLFKKINKDSEFEFIFFSKKDKHLTFEKYRTLVKFISKRASS DSKIKLVKNEQTLDISFSPEIEKVYRCTLIGKDAINTYMKKLSIANNHVIFRNLIKLSKK DDKIEILKKEKKSDQ TIDIDDI. YMR ARI. SSET K ISDEEYKKLIA IDE TQMDN IIFRYKER TSLYVYQNDKDYIRIDLTITKMDKKYNKI. NNAISNYELELETMCEKPKNEILEKMYN EIEVIFKVVQQSNFIITKSEANSVIENYKNLLAVPKQNERSLYGRQPISLEIQYVESLAN KYAVTDKADGERHFLVIFNNHVYLMNKNLDVKNTGIILKKEQERYNNSVIDGELIFL KNRHVFLTFDCLFNGGNDIRKEAKLEERLKNADIIINDCFVFNGQKGCKYLETEHHK EFNLNKTVEFIIRKQIKMTLDNLNUDIEIEKQFLI. VRRKYFIHSIGAKPWEIARYTAI WESYTGDPEIKCPYTLDGLIFQPNEQKYAVNKKDSRYDDYKWKPREKNSIDFYVEFL KDDDGNILTVYDNSYSKISDDPNAVDNAQERNANQTYRICRLHVGQIVGQLQAPVLF RENDNLYEAYIPLQNGEIRDIEGNILADKTVVEFYYNSMSNI. PNKFKWIPMRTRYDK TEMVMKYKQNYGNFVSVADKVWQSIVNPVLMTDFEDLAKGNNPDKNQYFYDKKI, DELRKRIGHDMIVSAAKEDKYFQLITKLGIPMRNFHNFIKSNIIFTFCHSMYQDNKQK SVLDLGCGRGGDQNRFYYAKVAFYVGVDYDRDALFNPLDSSTSRYNQQKRKPGFT KMFFIHGDQSNELDYESQFASI. GGMDFQNKQLIERFFSKDPAKRTTFDVINCQFAIHY SFKNDDSI. SNFKKNINNYLRNDGYVLITTFDGNAIRKIJ. KNKERFTQEYVNEEGEVKI LFDIVKKYDDVNDDVIMGTGNAIDVHMAWLFNEGVYQTEYLVDEKFIIEEFKRDCN LELVTMDSFENQFNILKEYLTEYVQYEADDRTREGILKRAGEFYKSNSVNDGCKKYT DLEKYYVFRKNIPKAKKQKGGNNDIYNTIIKYSISPLPGYDNNYSFINSIHHILKNHEIIDocket No. 10046-673W01

[0331] PKTLSPKTFCSDMGIDYENDINIDSNFKKIAKNIVIYHANNDGKHEKILNGLNILLAER NKDDEYTFKIIEKGKKITQNDNVIMLAKEGVMYAPIYQIDSDLRNGIFHMDDENIKQL LELQ

[0332] 8. SEQ ID NO: 8 - MACE VTKNKSENIRDIIDSDDISRVEEMVNNFRKNRNAEFEISVRKINYSNYIRISEYYVNNSS DIQQVTSEDVSIVEEDGNTYRISFENENLINDFESKYSNMKYGDIVKYILSENSGDDIEI IYKNRGSADRLSIEDLNLVIKLTEEVPVSNSNKPKLSGREKILYRYKNRYSFTIDDISR VDITDVKETSNIWELSKKISNYEIELEFTNKKINSNQVFEKIFDLLKIVQNTEIPIGIRES KQVLTDYQNLLGIKTSNHLDSRNVVSIETQHIVKFVPNRYArrDKADGERYFLFSNSN GVYELSTNMIVKKVNVPPEKKDFQNMEEDGEEIEIDGKEEFMVFDVVYHNNIDYRY DNNYTLTHRIIVINDIIDKAFNNLIPFTDYTDKYDNLEIDKIKEFYSNEIKTYWKKFNK

[0333] K1ENYSGLFVSRKLYFVPYGIDSSEVFMYADLVWKLCVYDQLTPYKLDGIIYTPIASP YMIKTSANF DSVPMEYKWKPPTQNSIDFYIKFDKDARGGEAIYYDNAVVRGEGRP YKICNEFVGENKSGEEKPIAFKVAGVEQKAFIYLTNDEAEDENGNVISDNTVVEFIFD NLKTDMDDPYKWIPIRTRYDKTESVQKYRKKYGNNLHIATRIWRTITNPVTEEIIAAL GNTLTYEKEMSKLTKMNESYNKQSFSYYQKNTSNAIGMRAFNNFIKSNMISTYCKE KQSVLDIGCGRGGDLIKFIRANIREYVGLDIDNNGLYVINDSAFNRYKNLKKTNQNV PPMTFINADARGLFNIEAQEKILPNMSESNKKEINNYLSNNKKYDAINCQFTLHYYES DDTSWNNFCQNVNSHIKDNGYLLITCFDGQLIYDKLKDKQKYSSSYTDNMGKKNIFF EINKIYSDEEIKSTGMAIDIYNSLISNPGVYQREYLVFPDFLKKSLKDNCGLELVETDM FYNIFNLYRNYFLINGGNFSTGEISGKRFNEIKDFYLALEGKSSSTTESDIAFASFKLAM

[0334] I. NRYYIFKKKTAINITEPSHIVSAVNKKTDLGKVLMPYFITNNMIIDYSLENNDVNKIY HFIRKKYSPVKPSVYLVRHTIVNNPMDGITFSRNKLEFVKIKKGTDPKILLVYKSPEKI FYPFYYQRFQNNDNSGDYFKNNLYLKDKG'rYLLDSNKIVNDLNILVNLSGKV

[0335] 9. SEQ ID NO: 9 - APMCE GTKLKKSNNDITIFSENEYNEIVEMLRDYSNGDNLEFEVSFKNINYI’NFMRnEHYINI rPENKIESNNYLDISLIFPDKNVYRVSLFNQEQIGEFriKFSKASSNDISRYIVSLDPSDD lEIVYKNRGSGKLIGIDNWAITIKSTEEIPLVAGKSKISKPKITGSERIMYRYKTRYSFTI NKNSRIDITDVKSSPIIWKI. MTVPSNYELELELINKIDINTLESELLNVFMIIQDTKIPISK AESDTVVEEYRNLLNVRQTNNLDSRNVISVNSNHIINFII’NRYAVTDKADGERYFLFS LNSGIYLLSINLTVKKLNIPVLEKRYQNMLIDGEYIKTTGHDLFMVFDVIFAEGTDYR YDNTYSLPKRIIIINNIIDKCFGNLIPFNDYTDKIINNI, ELDSIKTYYKSELSNYWKNFKDocket No. 10046-673W01

[0336] NRLNKSTDLFITRKLYLVPYGIDSSEIFMYADMIWKLYVYNELTPYQLDGIIYTPINSP YLIRGGIDAYDTII’MEYKWKPPSQNSIDFYIRFKKDVSGADAVYYDNSVERAEGKPY KICLLYVGLNKQGQEIPIQFKVNGVEQTANIYTKDGEATDINGNAINDNTVVEFVFDT LKIDMDDSYKWIPIRTRYDKTESVQKYHKRYGNNLQIANRIWKTITNPITEDIISSLGD PTTFNKEn’LLSDFRDTKYNKQALTYYQKNTSNAAGMRAFNNWlKSNMnTYCRDG SKVLDIGCGRGGDLIKHNAGVEFYVGIDIDNNGLYVINDSANNRYKNLKKTIQNIPP MYFINADARGEFTEEAQEKIEPGMPDENKSLINKYLVGNKYDTINCQFTIHYYLSDEL SWNNFCKNINNQLKDNGYLLITSFDGNLIHNKLKGKQKLSSSYTDNRGNKNIFFEINK IYSDTDKVGLGMAIDLYNSLISNPGTYIREYLVFPEFLEKSLKEKCGLELVESDLFYNI FNTYKNYFKKTYNEYGMTDVSSKKHSE1REFYLSLEGNANNDIEIDIARASFKLAML NRYYVFRKTSTINITEPSRIVNEENNRIDLGKFIMPYFRTNNMFIDEDNVDTDINRVYR NIRNKYRTTRPHVYLIKHNINENRLEDIYLSNNKLDFSKIKNGSDPKVLLIYKSPDKQF YPLYYQNYQSMPFDLDQ1YLPDKKKYLLDSDRIINDLNILINLTEKIKNIPQLS

[0337] 10. SEQ ID NO: 10 - GMCE SDLSQWFESNEVAPFLVEQPDDVEVELSFGRWKKEERQQVFQPGVTKAEFYNLFEFF SERSKIENPDFRREDFHTKEETMKVAGGGQRVNIRRITDLSSGKVSYQKKERNKIWD ERTWGIRIASSRENDVGTISDFVPSGFREKQRTRFTYIPNNSVFKGLILDLTKVVKSSF RNGVVYEVEIEKMPSVKKSASAMKNVVHRILQEMQSKNSLRDAEVIPVKQKEEVLE AFARLMDSKRQGDTSFLNKPRQVKIEDLFEPRLFAVTNKLDGERRLLWFGETGTYLV NPPSDIQKIGGIFYQYKDTVLDVELYGEETSKPKCYAFDCLFFMGKKQYNNNFDTRF AH VLEIEKHAGEELGIV SKK YFEVS1. SLGSFEGM A I’ATKDTRKT, Y1. EK ARTKKE. GGD LIJ DAVDAALDYAKVQKFRTDGI LQPREQGYKNEDTLKWKPAELMTIDFRLKRKS DNEYLLMMGGPRDKELNFLGSGPNRVSGTVKLPDKFLKEHQGRDFEGEILEFGFDK KKNIFFI’HRIRrDKDVPNYMTTVLSVWGNIFRGISLETMRGETLVLARRYNNILKQCL LTSASEAKTVLDIGSGRGGDIPKWQKGGMGMVWSIEPNKENIKEYEERAKAAKFTK YK NAKAQDTQAAIQFFGEDKSVDLSSAFYCiriYFGETKKDLDAECLTMSSLSKQG SLALFSFMDGDNIQRLLGKNKKFENSAFSVEKIGQWGKEAFGRKVRVDIKDADSMV KEQTEV’LVDIVNAKGLDGKVRI

[0338] 11. SEQ ID NO: 11 - BMCE SQLSEWFGSNEVSQFLTKQDANIEVELSFGRWRKTDRHQKFEPGVTKAEFYNLFESL DQISKKSPELFRREDVHTLEETMKSM’GQRANIRRIKNLSTGEVSFLMKDRSKVWDE RNWGI. RFAISREEEI. GVVPNFKAYGFREKQRTRFVFVGEKSAFKGI. TVDMTKVVRSDocket No. 10046-673W01

[0339] SFRDGVAFEVEVEKDSSVPKTAAAMTNSVKRIFEMMQSKSSCQTNEIISMKEKETAL EMFDRLMGSKRYGDTSFLNKI’RQVDADDLFEPRLFAVTNKLDGERRLLWFGSTGTY LVNPPFDIQK1AGVIAQYKDTVLDVELYGQELGDCSCYAFDCLWFMGQSQYNKNFD TRFSHVLEIEENVGQEARVIAKKYFEVSTSLGSFEGMPTATKDTRKLYLQKARTGKL GGDLLKEAVDRALEYAEKQKFRTDGLILQPREQGYKNEDTLKWKPADQMTIDFRLK RKSDNEFFLMMGGPKDRELKFEGTAAKR1PGTVVLPKEHESHPGHDFEGEILEFGFD I. EKSVFVPHRVRTDKDVPNYVKTVI. SVWGNIFRGVSI. ETLRGDTLVLPRKYNNKLK QCLLGIPKDVKTILDIGSGRGGDIQKWQHSGFSSVKMIEPNEENIREFERRAKITKFTG YKLFHGKAQETQR1VKFFGTEKVDFSSAFYCLGYFAETKKNLEDLCNTMSSMSSDKA MAAFAFMDGQAIREILGRKKKFENPAFSLEKEGEWTSKKYGNAVKVH1KDEGSMVK HQTEWLVDLEMLQDVMKGCGYELVFSRLADEGVSFLSTVGFEFVSLHRVAVFRKI

[0340] 12. SEQ ID NO: 12 - NCE SQQLSEWFGSNEVSQFLTKHEKNIEVELSFGRWVKTDRHQRFEPGVTKSEFYNLFDS LDQIAKKSPELFRREDIHTLF. ESAKTVAGQRGANVRRIKNLSTGEVSFLMKERNKIW DERSWGLRFALSREEELGVIPNFKASGFREKQRTRFVFVGEKSAFKGLTVDMTKVVK SSFRDGVVFEVEVEKDSSVI’KTAAAMTNSVKRILEMMQSKSSCQTNEIISSKEKEIAL EMFGKLVGSKRYGDTSFT. NKPRQVKVDDLFEPKLFAVTNKLDGERRLLWFSSLGTY LVNPPFDIQKIAGVFFQYKDTVLDVEFYGQELGDCSCYAFDCLFFSGQSQYHKNFDA RFSHVLEIEENVGQEAKVIAKKYFEVSVPLGSFEGLAPATKDTRKTYLEKAKTGKLG GDLLREAVDKALEYAEKQKFRTDGLVLQPREQGYKNEDTLKWKPAHQITIDFRLKR S AKNEFFI. MMGGPKD EIKFEG IAAKRTPG T V VI PKEFI JDSHPGHDFEGEII JEFGFDK GKSVFIPIIRVRTDKDVPNYIKTVI. SVWGDIFKGVSI. ETLRGETIALPRRYNNII. KQCL LGTSPGAKTDLDIGSGRGGDIQKWREGGFKMVWSIEPNEENIKEFESRARAAKFSSY RLLHAlGXQETQK.'kLKFFGDEKADMASAFYCLGYFAETKKSLEELCEriSSLSRIGAL AVFSFMDGQEIRSLLGKKKKFENSAFSLEKDGEWTSKKYGNAVKVHIKDADSMVKII QKEWTVDLEMLQDAMKKIIGYELI. FSRLADEGVSFLSTQGFEFVSLIIRVATFRKIE

[0341] 13. SEQ ID NO: 13 - EvoBMCE SQLSEWTGSNEVSQFLTKQDANIEVELSFGRWRKTDRHQKFEPGVTKAEFYNLFESL DQISKKSPELFRREDVHTLEETMKSAPGQRANIRRIKNLSTGEVSFLMKDRSKVWDE RNWGLRFAISREEELGVVPNFKAYGFREKQRTRFVFVGEKSAFKGLTVDMTKVVRS SFRDGVAFEVEVEKDSSVPKTAAAMTNAVKRIFEMMQSKSSCQTNEIISMKEKETAL EMFDRLMGSKRYGDTSFLNKPRQVDADDLFEPRLFAVTWKLDGERRLLWFGSTGTDocket No. 10046-673W01

[0342] YLVNPPFDIQKIAGVIAQYKDTVLDVELYGQELGDCSCYAFDCLWFMGQSQYNKNF DTRFSHVLEIEENVGQEARVIAKKYFEVSTSLGSFEGMPTATKDTRKLYLQKARTGK

[0343] 1 GGDI I JKEAVDRAI. EYAEKQKFRTDGLII VPREQGYKNEDTI. KWKPADQMTIDFRL KRKSDNEFFLMMGGPKDRELKFEGTAAKRIPGTVVLPKEFIESHPGHDFEGEILEFGF DLEKSVFVPHRVRTDKDVPNYVKTVLSVWGNIFRGVSLETLRGDTLVLPRKYNNKL KQCLLGIPKDVKTILDIGSGRGGDIQKWQHSGFSSVKMIEPNEENIREFERRAK1TKFT GYKEFHGKAQETQRIVKFFGTEKVDFSSAFYCLGYFAETKKNLEDECNTMSSMSSDK AMAAFAFMDGQAIREILGRKKKFENPAFSLEKEGEWTSKKYGNAVKVIIIKDEGSMV KHQTEWLVDLEMLQDVMKGCGYELVFSRLADEGVSFLSTVGFEFVSLHRVAVFRKI

[0344] 14. SEQ ID NO: 14 - SV40NES-ASFVCE (443) MFLEPPKKNRKVVASLDNLVARYQRCFNDQSLKNSTIELEIRFQQINFLLFKTVYEAL VAQEIPSTISHS1RCIKKVHHENHCREKILPSENLYFKKQPLMFFKFSEPASLGCKVSLA lEQPIRKFTLDSSVLVRLKNRTTFRVSELWKTELTIVKQLMGSEVSAKLAAFKTLLFDT PEQQTTKNMMTIJNPDGEYLYEIEIEYTGKPESETAADVIKIKNTVI. TLISPNIILMI. TA YHQAIEFIASHILSSEILLARIKSGKWGLKRLLPRVKSMTKANYMKFYPPVGYYVTDK ADGIRGIAVIQDrQIYVVADQLYSLGTTGIEPLKPriLDGEFMPEKKEFYGFDVIMYEG Nl I / rQQGFEFR IESESKGIKVEQ AFNIK AEMKPFISI. IS ADPN VI, T, K NFESIFKKK I’RPY SIDGIILVEPGNSYLNTNTFKWKPTW3NTI. DFLVRKCPESI. NVPEYAPKKGFSLHLLF VGISGELFKKLALNWCPGYTKLFPVTQRNQNYFPVQFQPSDFPLAFLYYHPDTSSFSN IDGKVLEMRCLKREINYVRWEIVKIREDRQQDLKTGGYFGNDFKTAELTWLNYMDP FSFEELAKGPSGMYFAGAKTGIYRAQTALISFIKQEIIQKISHQSWVIDLGIGKGQDLG RYEDAGVRHEVGIDKDQTAIAEEVYRKFSHATTRQHKHATNIYVIAIQDLAEPAKEI SEKVUQIYGFPKEGASSIVSNLHHYLMKNTQQVENLAVLCHKLLQPGGMVWFTTM LGEQVLELLHENRffiLNEVWEARENEVVKFAIKRLFKEDILQETGQEIGVLLPFSNGD FYNEYLVNTAHJKIFKIIHGFSLVQKQSFKDWIPEFQNFSKSLYKII. TEADKTWTSLF GFICLRKN

[0345] 15. SEQ ID NO: 15 - T7RNAP (WT) I. NTINIAKNDFSDIELAAIPFNTI. ADIIYGERLAREQLALEHESYEMGEARFRKMFERQ I AGEVADNAAAKPLITTIXPKMIARINDWTEEVKAKRGKRPTAFQFLQEIKPEAVA YF1 IK ITLACL IS ADNTTVQ A VAS AIGRA1EDEARFGRIRDLE AKHFKKN VEEQLNKR VGHVYKKAFMQVV’EADMLSKGLLGGEAWSSWHKEDSIHVGVRCIEMLIESTGMVS LHRQNAGVVGQDSETIELAPEYAEAIATRAGALAGISPMFQPCVVPPKPWTGITGGGDocket No. 10046-673W01

[0346] YWANGRRPLALVRTHSKKALMRYEDVYMPEVYKAINIAQNTAWKINKKVLAVAN

[0347] VrrKWKHCPVEDIPAIEREELPMKPEDIDMNPEALTAWKRAAAAVYRKDKARKSRRI SLEFMLEQANKFANHKAIWFPYNMDWRGRVYAVSMFNPQGNDMTKGLLTLAKGK PIGKEGYYWLKIHGANCAGVDKVPFPERIKFIEENHENIMACAKSPLENTWWAEQDS PFCFLAFCFEYAGVQHHGLSYNCSLPLAFDGSCSGIQHFSAMLRDEVGGRAVNLLPS ETVQDIYGIVAKKVNEILQADAINGTDNEVVTVTDENTGEISEKVKLGTKALAGQWL AYGVTRSVTKRSVMTLAYGSKEFGFRQQVI. EDTIQPAIDSGKGLMFTQPNQAAGYM AKLIWESVSVTVVAAVEAMNWLKSAAKLLAAEVKDKKTGEILRKRCAVHWVTPDG FPVWQEYKKPIQTRLNI. MFEGQFRLQFIINTNKDSEIDAHKQESGIAPNFVHSQDGSH 1. RKTVVW AHEKYGIESFALIHDSFGTIPADAANI FKA VRE TMVDT YESCDVLADEYD QFADQLHESQLDKMPALPAKGNLNLRDILESDFAFA

[0348] 16. SEQ ID NO: 16 - T7RNAP (443) LNT1NIAKNDFSDIELAAIPFNTLADHYGERLAREQLALEHESYEMGEARFRKMFERQ LKAGF. VADNAAAKPLITTLLPKMIARINDWFEEVKAKRGKRPTAFQFI. QEIKPEAVA YITIKTTLACLTSADKTTVQAVASAIGRAIEDEARFGRIRDLEAKIIFKKNVEEQLNKR VGHVYKKAFMQVVEADMLSKGLLGGEAWSSWHKEDSIHVGVRCIEMLIESTGMVS EHRQNAGVVGQDSET1ELAPEYAEAIATRAGAMAGISPMFQPCVVPPKPWTGITGGG YWANGRRPLAEVRTRSKKALMHYEDVYMPEVYKAINIAQNTAWKINKKVLAVAN VITKWKHCPVEDIPAIEREELPMKPEDIDMNPEALTAWKRAAAAVYRKDKARKSRRI SLEFMLEQANKFANHKAIWFPYNMDWRGRVYAVSMFNPQGNDMTKGLLTLAKGK PIGKEGYYWLKIHGANCAGVDKVPFPERIKFIEENHENIMACAKSPLENTWWAEQDS PFCH. AFCFEYAGVQHHGLSYNCSLPLAFDGSCSGIQHFSAMLRDEVGGRAVNIXPS ETVQDIYGIVAKKVNEILQADAINGTDNEVVTVTDENTGEISEKVKLGTKALAGQWL AYGV1RSVTKRSVMTLAYGSKEFGFRRQVLEDTIQPAIDSGKGLMFTQPNQAAGYM AKEIWESVSVTVVAAVEAMNWI. KSAAKLLAAEVKDKKTGEILRKRCAVHWVTPDG FPVWQEYKKPIQTRI. NLMH. GQFRI. QPTINTNKDSEIDARKQESGIAPNFVHSQDGSH LRKTVVWAHEKYGIESFALIHDSFGTIPADAANLFKAVRETMVDTYESCDVLADFYD QFADQLHESQLDKMPALPAKGNLNLRDILESDFAFA

[0349] 17. SEQ ID NO: 17 - EvoT7 LNTINIAKNDFSDIELAAIPFNTLADHYGERLAREQLALEHESYEMGEARFRKMFERS LKAGEVADNAAAKPLriTLLPKMIARINDWFEEVKAKRGKRPTAFQFLQEIKPEAVA YITIKTn. ACLTSADNTTVQAVASAIGRAIEDEARFGRIRDIFAKHFKKNVEEQI. NKRDocket No. 10046-673W01

[0350] VGHVYKKAFMQVVEADMLSKGLLGGEAWSSWHKEDSIHVGVRCIEMLIESTGMVS LHRQNAGVVGQDSETffiLAPEYAEAIATRAGALAGISPMFQPCVVPPKPWTGITGGG YWANGRRPLALVRTHSKKALMRYEDVYMPEVYKAINIAQNTAWKINKKVLAVAN VITKWKHCPVEDIPAIEREELPMKPEDIDMNPEALTAWKRAAAAVYRKDKARKSRRI SLEFMLELANKFANHKAIWFPYNMDWRGRVYAVSMFNPQGNDMTKGLLTLAKGK PIGKEGYYWLKIHGANCAGVDKVPFPERIKFIEENHENIMACAKSPLENTWWAEQDS PFCH. AFCFEYAGVQHHGESYNCSEPEAFDGSCSGIQHFSAMLRDEVGGRAVNIXPS ETVQDIYG IVAKKVNEILQKD AINGTDNE VVTVTDENTGEISEKSKLGTKAL AG QWL AYGVTRSVTKRSVMTLAYGSKEFGFRQQVLEDTIQPAIDSGKGLMFTQPNQAAGYM AKLIWESVSVTVVAAVEAMNF KSAAKLLAAEVKDKKTGEILRKRCAVHWVTPDG FPVWQEYKKPIQTRLNLMFLGQFRLQPTINTNKDSEIDAHKQESGIAPNEVHSLDGSH LRKTVVWAHEKYGIESFALIHDSFGTIPADAANLFKAVRETMVDTYESCDVLADFYD QFADQLHESQLDKMPALPAKGNLNLRDILESDFAFA

[0351] 18. SEQ ID NO: 18 - EvoT7 (443) LNTINIAKNDFSDIELAAIPFNTLADHYGERLAREQLALEHESYEMGEARFRKMFERS LKAGEVADNAAAKPLITTLLPKMIARINDWFEEVKAKRGKRPTAFQFLQEIKPEAVA YITIK TTLACI - TS ADKTTVQ A VAS AIGR AIEDEARFGRIRDLE AKHFKKN VEEQLNKR VGHVYKKAFMQVVEADMESKGELGGEAWSSWHKEDSIHVGVRCIEMLIESTGMVS LHRQNAGVVGQDSETIELAPEYAEAIATRAGAMAGISPMFQPCVVPPKPWTGITGGG YWANGRRPLALVRTRSKKALMHYEDVYMPEVYKAINIAQN'rAWKINKKVLAVAN V1TK KHCPVEDIPAIEREELPMKPEDIDMNPEALTAWKRAAAAVYRKDKARKSRRI SLEFMI. ELANKFANHKAIWFPYNMDWRGRVYAVSMFNPQGNDMTKGI. LTLAKGK PIGKEGYYV^TKIHGANCAGVDKVPFPERIKFIEENHENIMACAKSPLENTWWAEQDS PFCFLAFCFEYAGVQHHGLSYNCSLPLAFDGSCSGIQHFSAMLRDEVGGRAVNLLPS ETVQDIYGIVAKKVNFJEQKDAINGTDNEVVTVTDENTGFJSEKSKEGTKALAGQWL AYGVTRSVTKRSVMTLAYGSKEFGFRRQW. EDTIQPAIDSGKGLMFTQPNQAAGYM AKLIWESVSVTVVAAVEAMNFLKSAAKLLAAEVKDKKTGEILRKRCAVHWVTPDG FI WQEYKKI’IQTRLNLMFLGQFRLQPriN'INKDSEIDARKQESGIAPNFVHSLDGSH ERKTVVWAHF YGIESFAI HDSFGTIPADAANEFKAVRETMVDTYESCDVLADFYD QFADQLHESQI. DKMPALPAKGNLNI. RDILESDFAFA

Claims

Docket No. 10046-673W01CLAIMSWhat is claimed is:

1. An engineered T7 RNA polymerase (T7 RNAP), wherein said T7 RNAP comprises SEQ ID NO: 15, and further wherein SEQ ID NO: 15 comprises at least one substitution mutation selected from Q58S, Q404L, A584K, V609S, W698F, and / or Q786L, or any combination thereof.

2. The engineered T7 RNAP of claim 1, wherein SEQ ID NO: 15 comprises at least one additional mutation other than Q58S, Q404L, A584K, V609S, W698F, and / or Q786L.

3. The engineered T7 RNAP of claim 1 or 2, wherein the T7 RNAP is operably linked to a capping enzyme comprising SEQ ID NO: 11, or a variant thereof.

4. The engineered T7 RNAP of claim 3, wherein the variant of SEQ ID NO: 11 comprises at least one substitution mutation of S199A, N267W, and / or Q426V.

5. An engineered capping enzyme, wherein the engineered capping enzyme comprises SEQ ID NO: 11, and wherein SEQ ID NO: 11 comprises at least one substitution mutation comprising S199A, N267W, and / or Q426V.

6. The engineered capping enzyme of claim 5, wherein the engineered capping enzyme is operably linked to a T7 RNA polymerase comprising SEQ ID NO: 15, or a variant thereof.

7. The engineered capping enzyme of claim 6, wherein the variant of SEQ ID NO: 15 comprises at least one substitution comprising Q58S, Q404L, A584K, V609S, W698F, and / or Q786L.

8. An enzyme complex comprising a 1'7 RNA polymerase (1’7 RNAP) component and a capping enzyme component, wherein the T7 RNAP component comprises SEQ ID NO: 15, or a variant thereof, and further wherein the capping component comprises SEQ ID NO: 11, or a variant thereof, wherein a linker comprising at least 90% sequence identity to SEQ ID NO: 2 separates the T7 RNAP component and the capping enzyme component, wherein the variant of SEQ ID NO: 15 comprises at least one mutationDocket No. 10046-673W01selected from Q58S, Q404L, A584K, V609S, W698F, and / or Q786L, and wherein the variant of SEQ ID NO: 11 comprises at least one mutation selected from S199A, N267W, and / or Q426V.

9. The enzyme complex of claim 8, wherein the enzyme complex is a component in a transcription system, a post-transcription system, or a combination thereof.

10. The enzyme complex of claim 8 or 9, wherein the post-transcription system comprises protein translation.

11. A polynucleotide sequence encoding the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any one of claims 1-10.

12. A cell comprising the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any one of claims 1-10.

13. The cell of claim 12, wherein the cell is a eukaryotic cell or a prokaryotic cell.

14. The cell of claim 12 or 13, wherein the cell proliferates into a stable cell line or a transient cell line.

15. A cell comprising a polynucleotide sequence encoding the T7 RNA polymerase, the capping enzyme, or the enzyme complex of any one of claims 1-10.

16. A cell culture composition comprising the cell of any one of claims 12-15.

17. The cell culture composition of claim 16, wherein the cell is a tissue culture cell.

18. An in vitro system comprising one or more polynucleotide sequences encoding the enzyme complex of any one of claims 8-10 and a reporter gene, wherein the polynucleotide sequence is under the control of a T7 RNAP promoter.Docket No. 10046-673W0119. A method of producing a capped RNA transcript, the method comprising performing a transcription reaction utilizing the enzyme complex of any one of claims 8-10, wherein the method further comprises a composition comprising:a) one or more components needed to produce a methylguanylate cap; and b) a DNA template, wherein said DNA template is exposed to said composition under conditions such that the enzyme complex produces a capped RNA transcript.

20. A method of producing RNA, the method comprising:a) expressing the enzyme complex of any one of claims 8-10 in a cell, and b) isolating an amount of RNA from the cell that is increased relative to an otherwise identical control cell expressing the T7 RNA polymerase lacking the capping enzyme.