RNA polymerase variants and methods of their use

T7 RNA polymerase variants with targeted amino acid mutations address yield and impurity issues in mRNA manufacturing, enhancing transcription efficiency and reducing costs.

WO2026015073A1PCT designated stage Publication Date: 2026-01-15AGENCY FOR SCI TECH & RES
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/SG2025/050426
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-06-23
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Current T7 RNA polymerase technologies face challenges such as low RNA yields, 3'-inhomogeneity, formation of dsRNA impurities, limited nucleotide substrate specificity, and high costs, particularly when deviating from optimal DNA sequences, necessitating improved RNA polymerase variants for efficient mRNA manufacturing.

Method used

Development of T7 RNA polymerase variants with specific amino acid mutations, including R57H, R173S, A260S, P451L, D471N, G520E, A586T, S767G, and D859N, enhancing transcription efficiency and yield across various sequences and conditions.

Benefits of technology

The mutated T7 RNA polymerase variants exhibit improved RNA yields, increased transcription efficiency, and reduced impurities, facilitating more efficient mRNA production and lowering production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000004_0001
    Figure IMGF000004_0001
  • Figure IMGF000005_0001
    Figure IMGF000005_0001
  • Figure IMGF000006_0001
    Figure IMGF000006_0001
Patent Text Reader

Abstract

Various embodiments relate generally to polypeptides having RNA polymerase activity and to nucleic acids encoding those as well as methods of their use. More particularly, there are provided embodiments related to T7 RNA polymerase variants, fusion proteins comprising the T7 RNAP variants, nucleic acids encoding the same and methods of their use. Additionally, there are various embodiments related to split versions of the T7 RNA polymerase variants and methods of their use.
Need to check novelty before this filing date? Find Prior Art

Description

RNA POLYMERASE VARIANTS AND METHODS OF THEIR USECROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of priority of Singapore Patent Application No. 10202402036X filed 10 July 2024, the content of which being hereby incorporated by reference in its entirety for all purposes.TECHNICAL FIELD

[0002] The present invention lies in the technical field of enzyme technology and specifically relates to enzymes having RNA polymerase activity and to nucleic acids encoding those as well as methods of the manufacture of said enzymes. Further encompassed are methods and uses of these enzymes.BACKGROUND

[0003] mRNA therapeutics is a rapidly expanding field of medicine with significant applications in vaccine development and cancer immunotherapy. To enable the widespread adoption of mRNA medicine, robust protocols for mRNA manufacturing are essential. In fact, mRNA manufacturing and synthesis services market is expected to worth USD1.5B by 2035. In vitro transcription (IVT) using T7 RNA polymerase (T7 RNAP) is the primary process to synthesize mRNA. IVT yields and purity are heavily dependent on the initiation sequence and the sequence following the initiation site. Templates that deviate from the optimal sequence often sufferfrom low yields and 3’-inhomogeneity. Additional challenges associated with current T7 RNAP technology include formation of dsRNA impurities, limited nucleotide substrate specificity and high cost of co-transcriptional capping. Protein engineering efforts by academic research groups and companies have resulted in new T7RNAP variants that show improvement in one or more of the abovementioned properties.

[0004] Previous efforts to increase RNA yield rely heavily on (I) engineering the T7 promoter sequence, (ii) designing DNA template sequence to avoid secondary structures and (iii) optimizing various reaction conditions like ion concentrations and nucleotide triphosphate concentrations. These methods are laborious and time consuming. Provision of an engineered T7 RNAP that improves RNA yields across various sequences, compositions and reaction conditions will increase the ease of manufacturing RNA.

[0005] Given the competitive T7RNAP landscape, there is a need in the art to develop and provide RNA polymerase variants that address the ongoing challenges associated with T7 RNAP; including product homogeneity, stability and / or RNA yields. In particular, there is a need in the art for the provision of variants of T7 RNA polymerases with improved activity and increased RNA yields in methods of their use.SUMMARY

[0006] In one aspect, there is provided a polypeptide having RNA polymerase activity, comprising or consisting of : (I) an amino acid sequence set forth in SEQ ID NO:1 ;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in SEQ ID NO:1 ; or (iii) a functional fragment of (i) or (ii); wherein the polypeptide comprises one or more amino acid mutations at a position selected from R57, R173, A260, P451 , D471 , G520, A586, S767, D859 and combinations thereof, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

[0007] In various embodiments, the amino acid mutation is an amino acid substitution, and wherein: the amino acid residue R at position 57 is substituted with a basic amino acid residue, preferably H; and / or the amino acid residue R at position 173 is substituted with a polar amino acid residue, preferably S; and / or the amino acid residue A at position 260 is substituted with a polar amino acid residue, preferably S; and / or the amino acid residue P at position 451 is substituted with a non-polar amino acid residue, preferably a non-polar aliphatic amino acid, more preferably L; and / or the amino acid residue D at position 471 is substituted with a polar amino acid residue, preferably a polar amino acid with an amide side chain, more preferably N; and / or the amino acid residue G at position 520 is substituted with a polar amino acid residue, preferably a polar acidic amino acid, more preferably E, or a non-polar amino acid residue, preferably a non-polar aliphatic amino acid, more preferably A; and / or the amino acid residue A at position 586 is substituted with a non-polar amino acid residue, or a polar amino acid residue, more preferably T; and / or the amino acid residue S at position 767 is substituted with a polar amino acid residue, preferably G, or an aliphatic amino acid, preferably a non- bulky aliphatic amino acid, more preferably G or A; and / or the amino acid residue D at position 859 is substituted with a polar amino acid residue, preferably a polar amino acid with an amide side chain, more preferably N, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

[0008] In various embodiments, the amino acid residue Q at position 744 is invariable, and / or the amino acid residue S at position 430 is invariable, and / or the amino acid residue C at position 510 is invariable, and / or the amino acid residue F at position 880 is invariable, and / or the amino acid residue F at position 849 is invariable, optionally the amino acid residues at positions 878-883 are invariable, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1 .

[0009] In various embodiments, the one or more amino acid mutations are selected from R57H, R173S. A260S, P451 L D471 N, G520E, A586T, S767G. D859N and combinations thereof.

[0018] In various embodiments, the one or more amino acid mutations consist of: R173S, P451 L, D471 N and G520E; or A586T; or S767G; or A260S; or R57H; or R57H, A260S and D859N.

[0011] In various embodiments, the amino acid sequence (ii) and functional fragment (iii), comprise the amino acid sequence corresponding to residues 267-811 of SEQ ID NO:7.

[0012] In various embodiments, the polypeptide comprises or consists of the amino acid sequence set forth in SEQ ID NO:2; the amino acid sequence set forth in SEQ ID NO:3; the amino acid sequence set forth in SEQ ID NO:4; the amino acid sequence set forth in SEQ ID NO:5; the amino acid sequence set forth in SEQ ID NO:6; or the amino acid sequence set forth in SEQ ID NO:7.

[0013] In another aspect, there is provided a fusion protein comprising: a deaminase domain capable of catalyzing the deamination of nucleic acids; and a polypeptide disclosed herein: wherein the deaminase domain and the polypeptide are linked to form a single polypeptide chain.

[0014] In various embodiments, the deaminase domain is selected from the group consisting of cytidine deaminases, adenosine deaminases, and guanine deaminases.

[0015] In various embodiments, the fusion protein further comprises a linker peptide to link the deaminase domain to the polypeptide, such that the deaminase domain retains its enzymatic activity and the polypeptide retains its transcriptional activity.

[0016] In another aspect, there is provided a composition comprising the polypeptide disclosed herein and optionally an in vitro transcription (IVT) reagent, or a fusion protein disclosed herein.

[0017] In another aspect, there is provided a nucleic acid molecule encoding the polypeptide disclosed herein or fusion protein disclosed herein.

[0018] In various embodiments, the nucleic acid molecule is comprised in a vector preferably an expression vector, wherein said vector further comprises regulatory elements for controlling expression of said nucleic acid molecule, optionally the nucleic acid molecule is operably linked to a promoter suitable for expression in a host cell.

[0019] In another aspect, there is provided a host cell comprising the nucleic acid molecule disclosed herein.

[0020] In another aspect, there is provided a method for producing the polypeptide disclosed herein or fusion protein disclosed herein, comprising culturing a host cell disclosed herein under conditions that allow expression of the polypeptide or fusion protein, and isolating said polypeptide or fusion protein from the host cell or culture medium.

[0021] In another aspect, there is provided a cell-free method for producing the polypeptide disclosed herein or fusion protein disclosed herein, comprising subjecting the nucleic acid molecule disclosed herein to reaction conditions that allow the transcription and translation of the polypeptide or fusion protein, and isolating said polypeptide or fusion protein.

[0022] In another aspect, there is provided a method for RNA production or amplification to produce RNA transcript, comprising: contacting a nucleic acid template with the polypeptide disclosed herein under conditions that result in the production of RNA transcript.

[0023] In another aspect, there is provided a method of performing an in vitro transcription (IVT) reaction, comprising: contacting a nucleic acid template with the polypeptide disclosed herein in the presence of nucleoside triphosphates under conditions that result in the production of RNA transcript.

[0024] In another aspect, there is provided a method for heterologous expression of a protein of interest in a host cell, comprising: I) introducing into the host cell a nucleic acid sequence encoding the polypeptide disclosed herein, and a heterologous nucleic acid sequence encoding the protein of interest, wherein the heterologous nucleic acid sequence is operably linked to a T7 promoter; and ii) inducing expression of the polypeptide in the host cell under suitable conditions, wherein the induced expression of the polypeptide drives transcription of the heterologous nucleic acid sequence and expression of the protein of interest.

[0025] In various embodiments, the protein of interest is a recombinant protein selected from the group consisting of a therapeutic protein, an enzyme, and an antibody.

[0026] In various embodiments, the polypeptide disclosed herein is for use in RNA production or amplification to produce RNA transcript; or in performing an IVT reaction for producing RNA transcript; or heterologous expression of a protein of interest in a host cell, or detecting biomolecular interactions involving a target analyte.

[0027] In another aspect, there is provided a method of introducing one or more mutations in a target nucleic acid sequence in a host cell, comprising: introducing into the host cell a nucleic acid molecule encoding the fusion protein disclosed herein, preferably the nucleic acid molecule is a plasmid vector; and inducing expression of the fusion protein in the host cell to transcribe the target nucleic acid sequence, wherein during transcription the deaminase domain of the fusion protein introduces one or more mutations into the target nucleic acid sequence.

[0028] In various embodiments, the target nucleic acid sequence comprises a reporter gene for detecting the one or more mutations, and the target nucleic acid sequence is operably linked to a T7 promoter.

[0029] In various embodiments, the fusion protein disclosed herein is for use in introducing one or more mutations in a target nucleic acid sequence in a host cell.

[0030] In another aspect, there is provided a method for detecting biomolecular interactions involving a target analyte, comprising the steps of: providing a split version of the polypeptide disclosed herein, wherein the split version is linked to a detection system that facilitates reconstitution of the split version into a transcriptionally active polypeptide in the presence of the target analyte; contacting the split version with a sample suspected of containing the target analyte in the presence of a reporter gene construct comprising a reporter gene operably linked to a promoter recognized by the transcriptionally active polypeptide, wherein interaction with the target analyte induces reconstitution of the split version into the transcriptionally active polypeptide; and detecting the expression of the reporter gene as an indication of the biomolecular interaction involving the target analyte.

[0031] In various embodiments, the split version comprises two or more non-functional fragments that reconstitute into the transcriptionally active polypeptide upon interaction with the target analyte, wherein each non-functional fragment is operably linked to a binding or recognition partner specific to the target analyte, and wherein the interaction between the binding or recognition partner and the target analyte facilitates reconstitution of the non-functional fragments into the transcriptionally active polypeptide.

[0032] In various embodiments, the two or more non-functional fragments comprise a first nonfunctional fragment comprising an amino acid sequence corresponding to positions 1-179, 1 -510, 1 - 563, or 1 -601 of SEQ ID NO:1 , and a second non-functional fragment comprising an amino acid sequence corresponding to positions 180-883, 51 1 -883, 564-883, or 602-883 of SEQ ID NO:1.

[0033] In various embodiments, the two or more non-functional fragments comprise a first nonfunctional fragment comprising an amino acid sequence corresponding to positions 1 -67 of SEQ ID NO:1 , a second non-functional fragment comprising an amino acid sequence corresponding to positions 68-179 of SEQ ID NO:1 , a third non-functional fragment comprising an amino acid sequence corresponding to positions 180-601 of SEQ ID NO:1 , and a fourth non-functional fragment comprising an amino acid sequence corresponding to positions 602-883 of SEQ ID NO:1 .

[0034] In various embodiments, the expression of the reporter gene is quantified to determine the amount of the target analyte in the sample.

[0035] In another aspect, there is provided a split version of the polypeptide disclosed herein, for use in detecting biomolecular interactions.BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Various embodiments will be better understood with reference to the detailed descriptionwhen considered in conjunction with the non-limiting examples and the accompanying drawings.

[0037] FIG. 1 shows IVT templates (IVT-13, 22. 28, 24a, 25, 25b) used where +1 is the first nucleotide after the T7 promoter. Lengths are not drawn to scale. In each template the nucleotide sequence of the T7 promoter (TAATACGACTCACTATA: SEQ ID NO: 8) and nucleotides directly thereafter (and before either the broccoli aptamer, 5’UTR or 5’intron) are indicated as follows: IVT-13: TAATACGACTCACTATAGGG (SEQ ID NO: 9); IVT-22:TAATACGACTCACTATAAGGGGAAT (SEQ ID NO: 10); IVT-28:TAATACGACTCACTATAAGGAGACC (SEQ ID NO: 11); IVT- 24a:TAATACGACTCACTATAGGGAGACC (SEQ ID NO: 12); IVT-25:TAATACGACTCACTATA AGGAGACC (SEQ ID NO: 13); IVT-25b:TAATACGACTCACTATAAGGAGACC (SEQ ID NO: 14), where the GGG may be included to improve transcriptional yield and the AGG may be included to initiate transcription with co-capping agent.

[0038] FIG. 2 shows time-course broccoli fluorescence during IVT of template IVT-22 by the different T7 variants. The initial rate was estimated by the gradient of the tangent line between 6 to 12 min into the reaction. In the graph, the order of the lines from highest A.U to lowest is as follows: T7-295 > T7- 259 > T7-260 > T7-258 >T7>286>T7-184 > T7-256 . Calibrated RNA concentration was measured by RNA Tapestation of the IVT-28. IVT-28 is polyadenylated IVT-22 without the HHR-broccoli sequence.

[0039] FIG. 3 shows time-course broccoli fluorescence during IVT of template IVT-24a by the different T7 variants. The initial rate was estimated by the gradient of the tangent line between 6 to 12 min into the reaction. In the graph, the order of the lines from highest A.U to lowest is as follows: T7-295 > T7-260 / T7-259 > T7-258 > T7-256 > T7-184 / T7-286.

[0040] FIG. 4 shows time-course broccoli fluorescence during IVT of template IVT-25b by the different T7 variants. The initial rate was estimated by the gradient of the tangent line between 6 to 12 min into the reaction. Calibrated RNA concentration was measured by RNA Tapestation of IVT- 25. IVT-25 is IVT -25b without the HHR-broccoli sequence. In the graph, the order of the line from highest A.U to lowest is as follows: T7-295 > T7-256> T7-286 >T7-258 > T7-259 > T7-260 > T7-184.

[0041] FIG. 5 shows RNA Tapestation analyses of purified IVT-28 RNA. Expected size of IVT-28 is 1354-bp. RIN: RNA integrity score. RIN of >6 is generally accepted as acceptable RNA integrity.

[0042] FIG. 6 shows RNA Tapestation analyses of purified IVT-25 RNA. Expected size of IVT-25 is 4316-bp. RIN: RNA integrity score. RIN of >6 is generally accepted as acceptable RNA integrity.

[0043] FIG. 7A shows the broccoli fluorescence obtained from IVT of template IVT-13 with the indicated T7 RNAP. His-tagged T7 RNAP was bead-purified from 1 ml cultures and used to transcribe template IVT-13. To normalize the fluorescence values to protein abundance, purified proteins wererun on a PAGE-Urea gel and the abundance was quantified by the band intensity using Imaged. Values represent n=1 ; and FIG. 7B shows the broccoli fluorescence obtained from IVT of template IVT-24a with indicated T7 RNAP. His-tagged T7 RNAP was bead-purified from 1 ml cultures and used to transcribe template IVT-24a. To normalize the fluorescence values to protein abundance, purified proteins were run on a PAGE-Urea gel and the abundance was quantified by the band intensity using Imaged. Values represent n=1.

[0044] FIG. 8 shows an RNA Tapestation gel image of IVT-25 transcribed by the indicated T7 variants. His-tagged T7 RNAP variants were column-purified from 50 ml bacterial cultures, buffer- exchanged into a storage buffer, and protein concentrations were quantified by Nanodrop. 20 nM of protein and 20 nM of DNA template were used for IVT. The IVT reaction mixture was purified by SPRI beads before analyses by Tapestation. RIN= RNA integrity score. RIN of >6 is generally accepted as acceptable RNA integrity.

[0045] FIG. 9 shows the High-throughput sequencing assay to calculate the transcription fidelity of T7-295 variant and WT T7. Barcoded gene-specific primer used during reverse transcription eliminates mutation artefacts arising from PCR and sequencing.

[0046] FIG. 10A shows a template transcribed by WT T7 and T7-295 in the presence of AG capping reagent. Hammerhead ribozyme undergoes autocleavage to release a 17-nt long 5’ fragment This fragment will be capped or uncapped depending on the capping efficiency of the T7 polymerase; and FIG. 10B shows the expected PAGE gel analysis. Capping efficiency is calculated from the relative gel densitometries of the 5’ cleaved fragments.

[0047] FIG. 11A shows 10% PAGE-UREA gel of IVT reaction following transcription of ribozyme by WT T7 or T7-293. Gel was stained with SYBR Gold; and FIG. 11B shows mRNA capping efficiencies of WT T7 and T7-295 at different ratios of AG Clean cap to GTP concentration.DETAILED DESCRIPTION

[0048] The following detailed description refers to, by way of illustration, specific details and embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. Other embodiments may be utilized and structural and logical changes may be made without departing from the scope of the invention. Embodiments described below in context of the polypeptides, fusion proteins, nucleic acids, host cells are analogously valid for the respective methods, and vice versa. The various embodiments are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The singular terms "a," "an," and"the" include plural referents unless context clearly indicates otherwise. Similarly, the word "or" is intended to include "and" unless the context clearly indicates otherwise. The term "comprises" means "includes." In case of conflict, the present specification, including explanations of terms, will prevail. “About”, as used herein in connection with numerical values refers to the referenced numerical value ±10% or ±5%.

[0050] Through phage-assisted non-continuous evolution, the present inventors developed T7 RNA polymerase variants containing one or more mutations relative to its parental sequence that have been shown to confer superior transcription kinetics and RNA yield. The inclusion of a single or combination of mutations has been identified that contribute towards the provision of a more productive T7 RNA polymerase (T7 RNAP) relative to its parental sequence.

[0051] Accordingly, there is provided herein a novel polypeptide having RNA polymerase activity. In various embodiments, the polypeptide disclosed herein may have T7 RNA polymerase activity and may be referred to as a T7 RNA polymerase variant.

[0052] The term “T7 RNA polymerase” refers to an RNA polymerase from the T7 bacteriophage that catalyzes the formation of RNA from DNA in the 5' — > 3' direction. T7 RNA polymerase is highly specific for the T7 promoter. The polypeptides disclosed herein have T7 RNA polymerase activity, i.e. the polypeptides transcribe DNA specifically starting at a T7 promoter. The polypeptide (i.e. T7 RNA polymerase variant) may be derived from a Wild-type T7 RNA polymerase. The wild-type (WT) T7 RNA polymerase refers to the naturally occurring, unmodified form of RNA polymerase encoded by the gene 1 of bacteriophage T7. The amino acid sequence of wild-type T7 RNA polymerase is well- characterized and commonly used as a reference for engineered or mutant variants in research and biotechnology applications, whereby information may be accessed via online databases such as Uniprot: P00573. In various embodiments, the amino acid sequence of the wild-type T7 RNA polymerase is set forth in SEQ ID NO:1, as shown in Table 1 below. Accordingly, in various embodiments, the polypeptide disclosed herein is derived from the T7 RNA polymerase of SEQ ID NO: 1 and comprises one or more mutations.

[0053] Table 1 : Wild-Type T7 RNA Polymerase amino acid sequence, and domains indicated by shading, underlining and bold lettering as follows.

[0054] The wild-type T7 RNA polymerase (WT-T7 RNAP) comprises three major structural and functional domains. The N-terminal domain (amino acids 1-266 of SEQ ID NO:1 ) is primarily involved in promoter recognition and binding, facilitating the initiation of RNA transcription. The central catalytic domain (amino acids 267-811 of SEQ ID NO:1 ) contains the active site responsible for the core polymerase activity, including nucleotide binding and RNA chain elongation. The C-terminal domain (amino acids 812-883 of SEQ ID NO:1) contributes to the structural stability of the enzyme and is also implicated in interactions with the DNA template and incoming nucleotides.

[0055] In various embodiments, the amino acid residue D at position 537 may be invariable, and / or the amino acid residue D at position 812 may be invariable, and / or the amino acid residue D at position 813 is invariable, whereby numbering is relative to SEQ ID NOR . Basis and reasoning for these residues being invariable can be found in Osumi-Davis et al (J Mol Biol. 1992 Jul 5;226(1 ) :37- 45 https: / / pubmed.ncbi.nlm.nih.gov / 1619661 / ).

[0056] In various embodiments, the amino acid residue Q at position 744 may be invariable, and / or the amino acid residue S at position 430 is invariable, and / or the amino acid residue C at position 510 may be invariable, and / or the amino acid residue F at position 880 is invariable, and / or the amino acid residue F at position 849 may be invariable, whereby numbering is relative to SEQ ID NO:1. Basis and reasoning for these residues being invariable can be found in US Patent No. US9540670B2.

[0057] In various embodiments, the amino acid residues at positions 878-883 may be invariable, whereby numbering is relative to SEQ ID NO:1 . Residues at positions 878, 879, 880, 881 , 882, and / or 883 are located in the C-terminal region, which is important for enzymatic activity, particularly in promoter recognition, DNA binding, and formation of an active transcription complex (e.g. See Cheetham & Steitz (1999,), Nature, 399(6731), 80-83. https: / / doi.Org / 10.1038 / 19999).

[0058] The term “RNA polymerase activity” refers to the enzymatic function of RNA polymerases in catalysing the synthesis of RNA from a nucleic acid template, typically DNA, through the sequential incorporation of ribonucleoside triphosphates (rNTPs) into a growing RNA strand in a 5' to 3' direction. This activity includes initiation at a specific promoter sequence, elongation of the RNA transcript, and, in some cases, termination of transcription. In the context of T7 RNA polymerase, RNA polymeraseactivity is characterized by the enzyme’s ability to specifically recognize and bind to a T7 promoter sequence on a double-stranded DNA template and to transcribe downstream sequences with high efficiency and fidelity. Thus, T7 RNAP activity may refer to transcription initiation at T7-specific promoters and not general RNAP activity. T7 RNA polymerase is a single-subunit enzyme that exhibits strong promoter specificity and processivity, producing RNA transcripts without the need for additional transcription factors. RNA polymerase activity, including that of T7 RNA polymerase, may be measured using in vitro transcription assays by quantifying RNA products generated from a DNA template containing the appropriate promoter. Analytical methods include incorporation of labelled rNTPs, gel electrophoresis, qRT-PCR, and high-throughput sequencing to assess transcript yield, length, and error rates such as base substitutions, insertions, and deletions.

[0059] The polypeptides of the present invention preferably have enzymatic activity, in particular RNA polymerase activity. In various embodiments, this means that they can perform transcription with an efficiency of 60 % or more, preferably 70 % or more, more preferably 80% or more, relative to the transcription efficiency of a reference RNA polymerase enzyme, or parental polypeptide sequence. In various embodiments, the polypeptides of the present invention exhibit enhanced enzymatic activity, particularly RNA polymerase activity. This means that the polypeptides are capable of performing transcription with an efficiency of 110% or more, preferably 125% or more, and more preferably 150% or more, relative to the transcription efficiency of a reference RNA polymerase enzyme, such as a wild-type or parental polypeptide comprising the amino acid sequence of SEQ ID NO:1.

[0060] In various embodiments, the polypeptides disclosed herein have at least 50 %, more preferably at least 70, most preferably at least 90 % of the RNA polymerase activity of the enzyme having the amino acid sequence of SEQ ID NO:1 . In various embodiments, the polypeptides disclosed herein exhibit an increased RNA polymerase activity relative to the enzyme having the amino acid sequence of SEQ ID NO:1. In some embodiments, the polypeptides have at least 110%, more preferably at least 125%, and most preferably at least 150% of the RNA polymerase activity of the enzyme comprising SEQ ID NO:1 .

[0061] In various embodiments, the polypeptide disclosed herein may exhibit one or more improved enzyme activity. The term “improved enzyme property” refers to any enzyme property made better or more desirable for a particular purpose as compared to that property found in a reference enzyme. For the polypeptides disclosed herein, the comparison is generally made to a reference T7 RNA polymerase enzyme (i.e. WT T7 RNA polymerase) which does not contain the particular mutations which improve enzyme efficiency. However, in some embodiments, the reference T7 RNA polymerase can be another improved T7 RNA polymerase variant.

[0062] In various embodiments, the polypeptide disclosed herein may exhibit improved or increased transcription efficiency relative to its parent sequence (i.e. wild-type RNA polymerase) or a referenceenzyme. “Transcription efficiency” herein refers to RNA yield, and / or RNA quality, and / or rate of transcription. Highly efficient RNA polymerase may refer to an RNA polymerase that when used in an in vitro transcription reaction, improve the transcription efficiency. An efficient and highly active RNA polymerase results in increased production of specific target products, thereby lowering purification and production expenses. Moreover, it enables the creation of diverse RNA targets, surpassing the capabilities of existing polymerases.

[0063] In various embodiments, the polypeptide may be an isolated polypeptide. The term “isolated”, as used herein, relates to the polypeptide in a form where it has been at least partially separated from other cellular components it may naturally occur or associate with. The polypeptide may be a recombinant polypeptide, i.e. polypeptide produced in a genetically engineered organism that does not naturally produce said polypeptide.

[0064] In various embodiments, the polypeptides having RNA polymerase activity may be post- translationally modified, for example glycosylated. Such modification may be carried out by recombinant means, i.e. directly in the host cell upon production, or may be achieved chemically or enzymatically after synthesis of the polypeptide, for example in vitro.

[0065] In various embodiments, the polypeptide disclosed herein comprises or consists of the amino acid sequence set forth in SEQ ID NO:1, or a functional variant or fragment thereof, and comprises one or more amino acid mutations. In this regard, the polypeptide disclosed herein comprises one or more amino acid mutations, with the proviso that the polypeptide is not the wild-type T7 RNAP polymerase having the amino acid sequence set forth in SEQ ID NO:1.

[0066] The term “amino acid”, as used herein refers to natural and / or unnatural or synthetic amino acids, including both the D and L optical isomers, amino acid analogs (for example norleucine is an analog of leucine) and derivatives known in the art. The term “natural amino acid”, as used herein, relates to the 20 naturally occurring L-amino acids, namely Gly (G), Ala (A), Vai (V), Leu (L), lie (I), Phe (F), Cys (C), Met (M), Pro (P), Thr (T), Ser (S), Glu (E), Gin (Q), Asp (D), Asn (N), His (H), Lys (K), Arg (R), Tyr (Y) , and Trp (W). Generally, in the context of the present application, the polypeptides are shown in the N- to C-terminal orientation. All amino acid residues are generally referred to herein by reference to their one letter code and, in some instances, their three letter code. This nomenclature is well known to those skilled in the art and used herein as understood in the field. As a person skilled in the art would appreciate, amino acids can be categorized in different classes depending upon the chemical and physical properties of the amino acid residue. Amino acids may be grouped according to similarities in the properties of their side chains (in A. L. Lehninger, in Biochemistry, second ed., pp. 73-75, Worth Publishers, New York (1975)): (1 ) non-polar: Ala, Vai, Leu, lie, Pro, Phe, Trp, Met; (2) uncharged polar: Gly, Ser, Thr, Cys, Tyr, Asn, Gin; (3) acidic: Asp, Glu; and (4) basic: Lys, Arg, His. Alternatively, naturally occurring residues may be divided into groups based on common sidechain properties: (1 ) hydrophobic: Norleucine, Met, Ala, Vai, Leu, He; (2) neutral hydrophilic: Cys, Ser,Thr, Asn, Gin; (3) acidic: Asp, Glu; (4) basic: His, Lys, Arg; (5) residues that influence chain orientation: Gly, Pro; and (6) aromatic: Trp, Tyr, Phe. As one of skill in the art would appreciate, amino acids can be categorized in different classes depending upon the chemical and physical properties of the amino acid residue. Typically, hydrophobic amino acids can be further classified as having an aliphatic side chain or an aromatic side chain. Aliphatic amino acids and aromatic amino acids are known one of skill in the art. Typically, aliphatic amino acids have a side chain containing hydrogen and carbon atoms. Examples of aliphatic amino acids include alanine, isoleucine, proline, and valine. Typically, aromatic amino acids contain a side chain containing an aromatic ring. Examples of aromatic amino acids include phenylalanine, tyrosine, histidine and tryptophan.

[0067] The term “amino acid mutation”, as used herein, refers to any mutation such as substitution, deletion and also insertion of an amino acid residue at a position corresponding to the reference T7 RNA polymerase sequence, for example, the amino acid sequence set forth in SEQ ID NO:1. The mutation may be a conservative and / or non-conservative mutation, more particularly a conservative and / or non-conservative substitution. The term "conservative amino acid substitution" means the exchange (substitution) of one amino acid residue for another amino acid residue, where such exchange does not lead to a considerable change in the polarity or charge or size at the position of the exchanged amino acid, e.g. the exchange of a nonpolar amino acid residue for another nonpolar amino acid residue. Conservative amino acid substitutions in the context of the invention encompass, for example, G=A=S, l=V=L=M, D=E, N=Q, N=Q=S=T, K=R, K=R=H, Y=F=W, S=T, S=T=C, G=A=I=V=L=M=Y=F=W=P=S=T. In various embodiments, the amino acid mutation is an amino acid substitution. The amino acid mutation may also be a non-conservative mutation. The amino acid mutation may be an amino acid substitution or a non-conservative amino acid substitution.

[0068] In various embodiments, the polypeptides disclosed herein comprise at least one amino acid mutation at a position corresponding to position R57, R173, A260, P451 , D471 , G520, A586, S767, D859 and combinations thereof. The position numbering of the amino acid residues and mutations thereof are in accordance with the amino acid residue numbering of SEQ ID NO:1.

[0069] In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation at the position R173, P451 , D471 and G520. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation at the position A586. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation at the position S767. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation at the position A260. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation at the position R57. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation at the position R57, A260 and D859.

[0070] In various embodiments, the one or more mutations are amino acid substitutions selected from those listed in the below Table 2, with amino acid (AA) position numbering relative to WT-T7RNAP (SEQ ID NO:1 ).Table 2: Amino Acid Mutations

[0071] In various embodiments, the amino acid mutation may be an amino acid substitution, and the one or more mutations may comprise: the amino acid residue R at position 57 being substituted with a basic amino acid residue, preferably H; and / or the amino acid residue R at position 173 being substituted with a polar amino acid residue, preferably S; and / or the amino acid residue A at position 260 being substituted with a polar amino acid residue, preferably S; and / or the amino acid residue P at position 451 being substituted with a non-polar amino acid residue, preferably a non-polar aliphatic amino acid, more preferably L; and / or the amino acid residue D at position 471 being substituted with a polar amino acid residue, preferably a polar amino acid with an amide side chain, more preferably N; and / or the amino acid residue G at position 520 being substituted with a polar amino acid residue, preferably a polar acidic amino acid, more preferably E, or a non-polar amino acid residue, preferably a non-polar aliphatic amino acid, more preferably A; and / or the amino acid residue A at position 586 being substituted with a non-polar amino acid residue, or a polar amino acid residue, more preferably T; and / or the amino acid residue S at position 767 being substituted with a polar amino acid residue, preferably G, or an aliphatic amino acid, preferably a non-bulky aliphatic amino acid, more preferably G or A; and / or the amino acid residue D at position 859 being substituted with a polar amino acid residue, preferably a polar amino acid with an amide side chain, more preferably N.

[0072] In various embodiments, the one or more amino acid mutations comprise a substitution of the amino acid residue A at position 260 with an amino acid residue other than a high-helix propensity amino acid such as M, K, R, E or L.

[0073] In various embodiments, the one or more amino acid mutations may be selected from R57H, R1 / 3S, A260S, P451 L D4 / 1 N, G520E, A586T, S767G. D859N and combinations thereof.

[0074] In various embodiments, the polypeptide disclosed herein comprises amino acid mutations R173S, P451 L, D471 N and G520E. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation A586T. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation S767G. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation A260S. In various embodiments, the polypeptide disclosed herein comprises an amino acid mutation R57H. In various embodiments, the polypeptide disclosed herein comprises amino acid mutations R57H, A260S and D859N.

[0075] The present inventors identified six T7 RNA polymerase variants (namely T7-295, T7-259, T7-260, T7-258, T7-256, T7-286) that each contain at least one substitution or a combination of substitutions at positions 57, 173, 260, 451 , 471 , 520, 586, 767, 859 with respect to WT T7 RNA polymerase set forth in SEQ ID NO:1 .

[0076] Table 3: Lists amino acid sequences comprising amino acid substitutions and T7 RNAP variants identified in this invention using the amino acid sequence of wild-type T7 RNAP as reference (N- to C-terminal). The mutations are depicted in the amino acid sequences in bold and |bordered| for ease of reference.

[0077] In various embodiments, the polypeptide disclosed herein comprises or consists of the amino acid sequence set forth in any one of SEQ ID NO:1, 2, 3, 4, 5, 6 or 7 or a functional variant or fragment thereof, and comprises one or more amino acid mutations at a position corresponding to position R57H, R173S, A260S, P451 L, D471N, G520E, A586T, S767G, D859N and combinations thereof.

[0078] In various embodiments, the polypeptide disclosed herein comprises or consists of an amino acid sequence set forth in SEQ ID NO: 2 or a functional variant or fragment thereof, and comprises amino acid mutations R173S, P451 L, D471 N and G520E.

[0079] In various embodiments, the polypeptide disclosed herein comprises or consists of an amino acid sequence set forth in SEQ ID NO: 3 or a functional variant or fragment thereof, and comprises an amino acid mutation A586T.

[0080] In various embodiments, the polypeptide disclosed herein comprises or consists of an amino acid sequence set forth in SEQ ID NO: 4 or a functional variant or fragment thereof, and comprises an amino acid mutation S767G.

[0081] In various embodiments, the polypeptide disclosed herein comprises or consists of an amino acid sequence set forth in SEQ ID NO: 5 or a functional variant or fragment thereof, and comprises an amino acid mutation A260S.

[0082] In various embodiments, the polypeptide disclosed herein comprises or consists of an amino acid sequence set forth in SEQ ID NO: 6 or a functional variant or fragment thereof, and comprises an amino acid mutation R57H.

[0083] In various embodiments, the polypeptide disclosed herein comprises or consists of an amino acid sequence set forth in SEQ ID NO: 7 or a functional variant or fragment thereof, and comprises amino acid mutations R57H, A260S and D859N.

[0084] The term “functional variant", as used herein in relation to the polypeptide disclosed herein, relates to polypeptides that comprise or consist of an amino acid sequence that is at least 60%, 65%, 70%, 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 90.5%, 91 %, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.25%, or 99.5% identical or homologous to the amino acid sequence set forth in any one of SEQ ID NO:1-7 over their entire length, but retain the functionality of the reference sequence (i.e. polymerase activity). The functional variants comprise one or more amino acid mutations at a position corresponding to position R57, R173, A260, P451 , D471 , G520, A586, S767, D859 and combinations thereof. The term “functional variant", also encompasses variants that comprise the amino acid sequence set forth in any one of SEQ ID NO:1-7 but also comprise N- and / or C-terminal extensions of 1 or more amino acids, and comprise an amino acid mutation at a position corresponding to position R57, Ri 73, A260 P451 , D471 , G520, A586, S767, 0859 and combinations thereof of SEQ ID NO: 1. In various embodiments, the polypeptide disclosed herein has an amino acid sequence that shares at least 60, preferably at least 70, more preferably at least 80, most preferably at least 90 % sequence identity with the amino acid sequence set forth in any one of SEQ ID NO:1-7 over its entire length or has an amino acid sequence that shares at least 80, preferably atleast 90, more preferably at least 95% sequence homology with the amino acid sequence set forth in any one of SEQ ID NO:1-7 over its entire length, and comprise an amino acid mutation at a position corresponding to position R57, R173, A260, P451 , D471 , G520. A586, S767, D859 and combinations thereof.

[0085] The identity of amino acid sequences (or nucleotide sequences) is generally determined by means of a sequence comparison. This sequence comparison is based on the BLAST algorithm that is established in the existing art and commonly used (cf. e.g. Altschul et al. (1990) “Basic local alignment search tool”, J. Mol Biol 215:403-410, and Altschul et al (1997): “Gapped BLAST and PSI-BLAST: a new generation of protein database search programs”; Nucleic Acids Res , 25, p 3389- 3402) and is effected in principle by mutually associating similar successions of nucleotides or amino acids in the nucleic acid sequences and amino acid sequences, respectively. A tabular association of the relevant positions is referred to as an "alignment." Sequence comparisons (alignments), in particular multiple sequence comparisons, are commonly prepared using computer programs which are available and known to those skilled in the art. A comparison of this kind also allows a statement as to the similarity to one another of the sequences that are being compared. This is usually indicated as a percentage identity, i.e. the proportion of identical nucleotides or amino acid residues at the same positions or at positions corresponding to one another in an alignment. The more broadly construed term "homology", in the context of amino acid sequences, also incorporates consideration of the conserved amino acid exchanges, i.e. amino acids having a similar chemical activity, since these usually perform similar chemical activities within the protein. The similarity of the compared sequences can therefore also be indicated as a "percentage homology" or "percentage similarity." Indications of identity and / or homology can be encountered over entire polypeptides or genes, or only over individual regions. Homologous and identical regions of various nucleic acid sequences or amino acid sequences are therefore defined by way of matches in the sequences. Such regions often exhibit identical functions. They can be small, and can encompass only a few nucleotides or amino acids. Small regions of this kind often perform functions that are essential to the overall activity of the protein. It may therefore be useful to refer sequence matches only to individual, and optionally small, regions. Unless otherwise indicated, however, indications of identity and homology herein refer to the full length of the respectively indicated amino acid sequence (or nucleic acid sequence).

[0086] The term “functional fragment", as used herein in relation to the polypeptides disclosed herein, relates to polypeptides that differ from the amino acid sequence set forth in any one of SEQ ID NO:1- 7 by a deletion of one or more amino acids from its C- and / or N-terminus while retaining the polymerase activity. Said fragments preferably retain full functionality. In various embodiments, such fragment differs from the reference sequence and they may lack 1 -20 amino acids from their N- and / or C-terminus, for example 1 -15 amino acids or 1 -10 amino acids or 1 -5 amino acids, and encompasses an amino acid sequence that matches the initial molecule as set forth in SEQ ID NOs. 1-7 over a length of at least 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 860, 870, 880 continuously connected amino acids. The functional fragment comprises the one or more amino acidmutations at a position corresponding to position R57, R173, A260, P451 , D-+71 , G520, A586, S7S7, D859 and combinations thereof.

[0087] In various embodiments, the present invention thus also relates to functional fragments of the polypeptides described herein, with said fragments retain enzymatic activity, that is RNA polymerase activity. It is preferred that they have at least 50 %, more preferably at least 70, most preferably at least 90 % of the RNA polymerase activity of the initial molecule, preferably of the polypeptide having the amino acid sequence of SEQ ID NO:1.

[0088] In various embodiments, the functional fragment comprises the amino acid sequence corresponding to residues 267-811 of SEQ ID NO:1. In various embodiments, the functional fragment comprises the amino acid sequence corresponding to residues 812-883 of SEQ ID NO:1. In various embodiments, the functional fragment comprises the amino acid sequence corresponding to residues 267-883 of SEQ ID NO:1.

[0089] In various embodiments, the functional fragment comprises the one or more amino acid mutations R57H, R173S, A260S, P451 L, D471 N. G520E, A586T, S767G, D859N or combinations thereof.

[0090] In addition to the above-described mutations, polypeptides according to the embodiments disclosed herein can comprise amino acid modifications other than those described above (i.e. amino acid mutations (substitutions) at positions R57, R173, A260, P451 , D471 , G520, A586, S767, D859 of SEQ ID NO:1) and present in any one of the amino acid sequences set forth in SEQ ID NO:1-7. Such polypeptides are, for example, further developed by targeted genetic modification, i.e. by way of mutagenesis methods, and optimized for specific purposes or with regard to special properties (for example, with regard to their catalytic activity, stability, etc.). If such additional modifications are introduced into the polypeptides of the invention, these preferably do not affect, alter the polymerase activity of the reference T7 RNAP variants. This means that the above-defined features of these residues are not changed by these additional mutations beyond what is defined above. In addition, nucleic acids contemplated herein can be introduced into recombination formulations and thereby used to generate entirely novel RNA polymerases or other polypeptides.

[0091] Accordingly, in one aspect, the present invention relates to a polypeptide having RNA polymerase activity, comprising or consisting of:(i) an amino acid sequence set forth in any one of SEQ ID NO:1-7;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in any one of SEQ ID NO: 1-7; or(ill) a functional fragment of (I) or (ii);wherein the polypeptide comprises one or more amino acid mutations at a position selected from R57, R173, A260, P451. D471 , G520, A586, S767, D859 and combinations thereof, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

[0092] In various embodiments, the polypeptide disclosed herein may comprise or consist of:(i) an amino acid sequence set forth in SEQ ID NO:2;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in SEQ ID NO: 2; or(iii) a functional fragment of (i) or (ii); wherein the polypeptide comprises an amino acid mutation at a position corresponding to position Rl / 3, P451, D471 , and G520, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1 .

[0093] In various embodiments, the polypeptide disclosed herein may comprise or consist of:(i) an amino acid sequence set forth in SEQ ID NO:3;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in SEQ ID NO: 3; or(iii) a functional fragment of (I) or (ii); wherein the polypeptide comprises an amino acid mutation at a position corresponding to position A58S, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

[0094] In various embodiments, the polypeptide disclosed herein may comprise or consist of:(I) an amino acid sequence set forth in SEQ ID NO:4;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in SEQ ID NO: 4; or(iii) a functional fragment of (I) or (ii); wherein the polypeptide comprises an amino acid mutation at a position corresponding to position S767, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

[0095] In various embodiments, the polypeptide disclosed herein may comprise or consist of:(I) an amino acid sequence set forth in SEQ ID NO:5;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in SEQ ID NO: 5; or(iii) a functional fragment of (I) or (ii); wherein the polypeptide comprises an amino acid mutation at a position corresponding to position A260, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

[0096] In various embodiments, the polypeptide disclosed herein may comprise or consist of:(i) an amino acid sequence set forth in SEQ ID NO:6;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in SEQ ID NO: 6; or(iii) a functional fragment of (I) or (ii); wherein the polypeptide comprises an amino acid mutation at a position corresponding to position R57, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

[0097] In various embodiments, the polypeptide disclosed herein may comprise or consist of:(i) an amino acid sequence set forth in SEQ ID NO:7;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in SEQ ID NO: 7; or(iii) a functional fragment of (i) or (ii); wherein the polypeptide comprises an amino acid mutation at a position corresponding to position R57, A260 and D359, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

[0098] In another aspect, there is also provided a fusion protein comprising the polypeptide disclosed herein, operably linked to a deaminase domain capable of catalyzing the deamination of nucleic acids. The linking of the polypeptide and deaminase domain forms a single polypeptide chain. As used herein, the term "fusion protein" refers to a chimeric protein in which two or more distinct polypeptide sequences — originating from different proteins or functional domains — are joined together within a single polypeptide chain, such that each constituent domain or moiety retains at least part of its native biological activity. As used herein, the term "single polypeptide chain" refers to a contiguous sequence of amino acids translated from a single open reading frame (ORF), such that both the deaminase domain and the polypeptide are encoded by the same mRNA transcript and are synthesized as aunified molecule by a ribosome. In various embodiments, the fusion protein may refer to a single polypeptide chain comprising a first polypeptide comprising or consisting of the polypeptide disclosed herein operably linked to a second polypeptide comprising or consisting of a deaminase domain capable of catalyzing the deamination of nucleic acids.

[0099] The fusion of the polypeptide disclosed herein and the deaminase domain may include a linker. The resulting single polypeptide chain combines the transcriptional activity of the disclosed polypeptide with the nucleic acid-editing activity of the deaminase domain. This configuration enables site-specific base editing in the vicinity of transcriptionally active regions, particularly those targeted by the polypeptide through promoter recognition, such as T7 promoter-driven loci in genomic DNA or plasmids. In particular, the fusion protein disclosed herein may be used to introduce mutations into a target nucleic acid sequence.

[0100] In various embodiments, the C-terminus of the polypeptide may be linked to the N-terminus of the deaminase domain (C-to-N configuration), or vice versa (N-to-C configuration), depending on which orientation provides optimal functional activity and structural compatibility. In various embodiments, the N-terminus of the polypeptide may be linked to the N-terminus of the deaminase domain (N-to-N), or the C-terminus to the C-terminus (C-to-C), by employing loop-forming or branching linkers.

[0101] In various embodiments, the polypeptide disclosed herein may be indirectly linked to the deaminase domain via a linker. The linker may be a linker peptide that facilitates proper folding and function of the fusion protein, such that the deaminase domain retains its enzymatic activity and the polypeptide retains its transcriptional activity. As used herein, a linker refers to a peptide sequence that connects two functional domains within a fusion protein. The linker may be rigid or flexible, and may serve to: (i) facilitate correct spatial orientation and folding of each domain; (ii) minimize steric hindrance that may impede enzymatic activity; and / or (ill) preserve the functional independence of each domain. In various embodiments, the linker may comprise one or more Glycine-Serine (Gly- Ser) repeats, such as (G S)n, where n is an integer from 1 to 5. In other embodiments, the linker may comprise a rigid helical sequence, such as one or more repeats of EAAAK, to maintain fixed spatial separation between domains and prevent domain interference. In further embodiments, the linker may comprise or consist of a synthetic unstructured polypeptide composed of small polar, uncharged amino acids, such as Glycine (Gly), Serine (Ser), Alanine (Ala), Threonine (Thr), Glutamine (Gin), Proline (Pro), Glutamic acid (Glu), and Asparagine (Asn). Such linkers, including XTEN-like linkers, are designed to be flexible, hydrophilic, and resistant to folding into stable secondary structures, thereby enhancing solubility and reducing immunogenicity. In yet other embodiments, the linker may be cleavable under specific conditions (e.g., by a protease), thereby allowing controlled release or separation of the individual domains.

[0102] In various embodiments, the fusion protein comprises or consists of: a deaminase domain capable of catalyzing the deamination of nucleic acids; and a polypeptide disclosed herein, wherein the deaminase domain and the polypeptide disclosed herein are linked to form a single polypeptide chain.

[0103] In various embodiments, the deaminase domain may be selected from the group consisting of cytidine deaminases, adenosine deaminases, and guanine deaminases. In various embodiments, the deaminase domain is a cytidine deaminase, including but not limited to members of the APOBEC family (e.g., APOBEC3A, APOBEC3B, or activation-induced cytidine deaminase (AID)), or a doublestranded DNA-specific cytidine deaminase such as DddA (double-stranded DNA deaminase toxin A, UniProt P0DUH5). In various embodiments, the deaminase domain is an adenosine deaminase, such as ADAR1 (UniProt A0A3B3ISU1 ) or ADAR2 (UniProt C1JAR3), which catalyze the deamination of adenosine to inosine in RNA, resulting in A— >-G transitions. In various embodiments, the deaminase domain may also be derived from non-human orthologs, such as rat APOBEC (UniProt P38483). Each of these enzymes catalyzes the deamination of a target base, thereby enabling programmable base transitions (e.g., C^T, A^G) in DNA or RNA sequences, depending on the substrate specificity of the deaminase used.

[0104] In various embodiments, the deaminase domain may comprise or consist of an amino acid sequence set forth in any one of SEQ ID NO: 15-18 or a functional variant or fragment thereof.

[0105] Nucleic acid molecules encoding the polypeptide or fusion protein disclosed herein are also provided. All embodiments disclosed above in relation to the polypeptide and fusion protein similarly apply to the nucleic acid molecules and vice versa.

[0106] The nucleic acid molecules can be DNA molecules or RNA molecules. They can exist as an individual strand, as an individual strand complementary to said individual strand, or as a double strand. With DNA molecules in particular, the sequences of both complementary strands in all three possible reading frames are to be considered in each case. Also to be considered is the fact that different codons, i.e. base triplets, can code for the same amino acids, so that a specific amino acid sequence can be coded by multiple different nucleic acids. As a result of this degeneracy of the genetic code, all nucleic acid sequences that can encode one of the above-described T7 RNA polymerase variants or fusion protein are included. The skilled artisan is capable of unequivocally determining these nucleic acid sequences, since despite the degeneracy of the genetic code, defined amino acids are to be associated with individual codons. The skilled artisan can therefore, proceeding from an amino acid sequence, readily ascertain nucleic acids coding for that amino acid sequence. In addition, in the context of nucleic acids, according to the present invention one or more codons can be replaced by synonymous codons. This aspect refers in particular to heterologous expression of the enzymes contemplated herein. For example, every organism, e.g. a host cell of a production strain, possesses a specific codon usage. "Codon usage" is understood as the translation of the genetic code into amino acids by the respective organism. Bottlenecks in protein biosynthesis can occur if the codons located on the nucleic acid are confronted, in the organism, with a comparatively small number of loaded tRNA molecules. Also it codes for the same amino acid, the result is that a codon becomes translated in the organism less efficiently than a synonymous codon that codes for the same amino acid. Because of the presence of a larger number of tRNA molecules for the synonymous codon, the latter can be translated more efficiently in the organism. By way of methods commonly known today such as, for example, chemical synthesis or the polymerase chain reaction (PCR) in combination with standard methods of molecular biology or protein chemistry, a skilled artisan has the ability to manufacture, on the basis of known DNA sequences and / or amino acid sequences, the corresponding nucleic acids all the way to complete genes. Such methods are known, for example, from Sambrook, J., Fritsch, E. F., and Maniatis, T, 2001 , Molecular cloning: a laboratory manual, 3rd edition, Cold Spring Laboratory Press.

[0107] The term "degenerate variant" refers to a nucleotide sequence encoding a protein (for example, the polypeptide or fusion protein disclosed herein) that includes a nucleotide sequence that is degenerate as a result of the genetic code. There are twenty natural amino acids, most of which are specified by more than one codon. Therefore, all degenerate nucleotide sequences are includedas long as the resulting polypeptide or fusion protein disclosed herein encoded by the nucleotide sequence has the desired activity.

[0108] The nucleic acid molecules encoding the polypeptides or fusion protein described herein, as well as a circular DNA molecule containing such a nucleic acid, in particular a plasmid, vector, cosmid, bacterial artificial chromosome (BAC), bacteriophage, viral vector or hybrids thereof also form part of the present invention.

[0109] In various embodiments, the nucleic acid molecules may be comprised in a vector or is a vector. "Vectors" are understood for purposes herein as elements - made up of nucleic acids - that contain a nucleic acid contemplated herein as a characterizing nucleic acid region. They enable said nucleic acid to be established as a stable genetic element in a species or a cell line over multiple generations or cell divisions. In particular when used in bacteria, vectors are special plasmids, i.e. circular genetic elements. In the context herein, a nucleic acid as contemplated herein is cloned into a vector. Included among the vectors are, for example, those whose origins are bacterial plasmids, viruses, or bacteriophages, or predominantly synthetic vectors or plasmids having elements of widely differing derivations. Using the further genetic elements present in each case, vectors are capable of establishing themselves as stable units in the relevant host cells over multiple generations. They can be present extrachromosomally as separate units, or can be integrated into a chromosome resp. into chromosomal DNA. In various embodiments, the vector may be selected from the group of plasmids (i.e. bacterial plasmids), binary vectors, DNA vectors, mRNA vectors, retroviral vectors, lentiviral vectors, adenoviral vectors, transposon-based vectors, and artificial chromosomes.

[0110] In various embodiments, the vector may be an expression vector. Expression vectors encompass nucleic acid sequences which are capable of replicating in the host cells, by preference microorganisms, particularly preferably bacteria, that contain them, and expressing therein a contained nucleic acid. In various embodiments, the vectors described herein thus also contain regulatory elements that control expression of the nucleic acids encoding a polypeptide or fusion protein of the invention. Expression is influenced in particular by the promoter or promoters that regulate transcription. Expression can occur in principle by means of the natural promoter originally located in front of the nucleic acid to be expressed, but also by means of a host-cell promoter furnished on the expression vector or also by means of a modified, or entirely different, promoter of another organism or of another host cell. In the present case at least one promoter for expression of a nucleic acid as contemplated herein is made available and used for expression thereof. Expression vectors can furthermore be regulated, for example by way of a change in culture conditions or when the host cells containing them reach a specific cell density, or by the addition of specific substances, in particular activators of gene expression. One example of such a substance is the galactose derivative isopropyl-beta-D-thiogalactopyranoside (IPTG), which is used as an activator of the bacterial lactose operon (lac operon).

[0111] In various embodiments, the nucleic acid molecule disclosed herein may comprise an “expression construct" which refers to a functional unit built in the vector for the purpose of recombinantly expressing the polypeptide or fusion protein disclosed herein, when introduced into an appropriate host cell. The term "recombinant", as used herein (e.g. a recombinant polypeptide, a recombinant fusion protein, a recombinant nucleic acid, or the like), refers to any molecule which is prepared, expressed, created or isolated by recombinant means, and which is not naturally occurring. "Recombinant" can be used synonymously with "engineered" or "non-natural" and can refer to to an organism, microorganism, cell, nucleic acid molecule, or vector that includes at least one genetic alteration or has been modified by introduction of an exogenous nucleic acid molecule, wherein such alterations or modifications are introduced by genetic engineering (i.e., human intervention). Genetic alterations include, for example, modifications introducing expressible nucleic acid molecules encoding proteins, fusion proteins or enzymes, or other nucleic acid molecule additions, deletions, substitutions or otherfunctional disruption of a cell's genetic material. Additional modifications include, for example, non-coding regulatory regions in which the modifications alter expression of a polynucleotide, gene or operon.

[0112] In various embodiments, the nucleic acid molecule may be comprised in a bacterial plasmid or is a bacterial plasmid. The term “bacterial plasmid” as used herein refers to a circular DNA molecule capable of replication in a bacterial host cell. A bacterial plasmid may contain an appropriate origin of replication, which is a sequence of DNA sufficient to enable the replication of the plasmid in a host bacterial cell. A bacterial plasmid may also contain a selectable marker sequence, which encodes a selectable marker conferring cellular resistance to antibiotics such as ampicillin, kanamycin, chloramphenicol, and tetracycline.

[0113] In various embodiments, the nucleic acid molecule, or vector, or plasmid, containing the nucleic acid molecule, further comprises regulatory elements for controlling expression of said nucleic acid molecule.

[0114] The term "operably linked" as used herein refers to the relationship between two or more nucleotide sequences that interact physically or functionally. For example, a promoter or regulatory nucleotide sequence is said to be operably linked to a nucleotide sequence that codes for an RNA or a protein if the two sequences are situated such that the regulatory nucleotide sequence will affect the expression level of the coding or structural nucleotide sequence. “Regulatory nucleotide sequences" or “regulatory elements” as used herein refer to nucleotide sequences that influence the timing and level / amount of transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include promoters; translation leader sequences; introns; enhancers; stem-loop structures; repressor binding sequences; termination sequences; and polyadenylation recognition sequences. Particular regulatory sequences may be located upstream and / or downstream of a coding sequence operably linked thereto.

[0115] Another aspect of the invention relates to a host cell comprising the polypeptide, fusion protein or nucleic acid molecule disclosed herein. All embodiments disclosed above in relation to the polypeptide, fusion protein and nucleic acid molecule disclosed herein, similarly apply to the host cell, and vice versa.

[0116] All cells are in principle suitable as host cells, i.e. prokaryotic or eukaryotic cells. Those host cells can be manipulated in genetically advantageous fashion. The term "host ceil" as used herein refers to a living cell into which the polypeptide, fusion protein or nucleic acid molecule is to Lie or has been introduced. The living cell includes both a cultured cell and a cell within a living organism. In various embodiments, host cells can be engineered to incorporate a desired gene or expression construct on its chromosome or in its genome. The host cell may be any cell that is commonly used for expression, i.e. transcription and translation of the nucleic acid molecule for the production of the polypeptide or fusion protein disclosed herein. In particular, the term “host cell" relates to prokaryotes, eukaryotes, plants, insect cells or mammalian cells, cell lines and cell culture systems. Host cells include, without limitation, bacterial, microbial, plant or animal cells. In various embodiments, the host cell is represented by those host cells whose activity can be regulated on the basis of genetic regulation elements that are made available, for example, on the vector, but can also be present a priori in those cells. They can be stimulated to expression, for example, by controlled addition of chemical compounds that serve as activators, by modifying the culture conditions, or when a specific cell density is reached. This makes possible economical production of the polypeptides contemplated herein.

[0117] A nucleic acid molecule disclosed herein or a vector / plasmid containing said nucleic acid molecule may be transfected, transduced or transformed into the host cell, which all generally refers to the incorporation of the nucleic acid molecule into the host cell. Methods for the transformation, transduction or transfection of cells are established in the existing art and are sufficiently known to the skilled artisan. The host cells contemplated herein are cultured and fermented in a usual manner, for example in discontinuous or continuous systems. In the former case a suitable nutrient medium is inoculated with the host cells, and the product is harvested from the medium after a period of time to be ascertained experimentally. Continuous fermentations are notable for the achievement of a flow equilibrium in which, over a comparatively long period of time, cells die off in part but are also in part renewed, and the protein formed can simultaneously be removed from the medium.

[0118] Preferred host cells are prokaryotic or bacterial cells, such as E. coli cells. Bacteria are notable for short generation times and few demands in terms of culturing conditions. As a result, economical culturing methods for producing the polypeptides can be established. In addition, the skilled artisan has ample experience in the context of bacteria in fermentation technology. Gramnegative or Gram-positive bacteria may be suitable for a specific production instance, for a wide variety of reasons to be ascertained experimentally in the individual case, such as nutrient sources,product formation rate, time requirement, etc. In various embodiments, the host cells may be E.coli cells.

[0119] In various embodiments, the host cell may be a microbial cell that is GRAS-certified (Generally Recognized as Safe). In various embodiments, the microbial cell may be selected from bacterial and yeast species including, but not limited to, Lactobacillus spp., Bifidobacterium spp., Bacillus spp., Saccharomyces spp., Streptococcus thermophilus, Corynebacterium glutamicum, Kluyveromyces marxianus, Kluyveromyces lactis, and Propionibacterium freudenreichii.

[0120] Host cells contemplated herein can be modified in terms of their requirements for culture conditions, can comprise other or additional selection markers, or can also express other or additional proteins. They can, in particular, be those host cells that transgenically express multiple proteins or enzymes.

[0121] The host cell can, however, also be a eukaryotic cell, which is characterized in that it possesses a cell nucleus. A further embodiment is therefore represented by a host cell which is characterized in that it possesses a cell nucleus. In contrast to prokaryotic cells, eukaryotic cells are capable of post-translationally modifying the protein that is formed. Examples thereof are fungi such as Actinomycetes, or yeasts such as Saccharomyces or Kluyveromyces or insect cells, such as Sf9 cells. This may be particularly advantageous, for example, when the polypeptides or fusion proteins disclosed herein, in connection with their synthesis, are intended to experience specific modifications made possible by such systems. Among the modifications that eukaryotic systems carry out in particular in conjunction with protein synthesis are, for example, the bonding of low-molecular-weight compounds such as membrane anchors or oligosaccharides.

[0122] Host cells disclosed herein may be used to manufacture the polypeptides or fusion protein described herein.

[0123] Accordingly, a further aspect of the invention is therefore a method of producing / manufacturing a polypeptide or fusion protein as disclosed herein, comprising culturing a host cell contemplated herein under conditions that allow expression of the polypeptide or fusion protein; and isolating the polypeptide or fusion protein from the culture medium or from the host cell. Culture conditions and mediums can be selected by those skilled in the art based on the host organism used by resorting to general knowledge and techniques known in the art.

[0124] In this regard, the host cell may be readily manipulated in microbiological and biotechnological terms. This refers, for example, to easy culturability, high growth rates, low demands in terms of fermentation media, and good production and secretion rates for the polypeptides. The polypeptides can furthermore be modified, after their manufacture, by the cells producing them, forexample by the addition of sugar molecules, formylation, amination, etc. Post-translation modifications of this kind can functionally influence the polypeptide.

[0125] In various embodiments, the method may further comprise purifying the isolated polypeptide disclosed herein. “Purification” or “purifying” herein means the process of removing components from a host cell or culture, the presence of which is not desired. Purification is a relative term and does not require that all traces of the undesirable component be removed.

[0126] In yet another aspect of the invention, there is provided a cell-free method for producing a polypeptide or fusion protein as disclosed herein, comprising subjecting the nucleic acid molecule disclosed herein to reaction conditions that allow the transcription and translation of the polypeptide or fusion protein, and isolating said a polypeptide or fusion protein disclosed herein.

[0127] As used herein, the term "cell-free method" refers to an in vitro biochemical system that enables the synthesis of polypeptides in the absence of living cells, by providing the necessary molecular machinery for transcription and / or translation in a controlled, cell-free environment. Such systems may typically comprise cell extracts or reconstituted enzyme mixtures containing ribosomes, tRNAs, amino acids, nucleotides, energy sources (e.g. , ATP, GTP), cofactors, and other components required for gene expression. In various embodiments, the cell-free method may be derived from prokaryotic sources (e.g., E. coli lysates), eukaryotic sources (e.g., wheat germ, rabbit reticulocyte, or insect cell extracts), or synthetic transcription-translation (TX-TL) systems assembled from purified components. The nucleic acid molecule encoding the polypeptide or fusion protein may be introduced into the system to initiate expression, and the resulting polypeptide may be subsequently isolated and purified from the reaction mixture using standard biochemical techniques. Cell-free methods are advantageous for rapid, scalable, and high-throughput protein production, especially for toxic, unstable, or non-naturally occurring proteins that are difficult to express in living cells.

[0128] There is also provided a composition comprising the polypeptide, fusion protein, nucleic acid molecule encoding the polypeptide or fusion protein, or the host cell comprising the nucleic acid molecule disclosed herein.

[0129] In various embodiments, the composition may comprise one or more additional components or agents that enhance the function, delivery, or stability, of the polypeptide, fusion protein, nucleic acid molecule, or host cell for one or more of the uses described below. Accordingly, the compositions disclosed herein may be formulated and adapted for use in a wide range of applications.

[0130] In various embodiments, the composition may comprise one or more in vitro transcription (IVT) reagents. The in vitro transcription reagent may be selected from a transcription reagent, a deoxyribonucleic acid (DNA), nucleoside triphosphates, and a cap analog.

[0131] There is also provided a kit comprising the polypeptide, fusion protein, nucleic acid molecule encoding the polypeptide or fusion protein, or the host cell comprising the nucleic acid molecule disclosed herein. The kit may be used for any one of the methods of use disclosed herein.❖ Methods of Use

[0132] All embodiments disclosed herein in relation to the polypeptides, fusion proteins, nucleic acids, host cells and kits are similarly applicable to the uses and methods described herein and vice versa.

[0133] RNA polymerases are essential enzymes that catalyze the synthesis of RNA from a DNA template (i.e. transcription). In both biological and synthetic systems, RNA polymerases are widely employed for the in vitro or in vivo production of RNA molecules, including but not limited to messenger RNA (mRNA), non-coding RNAs, and RNA-based therapeutics.

[0134] Accordingly, the present invention relates to the use of polypeptides described herein for RNA production or amplification to produce RNA transcript.

[0135] In various embodiments, there is provided a method for RNA production or amplification to produce RNA transcript, comprising contacting a nucleic acid template with the polypeptide disclosed herein under conditions that result in the production of RNA transcript.

[0136] As used herein, the term "nucleic acid template" refers to a polynucleotide molecule — such as DNA or RNA — that contains a sequence encoding one or more regions to be transcribed into RNA by the polypeptide disclosed herein, which possesses RNA polymerase activity. In various embodiments, the nucleic acid template comprises a double-stranded DNA (dsDNA) or singlestranded DNA (ssDNA) molecule that includes a promoter sequence recognizable by the RNA polymerase (e.g., a T7 promoter), followed by a downstream transcription unit containing the target sequence for RNA synthesis. The template may be linear or circular, and may be present as a plasmid. In various embodiments, the nucleic acid template may be a DNA template.

[0137] The conditions that result in the production of RNA transcript may refer to suitable reaction parameters and components that support the transcriptional activity of the polypeptide. These conditions would be readily understood by a person skilled in the art and may include, but are not limited to: the presence of a suitable reaction buffer (e.g., containing Tris-HCI, Mg2+, and DTT), an appropriate concentration of ribonucleoside triphosphates (rNTPs), and optimal temperature and pH for enzymatic activity. The transcription reaction may also include RNase inhibitors, cofactors, or crowding agents (e.g., PEG) to enhance yield or stability. The precise conditions may be optimized depending on the sequence, structure, and activity profile of the specific RNA polymerase variant used.

[0138] In various embodiments, the method is an in vitro method.

[0139] The term "RNA" (or “ribonucleic acid”) as used herein relates to a molecule which comprises ribonucleotide residues. The term ''ribonucleotide" refers to a nucleotide containing ribose as its pentose component. The term "RNA" comprises double -stranded RNA, single stranded RNA, isolated RNA, synthetic RNA, recombinantly generated RNA, ribo-oligonucleotides (shorter RNA sequences generally in the range of 3 to 40 nucleotides), self-amplifying RNA (“saRNA”) also referred to as self-replicating RNA or self-amplifying / replicating mRNA (“SAM”), and modified RNA which differs from naturally occurring RNA by addition, deletion, substitution and / or alteration of one or more nucleotides. Nucleotides in RNA molecules can comprise non-standard nucleotides, such as non- naturally occurring nucleotides or chemically synthesized nucleotides or deoxynucleotides. "mRNA" (or messenger RNA") as used herein means "messenger-RNA" and relates to a transcript which is generated by using a DNA template and encodes a peptide or protein. Typically, mRNA comprises a protein coding region flanked by a 5'-UTR and a 3'-UTR. The term "antisense- RNA" relates to singlestranded RNA comprising ribonucleotide residues, which are complementary to the mRNA. The term "siRNA" means "small interfering RNA", which is a class of double-stranded RNA-molecules comprising about 20 to about 25 base pairs.

[0140] The term "RNA transcript" refers to the RNA product synthesized from the nucleic acid template by the disclosed polypeptide under transcription-permissive conditions. The RNA transcript may be messenger RNA (mRNA), non-coding RNA (ncRNA), or any synthetic or functional RNA, and may vary in length and composition depending on the template and transcriptional machinery used. In various embodiments, the RNA transcript may comprise untranslated regions (UTRs), coding sequences (CDS), or structural elements such as aptamers or riboswitches. The RNA transcript may be used directly, further processed, or incorporated into downstream applications such as translation, reverse transcription, or diagnostic detection.

[0141] In various embodiments, the RNA transcript ® RNA (e g , mRNA or seif- replicating RNA) that encodes a protein (e.g., a therapeutic protein). Thus, the RNA transcripts produced using the polypeptides disclosed herein may be used in a myriad of applications. For example, the RNA transcripts may be used to produce proteins of interest, e.g., therapeutic proteins, vaccine antigen, and the like. In various embodiments, the RNA transcripts are therapeutic RNAs. A therapeutic mRNA is ar: mRNA that encodes a therapeutic protein (tire term 'protein’ encompasses peptides) Therapeutic proteins mediate a variety of effects in a host cell or in a subject to treat a disease or ameliorate the signs and symptoms of a disease. An RNA transcript produced may encode one or more biologies. A biologic is a polypeptide-based moiecuie that may be used io treat, -cure, mitigate, prevent, or diagnose a serious or iite-ihreatening disease or medical condition In various embodiments, RNA transcripts are used as guide RNA (gRNA) for gene targeting. In various embodiments, RNA transcripts (e.g., mRNA) are used tor in vitro translation and micro injection, insome embodiments, SNA transcripts are used for RNA amplification, in some embodiments, RNA transcripts are used as anti-sense RNA for gene expression experiments.

[0142] In various embodiments, the nucleic acid template may comprise a linear or circular DNA molecule that encodes an RNA transcript to be synthesized by the polypeptide disclosed herein. The nucleic acid template may comprise a T7 promoter or other RNA polymerase recognition site upstream of the transcribed region, and may further comprise one or more regulatory, coding, or functional elements to enhance transcriptional output, translation efficiency, stability, or detection of the resulting RNA transcript. In various embodiments, the nucleic acid template comprises a T7 promoter operably linked to a gene of interest, and optionally one or more regulatory, coding, or functional elements. In various embodiments, the T7 promoter comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 8, or a variant thereof.

[0143] In various embodiments, variants of a nucleotide sequence may refer to nucleotide sequences that is at least 60%, 65%, 70%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 90.5%, 91 %, 91 .5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.25%, or 99.5% identical to the reference nucleotide sequence over their entire length, but retains the functionality of the reference sequence.

[0144] In various embodiments, the “gene of interest" may refer to a nucleotide sequence that is capable of being transcribed into an RNA molecule, which may optionally be translated into a protein of interest. In various embodiments, the gene of interest may be a nucleotide sequence that (i) encodes a protein of interest, or (ii) is transcribed into an RNA of interest.

[0145] The “protein of interest" may include, but is not limited to, viral antigens (e.g., the receptorbinding domain (RBD) of the SARS-CoV-2 spike protein, including variants such as the delta variant), structural viral proteins (e.g., Coxsackievirus B3 (CVB3) capsid proteins), full-length or truncated spike proteins, enzymes, cytokines, antibodies or fragments thereof, or other bioactive peptides. In addition, the term “RNA of interest” may refer to a functional RNA molecule that is transcribed from the nucleic acid template and exerts its biological, structural, or regulatory role at the RNA level without requiring translation into a protein. The RNA of interest may include, but is not limited to, guide RNAs (gRNAs), small interfering RNAs (siRNAs), microRNAs (miRNAs), messenger RNAs (mRNAs) that serve as templates for translation, or non-coding RNAs involved in gene regulation, splicing, or localization.

[0146] In various embodiments, the gene of interest may be a nucleotide sequence that encodes a SARS-CoV-2 receptor-binding domain (RBD) delta variant, a Coxsackievirus B3 (CVB3)-derived internal ribosome entry site, or a full-length SARS-CoV-2 spike protein. In various embodiments, the gene of interest may encode a reporter protein, such as green fluorescent protein (GFP) or any of its spectral variants (e.g., EGFP, mCherry, YFP, CFP), or a luciferase enzyme (e.g., firefly, Renilla, orNanoLuc luciferases and their derivatives). In various embodiments, the gene of interest may encode a therapeutic or clinically relevant protein, including ornithine transcarbamylase, coagulation Factor IX, or erythropoietin (EPO). The selection of the gene of interest may vary depending on the desired application, such as for use in diagnostics, therapeutics, reporter assays, or vaccine development.

[0147] In various embodiments, the gene of interest may be a nucleotide sequence that encodes a SARS-CoV-2 RBD delta variant, a Coxsackievirus B3 (CVB3)-derived internal ribosome entry site, or a Sars-CoV-2-full length spike protein.

[0148] In various embodiments, the nucleotide sequence that encodes a SARS-CoV-2 RBD delta comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 30, or a variant thereof. In various embodiments, the nucleotide sequence that encodes a CVB3-derived internal ribosome entry site comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 31 , or a variant thereof. In various embodiments, the nucleotide sequence that encodes a Sars-CoV-2-full length spike protein comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 32, or a variant thereof.

[0149] In various embodiments, the one or more regulatory, coding, or functional elements may comprise a nucleotide sequence encoding a 5' untranslated region (5'UTR) and a 3'UTR, a 5’ intron and a 3' intron, a polyadenylation sequence or polyA tail to enhance stability and translational efficiency, a nucleotide sequence encoding a self-cleaving ribozyme such as a hammerhead ribozyme (HHR) to generate precise 3’ ends, a luminescent reporter gene, and / or a fluorogenic RNA aptamer to enable fluorescence-based detection of transcript production.

[0150] In various embodiments, the nucleic acid template comprises a nucleotide sequence encoding a 5'UTR and a nucleotide sequence encoding a 3'UTR to enhance translation or stability. The nucleotide sequence encoding the 5’UTR may be positioned upstream of the gene of interest, and the nucleotide sequence encoding the 3’UTR may be positioned downstream of the gene of interest. In various embodiments, the nucleotide sequence encoding a 5'UTR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 27, 28, 38, or a variant thereof. In various embodiments, the nucleotide sequence encoding a 3'UTR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 34, 39, or a variant thereof.

[0151] The term "downstream" as used herein refers to a nucleotide sequence that is located 3’ to a reference nucleotide sequence. In particular, downstream nucleotide sequences generally relate to sequences that follow the starting point of transcription. For example, the translation initiation codon of a gene may be located downstream of the start site of transcription The term “immediately downstream” may be used to specify that the nucleotide sequence immediately follows and is directly next to the reference nucleotide sequence in the 3’ direction, with no other intervening nucleotide sequence or genetic element therein between. The term "upstream" as used herein refers to a nucleotide sequence that is located 5' to a reference nucleotide sequence. In particular, upstreamnucleotide sequences generally relate to sequences that are located on the 5' side of a coding sequence or starting point of transcription. For example, most promoters are located upstream of the start site of transcription The term “immediately upstream” may be used to specify that the nucleotide sequence immediately precedes and is directly next to the reference nucleotide sequence in the 5’ direction, with no other intervening nucleotide sequence or genetic element therein between.

[0152] In various embodiments, the nucleic acid template comprises a 5’ intron sequence and a 3’ intron sequence The 5’ intron sequence may be positioned upstream of the gene of interest, and the 3’ intron sequence may be positioned downstream of the gene of interest. As used herein, a ‘5’ intron’ refers to an intronic nucleotide sequence located at the 5' end of a transcription unit (i.e. comprising the gene of interest), typically upstream of or within the initial portion of the coding sequence. The 5' intron is transcribed as part of the primary RNA transcript and is capable of being spliced out by the host cell's splicing machinery Inclusion of a 5' intron has been shown to enhance gene expression by promoting efficient mRNA processing, nuclear export, transcript stability, and / or translation efficiency in eukaryotic systems. A ‘3’ intron’ refers to an intronic nucleotide sequence located at the 3' end of the transcription unit, and transcribed as part of the primary transcript. The 3’ intron is likewise spliced out during RNA processing The presence of a 3' intron may contribute to enhanced mRNA stability, regulated transcript processing, and improved expression by facilitating splicing- associated regulatory mechanisms. In various embodiments, the 5' intron sequence comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 29, or a variant thereof. In various embodiments, the 3’intron sequence comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 35, or a variant thereof.

[0153] In various embodiments, the nucleic acid template comprises a polyadenylation sequence or poly(A) tail sequence. In various embodiments, the polyadenylation sequence or poly(A) tail sequence is positioned downstream of the gene of interest. In various embodiments, when the nucleic acid template comprises a nucleotide sequence encoding a 3'UTR, the polyadenylation sequence or poly(A) tail sequence is positioned downstream of the nucleotide sequence encoding the 3'UTR. The term "polyA tail" is a region of mRNA that is downstream, e.g. .directly downstream (i.e., 3'), from a 3' UTR that contains multiple, consecutive adenosine monophosphates. A polyA tail may contain 10 to 300 adenosine monophosphates. For example, a polyA tail may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 350, 400, 450, 500, 550, or 600 adenosine monophosphates. In some embodiments, a polyA tail contains 50 to 250 adenosine monophosphates. In a relevant biological setting the poly(A) tail functions to protect mRNA from enzymatic degradation, e.g., in the cytoplasm, and aids in transcription termination, export of the mRNA from the nucleus, and translation. In various embodiments, the poly(A) tail sequence comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 37, or a variant thereof.

[0154] In various embodiments, the nucleic acid template comprises a nucleotide sequence encoding a self-cleaving ribozyme. In various embodiments, the nucleotide sequence encoding a selfcleaving ribozyme is positioned downstream of the gene of interest to define the 3' end of the transcript or to separate functional RNA domains. As used herein, the term “self-cleaving ribozyme” refers to a structured RNA molecule that possesses intrinsic catalytic activity enabling it to cleave its own phosphodiester backbone at a specific site within the RNA sequence. Self-cleaving ribozymes are commonly used in synthetic biology and in vitro transcription (I VT) systems to generate precise 3' ends of RNA transcripts or to enable modular processing of multi-component RNAs. Non-limiting examples of self-cleaving ribozymes include the hammerhead ribozyme (HHR), which cleaves at a specific conserved motif and is widely used in IVT constructs to produce homogeneous RNA ends; the hepatitis delta virus (HDV) ribozyme; the hairpin ribozyme; and the twister and pistol ribozymes. In various embodiments, the nucleotide sequence encoding a self-cleaving ribozyme comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 36, or a variant thereof.

[0155] In various embodiments, the nucleic acid template comprises a luminescent reporter gene. The luminescent reporter gene may be a gene encoding firefly luciferase (FLuc), a gene encoding renilla luciferase (Rluc), a gene encoding Cypridina Luciferase (Clue), a gene encoding Gaussia Luciferase (Glue), or a gene encoding NanoLuc Luciferase (NanoLuc). The luminescent reporter gene may be included to enable detection and quantification of RNA production and / or protein expression through luminescent signal generation. The luminescent reporter gene may be operably linked to the gene of interest. In various embodiments, the luminescent reporter gene is positioned downstream of the gene of interest. In various embodiments, the gene encoding NanoLuc Luciferase comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 33, or a variant thereof.

[0156] In various embodiments, the nucleic acid template comprises a fluorogenic RNA aptamer to enable fluorescence-based detection of transcript production. In various embodiments, the fluorogenic RNA aptamer is a broccoli aptamer, a structured RNA that binds to a fluorogenic dye (e.g . , DFHBI) and emits fluorescence upon binding, allowing visualization of RNA synthesis in vitro or in cells. In alternative embodiments, members of the broader fluorogenic RNA aptamer family may be used in place of or in addition to the broccoli aptamer, including but not limited to Spinach, Spinach2, Mango, Corn, and Pepper aptamers, each offering distinct spectral and structural properties for multiplexed or application-specific use. In various embodiments, the fluorogenic RNA aptamer may be positioned at the 3’ end of the nucleic acid template, and downstream of the other elements. In various embodiments, the broccoli aptamer comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 25 or 26, or a variant thereof.

[0157] In various embodiments, the nucleic acid template may comprise intervening nucleotide sequences that may be termed as a spacer sequence that separate elements from each other The spacer sequence may be termed as 5' spacer sequence or a 3’ spacer sequence dependent on the positional relationship to an element in the nucleic acid template. The spacer sequence may be ashort spacer sequence that is less than or equal to 30 nucleotides (<30 nt) in length, or a long spacer sequence that is more than 30 nucleotides (>30 nt) in length. In various embodiments, the spacer sequence is a short spacer sequence of at least 2, 3, 4, 5, 6, 7, 3, 9, 10, 11. 12, 13, 14, 15, 16, 17, 18, 19, 20, 25 or 30 nucleotides in length. In various embodiments, a spacer sequence is included to separate the T7 promoter sequence from other elements In various embodiments, a spacer sequence is included to separate the nucleotide sequence encoding a self-cleaving ribozyme and fluorogenic RNA aptamer from other elements.

[0158] In various embodiments, the nucleic acid template may comprise intervening nucleotide sequences between functional elements derived from cloning vector backbones, such as multiple cloning sites, antibiotic resistance markers, or other non-transcribed vector components.

[0159] In various embodiments, the nucleic acid template as used in the method may be referred to as an in vitro transcription (IVT) template, and the polypeptide disclosed herein may be used in an IVT method.

[0160] Accordingly, in a still further aspect, the present invention relates to the use of polypeptides described herein for performing an in vitro transcription (IVT) reaction.

[0161] In various embodiments, there is provided a method of performing an in vitro transcription (IVT) reaction, comprising contacting a nucleic acid template with the polypeptide disclosed herein in the presence of nucleoside triphosphates under conditions that result in the production of RNA transcript.

[0162] In various embodiments, the conditions may comprise incubation at 30-42°C for 1-4 hours. Suitable conditions to carry out an IVT reaction are readily know to those skilled in the art, and may be suitably selected.

[0163] The term “In vitro transcription (IVT)” refers to a procedure that allows for template -directed synthesis of RNA molecules of any sequence from short oligonucleotides to those of several kilobases in pg to mg quantities. It is based on the engineering of a template that includes a bacteriophage promoter sequence (e.g. from the T7 coliphage) upstream of the sequence of interest (i.e. gene of interest) followed by transcription using the corresponding RNA polymerase.

[0164] In various embodiments, the nucleoside triphosphates may comprise modified nucleoside triphosphates, preferably the modified nucleoside triphosphates comprise a modified nucleobase selected from pseudouridine, 1 -methylpseudouridine, 1 - ethylpseudouridine, 2-thiouridine, 4'- thiouridine, 2-thio-l -methyl- 1 -deaza-pseudouridine, 2 -thio- 1 -methyl -pseudouridine, 2-thio-5 -aza- uridine, 2-thio-dihydropseudouridine, 2-thio- dihydrouridine, 2-thio-pseudouridine, 4-methoxy-2-thio- pseudouridine, 4-methoxy- pseudouridine, 4-thio-l -methyl -pseudouridine, 4-thio-pseudouridine, 5 -aza-uridine, dihydropseudouridine, 5 -methyluridine, 5 -methoxyuridine (mo5U), 2’-O-methyl uridine, Phosphorothioate (PS), 2’-O-Methyl (2’-OMe), 2’-O-Methoxyethyl (2’-MOE), 2’-Fluoro (2’-F), and Locked Nucleic Acid (LNA).

[0165] In various embodiments, the method further comprises the use of a cap analog, and the contacting step is carried out in the presence of the cap analog. In various embodiments, the cap analog may be a dinucleotide cap analog, a trinucleotide cap analog, or a tetranucleotide cap analog.

[0166] In various embodiments, the nucleic acid template may comprise from 5’ to 3’ direction: the T7 promoter and tluorogenic RNA aptamer. This template corresponds to IVT-13 as shown in FIG. 1. In various embodiments, the nucleic acid template comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 19, or a variant thereof.

[0167] In various embodiments, the nucleic acid template may comprise from 5’ to 3’ direction: the T7 promoter, a nucleotide sequence encoding a 5'UTR, a gene of interest, a nucleotide sequence encoding a 3’UTR, a nucleotide sequence encoding a self-cleaving ribozyme, and a fiuorogenic RNA aptamer. This template corresponds to IVT-22 as shown in FIG. 1. In various embodiments, the nucleic acid template comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 20, or a variant thereof.

[0168] In various embodiments, the nucleic acid template may comprise from 5’ to 3’ direction: the T7 promoter, a nucleotide sequence encoding a 5'UTR, a gene of interest, a nucleotide sequence encoding a 3’UTR, and a poly(A) tail sequence. This template corresponds to IVT-25 and iVT-28 as shown in FIG. 1. In various embodiments, the nucleic acid template comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 21 or 23, or a variant thereof.

[0169] In various embodiments, the nucleic acid template may comprise from 5’ to 3’ direction: the T7 promoter, a 5’ intron sequence, a gene of interest, a luciferase reporter gene, 3’ intron sequence, a nucleotide sequence encoding a self-cleaving ribozyme, and a fluorogeoic RNA aptamer. This template corresponds to IVT-24a as shown in FIG. 1. In various embodiments, the nucleic acid template comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 22, or a variant thereof.

[0170] In various embodiments, the nucleic acid template may comprise from 5’ to 3’ direction: the T7 promoter, a nucleotide sequence encoding a 5'UTR, a gene of interest, a nucleotide sequence encoding a 3’UTR, a poly(A) tail sequence, a nucleotide sequence encoding a self-cleaving ribozyme, and a fluorogenic RNA aptamer. This template corresponds to iVT-25b as shown in FIG. 1. In various embodiments, the nucleic acid template comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 24, or a variant thereof.

[0171] In various embodiments, the polypeptide disclosed herein exhibits an increased RNA polymerase activity compared to the WT T7 RNA polymerase (SEQ ID NO:1 ), resulting in enhanced yield of the RNA transcript. In various embodiments, the RNA transcript yield may be increased by at least 1.1 fold, at least 1.2 fold, at least 1.5 fold, at least 2-fold, at least 3-fold, at least 5-fold, at least 10- fold, at least 20-fold, at least 50-fold, or at least 100-fold, compared to a corresponding wild-type T7 RNA polymerase which has the sequence of SEQ ID NO: 1 .

[0172] In a still further aspect, the present invention relates to the use of polypeptides disclosed herein for heterologous expression of a protein of interest in a host cell.

[0173] Accordingly, there is provided a method for heterologous expression of a protein of interest in a host cell, comprising: introducing into the host cell a nucleic acid sequence encoding the polypeptide disclosed herein, and a heterologous nucleic acid sequence encoding the protein of interest, wherein the heterologous nucleic acid sequence is operably linked to a T7 promoter; and inducing expression of the polypeptide in the host cell under suitable conditions wherein the induced expression of the polypeptide drives transcription of the heterologous nucleic acid sequence and expression of the protein of interest.

[0174] The term “heterologous nucleic acid sequence”, as used herein, refers to a nucleic acid sequence that originates from a different species, organism, cell type, or genomic context relative to the host cell or vector into which it is introduced. A heterologous sequence is not naturally found at the integration site or under the control of the same regulatory elements in the host genome. The heterologous nucleic acid sequence may encode the protein of interest, regulatory elements, and may be operably linked to a promoter to enable its expression in the host system.

[0175] In various embodiments, the heterologous nucleic acid sequence comprises a gene of interest that encodes the protein of interest. In various embodiments, the gene of interest is operably linked to a T7 promoter such that expression of the polypeptide drives direct transcription of the gene of interest and expression of the protein of interest. In various embodiments, the heterologous nucleic acid sequence comprises a gene of interest that encodes the protein of interest and a regulatory element operably linked to a T7 promoter, such that expression of the polypeptide drives expression of the regulatory element, which in turn regulates expression of the protein of interest. Thus, expression of the protein of interest may be induced either directly or indirectly via a regulatory intermediate.

[0176] In various embodiments, the step of introducing refers to the delivery of one or more nucleic acid sequences into a host cell using any suitable method known in the art. This may include physical, chemical, or biological techniques such as transformation, transfection, electroporation, microinjection, or viral vector-mediated delivery. The nucleic acid sequences may be in the form of plasmid vectors, linear DNA fragments, or episomal constructs. The introduction step enables thehost cell to express the RNA polymerase polypeptide and the protein of interest under transcriptional control of the T7 promoter.

[0177] In various embodiments, the step of inducing expression refers to the initiation or enhancement of transcription and / or translation of the polypeptide and / or protein of interest encoded by the introduced nucleic acid sequence. This may involve providing an inducer molecule (e.g., IPTG in lac-inducible systems), altering culture conditions (e.g., temperature shift, medium composition), or activating a promoter that controls expression of the RNA polymerase or a regulatory element. Expression of the RNA polymerase results in transcription of the T7 promoter-controlled gene or T7 promoter controlled regulatory element, thereby enabling downstream expression of the protein of interest. The use of a T7 promoter operably linked to the protein of interest coding sequence or regulatory element allows for selective and robust transcription by the expressed polypeptide having RNA polymerase activity. The method may be applied for transient or stable expression depending on the host cell system and intended use.

[0178] In various embodiments, the method further comprises purifying the protein of interest from the host cell. In various embodiments, the heterologous nucleic acid sequence further comprises a nucleotide sequence encoding a tag for purification.

[0179] In various embodiments, the protein of interest is a recombinant protein selected from the group consisting of therapeutic proteins, enzymes, and antibodies.

[0180] In various embodiments, separate nucleic acid molecules may be introduced into the host cell, one comprising the nucleic acid sequence encoding the polypeptide disclosed herein and the other comprising the heterologous nucleic acid sequence encoding the protein of interest. That is, a first nucleic acid molecule comprising the nucleic acid sequence encoding the polypeptide disclosed herein, and a second nucleic acid molecule comprising the heterologous nucleic acid sequence encoding the protein of interest may be introduced into the host cell.

[0181] In a still further aspect, the present invention relates to the use of the fusion protein disclosed herein for introducing one or more mutations in a target nucleic acid sequence in a host cell.

[0182] Accordingly, there is provided a method of introducing one or more mutations in a target nucleic acid sequence in a host cell, comprising: introducing into the host cell a nucleic acid molecule encoding the fusion protein disclosed herein: and inducing expression of the fusion protein in the host cell to transcribe the target nucleic acid sequence.

[0183] As used herein, the term “target nucleic acid sequence” refers to a nucleotide sequence of interest within the host cell that is subject to transcription and mutation, and may include coding or non-coding DNA, plasmid-encoded genes, or chromosomally integrated sequences. In variousembodiments, the target nucleic acid sequence may be endogenous to the host cell (i.e., naturally present in the genome), or may be exogenous and introduced into the host cell, for example, via transformation, transfection, electroporation, or viral transduction. In such embodiments, (I) a first nucleic acid molecule encoding the fusion protein disclosed herein, and (ii) a second nucleic acid molecule comprising the target nucleic acid sequence, optionally operably linked to a promoter sequence for transcriptional control may be introduced into a host cell.

[0184] In various embodiments, the target nucleic acid sequence may be operably linked to a T7 promoter such that expression of the fusion protein drives transcription of the target nucleic acid sequence and leads to or facilitates the introduction of one or more mutations in the target nucleic acid sequence. In various embodiments, the host cell comprises a nucleic acid molecule comprising the target nucleic acid sequence operably linked to a T7 promoter.

[0185] In various embodiments, the step of introducing the nucleic acid molecules into the host cell may be carried out by any method known in the art, including but not limited to: DNA transfection, transformation, or electroporation; viral transduction; nanoparticle-mediated delivery; biolistic particle delivery; ultrasound-mediated delivery; microinjection; or direct polypeptide delivery methods such as cell-penetrating peptides, liposome encapsulation, or protein transduction domains. The purpose of this introduction is to enable intracellular production of the fusion protein, which in turn facilitates transcription of a target nucleic acid sequence and the introduction of mutations therein.

[0186] In various embodiments, the step of inducing expression refers to initiating or enhancing the transcription and / or translation of the fusion protein encoded by the introduced nucleic acid molecule. This step may involve application of chemical inducers (e.g., IPTG), environmental triggers (e.g., temperature or pH shifts), or activation of inducible promoters (e.g., lac or arabinose-inducible systems). Upon expression, the fusion protein catalyzes transcription of the target nucleic acid sequence, during which mutations may be introduced, via enzymatic modification (e.g., base deamination by the fused deaminase domain).

[0187] In various embodiments, the host ceil may be as defined above, and generally refer to any cell that is capable of receiving and expressing a nucleic acid molecule introduced into it. Preferred host cells include prokaryotic and eukaryotic cells that are amenable to genetic manipulation and expression of recombinant proteins. In various embodiments, the host cell is a bacterium, such as an E.coli strain, which is compatible with T7 promoter-based expression systems. In other embodiments, the host cell is a eukaryotic cell such as a yeast, insect cell, or mammalian cell, particularly when eukaryotic post-translational modifications are desired.

[0188] In various embodiments, the nucleic acid molecule is a plasmid vector.

[0189] In various embodiments, the method further comprises selecting or identifying the one or more mutations in the target nucleic acid sequence of the host cell. The one or more mutations may be identified by sequencing the target nucleic acid sequence. Identification may be performed by analysing the resulting nucleotide changes using nucleic acid sequencing techniques, such as Sanger sequencing or next-generation sequencing (NGS), which allow detection of point mutations, insertions, or deletions introduced during transcription and / or replication.

[0190] In various embodiments, the one or more mutations are introduced by the activity of the deaminase domain, which enables direct chemical modification of nucleobases, for example, deamination of cytidine to uridine (or deoxycytidine to deoxyuridine) or adenosine to inosine in the target nucleic acid sequence. Such modifications may occur either during transcription (e.g., co- transcriptional RNA editing) or on exposed single-stranded DNA regions, such as transcription bubbles. These base conversions may subsequently be resolved into permanent mutations following reverse transcription, replication, or repair, thereby introducing site-specific or localized mutations into a target nucleic acid molecule in a programmable or stochastic manner.

[0191] In various embodiments, the one or more mutations may be introduced by error-prone transcription of the fusion protein disclosed herein, in addition to the activity of the deaminase domain. As used herein, “error-prone transcription" refers to the synthesis of RNA with a higher-than-normal rate of nucleotide misincorporation due to the intrinsic or engineered infidelity of the RNA polymerase, which can result in mutation propagation when such RNA intermediates are reverse-transcribed or influence DNA repair or replication mechanisms in the cell.

[0192] In various embodiments, the target nucleic acid sequence comprises a reporter gene for detecting the one or more mutations, and the target nucleic acid sequence may be operably linked to a T7 promoter. The reporter gene serves as a detectable or quantifiable marker to facilitate the identification of transcriptional activity and / or mutational changes. As used herein, the “reporter gene” refers to a gene whose expression produces an easily measurable phenotype or signal (e.g., luciferase, GFP, p-galactosidase), enabling detection of mutational effects. The target nucleic acid sequence may be operably linked to a T7 promoter, meaning that the reporter gene is placed downstream of the T7 promoter such that it can be transcribed by the fusion protein disclosed herein in the host cell.

[0193] In various embodiments, the method further comprises introducing into the host cell a mutagenic agent to increase the mutation rate in the target nucleic acid sequence. As used herein, “mutagenic agent” includes any chemical, biological, or physical factor capable of inducing mutations, such as ethyl methanesulfonate (EMS), or ultraviolet (UV) irradiation. The combination of mutagenic agents may elevate the mutation frequency for directed evolution or functional screening applications.

[0194] In previous studies, it has been revealed that T7 RNAP can be divided or split into multiple fragments that may be co-expressed and reconstituted to function (see: Segall-Shapiro et al., Nature Chemical Biology, 2015, PMC4299498). For example, the DNA-binding loop responsible for promoter recognition is located within a C-terminal 285-amino acid fragment, referred to as the o fragment’ This o fragment can be modularly exchanged to alter promoter specificity. When paired with a complementary 601 -amino acid ‘core fragment’ that retains catalytic function, the reconstituted polymerase can be directed to transcribe different target genes depending on the o fragment used This modularity enables engineered systems in which promoter specificity and enzymatic activity are independently programmable through fragment reassociation.

[0195] In this regard, split T7 RNA polymerase fragments can be used to detect biomolecular interactions. When the biomolecules interact, the proximity-induced reassociation of the fragments restores RNA polymerase activity. In particular, the reconstituted RNA polymerase may then transcribe a reporter gene from a T7 promoter, producing a detectable output. The presence or level of this detectable output serves as a direct readout of the underlying biomolecular interaction, enabling its use in biosensing and diagnostic applications.

[0196] As used herein, the term “biomolecular interactions” refers to any physical or functional association between partners of interest, which may include two or more biological molecules, such as proteins, nucleic acids, lipids, carbohydrates, or small molecules. These interactions may include, but are not limited to, protein-protein interactions, protein-nucleic acid interactions (e.g., DNA- protein or RNA-protein), nucleic acid-nucleic acid interactions (e.g., hybridization), protein-small molecule interactions (e.g., ligand binding), and antibody-antigen binding. Biomolecular interactions may be transient or stable, direct or mediated through one or more intermediates, and may occur in vitro or in vivo. Such interactions are often specific and can modulate biological function, making them useful targets or indicators in biosensing, therapeutic, or diagnostic applications.

[0197] In various embodiments, the partners of interest of the biomolecular interaction comprise a pair of molecules wherein one molecule is a target analyte and the other is a binding or recognition partner specific for the target analyte.

[0198] Accordingly, in a still further aspect, the present invention relates to the use of a split version of the polypeptide disclosed herein, for detecting biomolecular interactions involving a target analyte, or for detecting the presence, and optionally quantity, of the target analyte.

[0199] The term “split version”, as used herein in the context of the polypeptide disclosed herein refers to a modified form of the polypeptide that is engineered to be divided into two or more separate fragments (i.e. non-functional fragments). These fragments are individually inactive with regard to RNA polymerase activity but can reconstitute into a functional, transcriptionally active enzyme under specific conditions. This reconstitution typically occurs when the fragments are brought together by aspecific biomolecule interaction or binding event, such as the presence of a target analyte in a biosensor application.

[0200] In various embodiments, the split version may comprise and be referred to as split nonfunctional fragments of the polypeptide disclosed herein. “Reconstitution” refers to the physical or functional reassociation of protein fragments into a transcriptionally active polypeptide, more particularly a polypeptide having RNA polymerase activity, even more preferably a polypeptide having T7 RNA polymerase activity.

[0201] In various embodiments, the reporter gene encodes a detectable marker selected from the group consisting of fluorescent proteins (e.g., GFP, mCherry), luminescent proteins (e.g., luciferase), enzymatic reporters, chromogenic proteins (e.g., p-galactosidase) or RNA-based reporters such as aptamers or barcoded transcripts. In various embodiments, the expression of the reporter gene is quantified to determine the amount (e.g. concentration) of the target analyte in the sample. The expression of the reporter gene may be detected and quantified using standard molecular or biochemical techniques appropriate to the nature of the reporter. For fluorescent or luminescent reporters, signal intensity may be measured using a fluorometer, luminometer, or plate reader. For RNA-based reporters (e.g., fluorescent aptamers or barcoded transcripts), expression levels may be assessed by quantitative reverse transcription PCR (qRT-PCR), RNA sequencing, or fluorescence microscopy. Protein-based reporters may also be quantified by enzyme-linked immunosorbent assay (ELISA), Western blotting, or colorimetric assays. The level of reporter expression may serves as a direct or proportional readout of the target analyte presence or interaction. The resulting signal from the reporter gene may be detected qualitatively as an indicator of presence or absence, or quantitatively to assess the concentration or strength of the target analyte interaction.

[0202] Accordingly, there is provided a method for detecting biomolecular interactions involving a target analyte, comprising the steps of: providing a split version of the polypeptide disclosed herein, wherein the split version is linked to a detection system that facilitates reconstitution of the split version into a transcriptionally active polypeptide in the presence of the target analyte; contacting the split version with a sample suspected of containing the target analyte in the presence of a reporter gene construct comprising a reporter gene operably linked to a promoter recognized by the transcriptionally active polypeptide, wherein interaction with the target analyte induces reconstitution of the split version into the transcriptionally active polypeptide; and detecting the expression of the reporter gene as an indication of the biomolecular interaction involving the target analyte.

[0203] In another embodiment, there is also provided a method for detecting the presence and optionally the quantity of a target analyte in a sample, the method comprising the steps of:providing a split version of the polypeptide disclosed herein, wherein the split version is linked to a detection system that facilitates reconstitution of the split version into a transcriptionally active polypeptide in response to interaction with the target analyte; contacting the split version with the sample, in the presence of a reporter gene construct comprising a reporter gene operably linked to a promoter recognized by the transcriptionally active polypeptide, wherein the presence of the target analyte in the sample induces reconstitution of the split version into the transcriptionally active polypeptide; detecting expression of the reporter gene as an indication of the presence of the target analyte, and optionally quantifying the reporter gene expression to determine the quantity or concentration of the target analyte in the sample.

[0204] In various embodiments, the method further comprises providing the reporter gene construct and contacting the sample with the reporter gene construct. In various embodiments, the promoter is a T7 promoter. Accordingly, the contacting step may involve adding the split into a reaction mixture or cellular environment in the presence of a promoter-driven (e.g. T7 promoter driven) reporter gene construct.

[0205] In various embodiments, the method comprises inducing expression of the reporter gene by the active transcriptionally active polypeptide, if reconstituted.

[0206] In various embodiments, the split version comprises two or more non-functional fragments of the polypeptide disclosed herein.

[0207] The term “non functional fragment”, as used herein in relation to the polypeptides disclosed, refers to a fragment derived from the polypeptide having RNA polymerase activity disclosed herein, which, on its own, lacks RNA polymerase activity, either due to deletion, truncation, or the absence of critical catalytic or structural domains required for transcriptional function. A non-functional fragment may correspond to a portion of the full-length polypeptide that, when isolated or expressed independently, is insufficient to support transcription from a T7 promoter, but may regain function when reassociated or combined with one or more complementary non-functional fragments. The two or more non-functional fragments may reassociate under analyte-dependent conditions to form a functional, transcriptionally active complex.

[0208] In various embodiments, the non-functional fragments may be provided as recombinantly expressed and purified polypeptides, or alternatively, as nucleic acid molecules encoding the respective fragments for in vitro or in vivo expression.

[0209] In various embodiments, the detection system may comprise any molecular configuration, construct, or combination of components that enables the specific binding or recognition of a target analyte. The detection system serves to mediate or facilitate the reassociation of split T7 RNApolymerase fragments, thereby linking biomolecule interaction involving the target analyte to transcriptional activation of a reporter gene. For example, each fragment of the split RNA polymerase may be fused to a binding or recognition partner that interacts specifically with the target analyte. When the target analyte is present, upon interaction of a target analyte with its corresponding binding or recognition partner, the non-functional fragments are brought into close proximity through molecular recognition events, leading to their reassociation and restoration of RNA polymerase activity (i.e. transcriptional activity). The reconstitution of the active RNA polymerase, may initiate transcription from a T7 promoter of the reporter gene construct, ensuring that transcription is exclusive to the biosensor pathway, whereby the expression of the reporter gene produces a detectable signal, indicating the biomolecule interaction and presence of the target analyte.

[0210] In various embodiments, the detection system comprises one or more binding or recognition partners capable of specifically interacting with the target analyte. In various embodiments, the binding or recognition partner may be fused or operably linked to one or more of the non-functional fragments, such that the interaction or binding event involving the target analyte results in the physical colocalization or conformational stabilization of the fragments, thereby restoring transcriptional activity.

[0211] In various embodiments, one or more of the non-functional fragments may be linked, directly or indirectly, to the binding or recognition partners, thereby enabling reconstitution of transcriptional activity upon target analyte binding.

[0212] As used herein, the term “binding or recognition partner” refers to any molecule, domain, or structural element that specifically associates with a target analyte through molecular recognition. The binding or recognition partner may include, but is not limited to, antibodies, antibody fragments, aptamers, peptide ligands, nucleic acid-binding proteins, receptors, small molecule-binding domains, protein-protein interaction domains, nucleic acid-based elements such as toehold switches or strand displacement assemblies, or other affinity elements. These partners are capable of forming a specific and detectable interaction with the target analyte and may be fused or operably linked to a nonfunctional fragment of a split RNA polymerase to mediate proximity-induced or conformationally induced reconstitution of polymerase activity upon analyte binding. In various embodiments, the binding or recognition partner may be more simply termed as a binding partner that specifically and selectively binds to the target analyte.

[0213] This linking of the fragments to the binding or recognition partners may be achieved through genetic fusion, chemical conjugation, or non-covalent interactions, depending on the design. Not all fragments are required to be directly linked to the binding or recognition partners. In various embodiments, only one of the fragments may be linked to a binding or recognition partner or detection moiety, while the other fragment is freely diffusible or otherwise recruited upon target analyte-induced complex formation. In other embodiments, both or all fragments may be independently linked to the same or different binding or recognition partners, enabling target analyte-dependent multivalentassembly. The design of the detection system, and selection of the binding or recognition partners may vary depending on the nature of the target analyte, and may include heterodimerization modules, scaffolds, or bridging elements that mediate analyte-responsive reassociation of the polymerase fragments.

[0214] In various embodiments, binding or recognition partners operably linked to different nonfunctional fragments are configured to bind to distinct or different binding sites on the target analyte. This ensures that each partner can simultaneously interact with the target analyte without steric interference or competitive binding. Such design facilitates the co-localization of the fragments upon target analyte binding, thereby promoting their proximity-induced reconstitution into a transcriptionally active RNA polymerase. In various embodiments, the binding or recognition partners may include antibodies or aptamers targeting different epitopes of a protein target analyte, complementary oligonucleotides recognizing distinct sequences of a nucleic acid target analyte, or small moleculebinding domains recognizing partially or non-overlapping functional groups of a metabolite target analyte. In various embodiments, the reconstitution is driven by the binding of the binding or recognition partners (e.g., antibody fragments, aptamers, peptides) to distinct, non-overlapping epitopes or binding sites on the target analyte. In various embodiments, the reconstitution is driven by the binding of the binding or recognition partners to partially overlapping epitopes or binding sites on the target analyte.

[0215] In various embodiments, the split version comprises two or more non-functional fragments of the polypeptide that reconstitute into an active RNA polymerase upon interaction with a target analyte, wherein each non-functional fragment is operably linked to a binding or recognition partner specific to the target analyte, and wherein the interaction between the binding or recognition partner and the target analyte facilitates reconstitution of the non-functional fragments into a transcriptionally active enzyme.

[0216] In various embodiments, the non-functional fragments may be provided separately or premixed, depending on the detection format, and may be provided linked with the binding or recognition partners. In various embodiments, all components, including the split fragments, a nucleic acid construct containing a T7 promoter operably linked to a reporter gene (i.e. reporter gene construct), and necessary cofactors (e.g., nucleotides, magnesium ions, transcription buffer), may be comprised in a reaction mixture. In various embodiments, the non-functional fragments may be expressed from one or more expression vectors introduced into a host cell, optionally under the control of inducible or constitutive promoters. Thus, the providing step ensures that all molecular components necessary for the target analyte-dependent transcriptional activation are present and configured for subsequent contact with the sample.

[0217] The term “sample”, as used herein, refers to any material or composition that is suspected of containing a target analyte and is suitable for use in the detection method described herein. Thesample may be of biological, environmental, clinical, industrial, or synthetic origin, and may be in solid, liquid, or gaseous form, or a combination thereof. In various embodiments, the sample may comprise a biological fluid such as blood, serum, plasma, urine, saliva, cerebrospinal fluid, tears, or lymph; a tissue lysate or cellular extract; a whole cell population or microbial culture; or a purified or semipurified preparation such as a protein solution, nucleic acid preparation, or small molecule library. The sample may also comprise environmental samples such as water, soil, air, or food, or industrial or pharmaceutical formulations to be screened for target compounds. The sample may be unmodified or subjected to preliminary processing steps (e.g., filtration, lysis, centrifugation, dilution) to facilitate compatibility with the detection system.

[0218] As used herein, the term “target analyte" refers to any molecule, molecular complex, or biological entity that is suspected to be present in a sample and whose presence, and interaction is to be detected using the methods described herein. The target analyte may be of biological, environmental, clinical, industrial, or synthetic origin, and includes, but is not limited to, proteins, peptides, nucleic acids (DNA or RNA), carbohydrates, lipids, small molecules, metabolites, toxins, pathogens (e.g., bacteria, viruses), whole cells, or molecular assemblies. The target analyte specifically interacts with the binding or recognition partner operably linked to a non-functional fragment of a split RNA polymerase, such that binding of the target analyte facilitates or induces reconstitution of transcriptional activity. The target analyte may be present in complex mixtures or purified forms, and may be endogenous or exogenous to the sample under investigation. In various embodiments, the target analyte may be a protein (e.g., detection of protein-protein interactions such as p53-MDM2), a nucleic acid (e.g., a specific RNA or DNA strand that triggers toehold-mediated strand displacement to bring transcriptional components together), or a small molecule or metabolite, such as theophylline, cyclic AMP (cAMP), or metal ions.

[0219] In various embodiments, the split version may comprise two fragments, a first fragment may comprise or consist of an amino acid sequence corresponding to residues at positions 1 -179, 1 -510, 1 -563, or 1 -601 of SEQ ID NO:1, and a second fragment may comprise or consist of an amino acid sequence corresponding to residues 180-883, 511 -883, 564-883, or 602-883 of SEQ ID NO:1. In various embodiments, the two fragments comprise: a first fragment comprising or consisting of an amino acid sequence corresponding to residues at positions 1 -179 of SEQ ID NO:1, and a second fragment comprising or consisting of an amino acid sequence corresponding to residues 180-883 of SEQ ID NO:1 ; or a first fragment comprising or consisting of an amino acid sequence corresponding to residues at positions 1 -510 of SEQ ID NO:1, and a second fragment comprising or consisting of an amino acid sequence corresponding to residues 511 -883 of SEQ ID NO:1 ; or a first fragment comprising or consisting of an amino acid sequence corresponding to residues at positions 1 -563 of SEQ ID NO:1, and a second fragment comprising or consisting of an amino acid sequence corresponding to residues 564-883 of SEQ ID NO:1 ; ora first fragment comprising or consisting of an amino acid sequence corresponding to residues at positions 1 -601 of SEQ ID NO:1, and a second fragment comprising or consisting of an amino acid sequence corresponding to residues 602-883 of SEQ ID NO:1 ; or

[0220] In various embodiments, the split fragments may comprise or consist of an amino acid sequence corresponding to residues at positions 1 -67, 68-179, 180-601 , and / or 602-883 of SEQ ID NO:1.

[0221] In various embodiments, the split version of the polypeptide disclosed herein may comprise two non-functional fragments corresponding to amino acid positions 1 -601 , and 602-883 of SEQ ID NO:1. That is, a first non-functional fragment comprising an amino acid sequence corresponding to positions 1 -601 of SEQ ID NO:1 , and a second non-functional fragment comprising an amino acid sequence corresponding to positions 602-883 of SEQ ID NO:1 .

[0222] In various embodiments, the split version of the polypeptide disclosed herein may comprise or consist of four non-functional fragments of the polypeptide that may be used to form a split version. In various embodiments, the split version of the polypeptide disclosed herein may comprise four nonfunctional fragments corresponding to amino acid positions 1-67, 68-179, 180-601 , and 602-883 of SEQ ID NO:1. That is, a first non-functional fragment comprising or consisting of an amino acid sequence corresponding to positions 1 -67 of SEQ ID NO:1 , a second non-functional fragment comprising or consisting of an amino acid sequence corresponding to positions 68-179 of SEQ ID NO:1 , a third non-functional fragment comprising or consisting of an amino acid sequence corresponding to positions 180-601 of SEQ ID NO:1 , and a fourth non-functional fragment comprising or consisting of an amino acid sequence corresponding to positions 602-883 of SEQ ID NO:1 .

[0223] It will be appreciated that the polypeptide disclosed herein may be split at alternative positions, either naturally tolerated or experimentally optimized, to improve structural integrity, reduce background transcriptional activity, or enhance the efficiency of reconstitution. Suitable split points may include residues proximate to positions 179-180, which demarcate the interface between the N- terminal and central domains, positions 510-51 1 , positions 563-564, or positions near 600-602, corresponding to the boundary of the core and o fragment. These design strategies facilitate modular biosensor construction with tuneable responsiveness.

[0224] In various embodiments, expression of the reporter gene may be quantified to determine the concentration of the target analyte in the sample, enabling both qualitative detection and quantitative analysis.

[0225] While various embodiments herein describe the use of split versions of the RNA polymerase polypeptide for conditionally controlled transcriptional activity, it is understood that unsplit, full-lengthversions of the polypeptide disclosed herein may also be employed in alternative embodiments where constitutive or inducible activity may be desirable.

[0226] Accordingly, in a still further aspect, the present invention relates to the use of the polypeptide disclosed herein, for detecting biomolecular interactions involving a target analyte, or for detecting the presence, and optionally quantity, of the target analyte.

[0227] In various embodiments, there is provided a method for detecting biomolecular interactions involving a target analyte, the method comprising the steps of: providing a polypeptide disclosed herein linked to a detection system such that the transcriptional activity of the polypeptide is modulated in response to interaction with the target analyte; contacting the polypeptide with a sample suspected of containing the target analyte, in the presence of a reporter gene construct comprising a reporter gene operably linked to a promoter recognised by the polypeptide, wherein interaction with the target analyte modulates the transcriptional activity of the RNA polymerase; and detecting the expression of the reporter gene as an indication of the biomolecular interaction involving the target analyte.

[0228] In various embodiments, there is also provided a method for detecting the presence and optionally the quantity of a target analyte in a sample, the method comprising the steps of: providing a polypeptide disclosed herein linked to a detection system such that the transcriptional activity of the polypeptide is modulated in response to interaction with the target analyte; contacting the polypeptide with a sample suspected of containing the target analyte, in the presence of a reporter gene construct comprising a reporter gene operably linked to a promoter recognised by the polypeptide; detecting expression of the reporter gene as an indication of the presence of the target analyte, and optionally quantifying the reporter gene expression to determine the quantity or concentration of the target analyte in the sample.

[0229] In various embodiments, the polypeptide may be functionally regulated by one or more target analyte-responsive elements, such that the presence of a target analyte induces a conformational or steric change, interaction-induced release, or promoter accessibility that activates or suppresses transcriptional activity, thereby enabling detection of the biomolecular interaction through reporter gene expression.

[0230] The invention is further illustrated by the following non-limiting examples and the appended claims.EXAMPLESMaterials and Methods

[0231] General cloning: PCR was performed using Phusion Hot Start II DNA polymerase (Thermo Fisher Scientific) and purified using Zymo-Spin IC columns (Zymo Research). All plasmids were constructed by Gibson assembly using the HiFi DNA Assembly Master Mix (New England Biolabs). T7 RNAP sequences were amplified from bacteriophage encoding the evolved T7 variant. Coding and non-coding sequences were synthesized as gene blocks (Genscript). Maehl chemically competent E. coli cells (Thermo Fisher Scientific) were used for plasmid construction. Plasmids were purified using QIAprep Spin Miniprep Kit (Qiagen) or QIAGEN Plasmid Plus Midi Kit (Qiagen).

[0232] Protein purification: Unless otherwise stated, plasmids were transformed into XJB (DE3) chemically competent E.coli cells (Zymo research) and plated on Luria-Bertani (LB) agarose plates (1.5% w / v, except as noted) containing 50 pg / ml kanamycin and left to shake in an incubator (200 rpm) at 37°C overnight

[0233] Bead-clean-up: Colonies were isolated and inoculated into 1 ml of LB Broth containing 50 pg / ml of kanamycin and overnight autoinduction media (Merck) within a 96-well deep-well plate. The plates were sealed with breathable seals and left to grow overnight in at 37°C with shaking at 200 rpm. Bacteria pellet was collected using a centrifuge (10 minutes, 6,000 g) and resuspended in 300 pl of lysis buffer (50 mM sodium phosphate buffer pH 7.0, 300 mM sodium chloride, 10 mM imidazole, 0.03% Triton X-100) and 15 pl of DNAse I (Vazyme). The suspensions were passed through a freezethaw cycle to ensure thorough lysis. HisPur™ Ni-NTA magnetic beads were thoroughly vortexed, were buffer-exchanged with lysis buffer twice then resuspended in the original volume of beads. 15 pl of washed beads were added to each bacterial lysate then transferred column-wise to a 96-well roundbottom plate (Costar). The beads were then washed with 600 pl of wash buffer (50 mM sodium phosphate buffer pH 7.0, 300 mM sodium chloride 50 mM imidazole) then eluted into 80 pl of elution buffer (50 mM sodium phosphate buffer at pH 7.0, 300 mM sodium chloride, and 500 mM imidazole).

[0234] Large scale column cleanup: A single colony from the transformation was inoculated overnight at 37°C in 1 ml LB broth containing 50 pg / mL kanamycin to generate a starter culture. The starter culture was diluted 200-fold in a fresh LB broth (50 ml) containing 50 pg / mL kanamycin, 1 .5 M L-arabinose and 0.5 M magnesium chloride. Subcultures were grown at 37°C until the optical density at 600 nm (QD600) reached -0.4-0 5 Next, 100 pM isopropyl p-d-1 -thiogalactopyranoside (IPTG) was then added to induce the expression of proteins. Following this, the cultures were incubated at 200 rpm, 37°C for 4 hours. Cells were harvested from this culture using a centrifuge (15 minutes, 8,000 g) maintained at 4°C, before resuspending and freezing the resulting pellet at -80°C in 750 pl of lysis buffer (50 mM sodium phosphate buffer at pH 7.0, 300 mM sodium chloride, 10 mM imidazole, and 0.03% Triton X-100). To ensure protein stability, subsequent purification steps were also carried out at 4°C.

[0235] The frozen pellet was thawed and sonicated to release proteins from the cells. The supernatant obtained from centrifugation of lysates at 13,500* g was loaded on nickel resin (PureCube Biotech) and incubated overnight to capture the His-tagged proteins. 20 ml of wash buffer (50 mM sodium phosphate buffer, 300 mM sodium chloride, and 50 mM imidazole) was then used to wash the resin, before eluting the bound proteins with 2 ml of elution buffer (50 mM sodium phosphate buffer at pH 7.0, 300 mM sodium chloride, and 500 mM imidazole). Lastly, the proteins were exchanged into 50 mM sodium phosphate at pH 7.0 with 50% glycerol to allow for long-term storage at -80°C

[0236] In-vitro transcription: To generate DNA templates for in-vitro transcription, plasmids were linearized with the appropriate restriction enzyme (New England Biolabs) and purified using ZymoSpin V columns or Zymo-Spin V columns (Zymo Research). Transcription reactions were performed in a transcription buffer containing 40 mM Tris-HCI pH 7.9, 19 mM MgCl2, 5 mM DTT and 1 mM spermidine. 4 mM of each ribonucleotide triphosphate were added, and supplemented with Murine RNase inhibitor (1000 units / ml) and inorganic pyrophosphatase (4.15 units / mL) (New England Biolabs). For short templates like IVT-13, 500 ng of DNA template was used. For the remaining templates, the final template concentration was between 10 nM to 40 nM. For the bead-purified protein solutions, 5 pl of protein solution was used. For the column-purified proteins quantified by Nanodrop, the final protein concentration was equimolar to the template concentration. IVT reactions were conducted at 37°C in a PCR block for 2 h. For time-course assays, reactions were conducted at 37°C for up to 14 h. Nucleotide sequences of the IVT templates shown in FIGs.1-8 are provided in Table 5 below, along with individual nucleotide sequences of the elements.

[0237] Broccoli fluorescence assay:T7 transcription was performed as described earlier in 10 pl reaction volumes with the exception of a different transcription buffer. The 1 x transcription buffer contained 40 mM HEPES buffer pH 7.5, 19 mM MgCl2, 100 mM KCI, 5 mM DTT, 1 mM spermidine and 0.1 mM of DFHBI (Medchemexpress). Reactions were performed in 384-well black plates (Corning #3571 ). To determine the fluorescence directly during the transcription, measurements were performed in a Varioskan LUX Multimode microplate reader (Thermo Fisher Scientific), which was thermostatted at 37°C. To ensure simultaneous initiation of the transcription, the mastermix (containing buffer, template and DFHBI dye) and T7 protein were pipetted separately. To prevent evaporation during the transcription, the 384-well plates were covered with sealing tape. The following settings were applied for fluorescence measurements: ex=479 nm em=532 nm with a kinetic loop of 1 -minute interval for fluorescence reading.

[0238] RNA Tapestation: RNA produced from IVT reaction was purified using the RNACIean XP beads (Beckman Coulter) as per manufacturer's instructions RNA integrity was analyzed via the 4150 TapeStation system (Agilent) using the high-sensitivity RNA ScreenTape (Agilent) as per manufacturer's instructions.Results and DiscussionExample 1: Engineering T7 RNA Polymerase Variants

[0239] Six functional T7 RNA polymerase (T7 RNAP) variants (T7-256, T7-258, T7-259, T7-260, T7- 286 and T7-295) were engineered containing mutations at positions that have not been previously reported in the literature, as shown in Table 6 below.

[0240] Table 6: Genotypes of T7 variants (Mutated amino acid residues are in bold)

[0241] The kinetics and RNA yields of these variants were measured through in-vitro transcription (IVT) of three DNA templates that vary in length and sequence composition (FIG. 1). Compared to wild-type T7 RNAP, these variants demonstrated between 1.3- to 2.3-fold improvement in initial transcription rate, and between 1.2- to 3.9-fold increase in RNA yield, as measured by RNA Tapestation (FIG.2, 3, 4 ).

[0242] In addition, the RNA integrity of IVT products produced by these variants were not compromised, as indicated by the comparable RNA integrity (RIN) scores between the variants and wild-type T7 RNAP (FIG. 5, 6).

[0243] To illustrate the positional importance of R173, P451 , D471 , G520, A586, S767, A260, R57, and D859, refer to FIG. 7A and 7B, and the below Table 7 that summarises the positional mutants and their effects in IVT yield.

[0244] Table 7: Positional Mutations in the T7 RNA polymerase

[0245] T7 variants that resulted in broccoli fluorescence comparable to or higher than wild-type T7 RNAP were selected for further corroborative analyses for mRNA integrity and calibrated yield quantification through RNA TapeStation analysis (FIG. 8). The results show that single amino acid substitutions at positions R173, P451 , G520 and S767 maintained or improve RNA yields. Point mutations at A260 and A586 mostly decreased RNA yields, suggested that mutations at these positions alone are insufficient to improve RNA yields and required the synergistic effects of other mutations.

[0246] The calibrated RNA concentrations and size of the RNA products analyzed by Tapestation in FIG. 8 are shown in the below Table 8.

[0247] Table 8: RNA yield of T7 RNAP Variants

[0248] Six T7 RNA polymerase variants that each contain at least one substitution or a combination of substitutions at positions 57, 173, 260, 451 , 471 , 520, 586, 767, 859 (With respect to WT T7 RNA polymerase) were ranked relative to the performance of each based on initial rate data, as indicated in Table 9.

[0249] Table 9: Ranking of T7 RNAP variants

[0250] The ranking in Table 9 is based on descending RNA yields (ie. 1 has the highest RNA yield among the six variants. Depending on the application of the RNAP, different variants might be used. For example, in the case of heterologous protein expression, a highly active T7 variant might be toxic to the cell. As such, a less active T7 variant that gives lower RNA yields might be favoured.Example 2: High-throughput sequencing ofeGFP cDNA to determined transcription fidelity of T7-295

[0251] The variant T7-295 was further investigated and shown to increase RNA yield while maintaining transcription fidelity. In particular, the variant T7-295 was shown to have similar transcription fidelity and co-capping efficiency compared to WT-T7. Thus, the mutations did not negatively impact these properties.

[0252] FIG. 9 illustrates High-throughput sequencing of eGFP cDNA used to determine transcription fidelity of T7-295. To determine the transcription fidelity of the T7 variants, IVT was performed on a reporter gene and reverse-transcribed the RNA with a barcoded gene-specific primer. The barcoded cDNA helps to eliminate potential artefacts from sequencing or PCR, and the reads are aligned to the reference and the frequency of point mutations and indels is calculated using the formula shown in FIG. 9.

[0253] The error rate of reverse transcriptase and Phusion polymerase is approximately 1 in 100,000 bases, which is an order of magnitude lower than the error rate of T7 RNA polymerase. Therefore, the observed error frequency is likely attributable to misincorporation events by the RNA polymerase. From approximately 2 million sequencing reads, the total number of point mutations, insertions, and deletions was determined and divided by the total number of bases transcribed to calculate the overall error frequency. A high fidelity of the T7 RNAP means less likelihood of introducing mutations into RNA during transcription.

[0254] The mutation frequency per base for WT T7 and T7-295 variant is shown in the below Table 10.

[0255] The T7-295 variant was shown to increase RNA yield while maintaining transcription fidelity relative to the WT T7. Further, the mutation frequency of T7-295 was observed to be similar to that of WT T7, with an error rate of approximate 1 in 1000 base pairs for this eGFP template.Example 3: Gel-shift assay for quantitative assessment of capping efficiency

[0256] To assess if the variant maintained its co-transcriptional capping abilities a gel shift assay that provides an easy and quantitative assessment of capping efficiency was used. A template was transcribed by WT T7 and T7-295 in the presence of AG capping reagent. Hammerhead ribozyme undergoes autocleavage to release a 17-nt long 5’ fragment and this fragment will be capped or uncapped depending on the capping efficiency of the T7 polymerase (FIG. 10A).

[0257] As shown in FIG. 10B, the capped transcript can be distinguished from its uncapped version on a PAGE gel. The relative gel band intensities (densitometries) of each product serve as a proxy for capping efficiency. In this regard, the use of the T7 polymerase variant provides a faster and simpler method for quantifying capping efficiency without the need for expensive and time-consuming techniques such as LC-MS.Example 4: Co-transcriptional capping efficiency of T7-295 is comparable to WT T7

[0258] An IVT reaction following transcription of ribozyme by WT T7 or T7-293 was carried out to investigate the co-transcriptional capping efficiency of the variant T7-295.

[0259] As shown in FIG. 11A and 11 B, the co-transcriptional capping efficiency of T7-295 is shown to be comparable to WT, more than 80% capping in the presence of excess cap. Advantageously, the T7 RNAP variants result in higher RNA yields compared to WT T7 and yet is able to maintain a level of co-capping efficiency comparable to WT T7.

[0260] The invention has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the invention. This includes the generic description of the invention with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein. Other embodiments are within the following claims.

[0261] One skilled in the art would readily appreciate that the present invention is well adapted to carry out the objects and obtain the ends and advantages mentioned, as well as those inherent therein. Further, it will be readily apparent to one skilled in the art that varying substitutions and modifications may be made to the invention disclosed herein without departing from the scope and spirit of the invention. The polypeptides, fusion proteins, nucleic acid molecules, methods, and uses described herein are presently representative of preferred embodiments are exemplary and are not intended as limitations on the scope of the invention. The listing or discussion of a previously published document in this specification should not necessarily be taken as an acknowledgement that the document is part of the state of the art or is common general knowledge.

[0262] The invention illustratively described herein may suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein. Thus, it should be understood that although the present invention has been specifically disclosed by exemplary embodiments and optional features, modification and variation of the inventions embodied therein herein disclosed may be resorted to by those ski lled in the art, and that such modifications and variations are considered to be within the scope of this invention.

[0263] The content of all documents and patent documents cited herein is incorporated by reference in their entirety.

Claims

CLAIMS1 . A polypeptide having RNA polymerase activity, comprising or consisting of:(i) an amino acid sequence set forth in SEQ ID NO:1 ;(ii) an amino acid sequence that shares at least 65%, preferably at least 75%, even more preferably at least 85%, most preferably at least 95% sequence identity with, or at least 80%, preferably at least 90%, more preferably at least 95% sequence homology with, the amino acid sequence as set forth in SEQ ID NO:1 ; or(iii) a functional fragment of (i) or (ii); wherein the polypeptide comprises one or more amino acid mutations at a position selected from R57, R173, A260, P451 , D471 , G520, A586, S767, D859 and combinations thereof, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1 .

2. The polypeptide of claim 1 , wherein the amino acid mutation is an amino acid substitution, and wherein: the amino acid residue R at position 57 is substituted with a basic amino acid residue, preferably H; and / or the amino acid residue R at position 173 is substituted with a polar amino acid residue, preferably S; and / or the amino acid residue A at position 260 is substituted with a polar amino acid residue, preferably S; and / or the amino acid residue P at position 451 is substituted with a non-polar amino acid residue, preferably a non-polar aliphatic amino acid, more preferably L; and / or the amino acid residue D at position 471 is substituted with a polar amino acid residue, preferably a polar amino acid with an amide side chain, more preferably N; and / or the amino acid residue G at position 520 is substituted with a polar amino acid residue, preferably a polar acidic amino acid, more preferably E, or a non-polar amino acid residue, preferably a non-polar aliphatic amino acid, more preferably A; and / or the amino acid residue A at position 586 is substituted with a non-polar amino acid residue, or a polar amino acid residue, , more preferably !; and / or the amino acid residue S at position 767 is substituted with a polar amino acid residue, preferably G, or an aliphatic amino acid, preferably a non-bulky aliphatic amino acid, more preferably G or A; and / or the amino acid residue D at position 859 is substituted with a polar amino acid residue, preferably a polar amino acid with an amide side chain, more preferably N, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1 .

3. The polypeptide of claim 1 or 2, wherein the amino acid residue Q at position 744 is invariable, and / or the amino acid residue S at position 430 is invariable, and / or the amino acid residue C atposition 510 is invariable, and / or the amino acid residue F at position 880 is invariable, and / or the amino acid residue F at position 849 is invariable, optionally the amino acid residues at positions 878- 883 are invariable, wherein position numbering is relative to the amino acid sequence set forth in SEQ ID NO:1.

4. The polypeptide of any one of claims 1-3, wherein the one or more amino acid mutations are selected from R57H, R173S, A260S, P451 L D471 N, G520E, A586T, S767G, D859N and combinations thereof.

5. The polypeptide of claim 4, wherein the one or more amino acid mutations consist of: R173S, P451 L, D471 N and G520E; or A586T; or S767G; or A260S; or R57H; or R57H, A260S and D859N.

6. The polypeptide of any one of claims 1-5, wherein the amino acid sequence (ii) and functional fragment (iii), comprise the amino acid sequence corresponding to residues 267-811 of SEQ ID NOT.

7. The polypeptide of any one of claims 1-6, wherein the polypeptide comprises or consists of: the amino acid sequence set forth in SEQ ID NO:2; the amino acid sequence set forth in SEQ ID NO:3; the amino acid sequence set forth in SEQ ID NO:4; the amino acid sequence set forth in SEQ ID NO:5; the amino acid sequence set forth in SEQ ID NO:6; or the amino acid sequence set forth in SEQ ID NOT.

8. A fusion protein comprising: a deaminase domain capable of catalyzing the deamination of nucleic acids; and a polypeptide of any one of claims 1-7; wherein the deaminase domain and the polypeptide are linked to form a single polypeptide chain.

9. The fusion protein of claim 8, wherein the deaminase domain is selected from the group consisting of cytidine deaminases, adenosine deaminases, and guanine deaminases.

10. The fusion protein ofclaim 8, further comprising a linker peptide to link the deaminase domain to the polypeptide, such that the deaminase domain retains its enzymatic activity and the polypeptide retains its transcriptional activity.

11. A composition comprising the polypeptide of any one of claims 1-7 and optionally an in vitro transcription (IVT) reagent, or a fusion protein of any one of claims 8-10.

12. A nucleic acid molecule encoding the polypeptide of any one of claims 1-7 or fusion protein of any one of claims 8-10.

13. The nucleic acid molecule of claim 12, wherein the nucleic acid molecule is comprised in a vector preferably an expression vector, wherein said vector further comprises regulatory elements for controlling expression of said nucleic acid molecule, optionally the nucleic acid molecule is operably linked to a promoter suitable for expression in a host cell.

14. A host cell comprising the nucleic acid molecule of claims 12 or 13.

15. A method for producing the polypeptide of any one of claims 1-7 or fusion protein of any one of claims 8-10, comprising culturing a host cell of claim 14 under conditions that allow expression of the polypeptide or fusion protein, and isolating said polypeptide or fusion protein from the host cell or culture medium.

16. A cell-free method for producing the polypeptide of any one of claims 1 -7 or fusion protein of any one of claims 8-10, comprising subjecting the nucleic acid molecule of claim 12 to reaction conditions that allow the transcription and translation of the polypeptide or fusion protein, and isolating said polypeptide or fusion protein.

17. A method for RNA production or amplification to produce RNA transcript, comprising: contacting a nucleic acid template with the polypeptide of any one of claims 1-7 under conditions that result in the production of RNA transcript.

18. A method of performing an in vitro transcription (IVT) reaction, comprising: contacting a nucleic acid template with the polypeptide of any one of claims 1-7 in the presence of nucleoside triphosphates under conditions that result in the production of RNA transcript.

19. A method for heterologous expression of a protein of interest in a host cell, comprising: i) introducing into the host cell a nucleic acid sequence encoding the polypeptide of any one of claims 1-7, and a heterologous nucleic acid sequence encoding the protein of interest; and ii) inducing expression of the polypeptide in the host cell under suitable conditions, wherein the induced expression of the polypeptide drives transcription of the heterologous nucleic acid sequence and expression of the protein of interest.

20. The method of claim 20, wherein the protein of interest is a recombinant protein selected from the group consisting of a therapeutic protein, an enzyme, and an antibody.

21. The polypeptide of any one of claims 1 -7 for use in RNA production or amplification to produce RNA transcript; or in performing an IVT reaction for producing RNA transcript; or heterologous expression of a protein of interest in a host cell, or detecting biomolecular interactions involving a target analyte; or detecting the presence, and optional quantity, of a target analyte.

22. A method of introducing one or more mutations in a target nucleic acid sequence in a host cell, comprising: introducing into the host cell a nucleic acid molecule encoding the fusion protein of any one of claims 8-10, preferably the nucleic acid molecule is a plasmid vector; and inducing expression of the fusion protein in the host cell to transcribe the target nucleic acid sequence, wherein during transcription the deaminase domain of the fusion protein introduces one or more mutations into the target nucleic acid sequence.

23. The method of claim 21 , wherein the target nucleic acid sequence comprises a reporter gene for detecting the one or more mutations, and the target nucleic acid sequence is operably linked to a T7 promoter.

24. The fusion protein of any one of claims 8-10 for use in introducing one or more mutations in a target nucleic acid sequence in a host cell.

25. A method for detecting biomolecular interactions involving a target analyte, comprising: providing a split version of the polypeptide according to any one of claims 1-7, wherein the split version is linked to a detection system that facilitates reconstitution of the split version into a transcriptionally active polypeptide in the presence of the target analyte; contacting the split version with a sample suspected of containing the target analyte in the presence of a reporter gene construct comprising a reporter gene operably linked to a promoter recognized by the transcriptionally active polypeptide, wherein interaction of the split version with the target analyte induces reconstitution of the split version into the transcriptionally active polypeptide; and detecting the expression of the reporter gene as an indication of the biomolecular interaction involving the target analyte.

26. The method of claim 25, wherein the split version comprises two or more non-functional fragments that reconstitute into the transcriptionally active polypeptide upon interaction with the target analyte, wherein each non-functional fragment is operably linked to a binding or recognition partner specific to the target analyte, and wherein the interaction between the binding or recognition partner and the target analyte facilitates reconstitution of the non-functional fragments into the transcriptionally active polypeptide.

27. The method of claim 26, wherein the two or more non-functional fragments comprise a first non-functional fragment comprising an amino acid sequence corresponding to positions 1-179, 1- 510, 1-563, or 1-601 of SEQ ID NO:1 , and a second non-functional fragment comprising an amino acid sequence corresponding to positions 180-883, 511-883, 564-883, or 602-883 of SEQ ID NO:1.

28. The method of claim 26, wherein the two or more non-functional fragments comprise a first non-functional fragment comprising an amino acid sequence corresponding to positions 1 -67 of SEQ ID NO:1 , a second non-functional fragment comprising an amino acid sequence corresponding to positions 68-179 of SEQ ID NO:1 , a third non-functional fragment comprising an amino acid sequence corresponding to positions 180-601 of SEQ ID NO:1 , and a fourth non-functional fragment comprising an amino acid sequence corresponding to positions 602-883 of SEQ ID NO:1.

29. The method of any one of claims 25-28, wherein the expression of the reporter gene is quantified to determine the amount of the target analyte in the sample.

30. A split version of the polypeptide of any one of claims 1-7, for use in detecting biomolecular interactions.

Citation Information

Patent Citations

  • T7 RNA polymerase mutant which does not generate immunogenic byproducts and has high transcriptional activity and application of T7 RNA polymerase mutant

    CN118374471A

  • RNA polymerase mutant with improved functions

    US20110136181A1

  • T7 RNA polymerase variants

    WO2019005539A1

  • T7 RNA polymerase variants for RNA synthesis

    WO2023031788A1

  • RNA polymerase variant, and preparation method therefor and use thereof in RNA synthesis

    WO2024131998A2