Engineered t7 RNA polymerases for improved RNA production
The EVOLVE-Pro model efficiently evolves T7 RNA polymerases with few-shot active learning, overcoming limitations of existing PLMs by achieving substantial activity enhancements in protein variants for mRNA and circular RNA production.
Patent Information
- Application Number
- PCT/US2025/026131
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-20
- Filing Date
- 2025-04-24
- Publication Date
- 2026-01-22
AI Technical Summary
Existing protein language models (PLMs) struggle to significantly improve protein activity beyond wild-type levels due to limited training data and evolutionary constraints, requiring extensive experimental testing for variant evaluation.
The EVOLVE-Pro model, an ensemble approach combining a foundational protein language model with a top-layer discrimination model, employs few-shot active learning to efficiently nominate high-activity protein variants, such as T7 RNA polymerases, by iteratively selecting and testing mutants in an active learning framework.
EVOLVE-Pro achieves 2- to 515-fold improvements in protein activity with minimal experimental effort, demonstrating enhanced translation efficiency and reduced immunogenicity of T7 RNA polymerases, suitable for mRNA production and circular RNA therapies.
Smart Images

Figure IMGF000019_0001 
Figure IMGF000020_0001 
Figure IMGF000021_0001
Abstract
Description
ENGINEERED T7 RNA POLYMERASES FOR IMPROVED RNA PRODUCTIONCROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial Nos. 63 / 672,398 and 63 / 722,658 filed respectively on July 17, 2024, and November 20, 2024. The content of the above-referenced patent applications is incorporated by reference in its entirety herein.SEQUENCE LISTING
[0002] The instant application contains a Sequence Listing, which has been submitted electronically in XML file format, and is hereby incorporated by reference in its entirety. Said XML copy, created on April 8, 2025, is named 761870_SL.xml and is 111,365 bytes in size.BACKGROUND
[0003] Billions of years of evolutionary pressures have shaped protein diversity, filtering potential design space for diverse biological activities. There is emerging evidence that these sequences embody a fundamental language of biology that can be modeled with deep learning and offer unique insights into the evolutionary processes that have sculpted life on our planet. Protein language models (PLMs) learn the grammar of protein diversity by training to complete masked amino acids in a large protein sequence database, resulting in novel representations of biology. PLMs, with their internal representations of the protein evolutionary landscape, have been used to nominate variants with improved activity with limited success. Generative PLMs, such as ESM3, ProtGPT2, and ProGen, can design novel proteins, but these de novo-designed variants typically only reach wild-type level activity after extensive rounds of experimental testing. This inability of PLMs to drastically improve upon protein activity in zero-shot is partially driven by their inability to generalize to new contexts due to evolutionary constraints as well as limited training data. Active learning methods that utilize context-specific data in combination with deep learning models, including machine learning-directed protein evolution (MLDE) methods, have effectively improved diverse proteins, but typically require extensive efforts to experimentally evaluate many variants. Previous attempts combined protein representation models with active learning to simplify the evolution process, but these approaches have not generalized well beyond proof-of-concept demonstrations like GFP engineering.
[0004] Accordingly, there exists a need for novel protein variants nominated by means of PLM and methods of using same for biological applications.SUMMARY
[0005] In one aspect, the disclosure provides a non-natural T7 RNA polymerase comprising an amino acid sequence with one or more mutations relative to a wild-type T7 RNA polymerase having an amino acid sequence set forth in SEQ ID NO: 1, wherein the one or more mutations are located at one or more amino acids at positions 3, 5, 10, 12, 25, 47, 89, 105, 134, 147, 152, 177, 225, 256, 229, 241, 249, 273, 279, 281, 370, 371, 469, 531, 567, 643, 668, 683, 735, 738, 792, 797, 800, and / or 822 relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
[0006] In certain embodiments, the one or more mutations comprise 3M, 5D, 51, 5M, 10K,12N, 25N, 47A, 89R, 1051, 134T, 147K, 152N, 177L, 225E, 2291, 241W, 249C, 256P, 273H, 279S, 281L, 370V, 370W, 371H, 469Q, 531T, 567R, 643A, 668E, 643G, 643K, 643N, 643R, 643S, 643T, 683K, 735S, 738N, 792H, 797K, 800K, and / or 822S relative to the wildtype T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
[0007] In certain embodiments, the one or more mutations are located at two or more amino acids at positions: 47 and 643; 531, 370, and 643; 738 and 643; 3 and 643; 105 and 643; 469, 47, and 643; 47, 738, and 643; 738, and 643; 3, 105, and 643; 177, 47, and 643; 3, 47, and 643; 668, 47, and 643; 134, 47, and 643; 12, 47, and 643; 822, 47, and 643; 12 and 643; 567, 47, and 643; 668, 3, and 643; or 370, 47, and 643, relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
[0008] In certain embodiments, the one or more mutations comprise: 47A and 643G; 53 IT, 370V, and 643G; 738N and 643G; 3M and 643G; 1051 and 643G; 469Q, 47A, and 643G; 47A, 738N, and 643G; 738N, and 643G; 3M, 1051, and 643G; 177L, 47A, and 643G; 3M, 47A, and 643G; 668E, 47A, and 643G; 134T, 47A, and 643G; 12N, 47A, and 643G; 822S, 47A, and 643G; 12N and 643G; 567R, 47A, and 643G; 668E, 3M, and 643G; or 370W, 47A, and 643 G, relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
[0009] In certain embodiments, the one or more mutations are located at one or more amino acids at positions 370, 567, 643, 738, 797, and / or 800 relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
[0010] In certain embodiments, the amino acid mutation at position 370 is 370V or 370W.
[0011] In certain embodiments, the amino acid mutation at position 567 is 567R.
[0012] In certain embodiments, the amino acid mutation at position 643 is 643A, 643K,643N, 643R, 643S, 643T, or 643G.
[0013] In certain embodiments, the amino acid mutation at position 738 is 738N.
[0014] In certain embodiments, the amino acid mutation at position 797 is 797K.
[0015] In certain embodiments, the amino acid mutation at position 800 is 800K.
[0016] In certain embodiments, the non-natural T7 RNA polymerase amino acid sequence is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 2-61.
[0017] In another aspect, the disclosure provides a nucleic acid encoding a non-natural T7 RNA polymerase.
[0018] In another aspect, the disclosure provides a host cell comprising a non-natural T7 RNA polymerase.
[0019] In another aspect, the disclosure provides a method for producing an RNA sequence in vitro, the method comprising combining a non-natural T7 RNA polymerase, a nucleic acid template comprising a T7 promoter sequence, and NTPs.
[0020] In certain embodiments, the nucleic acid template comprises plasmid DNA, linear DNA, or mRNA.
[0021] In certain embodiments, the method further comprises a circularization step to promote covalently closed circular RNA production.
[0022] In certain embodiments, the produced RNA sequence is less immunogenic relative to a reference RNA sequence produced in vitro comprising a wild-type RNA polymerase comprising a sequence as set forth in SEQ ID NO: 1.
[0023] In certain embodiments, the non-natural T7 RNA polymerase comprises increased activity relative to a wild-type RNA polymerase comprising a sequence as set forth in SEQ ID NO: 1.
[0024] In certain embodiments, the increased activity is increased translation efficiency.BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Aspects, features, benefits, and advantages of the embodiments described herein will be apparent with regard to the following description, appended claims, and accompanying drawings.
[0026] Fig. 1A shows a schematic describing the protein evolution model EVOLution Via Language model-guided Variance Exploration for proteins (EVOLVE-Pro), and methods of using same. With the EVOLVE-Pro method, proteins of interest go through iterative rounds of low-N screening. A foundational PLM generates embeddings for all mutants of a protein and the average embedding by pooling across all residues is used as input for the top layer model. Each mutant’s activity is experimentally determined and used to train a domain expert top layer model with PLM embedding as input. The top layer model then nominates the top- N mutants for the next round of testing and the weights are updated iteratively in an active learning format. Fig.lA discloses SEQ ID NOS 63-66, 64-65, and 63, respectively, in order of appearance. Fig. IB shows the benchmarking of foundational models across a panel of 12 comprehensive deep mutational scanning (DMS) datasets. Each point is a unique protein and its DMS data. ESM2-15B has the highest average percent success in high activity variants prediction. Fig. 1C shows the comparison between EVOLVE-Pro in active learning format, in zero-shot pretraining format, and an existing zero-shot prediction method using protein language model across 12 DMS datasets. Each point is a unique protein using its DMS data. Fig. ID shows the performance over 10 rounds of EVOLVE-Pro with 16 mutants per round, compared to two different non-language model encoding schemes (One-hot encoding and integer encoding). Model performance is benchmarked on four datasets and compared to zero-shot ESM2 nomination success rate and background random sampling. Error bar represents the standard deviation for n=10 random simulations.
[0027] Fig. 2 shows a table summarizing the parameters grid search for EVOLVE-Pro application.
[0028] Fig. 3 shows a table describing 12 DMS datasets for EVOKV-Pro application.
[0029] Fig. 4 shows a summary of parameter grid searches for EVOLVE-Pro with an ESM-2 15B foundational model, showing the random forest regressor combined with the top 10 active learning selection strategy returned the highest average binary top fitness success rate across 12 DMS datasets.
[0030] Fig. 5 shows the Evolve-Pro model characterization using H3N2 hemagglutinin, GPCT, HIV envelope protein, and infA. The performance is evaluated over 10 rounds of EVOLVE-Pro with 16 mutants per round, compared to two different non-language model encoding schemes (One-hot encoding, integer encoding). Model performance is benchmarked on 8 additional datasets and compared to both the zero-shot ESM2 nomination success rate and background random sampling (1). Error bars represent the standard deviations for 10 random simulations.
[0031] Fig. 6A shows a schematic of the strategy for high throughput T7 RNA polymerases mutant testing and evolution policy setup for evolving a high fidelity T7 RNAP. Fig. 6B shows the screening T7 RNAP mutations by IVTT. In vitro transcription / translation (IVTT) coupled kit for production of T7 polymerase under an SP6 promoter was used. Different variants of T7 pol were cloned under the SP6 promoter and then the IVTT kit was used to make functional T7 polymerase in the test tube. Subsequently, the produced T7 pol was used to produce Luciferase mRNA by coincubation with a DNA template containing T7 promoter and Clue. The concentration of the produced Luciferase mRNA was then measured to check for T7 pol yield and transfection into BJ Fibroblast cell line was performed. 24 hours after transfection, the cell media was harvested to check for luciferase translation by the cells and check for IFNB signaling pathway activation by qPCR. These three parameters then were used to assess T7 polymerase variant’s function. Fig. 6C-E show the results from the T7 round 1. The three graphs depict the T7 pol mutant’s yield of mRNA, translation by the BJ fibroblast in terms of the luciferase protein output and Interferon beta gene activation by qPCR. 11 mutants that were randomly designed were compared with the wild type T7 polymerase side by side for all three parameters. Fig. 6F-H show the engineering of T7 pol. The model utilized the data generated in round 1 to predict for better mutants in round2 and round 3. The performance of each mutants is showed by projecting their parameters onto 2d graph where x-axis shows immunogenicity as measured by IFNB gene activation and y-axis shows increase in luciferase protein translation. Fig. 61 shows that the position 643 is key to reduce immunogenicity and improved translation. In round 4 prediction by the model, multiple mutations at position 643 in the T7 pol is nominated. In particular mutating position 643 to G yield an ultra-high fidelity T7 pol mutants that has 0 activation of IFNB in B J fibroblast cell line and more than 20 fold boost in luciferase protein translation. Fig. 6J shows the engineering of T7 RNAP over six rounds of EVOLVE-Pro. Data shows the top 10 mutants from current and preceding rounds, as measured by fold improvement oftranscription fidelity over wild-type. Fig. 6K shows the performance of T7 mutants from six EVOLVE-Pro rounds and previously engineered G47A / 884G SOTA T7 RNAP in Clue mRNA translation and immunogenicity in BJ Fibroblast cells. Fig. 6L shows the validation of epT7 for production of 6 mRNA sequences ranging from 513 nt to 6496nt. Purified WT or mutant RNAP is used to produce these sequences, and they were transfected into BJ fibroblast cells for either protein translation readout or targeted IFNB 1 gene expression analysis using qPCR 24 hours after transfection. A two-sided Student’s t-test was run between WT and each evolved T7 RNAP (**, p<0.01, ***, p<0.001, ****, p<0.0001). Error bars represent standard deviation with n=3 biological replicates. Fig. 6M shows dsRNA ELISA used to analyze the amount of dsRNA during transcription of a 1662 nt Cypridina luciferase mRNA. 500 ng of post-transcription product is used as input for the dsRNA ELISA. A two-sided Student’s t-test was run between WT and each evolved T7 RNAP (**** p<0.0001). Error bars represent standard deviation with n=3 biological replicates. Fig. 6N shows the mapping of the top mutations on the T7 RNAP structure (PDB: 3E2E). The active site is indicated by a red circle. Fig. 60 shows an heatmap showing most common T7 RNAP mutations explored by EVOLVE-Pro over rounds of evolution. Any position explored more than once is shown on a cumulative basis across rounds. Fig. 6P shows a scatter plot comparing the predicted ESM-2 protein fitness score versus experimentally measured T7 RNAP transcription fidelity scaled score across evolution rounds. The correlation and linear regression line are shown in the plot. Fig. 6Q and Fig. 6R show a comparison of the T7 RNAP latent space with either predicted ESM-2 protein fitness (masked marginal score) or EVOLVE-Pro protein activity fold improvement. Fig. 6S shows a kernel density estimate of protein fitness as predicted by ESM-2 versus protein function as predicted by EVOLVE-Pro. The correlation and linear regression line are shown in red and the R square of correlation is reported.
[0032] Fig. 7A shows a schematic of circular RNA production. Fig. 7B shows the validation of epT7 produced circRNA on four different template sequences compared to both T7E643G and wild-type T7. Translation of each protein is measured in HEK293FT cells 48 hours after transfection. A two-sided Student’s t-test was run between WT and each evolved T7 RNAP (***, p<0.001, ****, p<0.0001). Error bars represent standard deviation with n=3 biological replicates. Fig. 7C shows the comparison of RNA quality for nanoluc and eGFP circRNA produced by epT7 compared to wild-type T7 via gel electrophoresis at different steps in the production process: post- initial IVT and post-RNaseR processing. Fig. 7D shows thecomparison of dsRNA content for nanoluc circRNA produced by epT7 compared to wildtype T7 using either 2 hours of IVT or 12 hours of IVT. 500 ng post-RNAseR cleaned-up samples are used as input for dsRNA ELISA. A two-sided Student’s t-test was run between WT and each evolved T7 RNAP (**, p<0.01). Error bars represent standard deviation with n=3 biological replicates. Fig. 7E shows the TapeStation RNA integrity analysis of either epT7 or WT-produced circular nanoluc RNA. epT7 shows reduced concatemer production. Fig. 7F shows the comparison of purified nanoluc circRNA yield by epT7 compared to wildtype T7 after the initial RNaseR clean-up. The panel on the left shows the raw mass percentage left after the cleanup. The panel on the right shows the purity of the circular RNA in the post clean-up reaction as determined by quantification using a TapeStation analysis. A two-sided Student’s t-test was run between WT and epT7 (**, p<0.01, ****, p<0.0001). Error bars represent standard deviation with n=3 biological replicates. Fig. 7G shows a schematic of the in vivo mRNA assay for measuring mRNA expression in the liver via non- invasive luminescent imaging. Fig. 7H shows the in vivo luminescent signal detected 24 hours post-injection in mice injected with mRNA produced by either epT7 or wild-type T7 or PBS controls. A two-sided Student’s t-test was run between WT, wild-type T7 RNAP, and epT7 (*, p<0.05). Error bars represent standard deviation with n=3 biological replicates. Fig. 71 shows the time-course of in vivo luminescent signal detected up to 96 hours post-injection of LNP-mRNA produced by either epT7 or wild-type T7, or PBS controls. A two-sided paired Student’s t-test was run between WT, wild-type T7 RNAP, and epT7 (*, p<0.05) for each time point. Error bars represent the standard error of mean with n=3 biological replicates. Fig. 7J shows a schematic showing the evolution of higher activity variants with EVOLVE-Pro. The mutagenesis landscape of proteins is often conceptualized as a complex terrain with numerous potential paths. Shown here is a gray road that conceptualizes the protein mutagenesis landscape where traversing upwards results in higher protein function and traversing downwards reduces protein fitness. Traditional frameworks of evolutionary plausibility attempt to navigate this terrain based on natural selection, which is constrained by historical and environmental factors.
[0033] Fig. 8A shows an individual mutant’s fold improvement in transcription fidelity in producing Cypridina luciferase (Clue) mRNA across 6 rounds of evolution. Error bars represent standard deviation of three biological replicates. Fig. 8B shows a comparison of each round’s top-performing mutant with G47A / 884G (previous SOTA) T7 RNAP in their production of Clue mRNA. Two metrics are evaluated by assaying the protein translation andIFNB1 gene expression in BJ fibroblasts. Fold improvement is scaled to WT. Fig. 8C shows a comparison of the yield of mRNA across 4 different template sequences between WT, G47A / 884G, E643G, and epT7. Error bars represent standard deviation of three technical replicates. Fig. 8D shows a gel electrophoresis analysis of 3 different mRNAs produced by different T7 RNAP variants: WT, G47A / 884G, E643G, and epT7. Fig. 8E shows a TapeStation analysis of Clue mRNA produced by different T7 RNAP variants: WT, G47A / 884G, E643G, and epT7. Fig. 8F shows a TapeStation analysis of prime editor mRNA produced by different T7 RNAP variants: WT, G47A / 884G, E643G, and epT7.
[0034] Fig. 9A shows a translation and immunogenicity comparison of the previously engineered G47A / 884G SOTA T7 RNAP with WT, E643G, and epT7 across 4 different template sequences in BJ Fibroblast cells. mRNAs were transfected into BJ fibroblast cells for either protein translation readout or targeted IFNB 1 gene expression analysis using qPCR 24 hours after transfection. Error bars represent standard deviation of three biological replicates. Fig. 9B shows a comparison of editing activity of Cas9 mRNA produced by WT, G47A / 884G, E643G, and epT7 on endogenous ENO1 genomic loci in HEK293FT cells and BJ Fibroblast cells. Indel rates are quantified 48 hours after transfection of Cas9 mRNA and synthetic guide RNA targeting ENO1. Error bars represent standard deviation of three biological replicates.
[0035] Fig. 10A shows a validation of epT7 produced circRNA sequences compared to G47A / 884G, T7E643G, and wild-type T7 on Nanoluc encoded circRNAs, Fig. 10B shows a validation of epT7 produced circRNA sequences compared to G47A / 884G, T7E643G, and wild-type T7 on Glue encoded circRNAs, Fig. 10C shows a validation of epT7 produced circRNA sequences compared to G47A / 884G, T7E643G, and wild-type T7 on EGFP encoded circRNAs, and Fig. 10D shows a validation of epT7 produced circRNA sequences compared to G47A / 884G, T7E643G, and wild-type T7 on Clue encoded circRNAs. Translation of each protein is measured in HEK293FT cells 48 hours after transfection. Fig. 10E and Fig. 10F show representative fluorescence images for circular GFP RNA produced by wild-type T7 RNAP or epT7 at 24 hours (Fig. 10E) and 72 hours (Fig. 10F) posttransfection in HEK293 FT cells. Scale bar, 100 pm. Fig. 10G shows a TapeStation analysis of pre and post-RNaseR cleaned-up nanoluc circular RNA produced by either wild-type T7 or epT7. Fig. 10H shows time course kinetics of firefly luciferase luminescence (log scale) at 4 different time points after injection of LNP encapsulated mRNA made by either wild-type T7 or epT7 compared to background PBS control.DETAILED DESCRIPTION
[0036] It will be appreciated that for clarity, the following discussion will describe various aspects of embodiments of the applicant’s teachings. It should be noted that the specific embodiments are not intended as an exhaustive description or as a limitation to the broader aspects discussed herein. One aspect described in conjunction with a particular embodiment is not necessarily limited to that embodiment and can be practiced with any other embodiment(s). Reference throughout this specification to “one embodiment,” “an embodiment,” “an example embodiment,” means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrases “in one embodiment,” “in an embodiment,” or “an example embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular feature, structures or characteristics may be combined in any suitable manner, as would be apparent to a person skilled in the art from this disclosure, in one or more embodiments. Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the disclosure. For example, in the appended claims, any of the claimed embodiments can be used in any combination.General Definitions
[0037] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Definitions of common terms and techniques in molecular biology may be found in Molecular Cloning: A Laboratory Manual, 2nd edition (1989) (Sambrook, Fritsch, and Maniatis); Molecular Cloning: A Laboratory Manual, 4th edition (2012) (Green and Sambrook); Current Protocols in Molecular Biology (1987) (F.M. Ausubel et al. eds.); the series Methods in Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (1995) (M.J. MacPherson, B.D. Hames, and G.R. Taylor eds.): Antibodies, A Laboratory Manual (1988) (Harlow and Lane, eds.): Antibodies A Laboratory Manual, 2nd edition 2013 (E.A. Greenfield ed.); Animal Cell Culture (1987) (R.I. Freshney, ed.); Benjamin Lewin, Genes IX, published by Jones and Bartlet, 2008 (ISBN 0763752223); Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994 (ISBN0632021829); Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 9780471185710); Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, N.Y. 1994), March, Advanced Organic Chemistry Reactions, Mechanisms and Structure 4th ed., John Wiley & Sons (New York, N.Y. 1992); and Marten H. Hofker and Jan van Deursen, Transgenic Mouse Methods and Protocols, 2nd edition (2011).
[0038] As used herein, the singular forms “a,” “an,” and “the” include both singular and plural forms unless the context clearly dictates otherwise. Thus, for example, reference to “a cell” includes a plurality of such cells.
[0039] As used herein, the term “optional” or “optionally” means that the subsequent described event, circumstance or substituent may or may not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0040] The recitation of numerical ranges by endpoints includes all numbers and fractions subsumed within the respective ranges, as well as the recited endpoints.
[0041] As used herein, the term “about” or “approximately” refers to a measurable value such as a parameter, an amount, a temporal duration, and the like, are meant to encompass variations of and from the specified value, such as variations of + / -10% or less, + / -5% or less, + / -1% or less, + / -0.5% or less, and + / -0.1% or less of and from the specified value, insofar such variations are appropriate to perform in the disclosure. It is to be understood that the value to which the modifier “about” or “approximately” refers is itself also specifically disclosed.
[0001] It is noted that all publications and references cited herein are expressly incorporated herein by reference in their entirety. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present disclosure is not entitled to antedate such publication. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.Evolve-pro for Polymerase Enhancement
[0042] The EVOLVE-Pro model is used herein to enhance polymerases, such as T7 RNA polymerases. EVOLVE-Pro is a frontier multi-modal protein design model, which evolves high-activity protein variants with few-shot learning and minimal experimental testing to achieve accurate predictions of sequence-to-function for general properties. This performance comes from an ensemble approach, combining evolutionary-scale protein foundation models with a top-layer discrimination model to learn a protein’s functional landscape and guide the directed evolution process in silico. By applying EVOLVE-Pro in a few-shot, active learning framework, protein sequences with significantly higher activity can be efficiently nominated in a generalizable fashion with minimal effort. The modularity of the EVOLVE-Pro architecture allows this framework to scale with larger parameter PLMs. Moreover, EVOLVE-Pro prompting only requires protein sequences to be evolved without any structural information, expert knowledge, or prior data. As EVOLVE-Pro is multi-modal, multiple protein features of any type or data class can be simultaneously engineered, opening up vast possibilities for its use in biology and medicine.
[0043] EVOLVE-Pro was benchmarked in silico across a panel of 12 different proteins, showing state-of-the-art performance, and then apply the final model for mRNA manufacturing with a T7 RNA polymerase. EVOLVE-Pro yields mutants with 2- to 515-fold improvement over initial proteins. Improvement in vivo mRNA performance generated from an EVOLVE-Pro evolved T7 polymerase was observed. Analyzing nominated mutations, EVOLVE-Pro explores disparate sites, and the learned functional landscape is separate, and often negatively correlated with, the fitness inferred by the underlying protein language model. Lastly, EVOLVE-Pro’s utility is showcased in nominating multi-mutant protein designs out of a vast sequence space that enables final mutants that are much more active than naturally observed proteins. EVOLVE-Pro establishes the capabilities of few-shot active learning with protein language models for optimizing proteins for diverse activities.Development and Benchmarking of the EVOLVE-Pro Model
[0044] An ensemble model was designed to establish EVOLVE-Pro (Fig. 1A-D). The ensemble model involves: 1) a foundational protein language model to encode protein sequences into an information-rich latent space, and 2) a top-layer discrimination model to learn protein functional grammar in this evolutionary landscape and rank protein sequencesaccording to a designed policy framework, and 3) an active learning framework using top layer discrimination model to nominate the next set of protein variants for experimental evaluation. This cycle is performed iteratively to evolve defined protein activities until they reach desired levels (Fig. 1A).
[0045] EVOLVE-Pro was optimized across five parameters: 1) the strategy employed for the first round mutant selection, 2) the top layer discrimination model that learns the fitness landscape, 3) the active learning strategy for selecting mutants for the next round, 4) the evolution policy, and 5) the embedding vector transformation (Fig. 2). To perform a grid search across this space, a panel of twelve unique deep mutagenesis scanning (DMS) datasets for in silico validation was curated (Fig. 3). These twelve proteins represent diverse functions, including viral spike proteins, RNA-guided nucleases, lactases, and kinases, ensuring that the resulting model will be as generalizable as possible for learning diverse protein activity landscapes in PLM latent space.
[0046] ESM-2 protein language model was first assessed because of its large training data and available model size of >200M proteins and 15B parameters, respectively. Using the ESM-2 15B parameter model, the grid search found the optimal strategy was: 1) selecting a random set of first-round variants, 2) employing a random forest regressor discriminatory model to predict protein function, 3) using residue pooled average embeddings, and 4) using a top-N selection strategy in each round of evolution (Fig. 4). This policy nominated a high frequency of gain-of-function protein variants in only 5 rounds (Fig. 4A. Since the focus was on the percent of activity passing a threshold as the evaluation metric in the grid search, increasing function during in silico evolution was next evaluated. Both the median activity and the activity of the nominated top mutant were found to increase monotonically from round to round across all DMS datasets, further validating the model’s performance in this low-N active learning setting.
[0047] In general, 16 mutants per round of evolution for 10 rounds identified top mutants with fitness in the 50th percentile for eleven of the twelve DMS datasets. To understand how the number of variants per round affected performance, between 10 and 100 variants per round were tested, finding that larger rounds increased prediction accuracy. This performance trade off indicates that EVOLVE-Pro can be used for both extremely low-N evolution (<20 mutants per round) for rapid and cheap experimental characterization and medium-N (-100 mutants per round) for quicker and more efficient evolution with fewer rounds.
[0048] After optimizing the top layer model and learning strategies, the PLM was optimized, comparing ESM-2 15B to a panel of foundational models. Using the optimal parameters from the grid search, performance was benchmarked against smaller versions of ESM-2 and ESM- 1, UniRep, ProtT5, ProteinBERT, Ankh, one-hot encoding, and integer encoded protein representations. ESM-2 15B parameter model outperformed all the other models for identifying the highest fitness proteins for all datasets except two, confirming its final selection for the EVOLVE-Pro latent space model (Fig. IB). Importantly, large parameter PLMs showed a significant boost in prediction accuracy compared to non-language modelbased architectures, indicative of the powerful feature extraction present in transformer-based models (Fig. IB).
[0049] EVOLVE-Pro ’s performance was then benchmarked relative to other PLM -based engineering approaches. As many methods require pre-training a discriminatory model on thousands of variants, tested versions of EVOLVE-Pro augmented with various amounts of pre-training were tested (Fig. 1C). Reinforcement learning drastically reduced the overall number of mutants required: EVOLVE-Pro with only 5 rounds of evolution (16 mutants per round) was equivalent in performance to EVOLVE-Pro pre -trained with 160 mutants, while 10 rounds of evolution (16 mutants per round) was equivalent to pre-training with 500 mutants. Moreover, EVOLVE-Pro significantly outperformed zero-shot prediction methods. This comparison confirms that the few-shot nature of EVOLVE-Pro allows for efficient directed evolution with minimal effort and low-N testing per round (Fig. 1C).
[0050] Lastly, the per-round evolution improvement for EVOLVE-Pro compared to one-hot and integer encoding and zero-shot prediction were evaluated, finding that by round 5 variants with significantly enhanced fitness could universally be found (at 16 mutations per round) (Fig. ID and Fig. 5). Moreover, in many cases, the one-hot and integer encoding frameworks saturated much earlier in the evolution process and never reached the fitness levels achieved by EVOLVE-Pro. Interestingly, for some proteins a non-linear increase in protein fitness after round 3 was observed, suggesting greater gains in mapping the protein fitness landscape as EVOLVE-Pro evolution proceeds.Evolving T7 RNA polymerase for highly pure and efficient RNA production
[0051] T7 RNA polymerase mutants were generated using the large language program EVOLVE-Pro and evaluated (Fig. 6A-S, Fig. 7A-J, Fig. 8A-F, Fig. 9A-B, and Fig. 10A-H).a. RNA production via T7 RNA polymerase
[0052] T7 RNA polymerase (RNAP) plays a critical role in RNA production for mRNA therapies, mRNA vaccines, cell engineering, and basic scientific studies. As mRNA production has numerous features characterizing its potency and quality, a multi-objective optimization function was designed to evolve a high-fidelity T7 RNAP for mRNA production with these three parameters: 1) RNA yield measured via UV-vis spectrophotometry, 2) mRNA translation in a dsRNA sensitive cell line measured via luciferase translation, and 3) RNA purity measured via immunogenicity in BJ fibroblast cells by IFN-beta RNA production (Fig. 6A). These features were weighted in the EVOLVE-Pro objective function by 20%, 40%, and 40%, respectively to prioritize the higher fidelity and lower immunogenicity aspects of this enzyme for clinical applications. To facilitate high throughput variant testing, SP6 in vitro transcription-translation coupled reaction kits was relied on to generate mutant T7 RNAP in a one-pot reaction and subsequently use the produced T7 RNAP to produce co-transcriptionally capped Cypridina luciferase mRNA for downstream in vitro testing.
[0053] During the initial two rounds of evolution, improvements were observed but were only in the 2-4 fold improvement range. However, by rounds 3 and 4, significant improvements started to be observed in all features, especially in translation and immunogenicity fold changes over the wild-type T7 RNAP (Fig. 6J, Fig. 6K, and Fig. 8A). By the end of round 4, one T7 RNAP mutant was observed, E643G, that could generate luciferase mRNA that produced 34x more translated luciferase and -98% less immunogenicity (Fig. 6K). E643G was benchmark against the previously engineered state of the art G47A / 884G mutant T7 RNAP that has markedly reduced immunogenic byproduct in our IVTT assay. The E643G mutant was found to produce 7-fold higher translation in cells and ~2-fold less INFB inflammation in BJ fibroblasts (Fig. 8B).
[0054] Given the plethora of mutants tested in the first 4 rounds along with the existing G47A mutation that is known to reduce dsRNA formation, EVOLVE-Pro was then used to generate multi-mutants that involved the combination of up to 7 previously tested mutations. Typically, in rational mutagenesis for higher activity mutants, rational combinations of single beneficial mutations are combined according to their spatial location under the assumption of synergistic effects of these mutations. Here, EVOLVE-Pro’s ability was relied on to learn the activity landscape to nominate multi-mutants in an unbiased fashion. Surprisingly, the top nominated mutant in Round 5 was found to be a combination of previously reported G47Aand the best single mutant E643G as normal rational mutagenesis would do. However, it is worth noting that this combination resulted in a protein with worse performance than E643G alone (Fig. 8B). This points to the vast unknown epistatic effect between residues at different spatial positions in a protein and the utility of EVOLVE-Pro in nominating multi-mutants. By round 6, EVOLVE-Pro was able to nominate one particular variant that had ~57x more translation from luciferase mRNA and ~515x less immunogenicity than the original wildtype T7 RNAP (Fig. 6K). Moreover, this variant was substantially more effective at translation and less immunogenic than the G47A / 884G mutant. This multi-mutant, T7 RNAPT3M / G47A / E643G, was chosen as the final EVOLVE-Pro evolution candidate and termed enhanced T7 RNAP, or epT7.
[0055] Given the high throughput testing of mutant T7 RNAPs in the IVTT reaction, it was hypothesized that the unoptimized IVT buffer could change these mutant’s mRNA production and the performance of top mutants in clinically relevant IVT settings with NEB’s HiScribe transcription kit, followed by Vaccinia cap-1 capping and polyA tailing were compared. The top performing single mutant (E643G), previously reported state-of-the-art mutant (G47A / 884G), and our epT7 (T3M / G47A / E643G) along with WT were therefore purified to compare their performance. The production of six different mRNA sequences, ranging in size from 500 nt to 6,500 nt, between epT7, T71 64'*', and wild-type T7 were compared. Consistent with the IVTT-based experiment, it was found that epT7 and E643G produced significantly higher mRNA in a 2-hour transcription scheme than both wild-type T7 RNAP and the G47A / 884G variant (Fig. 8C). Analysis of the 3 different mRNA products by both gel-electrophoresis and TapeStation confirmed the presence of a single on-target product across all four enzymes (Fig. 8D, Fig. 8E, and Fig. 8F). Looking at the translation and immunogenicity aspects of the mRNAs produced by these enzymes, it was found that in all cases epT7 produced mRNA had 4 to 120 fold higher translation than wild type and 4 to 256 fold lower immunogenicity (Fig. 6L, Fig. 9A). Functional testing of SpCas9 mRNA also shows significantly higher editing from epT7’s produced mRNA in two separate cell lines (Fig. 9B). These results validate that the EVOLVE-Pro derived epT7 mutants are not buffer or template-specific and are genuinely improving the quality of mRNA produced by the polymerase. The mechanism of the epT7 performance enhancements was investigated by investigating the quality of the RNA. Using an established ELISA for dsRNA, it was found that the dsRNA in the ep T7 -produced mRNA was 5 -fold lower than wild-type T7-producedRNA and it performed equally well as the RNA produced by the state-of-the-art G47A / 884G mutant (Fig. 6M).
[0056] Previous efforts to reduce dsRNA production relied on adding a glycine residue at the C -terminal “foot” region of the enzyme (884G insertion). The instant model revealed the functional importance of E643 in transcription and, surprisingly, mutating this residue rendered the same effect as 884G (Fig. 6N, Fig. 8B). Indeed, analysis of the T7 RNAP structure reveals that E643 is close to the DNA template, suggesting that E643G improves template binding and RNA production (Fig. 6N). However, E643K / E643R did not improve the fidelity of transcription (Fig. 8A), suggesting that these bulky residues sterically clash with the template DNA. Therefore, it is likely that EVOLVE-Pro is able to identify a unique mechanism, interrogate the effect of this mechanism, and determine the right balance biochemically to mutagenize, thereby producing a novel, SOTA T7 RNAP variant that has never been described before. G47A has been previously reported to increase helix formation, and EVOLVE-Pro took advantage of this helix-favoring mutation in our multi-mutant generation. The third mutated residue in epT7 is in a disordered region (T3M), suggesting a role independent of DNA template binding. T3M might be involved in improving protein stability or other aspects that can modulate the polymerase function. These residues further highlight the insightfulness of EVOLVE-Pro to identify novel mutants that one would not test via rational mutagenesis. Further analysis of EVOLVE-Pro’s residue exploration in evolution revealed that E643 was found first in round 3 with the most beneficial mutation being E643N (Fig. 60). The model quickly gained an understanding of the functional importance of this residue and zoomed into this region by exploring it 5 more times in round 4, yielding E643G the best single mutant (Fig. 60). This trend is similar to the evolution of proteins reported above, where a beneficial mutation at a certain residue is capitalized by the model in the next round by exploring additional mutations around that region.
[0057] The relationship between the function (observed data) and fitness (pMMS) for T7 RNAP was then calculated and a negative correlation of 0.13 in this case was found, denoting the lack of association between the two metrics. EVOLVE-Pro successfully navigated through this divergence by selecting mutants with higher activity but not fitness in later rounds (Fig. 6P). Lastly, the global evolutionary landscape of epT7 and EVOLVE-Pro’s mutational trajectory were investigated. At a high level, as with the previous proteins evolved, the activity map learned by EVOLVE-Pro diverged from the fitness map predicted by ESM-2, showing that fitness predictions would not be able to predict the mutants that wereultimately discovered to improve protein activity and other parameters (Fig. 6Q, Fig. 6R, and Fig. 6S). b. Circular RNA production with epT7
[0058] Circular RNA has emerged as a promising therapeutic modality for protein replacement therapy thanks to its enhanced stability and prolonged expression of proteins. Since significantly lower dsRNA production and higher fidelity of transcription with epT7 was observed, we hypothesized that epT7 would enhance circular RNA production since the use of RNAse R during post-IVT processing typically enriches for both circular RNA and dsRNA species that are immunogenic (Fig. 7A). EpT7 was thus applied to the circularization of four different RNA sequences, finding that the translation obtained by circRNA from epT7 is 3 to 30 fold higher than RNA produced by WT T7 RNAP (Fig. 7B and Fig. 10A-D). To better understand the mechanism behind this enhanced translation, gel electrophoresis was performed of both pre and post-RNaseR products to check for the integrity of the IVT (Fig. 7C). Reduced byproducts were noticed in both eGFP-circRNA and nanoLuc-circRNA, showing higher fidelity of transcription by epT7. To confirm the higher stability of circular- eGFP RNA, both wild-type T7 and epT7’s produced circRNA were transfected in HEK293FT cells and imaged 24 hours and 72 hours post -transfection (Fig. 10E-F). Higher GFP fluorescence was observed from epT7 than wild-type T7 RNAP and stable expression of GFP was observed at 72 hours similar to previously reported. DsRNA ELISA was then used to detect the amount of dsRNA left in the product after RNAse R cleanup. Consistent with the hypothesis, there is a large increase in dsRNA percentage at around 1.5% from WT T7’s produced dsRNA (Fig. 7D). This dsRNA ratio is significantly reduced to 0.2% using epT7, highlighting the fidelity of epT7 during long transcription that is needed to accommodate circular RNA production (Fig. 7D). TapeStation was used to quantify the ratio of circular RNA in the original IVT products and significantly higher circular RNA production at around 13% efficiency was found, which was ~2.4 fold higher than the efficiency of WT T7 RNAP, higher circRNA purity, and lower concatemer production (Fig. 7E-F, Fig. 10G). c. mRNA for in vivo bioluminescent imaging
[0059] Given the high fidelity of epT7, the performance of epT7 was compared with WT T7 RNAP in producing 100% N1-Mcthylpsciidoiiridiric-5'-Triphosphatc-inodil'icd firefly luciferase mRNA that is commonly used for in vivo deep tissue imaging (Fig. 7G). This production process, including the modified bases, mimics the clinical production oftherapeutic mRNAs, allowing for a translationally relevant evaluation of epT7’s mRNA production. The produced mRNA was packaged with lipid nanoparticles (LNPs) that traffic to the liver for bio luminescent imaging. After 24 hours post-injection of the LNP formulations, ~ 10-fold higher luminescence was observed for the epT7 -produced mRNA compared to mRNA produced by WT T7 RNAP (Fig. 7H). Moreover, the kinetics of both mRNA formulations was tracked for 96 hours and consistently higher translation was found with the epT7-produced Flue mRNA for a longer period of time (Fig. 71, Fig. 10H).Sequences Table 1 : T7 RNA polymerase wild type and mutants.EXAMPLESExample 1. Measurement of luciferase activity
[0060] Media containing secreted or intracellular luciferase was harvested 48 hours after transfection unless otherwise noted. 20pL of media is used to measure secreted luciferase activity using Targeting Systems Cypridinia and Targeting systems Gaussia luciferase assay kits (Targeting Systems) on a Biotek Synergy 4 plate reader with an injection protocol. All replicates were performed as biological replicates. Intracellular Nanoluc and firefly luciferase were measured by lysing the cell in the luciferase assay mix (Promega) according to themanufacturer’s protocol. 5 minutes after lysis at room temperature, the signal is read out using a Biotek Synergy 4 plate reader.Example 2. Quantification of protein expression
[0061] Two days after the transfection of HEK293FT or BJ Fibroblast cells, the Nano-Gio HiBiT Lytic Detection System (Promega) was used for the quantification of the HiBiT tags, in cell lysates. For the preparation of the Nano-Gio HiBiT Lytic Reagent, the Nano-Gio HiBit Lytic Buffer (Promega) was mixed with Nano-Gio HiBiT Lytic Substrate (Promega) and the LgBiT Protein (Promega) according to the manufacturer’s protocol. The volume of Nano-Gio HiBiT Lytic Reagent added was equal to the culture medium present in each well, and the samples were placed on an orbital shaker at 600 rpm for 3 minutes. After incubation of 10 minutes at room temperature, the readout took place with 125 gain and 2 seconds integration time using a plate reader (Biotek Synergy Neo 2). The control background was subtracted from the final measurements.Example 3. Harvest of total RNA and quantitative PCR
[0062] For gene expression experiments in mammalian cells, cell harvesting and reverse transcription for cDNA generation were performed using a previously described modification of the commercial Cells-to-Ct kit (Thermo Fisher Scientific) 48 h after transfection. Transcript expression was then quantified with qPCR using Fast Advanced Master Mix (Thermo Fisher Scientific) and TaqMan qPCR probes (Thermo Fisher Scientific) with GAPDH control probes (Thermo Fisher Scientific). All qPCR reactions were performed in 10-j.il reactions with two technical replicates in a 384-well format and read out using a LightCycler 480 Instrument II (Roche). For multiplexed targeting reactions, readout of different targets was performed in separate wells. Expression levels were calculated by subtracting housekeeping control (GAPDH) cycle threshold (Ct) values from target Ct values to normalize for total input, resulting in ACt levels. Relative transcript abundance was computed as 2-ACt. All replicates were performed as biological replicates.Example 4. T7 RNA polymerase purification
[0063] Plasmids for overexpressing Twin-Strep-tagged SUMO (Small Ubiquitin-like Modifier)-fused WT or mutant T7 polymerase (pET-6xHis-thrombin-Twin-Strep-tag-SUMO- T7) (“6xHis” disclosed as SEQ ID NO: 62) were transformed into BL21 T7 expression E. coli strain (NEB, C2566H). The transformed cells were inoculated into 1.2 liters of Terrific Broth (TB) with 100 pg / ml ampicillin using 12 ml of an overnight culture of T7 Express cells containing the T7 polymerase expression construct. Cultures were grown at 37 °C until the cell density reached OD600 ~0.6, then protein overexpression was induced by adding 0.2 mM isopropyl (3-D-thiogalactoside (IPTG) and incubating for 24 hours at 16 °C. Cells were harvested by centrifugation at 4,000 g for 15 minutes and stored at -80 °C until purification.
[0064] The cell pellet was resuspended in lysis buffer (50 mM Tris-HCl, pH 8.0, 500 mM NaCl, 1 mM DTT) containing Protease Inhibitor (Roche Complete ULTRA, EDTA-free), lysozyme (Thermo Fisher Scientific), and Benzonase (Millipore) to degrade nucleic acids after lysis. Cells were lysed using an ultrasonic homogenizer under ice-cooling, followed by clarification through 60 minutes of centrifugation at 10,000 g. The lysate was incubated with Strep-Tactin®XT resins (iba) at 4 °C for 1 hour with orbital shaking, then applied to a gravity flow chromatography column equilibrated with wash buffer (50 mM Tris-HCl, pH 8.0, 500 mM NaCl, 1 mM DTT, Protease Inhibitor). After washing, the protein was eluted with the same cleavage buffer (50 mM Tris-HCl, pH 8.0, 500 mM NaCl, 0.1% Triton® X-100, 1 mM DTT, SUMO protease) following overnight cleavage of the SUMO-fusion by SUMO protease (Sigma- Aldrich) at 4 °C. Proteins were concentrated to 500 ml using Amicon® Ultra Centrifugal Filters (30 kDa molecular-mass cut-off, Millipore) before loading onto a gel filtration column (Superdex® 200 Increase 10 / 300 GL) via FPLC (AKTA Pure). Gel filtration fractions were analyzed by SDS-PAGE, and those containing T7 polymerase were pooled.
[0065] Proteins were quantified using the CBQCA Protein Quantitation Kit (Thermo Fisher Scientific) per the manufacturer’s instructions, buffer exchanged into storage buffer (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, 1 mM DTT, 0.1 mM EDTA, 50% Glycerol, 0.1% Triton® X-100), and stored at -20 °C until use.Example 5. Linear mRNA production
[0066] IVT templates for linear mRNA synthesis were prepared either by PCR amplification for 35 cycles or by linearizing plasmid DNA with Pmel (NEB, R0560L) overnight. Theresulting products were purified using silica columns (Qiagen, 28006 for PCR products or 28115 for plasmid digestion products) before RNA synthesis. Linear mRNA was synthesized in vitro at 37°C for 2 hours using 1 pM purified T7 polymerase with either 1000 ng of linearized plasmid template or 500 ng of PCR-amplified DNA templates per 20 pL of IVT reaction. Equimolar concentrations of NTPs (NEB, N0466L) were used, and the reactions for wild-type and mutant T7 polymerase were performed under identical conditions. For in vivo tests, mRNA was synthesized in vitro by T7 RNAP -mediated transcription at 37°C for 4 hours using 100% substituted N1 -methylpseudouridine -triphosphate (TriLink, N-1019) with either 1 pM purified epT7 polymerase or 2 pL of commercially available WT T7 RNA Polymerase (NEB, M0251S) per 20 pL following the manufacturer’s instructions. DNase I (NEB, M0303L) was used to remove the DNA template, terminating transcription. The mRNA was then purified using a silica column (NEB, T2040L) and quantified using a Nanodrop One Microvolume UV spectrophotometer (Thermo Fisher, ND-ONE-W) before further enzymatic reactions. The Cap 1 structure was added to the 5' end using Vaccinia capping enzyme (NEB, M2080) and mRNA Cap 2'-O-Methyltransferase (NEB, M0366). The mRNA was purified again using a silica column (NEB, T2040L), followed by the addition of AMP from ATP to the 3' end using E. coli Poly(A) Polymerase (NEB, M0276L). All mRNAs were column purified (NEB, T2040L) and eluted with 1 mM sodium citrate (Thermo Fisher, AM7001).Example 6. Circular RNA production
[0067] The construction of the plasmid template for circular mRNA synthesis has been previously described (Chen et al., 2022). The plasmid was digested with Notl restriction enzyme (Thermo Fisher, FD0593), and the resulting DNA product, serving as the transcriptional template for circular RNA, was column purified using a MinElute PCR Purification Kit (Qiagen, 28006). For each 20 pl IVT reaction, 500 ng of the purified transcriptional template was used. The IVT reactions were incubated at 37°C for 12 hours, followed by degradation of the DNA template with 2 pl of DNase I per 500 ng of transcriptional template for 30 minutes at 37°C. The remaining RNA was column purified before further enzymatic processing.
[0068] A separate circularization step was performed to promote covalently closed circular RNA formation from the un-circularized IVT product. This step included IX T4 RNA ligaseI buffer (NEB, B0216L), 2 mM GTP (NEB, N0450L), and 0.5 units of RNase inhibitor (NEB, M0314L) incorporated into 45 pg of silica column-purified RNA in a 50 pl reaction. The reaction mixture was heated at 55°C for 8 minutes and then silica column purified (NEB, T2040L). To isolate circular RNAs, the column-purified RNA was digested with 4 units of RNase R (Abeam, ab286929) per microgram of RNA for 15 minutes at 37°C. The samples were subsequently column purified, quantified using a Nanodrop One Microvolume UV spectrophotometer, and verified for complete digestion using the E-Gel Electrophoresis System or an Agilent TapeStation, following the manufacturer's instructions.Example 7. dsRNA ELISA
[0069] The dsRNA byproduct from in vitro transcription was detected using a dsRNA sandwich enzyme-linked immunosorbent assay (ELISA) that selectively identifies multispecies dsRNA molecules larger than 30-40 bp. This assay was performed with a multispecies dsRNA ELISA Kit (Novus Biologicals, NBP3-11368), using antibodies as previously described by Schonbom, J., et al. The KI (IgG2a) mouse monoclonal antibody was immobilized on 96-well Immulon 2 HB plates (Thermo Fisher Scientific, 3455) overnight at 4°C, then blocked with 1% BSA in PBS at 37°C for 2 hours. After triple washing, mRNA samples and Poly (EC) dsRNA standards were added to the plates and incubated for 1 hour at 37°C. Following another triple wash, the plates were incubated with monoclonal antibody K2 (IgM) at 37°C for 1 hour. After triple washing again, the plates were exposed to horseradish peroxidase (HRP)-conjugated F(ab’)2 fragment of goat anti-mouse secondary antibody at 37°C for 1 hour, followed by a final wash. Subsequently, 100 pL of TMB (3, 3', 5, 5' tetramethylbenzidine) substrate solution was added to each well, followed by 100 pL of 2M H2SO4. The absorbance was measured at 450 nm using a BioTek Synergy Neo2 Hybrid Multimode Reader (BioTek, BTNEO2).Example 8. mRNA analysis by Egel system and TapeStation
[0070] Linear RNA or isolated circular RNA was column purified and quantified using a NanoDrop One spectrophotometer. For RNA characterization via the E-gel system, 1200 ng of RNA samples and 4 pL ssRNA Ladder (NEB, N0362S) were denatured by a 1 : 1 dilution with formamide (Sigma, F7503-250ML). The samples were then loaded onto 2% E-Gel™ EX Agarose Gels with SYBR-GOLD II and run on the E-Gel Power Snap PlusElectrophoresis System (Thermo Fisher, G9101) using the settings: E-gel Category “11 wells,” E-gel Type “E-Gel™ EX 2%,” and Time “12 minutes” at room temperature. Images were captured using a Bio-Rad ChemiDoc Imaging System with the “SYBR-Gold” settings. For the characterization of circular RNA using the Agilent TapeStation, 2 pL of 10 ng / pL mRNA samples and RNA ladder were mixed with 1 pL of High Sensitivity RNA Sample Buffer. The samples were denatured at 72°C for 3 minutes and then cooled at 4°C for 2 minutes. They were subsequently loaded into the 4200 TapeStation instrument (Agilent, G2991BA) and analyzed using the 4200 TapeStation Controller Software, following the manufacturer’s instructions.Example 9. IVTT for T7 RNA polymerase production
[0071] For high throughput T7 RNAP generation, we used the TnT Quick Coupled Transcription / Translation System (Promega #2080) by cloning mutant T7 RNAP under an SP6 promoter to use for in vitro coupled transcription-translation. We incubated the input plasmid with the reaction mixture for 90 minutes according to the manufacturer’s protocol and subsequently used 2 pL of the reaction mixture as the input T7 RNA polymerase for in vitro transcription of downstream mRNA. For FVT, we used the NEB ARCA HiScribe co- transcriptional kit with the manufacturer-supplied Clue mRNA template. We followed the protocol by incubating the reaction mixture for 30 minutes at 37 °C. After IVT, 2 pL of DNase I was added to degrade the DNA template, and polyA polymerase was added for polyA tailing at 37 °C for 30 minutes. mRNA is purified using monarch mRNA cleanup column (#T2047, NEB) according to the manufacturer’s protocol and eluted in mRNA storage buffer (#AM7000, ThermoFisher). mRNA is used directly for transfection using messengerMax (#LMRNA001, ThermoFisher), or frozen at -80 °C for subsequent characterization.Example 10. LNP production
[0072] LNPs were prepared using a vortex mixing method(9). ALC-0315, DSPC, cholesterol, and DMG-PEG-2000 were dissolved in ethanol at 75, 10, 10, and 10 mg / ml, respectively, and mixed at a molar ratio of 50: 10:38.5: 1.5. The final volume was brought to 30 pl with ethanol. Separately, 10 pg of purified mRNA was diluted in 10 mM citrate buffer (pH 4) to a final volume of 90 pl. The lipid mixture (30 pl) was rapidly pipetted into the vortexing RNAsolution (90 JJ.1) at a 1 :3 volume ratio, and vortexed for an additional 20 seconds. The mixture was left at room temperature for 10-15 minutes, then dialyzed against PBS at 4°C overnight to remove ethanol and residual lipids, resulting in the final LNP formulation. RNA encapsulation efficiency was measured using the Thermo Fisher RiboGreen assay. LNP samples were treated with Tris-EDTA or Tris-EDTA + 1% Triton-X. Free RNA was detected using an RNA-binding fluorescent dye, RiboGreen, and fluorescence was measured with a plate reader. Encapsulation efficiency was quantified by comparing the fluorescence of treated samples to RNA standards.
[0073] Alternatively, LNPs were synthesized using the microfluidic organic-aqueous precipitation method. The organic / ethanol phase consisted of lipids Dlin-MC3-DMA, DSPC, Cholesterol, and DMG-PEG2k at a molar ratio of 50: 10:38.5: 1.5. The aqueous phase was prepared by diluting the RNA payload in a 10 mM citrate buffer at pH 3.0. The two phases were prepared at an ethanokaqueous volume ratio of 1 :2, and at an N:P ratio of 5 : 1. The phases were mixed through the NxGen microfluidic cartridge by the NanoAssemblr Ignite (Precision Nanosystems). The Ignite was set to: volume ratio- 2: 1; flow rate- 12 ml / min; waste volume- 0 mL. The resulting LNPs were dialyzed against PBS using 20K MWCO Slide- A-Lyzer™ MINI Dialysis cassettes (ThermoFisher Scientific) at 25°C for 90 min, with an exchange of the buffer reservoir at 45 min.References1. Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, N. Smetanin, R. Verkuil, O. Kabeli, Y. Shmueli, A. Dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido, A. Rives, Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123-1130 (2023).2. M. Heinzinger, K. Weissenow, J. G. Sanchez, A. Henkel, M. Mirdita, M. Steinegger, B. Rost, Bilingual Language Model for Protein Sequence and Structure, bioRxiv (2024)p. 2023.07.23.550085.3. A. Elnaggar, H. Essam, W. Salah-Eldin, W. Moustafa, M. Elkerdawy, C. Rochereau, B. Rost, Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, arXiv [cs.LG] (2023). http: / / arxiv.org / abs / 2301.06568.4. N. Brandes, D. Ofer, Y. Peleg, N. Rappoport, M. Linial, ProteinBERT: a universal deeplearning model of protein sequence and function. Bioinformatics 38, 2102-2110 (2022).5. Y. He, X. Zhou, C. Chang, G. Chen, W. Liu, G. Li, X. Fan, M. Sun, C. Miao, Q. Huang, Y. Ma, F. Yuan, X. Chang, Protein language models-assisted optimization of an uracil- N-glycosylase variant enables programmable T-to-G and T-to-C base editing. Mol. Cell 84, 1257-1270. e6 (2024).6. T. Hayes, R. Rao, H. Akin, N. J. Sofroniew, D. Oktay, Z. Lin, R. Verkuil, V. Q. Tran, J. Deaton, M. Wiggert, R. Badkundri, I. Shafkat, J. Gong, A. Derry, R. S. Molina, N. Thomas, Y. A. Khan, C. Mishra, C. Kim, L. J. Bartie, M. Nemeth, P. D. Hsu, T. Sercu, S. Candido, A. Rives, Simulating 500 million years of evolution with a language model, bioRxiv (2024)p. 2024.07.01.600583.7. N. Ferruz, S. Schmidt, B. Hocker, ProtGPT2 is a deep unsupervised language model for protein design. Nat. Commun. 13, 4348 (2022).8. A. Madani, B. Krause, E. R. Greene, S. Subramanian, B. P. Mohr, J. M. Holton, J. L. Olmos Jr, C. Xiong, Z. Z. Sun, R. Socher, J. S. Fraser, N. Naik, Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 41, 1099— 1106 (2023).9. J. A. Ruffolo, S. Nayfach, J. Gallagher, A. Bhatnagar, J. Beazer, R. Hussain, J. Russ, J. Yip, E. Hill, M. Pacesa, A. J. Meeske, P. Cameron, A. Madani, Design of highly functional genome editors by modeling the universe of CRISPR-Cas sequences, bioRxiv (2024)p. 2024.04.22.590591.10. K. K. Yang, Z. Wu, F. H. Arnold, Machine-leaming-guided directed evolution for protein engineering. Nat. Methods 16, 687-694 (2019).11. H. Lu, D. J. Diaz, N. J. Czarnecki, C. Zhu, W. Kim, R. Shroff, D. J. Acosta, B. R. Alexander, H. O. Cole, Y. Zhang, N. A. Lynd, A. D. Ellington, H. S. Alper, Machine learning-aided engineering of hydrolases for PET depolymerization. Nature 604, 662- 667 (2022).12. N. Thomas, D. Belanger, C. Xu, H. Lee, K. Hirano, K. Iwai, V. Polic, K. D. Nyberg, K. G. Hoff, L. Frenz, C. A. Emrich, J. W. Kim, M. Chavarha, A. Ramanan, J. J. Agresti, L. J. Colwell, Engineering of highly active and diverse nuclease enzymes by combining machine learning and ultra-high-throughput screening, bioRxiv (2024)p.2024.03.21.585615.13. Z. Wu, S. B. J. Kan, R. D. Lewis, B. J. Wittmann, F. H. Arnold, Machine leaming-assisted directed protein evolution with combinatorial libraries. Proc. Natl. Acad. Sci. U. A. 116, 8852-8858 (2019).14. B. J. Wittmann, Y. Yue, F. H. Arnold, Informed training set design enables efficient machine learning-assisted directed protein evolution. Cell Syst 12, 1026-1045. e7 (2021).15. S. Biswas, G. Khimulya, E. C. Alley, K. M. Esvelt, G. M. Church, Low-N protein engineering with data-efficient deep learning. Nat. Methods 18, 389-396 (2021).16. L. Brenan, A. Andreev, O. Cohen, S. Pantel, A. Kamburov, D. Cacchiarelli, N. S. Persky, C. Zhu, M. Bagul, E. M. Goetz, A. B. Burgin, L. A. Garraway, G. Getz, T. S. Mikkelsen, F. Piccioni, D. E. Root, C. M. Johannessen, Phenotypic Characterization of a Comprehensive Set of MAPK1 / ERK2 Missense Mutants. Cell Rep. 17, 1171-1183 (2016).17. P. Notin, A. W. Kollasch, D. Ritter, L. vanNiekerk, S. Paul, H. Spinner, N. Rollins, A. Shaw, R. Weitzman, J. Frazer, M. Dias, D. Franceschi, R. Orenbuch, Y. Gal, D. S. Marks, ProteinGym: Large-Scale Benchmarks for Protein Design and Fitness Prediction. bioRxiv, doi: 10.1101 / 2023.12.07.570727 (2023).18. T. Hino, S. N. Omura, R. Nakagawa, T. Togashi, S. N. Takeda, T. Hiramoto, S. Tasaka, H. Hirano, T. Tokuyama, H. Uosaki, S. Ishiguro, M. Kagieva, H. Yamano, Y. Ozaki, D. Motooka, H. Mori, Y. Kirita, Y. Kise, Y. Itoh, S. Matoba, H. Aburatani, N. Yachie, T. Karvelis, V. Siksnys, T. Ohmori, A. Hoshino, O. Nureki, An AsCasl2f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell 186, 4920-4935. e23 (2023).19. H. K. Haddox, A. S. Dingens, J. D. Bloom, Experimental Estimation of the Effects of All Amino-Acid Mutations to HIV’s Envelope Protein on Viral Replication in Cell Culture. PLoS Pathog. 12, el006114 (2016).20. E. D. Kelsic, H. Chung, N. Cohen, J. Park, H. H. Wang, R. Kishony, RNA Structural Determinants of Optimal Codons Revealed by MAGE-Seq. Cell Syst 3, 563-57 l.e6 (2016).21. M. A. Stiffler, D. R. Hekstra, R. Ranganathan, Evolvability as a function of purifying selection in TEM- 1 p-lactamase. Cell 160, 882-892 (2015).22. C. J. Markin, D. A. Mokhtari, F. Sunden, M. J. Appel, E. Akiva, S. A. Longwell, C. Sabatti, D. Herschlag, P. M. Fordyce, Revealing enzyme functional architecture viahigh-throughput micro fluidic enzyme kinetics. Science 373 (2021). A. O. Giacomelli, X. Yang, R. E. Lintner, J. M. McFarland, M. Duby, J. Kim, T. P. Howard, D. Y. Takeda, S. H. Ly, E. Kim, H. S. Gannon, B. Hurhula, T. Sharpe, A. Goodale, B. Fritchman, S. Steelman, F. Vazquez, A. Tshemiak, A. J. Aguirre, J. G. Doench, F. Piccioni, C. W. M. Roberts, M. Meyerson, G. Getz, C. M. Johannessen, D. E. Root, W. C. Hahn, Mutational processes shape the landscape of TP53 mutations in human cancer. Vat. Genet. 50, 1381-1387 (2018). E. M. Jones, N. B. Lubock, A. J. Venkatakrishnan, J. Wang, A. M. Tseng, J. M. Paggi, N. R. Latorraca, D. Cancilla, M. Satyadi, J. E. Davis, M. M. Babu, R. O. Dror, S. Kosuri, Structural and functional characterization of G protein-coupled receptors with deep mutational scanning. Elife 9 (2020). M. B. Doud, J. D. Bloom, Accurate Measurement of the Effects of All Amino-Acid Mutations on Influenza Hemagglutinin. Viruses 8 (2016). J. M. Lee, J. Huddleston, M. B. Doud, K. A. Hooper, N. C. Wu, T. Bedford, J. D. Bloom, Deep mutational scanning of hemagglutinin helps predict evolutionary fates of human H3N2 influenza variants. Proc. Natl. Acad. Sci. U. S. A. 115, E8276-E8285 (2018). A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma, R. Fergus, Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl. Acad. Sci. U. S. A. 118 (2021). E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, G. M. Church, Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods 16, 1315-1322 (2019). A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y. Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger, D. Bhowmik, B. Rost, ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning. IEEE Trans. Pattern Anal. Mach. Intell. 44, 7112-7127 (2022). B. L. Hie, V. R. Shanker, D. Xu, T. U. J. Bruun, P. A. Weidenbacher, S. Tang, W. Wu, J. E. Pak, P. S. Kim, Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol. 42, 275-283 (2024). A. Baum, D. Ajithdoss, R. Copin, A. Zhou, K. Lanza, N. Negron, M. Ni, Y. Wei, K.Mohammadi, B. Musser, G. S. Atwal, A. Oyejide, Y. Goez-Gazi, J. Dutton, E. Clemmons, H. M. Staples, C. Bartley, B. Klaffke, K. Alfson, M. Gazi, O. Gonzalez, E. Dick Jr, R. Carrion Jr, L. Pessaint, M. Porto, A. Cook, R. Brown, V. Ali, J. Greenhouse, T. Taylor, H. Andersen, M. G. Lewis, N. Stahl, A. J. Murphy, G. D. Yancopoulos, C. A. Kyratsous, REGN-COV2 antibodies prevent and treat SARS-CoV-2 infection in rhesus macaques and hamsters. Science 370, 1110-1115 (2020).32. C.-L. Hsieh, J. A. Goldsmith, J. M. Schaub, A. M. DiVenere, H.-C. Kuo, K. Javanmardi, K. C. Le, D. Wrapp, A. G. Lee, Y. Liu, C.-W. Chou, P. O. Byrne, C. K. Hjorth, N. V. Johnson, J. Ludes-Meyers, A. W. Nguyen, J. Park, N. Wang, D. Amengor, J. J. Lavinder, G. C. Ippolito, J. A. Maynard, I. J. Finkelstein, J. S. McLellan, Structure-based design of prefusion-stabilized SARS-CoV-2 spikes. Science 369, 1501-1505 (2020).33. C. Xin, J. Yin, S. Yuan, L. Ou, M. Liu, W. Zhang, J. Hu, Comprehensive assessment of miniature CRISPR-Casl2f nucleases for gene disruption. Nat. Commun. 13, 5623 (2022).34. Z. Wu, Y. Zhang, H. Yu, D. Pan, Y. Wang, Y. Wang, F. Li, C. Liu, H. Nan, W. Chen, Q. Ji, Programmed genome editing by a miniature CRISPR-Casl2f nuclease. Nat. Chem. Biol. 17, 1132-1138 (2021).35. X. Xu, A. Chemparathy, L. Zeng, H. R. Kempton, S. Shang, M. Nakamura, L. S. Qi, Engineered miniature CRISPR-Cas system for mammalian genome regulation and editing. Mol. Cell 81, 4333-4345.e4 (2021).36. B. P. Kleinstiver, A. A. Sousa, R. T. Walton, Y. E. Tak, J. Y. Hsu, K. Clement, M. M. Welch, J. E. Homg, J. Malagon-Lopez, I. Scarfo, M. V. Maus, L. Pinello, M. J. Aryee, J. K. Joung, Engineered CRISPR-Cas 12a variants with increased activities and improved targeting ranges for gene, epigenetic and base editing. Nat. Biotechnol. 37, 276-282 (2019).37. X. Kong, H. Zhang, G. Li, Z. Wang, X. Kong, L. Wang, M. Xue, W. Zhang, Y. Wang, J. Lin, J. Zhou, X. Shen, Y. Wei, N. Zhong, W. Bai, Y. Yuan, L. Shi, Y. Zhou, H. Yang, Engineered CRISPR-OsCasl2fl and RhCasl2fl with robust activities and expanded target range for genome editing. Nat. Commun. 14, 2046 (2023).38. L. Zhang, J. A. Zuris, R. Viswanathan, J. N. Edelstein, R. Turk, B. Thommandru, H. T. Rube, S. E. Glenn, M. A. Collingwood, N. M. Bode, S. F. Beaudoin, S. Lele, S. N. Scott, K. M. Wasko, S. Sexton, C. M. Borges, M. S. Schubert, G. L. Kurgan, M. S. McNeill, C.A. Fernandez, V. E. Myer, R. A. Morgan, M. A. Behlke, C. A. Vakulskas, AsCasl2a ultra nuclease facilitates the rapid generation of therapeutic cell medicines. Nat. Commun. 12, 3908 (2021).39. D. Y. Kim, J. M. Lee, S. B. Moon, H. J. Chin, S. Park, Y. Lim, D. Kim, T. Koo, J.-H.Ko, Y.-S. Kim, Efficient CRISPR editing with a hypercompact Casl2fl and engineered guide RNAs delivered by adeno-associated virus. Nat. Biotechnol. 40, 94-102 (2022).40. P. Pausch, B. Al-Shayeb, E. Bisom-Rapp, C. A. Tsuchida, Z. Li, B. F. Cress, G. J. Knott, S. E. Jacobsen, J. F. Banfield, J. A. Doudna, CRISPR-CasO from huge phages is a hypercompact genome editor. Science 369, 333-337 (2020).41. J. L. Doman, S. Pandey, M. E. Neugebauer, M. An, J. R. Davis, P. B. Randolph, A. McElroy, X. D. Gao, A. Raguram, M. F. Richter, K. A. Everette, S. Banskota, K. Tian, Y. A. Tao, J. Tolar, M. J. Osborn, D. R. Liu, Phage-assisted evolution and protein engineering yield compact, efficient prime editors. Cell 186, 3983-4002. e26 (2023).42. M. T. N. Yamall, E. I. loannidi, C. Schmitt-Ulms, R. N. Krajeski, J. Lim, L. Villiger, W. Zhou, K. Jiang, S. K. Garushyants, N. Roberts, L. Zhang, C. A. Vakulskas, J. A. Walker, A. P. Kadina, A. E. Zepeda, K. Holden, H. Ma, J. Xie, G. Gao, L. Foquet, G. Bial, S. K. Donnelly, Y. Miyata, D. R. Radiloff, J. M. Henderson, A. Ujita, O. O. Abudayyeh, J. S. Gootenberg, Drag-and-drop genome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases. Nat. Biotechnol., 1-13 (2022).43. J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, A. Rives, Language models enable zeroshot prediction of the effects of mutations on protein function, bioRxiv (202 l)p. 2021.07.09.450648.44. A. Dousis, K. Ravichandran, E. M. Hobert, M. J. Moore, A. E. Rabideau, An engineered T7 RNA polymerase that produces mRNA free of immunostimulatory byproducts. Nat. Biotechnol. 41, 560-568 (2023).45. Z. J. Kartje, H. I. Janis, S. Mukhopadhyay, K. T. Gagnon, Revisiting T7 RNA polymerase transcription in vitro with the Broccoli RNA aptamer as a simplified realtime fluorescent reporter. J. Biol. Chem. 296, 100175 (2021).46. R. Chen, S. K. Wang, J. A. Belk, L. Amaya, Z. Li, A. Cardenas, B. T. Abe, C.-K. Chen, P. A. Wender, H. Y. Chang, Author Correction: Engineering circular RNA for enhanced protein production. Nat. Biotechnol. 41, 293 (2023).47. S. R. Johnson, X. Fu, S. Viknander, C. Goldin, S. Monaco, A. Zelezniak, K. K. Yang, Computational scoring and experimental evaluation of enzymes generated by neural networks. Nat. Biotechnol., doi: 10.1038 / s41587-024-02214-2 (2024).48. V. R. Shanker, T. U. J. Bruun, B. L. Hie, P. S. Kim, Unsupervised evolution of protein and antibody complexes with a structure-informed language model. Science 385, 46-53 (2024).49. Y. Serrano, A. Ciudad, A. Molina, Are Protein Language Models Compute Optimal?, arXiv [q-bio.BM] (2024). http: / / arxiv.org / abs / 2406.07249.50. X. Cheng, B. Chen, P. Li, J. Gong, J. Tang, L. Song, Training Compute-Optimal Protein Language Models, bioRxiv (2024)p. 2024.06.06.597716.51. B. Chen, X. Cheng, P. Li, Y.-A. Geng, J. Gong, S. Li, Z. Bei, X. Tan, B. Wang, X. Zeng, C. Liu, A. Zeng, Y. Dong, J. Tang, L. Song, xTrimoPGLM: Unified lOOB-Scale Pretrained Transformer for Deciphering the Language of Protein, bioRxiv (2024)p.2023.07.05.547496.52. C. N. Bedbrook, K. K. Yang, J. E. Robinson, E. D. Mackey, V. Gradinaru, F. H. Arnold, Machine learning-guided channelrhodopsin engineering enables minimally invasive optogenetics. Nat. Methods 16, 1176-1184 (2019).53. J. W. Thornton, Resurrecting ancient genes: experimental analysis of extinct molecules. Nat. Rev. Genet. 5, 366-375 (2004).54. D. Ghosh, J. Cabrera, Enriched Random Forest for High Dimensional Genomic Data. IEEE / ACM Trans. Comput. Biol. Bioinform. 19, 2817-2828 (2022).55. A. Kirjner, J. Yim, R. Samusevich, S. Bracha, T. S. Jaakkola, R. Barzilay, I. R. Fiete, Improving protein optimization with smoothed fitness landscapes (2023). https: / / openreview.net / pdf / idmxlF2Zv8xO.56. K. Huang, R. Lopez, J.-C. Hutter, T. Kudo, A. Rios, A. Regev, Sequential Optimal Experimental Design of Perturbation Screens Guided by Multi-modal Priors, bioRxiv (2023)p. 2023.12.12.571389.57. P. M. Groth, M. H. Kerrn, L. Olsen, J. Salomon, W. Boomsma, Protein property prediction with uncertainties, arXiv [q-bio.BM] (2024). http: / / arxiv.org / abs / 2407.00002.58. J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, “ImageNet: A large-scalehierarchical image database” in 2009 IEEE Conference on Computer Vision and Pattern Recognition (IEEE, 2009), pp. 248-255.59. J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figumov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Zidek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli, D. Hassabis, Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021).60. M. Sourisseau, D. J. P. Lawrence, M. C. Schwarz, C. H. Storrs, E. C. Veit, J. D. Bloom, M. J. Evans, Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. J. Virol. 93 (2019).61. B. L. Hie, V. R. Shanker, D. Xu, T. U. J. Bruun, P. A. Weidenbacher, S. Tang, W. Wu, J. E. Pak, P. S. Kim, Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol. 42, 275-283 (2024).62. T. Hino, S. N. Omura, R. Nakagawa, T. Togashi, S. N. Takeda, T. Hiramoto, S. Tasaka, H. Hirano, T. Tokuyama, H. Uosaki, S. Ishiguro, M. Kagieva, H. Yamano, Y. Ozaki, D. Motooka, H. Mori, Y. Kirita, Y. Kise, Y. Itoh, S. Matoba, H. Aburatani, N. Yachie, T. Karvelis, V. Siksnys, T. Ohmori, A. Hoshino, O. Nureki, An AsCasl2f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell 186, 4920-4935. e23 (2023).63. A. J. Greaney, T. N. Starr, C. O. Barnes, Y. Weisblum, F. Schmidt, M. Caskey, C. Gaebler, A. Cho, M. Agudelo, S. Finkin, Z. Wang, D. Poston, F. Muecksch, T. Hatziioannou, P. D. Bieniasz, D. F. Robbiani, M. C. Nussenzweig, P. J. Bjorkman, J. D. Bloom, Mapping mutations to the SARS-CoV-2 RBD that escape binding by different classes of antibodies. Nat. Commun. 12, 4196 (2021).64. M. Sourisseau, D. J. P. Lawrence, M. C. Schwarz, C. H. Storrs, E. C. Veit, J. D. Bloom, M. J. Evans, Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. J. Virol. 93 (2019).65. A. Elnaggar, H. Essam, W. Salah-Eldin, W. Moustafa, M. Elkerdawy, C. Rochereau, B. Rost, Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, arXiv [cs.LG] (2023). http: / / arxiv.org / abs / 2301.06568.66. E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, G. M. Church, Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods 16, 1315-1322 (2019).67. J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, A. Rives, Language models enable zeroshot prediction of the effects of mutations on protein function, bioRxiv (2021) p. 2021.07.09.450648.68. L. Gieselmann, C. Kreer, M. S. Ercanoglu, N. Lehnen, M. Zehner, P. Schommers, J. Potthoff, H. Gruell, F. Klein, Effective high-throughput isolation of fully human antibodies targeting infectious pathogens. Nat. Protoc. 16, 3639-3671 (2021).69. X. Wang, S. Liu, Y. Sun, X. Yu, S. M. Lee, Q. Cheng, T. Wei, J. Gong, J. Robinson, D. Zhang, X. Lian, P. Basak, D. J. Siegwart, Preparation of selective organ-targeting (SORT) lipid nanoparticles (LNPs) using multiple technical methods for tissue-specific mRNA delivery. Nat. Protoc. 18, 265-291 (2023).70. K. Clement, H. Rees, M. C. Canver, J. M. Gehrke, R. Farouni, J. Y. Hsu, M. A. Cole, D. R. Liu, J. K. Joung, D. E. Bauer, L. Pinello, CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol. 37, 224-226 (2019).71. J. Schymkowitz, J. Borg, F. Stricher, R. Nys, F. Rousseau, L. Serrano, The FoldX web server: an online force field. Nucleic Acids Res. 33, W382-8 (2005).72. S. Bae, J. Park, J.-S. Kim, Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473— 1475 (2014).
Claims
1. What is claimed:
1. A non-natural T7 RNA polymerase comprising an amino acid sequence with one or more mutations relative to a wild-type T7 RNA polymerase having an amino acid sequence set forth in SEQ ID NO: 1, wherein the one or more mutations are located at one or more amino acids at positions 3, 5, 10, 12, 25, 47, 89, 105, 134, 147, 152, 177, 225, 256, 229, 241, 249, 273, 279, 281, 370, 371, 469, 531, 567, 643, 668, 683, 735, 738, 792, 797, 800, and / or 822 relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
2. The non-natural T7 RNA polymerase of claim 1, wherein the one or more mutations comprise 3M, 5D, 51, 5M, 10K,12N, 25N, 47A, 89R, 1051, 134T, 147K, 152N, 177L, 225E, 2291, 241W, 249C, 256P, 273H, 279S, 281L, 370V, 370W, 371H, 469Q, 531T, 567R, 643A, 668E, 643G, 643K, 643N, 643R, 643S, 643T, 683K, 735S, 738N, 792H, 797K, 800K, and / or 822S relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
3. The non-natural T7 RNA polymerase of claim 1, wherein the one or more mutations are located at two or more amino acids at positions:47 and 643;531, 370, and 643;738 and 643;3 and 643;105 and 643;469, 47, and 643;47, 738, and 643;738, and 643;3, 105, and 643;177, 47, and 643;3, 47, and 643;668, 47, and 643;134, 47, and 643;12, 47, and 643;822, 47, and 643;12 and 643;567, 47, and 643;668, 3, and 643; or370, 47, and 643, relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
4. The non-natural T7 RNA polymerase, wherein the one or more mutations comprise:47A and 643G;53 IT, 370V, and 643G;738N and 643G;3M and 643G;1051 and 643G;469Q, 47A, and 643G;47A, 738N, and 643G;738N, and 643G;3M, 1051, and 643G;177L, 47A, and 643G;3M, 47A, and 643G;668E, 47A, and 643G;134T, 47A, and 643G;12N, 47A, and 643G;822S, 47A, and 643G;12N and 643G;567R, 47A, and 643G;668E, 3M, and 643G; or370W, 47A, and 643G, relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO: 1.
5. The non-natural T7 RNA polymerase of claim 1, wherein the one or more mutations are located at one or more amino acids at positions 370, 567, 643, 738, 797, and / or 800 relative to the wild-type T7 RNA polymerase amino acid sequence set forth in SEQ ID NO:
6. The non-natural T7 RNA polymerase of claim 1, wherein the amino acid mutation at position 370 is 370V or 370W.
7. The non-natural T7 RNA polymerase of claim 1, wherein the amino acid mutation at position 567 is 567R.
8. The non-natural T7 RNA polymerase of claim 1, wherein the amino acid mutation at position 643 is 643A, 643K, 643N, 643R, 643S, 643T, or 643G.
9. The non-natural T7 RNA polymerase of claim 1, wherein the amino acid mutation at position 738 is 738N.
10. The non-natural T7 RNA polymerase of claim 1, wherein the amino acid mutation at position 797 is 797K.
11. The non-natural T7 RNA polymerase of claim 1 , wherein the amino acid mutation at position 800 is 800K.
12. The non-natural T7 RNA polymerase of claim 1, wherein the non-natural T7 RNA polymerase amino acid sequence is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 2-61.
13. A nucleic acid encoding the non-natural T7 RNA polymerase of claim 1.
14. A host cell comprising the non-natural T7 RNA polymerase of claim 1.
15. A method for producing an RNA sequence in vitro, the method comprising combining the non-natural T7 RNA polymerase of claim 1, a nucleic acid template comprising a T7 promoter sequence, and NTPs.
16. The method of claim 16, wherein the nucleic acid template comprises plasmid DNA, linear DNA, or mRNA.
17. The method of claim 16, further comprising a circularization step to promote covalently closed circular RNA production.
18. The method of claim 15, wherein the produced RNA sequence is less immunogenic relative to a reference RNA sequence produced in vitro comprising a wild-type RNA polymerase comprising a sequence as set forth in SEQ ID NO: 1.
19. The method of claim 15, wherein the non-natural T7 RNA polymerase comprises increased activity relative to a wild-type RNA polymerase comprising a sequence as set forth in SEQ ID NO: 1.
20. The method of claim 19, wherein the increased activity is increased translation efficiency.
Citation Information
Patent Citations
Improved t7 expression system
WO2009021191A2
Thermostable variants of t7 RNA polymerase
WO2017123748A1
Mutant polymerases and methods of using the same
WO2021228905A2
Custom bacterial strain for recombinant protein production
WO2023069900A1