Genetic engineering of proteins and protein design methods for miniature crispr nucleases
EVOLVE-Pro rapidly identifies highly active miniature CRISPR nucleases for efficient genome editing in mammalian cells, overcoming size limitations and experimental challenges, achieving high activity and minimal off-target effects.
Patent Information
- Application Number
- PCT/US2025/029698
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-20
- Filing Date
- 2025-05-16
- Publication Date
- 2026-01-22
AI Technical Summary
Existing CRISPR nucleases are too large for effective delivery in mammalian models and require extensive experimental testing to achieve high activity, limiting their application in gene editing and therapeutics.
Development of miniature CRISPR nucleases using a few-shot active learning framework with EVOLVE-Pro, a multi-modal protein design model, to rapidly identify highly active variants with minimal experimental testing, enabling efficient delivery and editing.
The miniature CRISPR nucleases demonstrate significantly improved activity, achieving up to 50% indel formation and efficient genome editing in mammalian cells, with minimal off-target effects, using a single AAV vector delivery.
Smart Images

Figure IMGF000049_0001 
Figure IMGF000012_0001 
Figure IMGF000013_0001
Abstract
Description
GENETIC ENGINEERING OF PROTEINS AND PROTEIN DESIGN METHODS FOR MINIATURE CRISPR NUCLEASESCross Reference to Related Applications
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial Nos. 63 / 672,398 and 63 / 722,658 filed respectively on July 17, 2024, and November 20, 2024. The entire content of the above-referenced patent applications is incorporated by reference herein.Sequence Listing
[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on May 15, 2025, is named 761817_083474-042PCl_SL.xml and is 92,808 bytes in size.Background
[0003] Cluster Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (Cas) nuclease systems are widely used as genome editing tools. Cas9 and Cas 12 are two examples of nucleases that are often used in CRISPR-Cas systems. These nucleases are generally more than 1000 amino acids long and can be guided by a guide RNA to edit a single stranded or double-stranded DNA target near a short sequence called protospacer adjacent motif (PAM).
[0004] However, while these nucleases offer flexibility, their size remains a significant barrier to their use. For example, gene editing and programmable gene activation and inhibition technologies based on these nucleases cannot be delivered in mouse models using common methods, such as adeno-associated vectors (AAV), because of the large size of the nuclease. Furthermore, the development of effective gene and cell therapies requires genome editing tools that can meet the demands for reduced payload sizes and efficient integration of diverse and large sequences, regardless of cell type or active repair pathways. In addition, CRISPR associated transposases, such as Cas 12k or type I-F directed Tn7 systems, allow for programmable integration in bacteria without the need for repair-pathway dependent editing, but have yet to be reconstituted in eukaryotic cells for mammalian genome editing. The difficulty in reconstitution of these systems can be due to the sheer number of proteins (4-7 proteins) that must be properly expressed and delivered to the nucleus for proper assemblyand DNA targeting. Prime editing was also reported for programmable gene editing independent of DNA repair pathways but is limited to base substitutions or small deletions and insertions (about < 50 bp).
[0005] Thus, there is a need for smaller and more compact CRISPR nucleases for gene editing, programmable gene activation and inhibition, and new applications. Smaller and more compact CRISPR nucleases can simplify delivery and extend application, and the additional space on such nucleases can enable fusion with effector domains.
[0006] One way to design and discover smaller and more compact CRISPR nucleases for gene editing is by the use of deep learning and active learning networks. Such networks rely on the fact that billions of years of evolutionary pressures have shaped protein diversity, filtering potential design space for diverse biological activities. There is emerging evidence that these sequences embody a fundamental language of biology that can be modeled with deep learning, which offers unique insights into the evolutionary processes that have sculpted life on our planet. Protein language models (PLMs) learn the grammar of protein diversity by training to complete masked amino acids in a large protein sequence database, resulting in novel representations of biology. PLMs, with their internal representations of the protein evolutionary landscape, have been used to nominate variants with improved activity with limited success. Generative PLMs, such as ESM3, ProtGPT2, and ProGen, can design novel proteins, but these de novo-designed variants typically only reach wild-type level activity after extensive rounds of experimental testing. This inability of PLMs to drastically improve upon protein activity in zero-shot is partially driven by their inability to generalize to new contexts due to evolutionary constraints, as well as limited training data. Active learning methods that utilize context-specific data in combination with deep learning models, including machine learning-directed protein evolution (MLDE) methods, have effectively improved diverse proteins, but typically require extensive efforts to experimentally evaluate many variants. Previous attempts combined protein representation models with active learning to simplify the evolution process, but these approaches have not been generalized well beyond proof-of-concept demonstrations, like GFP engineering.
[0007] Accordingly, there exists a need for novel protein variants nominated by means of PLM and methods of using same for biological applications. Here a few shot active learning framework is used to rapidly develop highly active programmable RNA-guided miniature CRISPR nucleases.Summary
[0008] In one aspect, the disclosure provides a composition comprising (a) a target specific miniature nuclease comprising an amino acid sequence at least 70% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs 2-60; and (b) a guide RNA (gRNA), wherein the target specific miniature nuclease and the gRNA form a complex capable of binding to, nicking, unwinding, and / or cleaving a DNA target.
[0009] In one embodiment, the target specific miniature nuclease comprises one or more point mutations, relative to wild-type sequence SEQ ID NO:1.
[0010] In one embodiment, the target specific miniature nuclease comprises a point mutation at position 333, relative to wild-type sequence SEQ ID NO: 1.
[0011] In one embodiment, the target specific miniature nuclease comprises a point mutation at position 333 encoding a valine, relative to wild-type sequence SEQ ID NO:1.
[0012] In one embodiment, the target specific miniature nuclease comprises a point mutation at position 178, relative to wild-type sequence SEQ ID NO: 1.
[0013] In one embodiment, the target specific miniature nuclease comprises a point mutation at position 178 encoding an arginine, relative to wild-type sequence SEQ ID NO: 1.
[0014] In one embodiment, the target specific miniature nuclease comprises a point mutation at position 454, relative to wild-type sequence SEQ ID NO: 1.
[0015] In one embodiment, the target specific miniature nuclease comprises a point mutation at position 454 encoding a proline, relative to wild-type sequence SEQ ID NO: 1.
[0016] In one embodiment, the target specific miniature nuclease comprises a point mutation at positions 333, 178, and 454, relative to wild-type sequence SEQ ID NO: 1.
[0017] In one embodiment, the target specific miniature nuclease comprises a point mutation at positions 333 encoding a valine, 178 encoding an arginine, and 454 encoding a proline, relative to wild-type sequence SEQ ID NO: 1.
[0018] In one aspect, the disclosure provides a nucleic acid molecule encoding the target specific miniature nuclease comprising (a) a target specific miniature nuclease comprising an amino acid sequence at least 70% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs 2-60; and (b) a guide RNA (gRNA), wherein the targetspecific miniature nuclease and the gRNA form a complex capable of binding to, nicking, unwinding, and / or cleaving a DNA target.
[0019] In one aspect, the disclosure provides a vector comprising the nucleic acid molecule comprising (a) a target specific miniature nuclease comprising an amino acid sequence at least 70% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs 2-60; and (b) a guide RNA (gRNA), wherein the target specific miniature nuclease and the gRNA form a complex capable of binding to, nicking, unwinding, and / or cleaving a DNA target.
[0020] In one aspect, the disclosure provides a cell comprising the vector comprising (a) a target specific miniature nuclease comprising an amino acid sequence at least 70% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs 2-60; and (b) a guide RNA (gRNA), wherein the target specific miniature nuclease and the gRNA form a complex capable of binding to, nicking, unwinding, and / or cleaving a DNA target.
[0021] In one embodiment, the target specific miniature nuclease and the gRNA are packaged into one or more AAV vector(s) for in vitro or in vivo delivery.
[0022] In one embodiment, the gRNA comprises a PCSK9 targeting guide.
[0023] In one aspect, the disclosure provides a method of gene editing comprising adding the composition comprising (a) a target specific miniature nuclease comprising an amino acid sequence at least 70% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs 2-60; and (b) a guide RNA (gRNA), wherein the target specific miniature nuclease and the gRNA form a complex capable of binding to, nicking, unwinding, and / or cleaving a DNA target to a DNA target, wherein the DNA target is bound to, nicked, unwound, and or cleaved.
[0024] In one aspect, the disclosure provides a method of gene editing comprising (a) adding to a DNA target, a composition comprising a miniature CRISPR nuclease and a guide RNA; wherein the miniature CRISPR nuclease and the guide RNA bind to the DNA target; and (b) cleaving the DNA target, wherein one or more mutations of the miniature CRISPR nuclease amino acid sequence is located in a WED domain of the miniature CRISPR nuclease amino acid sequence, or in an a-helix of a REC domain of the miniature CRISPR nuclease amino acid sequence, or in a C-terminus of a RuvC domain of the miniature CRISPR nuclease amino acid sequence.
[0025] In one embodiment, the mutation located in the WED domain is at position 333, relative to wild-type sequence SEQ ID NO: 1.
[0026] In one embodiment, the mutation located in the a-helix of a REC domain is at position 178, relative to wild-type sequence SEQ ID NO: 1.
[0027] In one embodiment, the mutation located at the C-terminus of a RUVC domain is at position 454, relative to wild-type sequence SEQ ID NO: 1.
[0028] In one aspect, the disclosure provides a non-natural miniature CRISPR nuclease comprising an amino acid sequence with one or more mutations relative to a wild-type miniature nuclease sequence of SEQ ID NO: 1, wherein the one or more mutations are located at amino acid positions 20, 40, 44, 45, 54, 78, 88, 140, 147, 155, 159, 163, 178, 186, 188, 189, 190, 191, 192, 194, 197, 333, 365, 425, 452, 454, and combinations thereof.Brief Description of the Drawings
[0029] Aspects, features, benefits, and advantages of the embodiments described herein will be apparent with regard to the following description, appended claims, and accompanying drawings.
[0030] Fig. 1. Evolution of highly active genome editing enzymes with EVOLVE-Pro. (A) Schematic of the evolution strategy with EVOLVE-Pro for engineering a miniature Casl2f. (B) Engineering of PsaCasl2f over four rounds of EVOLVE-Pro and a rational combination multi-mutant round. Data shows cumulative top 10 mutants from current and preceding rounds, as measured by fold improvement of indel activity at the endogenous RNF2 genomic locus. (C) Indel activities of WT PsaCasl2f, epPsaCasl2f, and a panel of published Casl2a and Casl2f nucleases on 10 different genomic targets across five genes (two guides per gene). The fold change on top of each guide denotes the relative fold increase of epPsaCasl2f compared to the average of the other published Casl2a and Casl2f nucleases. A one-way ANOVA is performed for each guide sequence shown (****, p<0.0001). (D) Next-generation sequencing quantified indel formation at murine proprotein convertase subtilisin / kexin type 9 (PCSK9) genomic loci by epPsaCasl2f, WT PsaCasl2f, and SpCas9. A one-way ANOVA is performed for each guide sequence shown (****, p<0.0001). (E) Schematic of the in vivo validation assay for EnPsaCasl2f editing at the murine PCSK9 locus for PCSK9 reduction.(F) Serum PCSK9 levels at three different time points from -2 days of injection to +14 days.The percent of control PCSK9 was calculated by normalizing to the control group with PBS injected. A two-sided Student’s t-test was run on each time point relative to -2 days’ baseline PCSK9 level (ns, non-significant, *, p<0.05). (G) Mapping of the top mutations on the AlphaFold3 model of PsaCasl2f. The RuvC active site is indicated by a red circle. (H) Heatmap showing most common PsaCasl2f mutations explored by EVOLVE-Pro over rounds of evolution. Any position explored more than once is shown on a cumulative basis across rounds. (I) Scatter plot comparing the predicted naive ESM-2 protein fitness (predicted masked marginal score) and scaled tested activity of nominated mutants across evolution, scatter points are colored by rounds in evolution. (J-K) Comparison of the PsaCasl2f embedding latent space with either predicted naive ESM-2 protein fitness landscape or EVOLVE-Pro protein function landscape. (L) A kernel density estimate plot of protein fitness as predicted by ESM-2 versus protein function as predicted by EVOLVE-Pro. The correlation and linear regression line are shown in red and the R square of the correlation is reported.
[0031] Fig. 2. Additional characterization of evolution of PsaCasl2f with EVOLVE-Pro. (A) Indel editing rate of individual nominated mutants from Evolve-Pro across 5 rounds of evolution relative to WT. Error bars represent standard deviation of three biological replicates. (B) Average indel size generated by PsaCasl2f, epPsaCasl2f, and other Casl2f and Casl2a systems across 10 tested genomic regions. Error bars represent standard deviation of three biological replicates. (C) Kinetics of in vitro cleavage of a target DNA by either enPsaCasl2 or WT PsaCasl2f. Error bars represent standard deviation of three biological replicates. (D) Serum PCSK9 level determined by ELISA at -2, 7, and 14 days post-injection of AAV-epPsaCasl2f or a PBS control group. Error bars represent standard deviation of three biological replicates. (E) Next-generation sequencing determined editing in the livers of either AAV-epPsaCasl2f or PBS-treated mice at the targeted PCSK9 site. Error bars represent standard deviation of 12 technical replicates. (F) Individual insertion and deletion frequency around the target spacer region in PCSK9 as quantified by CRISPREsso2. Error bars represent standard deviation of three biological replicates. (G) Specificity of epPsaCasl2f. Indel formation was examined on the top four off-target (OT) sites predicted by Cas-OFFinder for epPsaCasl2f with its PCSK9 targeting guide, showing minimal off-target editing at these OT sites. Error bars represent standard deviation of three biological replicates. Figure 2 discloses SEQ ID NOS 61-65, respectively, in order of appearance. (H) Aviolin plot of free energy changes for each mutant in evolution compared to WT grouped by evolution round.Detailed Description
[0032] Evolve-pro
[0033] EVOLVE-Pro is a frontier multi-modal protein design model. EVOLVE-Pro evolves high-activity protein variants with few-shot learning and minimal experimental testing, achieving accurate prediction of sequence-to-function for general properties. This performance comes from an ensemble approach, combining evolutionary-scale protein foundation models with a top-layer discrimination model to learn a protein’s functional landscape and guide the directed evolution process in silico. By applying EVOLVE-Pro in a few-shot, active learning framework, protein sequences with significantly higher activity can be efficiently nominated in a generalizable fashion with minimal effort. The modularity of the EVOLVE-Pro architecture allows this framework to scale with larger parameter PLMs. Moreover, EVOLVE-Pro prompting only requires protein sequences to be evolved without any structural information, expert knowledge, or prior data. As EVOLVE-Pro is multi-modal, multiple protein features of any type or data class can be simultaneously engineered, opening up vast possibilities for its use in biology and medicine.
[0034] We benchmark EVOLVE-Pro in silico for proteins showing state-of-the-art performance, then apply the final model for a miniature CRISPR nuclease. EVOLVE-Pro yields mutants with 2- to 515-fold improvement over initial proteins. We demonstrate in vivo liver editing with an EVOLVE- Pro-engineered miniature nuclease. Analyzing nominated mutations, EVOLVE-Pro explores disparate sites, and the learned functional landscape is separate, and often negatively correlated with, the fitness inferred by the underlying protein language model. Lastly, we showcase EVOLVE-Pro’s utility in nominating multi-mutant protein designs out of a vast sequence space that enables final mutants that are much more active than naturally observed proteins. EVOLVE-Pro establishes the capabilities of few-shot active learning with protein language models for optimizing proteins for diverse activities.
[0035] Development and benchmarking of the EVOLVE-Pro model
[0036] To establish EVOLVE-Pro, we designed an ensemble model that involves: 1) a foundational protein language model to encode protein sequences into an information-richlatent space, and 2) a top-layer discrimination model to learn protein functional grammar in this evolutionary landscape and rank protein sequences according to a designed policy framework, and 3) an active learning framework using top layer discrimination model to nominate the next set of protein variants for experimental evaluation. This cycle is performed iteratively to evolve defined protein activities until they reach desired levels.
[0037] Evolution of a miniature RNA-guided CRISPR nuclease
[0038] Programmable RNA-guided nucleases have diverse applications in basic biology, therapeutics, and diagnostics. However, commonly used nucleases, such as the Cas9 from Streptococcus pyogenes (SpCas9) are too large to effectively be packaged in common viral vectors such as adeno-associated viral (AAV), and more compact high- efficiency nucleases, such as the Cas9 from Staphylococcus aureus (SaCas9) still preclude the use of larger regulatory elements or protein fusions. Miniature Casl2f nucleases have compact sizes (<700 residues) but suffer from reduced efficiencies, requiring significant engineering for genome editing applications. Previous Casl2f engineering efforts relied on DMS or rationally designed mutations to increase the in vitro cleavage activity, requiring extensive screening to find the optimal variant. To accelerate miniature nuclease engineering, we tested whether EVOLVE-Pro could rapidly develop highly active Casl2f variants.
[0039] We selected the Casl2f from Pseudomonas aeruginosa (PsaCasl2f) SEQ ID NO: 1 : MPSETYITKTLSLKLIPSDEEKQALENYFITFQRAVNFAIDRIVDIRSSFRYLNKNEQFP AVCDCCGKKEKIMYVNISNKTFKFKPSRNQKDRYTKDIYTIKPNAHICKTCYSGVAG NMFIRKQMYPNDKEGWKVSRSYNIKVNAPGLTGTEYAMAIRKAISILRSFEKRRRNA ERRIIEYEKSKKEYLELIDDVEKGKTNKIVVLEKEGHQRVKRYKHKNWPEKWQGISL NKAKSKVKDIEKRIKKLKEWKHPTLNRPYVELHKNNVRIVGYETVELKLGNKMYTI HFASISNLRKPFRKQKKKSIEYLKHLLTLALKRNLETYPSIIKRGKNFFLQYPVRVTVK VPKLTKNFKAFGIDRGVNRLAVGCIISKDGKLTNKNIFFFHGKEAWAKENRYKKIRD RLYAMAKKLRGDKTKKIRLYHEIRKKFRHKVKYFRRNYLHNISKQIVEIAKENTPTVI VLEDLRYLRERTYRGKGRSKKAKKTNYKLNTFTYRMLIDMIKYKAEEAGVPVMIID PRNTSRKCSKCGYVDENNRKQASFKCLKCGYSLNADLNAAVNIAKAFYECPTFRWE EKLHAYVCSEPDK for evolution with set indel formation at the endogenous RNF2 locus target site as the optimization metric (Fig. 1A). After four rounds of evolution of 12 single mutants per round, EVOLVE-Pro yielded point-mutants of PsaCasl2f with up to 4.9-fold improvement in indelformation. This top variant, PsaCasl2fK333v, had >40% indel efficiency at the RNF2 site (Fig. IB, Fig. 2A). To identify synergies between EVOLVE-Pro nominated mutants, we combined the top-performing variants from previous rounds in a fifth round. We evaluated a set of these multi-mutants and found that PsaCasl2fI178A / K333V / K454Pexhibited greater than 50% indel activity at the RNF2 locus, a 25% higher activity than any of the single mutants (Fig. 2A). Given its performance, we refer to the PsaCasl2fI178A / K333V / K454Pvariant as EVOLVE-Pro PsaCasl2f (epPsaCasl2f). Because these mutants synergized to produce an even more active enzyme, it suggests that the variants identified by EVOLVE-Pro are uniquely independent in mechanism, highlighting the insightful potential of the method.
[0040] To generalize epPsaCasl2fs improved activity, we evaluated the enzyme at 10 different targets across five endogenous genomic loci, comparing to WT PsaCasl2f and seven previously characterized Casl2 effectors, AsCasl2a, Casl2, UnCasl2fl, enAsCasl2f, OsCasl2f, RhCasl2f, and CasMINE We observed consistently higher epPsaCasl2f activity compared to WT PsaCasl2f on 9 of 10 tested targets (Fig. 1C). Moreover, epPsaCasl2f edited the 10 targets with a 23.3 ± 16.7% average indel rate, surpassing all tested miniature Casl2f effectors and AsCasl2a with 2.2- to 44-fold improvement. Interestingly, epPsaCasl2f generated an average deletion of 5-bp across the 10 tested targets, larger than the deletions generated by other orthologs (Fig. 2B). Consistent with our mammalian data, purified epPsaCasl2f exhibited higher biochemical DNA cleavage activity than WT PsaCasl2f (Fig. 2C). Together, these data demonstrate that epPsaCasl2f is a highly active, compact effector for mammalian genome editing that outperforms other small effectors.
[0041] We applied epPsaCasl2f for in vivo genome editing applications, using its compact size for single-vector viral delivery in vivo. We designed guides targeting a sequence 5' of exon 3 in the mouse PCSK9 gene (Fig. ID). The PCSK9 protein regulates blood low-density lipoprotein (LDL) by binding to LDL receptors, making it a valuable therapeutic target. We first tested the efficacy of epPsaCasl2f in a murine hepatocyte cell line (Hepa 1-6) by cotransfecting murine codon-optimized epPsaCasl2f and sgRNA targeting sequences 5' of exon3 in the PCSK9 gene. Analyses of epPsaCasl2f, WT PsaCasl2f, and Staphylococcus pyogenes Cas9 (SpCas9) revealed that epPsaCasl2f robustly edited PCSK9 in Hepal-6 cells with -40% indel formation (comparable levels to SpCas9 and 3 -fold higher than the WT PsaCasl2f) (Fig. ID).
[0042] After validation of epPsaCasl2f in Hepal-6 cells, we packaged both epPsaCasl2f and its sgRNA targeting PCSK9 in a single AAV2 / 8 vector (Fig. IE). AAV-epPsaCasl2f was administered at a titer of 1.5 x 1012viral genome copies per mouse via retro-orbital injection into 3-month-old C57BL / 6J mice. We tracked blood PCSK9 levels for 14 days post-injection of AAV and found a significant decrease to around 50% of the original levels after 14 days (Fig. IF, Fig. 2D). We then harvested the liver at day 15, isolated the genomic DNA, and performed next-generation sequencing to survey for indel formation at the PCSK9 target site (Fig. 2E-F). We found around 7% on-target indel formation in the AAV-epPsaCasl2f injected mice (Fig. 2E), demonstrating that epPsaCasl2f can be used for single-vector AAV- mediated genome editing. To survey off-targets, we used CRISPR-Off finder to predict the top four off-target cleavage sites generated by epPsaCasl2f and analyzed the guidedependent off- target cleavage in the liver. We only found detectable editing at one of the four sites with a maximum level of 0.27% indels, confirming minimal off-target cleavage triggered by epPsaCasl2f (Fig. 2G).
[0043] To understand the mechanisms of the beneficial mutations nominated by the EVOLVE-Pro, we used Alphafold3 to predict the structure of PsaCasl2f (Fig. 1G). The predicted structure provides insights into how the PLM -nominated mutations, including I178A / K333V / K454P, contribute to enhancing the DNA cleavage activity (Fig. 1G). The K333V mutation is located in the WED domain, suggesting that it could increase the binding to its RNA guide. The 1178 A mutation is located in the middle of the long a-helix in the REC domain and forms a hydrophobic core with 1245 and L248 in the adjacent a-helix. Given that alanine is a helix-forming residue, the I178A mutation may stabilize the a-helix in the REC domain and thus augment the cleavage activity. The K454P mutation is located at the C- terminus of an a-helix in the RuvC domain and forms hydrophobic interactions with A509 and V511 in the adjacent a-helix, suggesting that it also stabilizes the protein conformation.
[0044] We then looked at the model’s attention to particular residues in the protein by calculating the cumulative frequency of individual residues explored by the model and found that multiple residues are repeatedly nominated by the model, including G147 and E451 (Fig. 1H). We calculated the pMMS for each nominated mutant to understand the relationship between the base layer PLM’s fitness prediction and the actual measured protein activity (Fig. II). We found that there is a weak negative correlation between fitness and function in PsaCasl2’s local context but EVOLVE-Pro nominated for higher activity mutants toward the later rounds contrary to high fitness mutants recommended by the PLM base layer (Fig. II).We then further projected both base layer PLM’s fitness score and top layer random forest regressor’s activity score in the EMS2 latent space to better understand EVOLVE-Pro’s global mutational trajectory (Fig. 1 J-L). We found a weak positive correlation of 0.03 between fitness and activity, further denoting the necessity of a top-layer discrimination model to properly distinguish between high fitness and high activity (Fig. 11).
[0045] SequencesTable 1 : Casl2f nuclease mutants, SEQ ID NOs 2-60.EXAMPLESExample 1. Casl2f protein purification and preparation
[0046] The gene encoding PsaCasl2f (residues 1-586) with an N-terminal Hise-SUMO tag (“Hise” disclosed as SEQ ID NO: 66) was cloned into the pE-SUMO vector (LifeSensors). The mutations were introduced by a PCR-based method, and the sequences were confirmed by DNA sequencing. The Casl2f protein was expressed in Escherichia coli Rosetta2 (DE3) (Novagen) by induction with 0.25 mM isopropyl (3-D-thiogalactopyranoside (Nacalai Tesque) at 20°C overnight. The E. coli cells were lysed by sonication in buffer A (20 mM Tris-HCl, pH 8.0, 20 mM imidazole, 1 M NaCl, 3 mM 2-mercaptoethanol, 10% glycerol) and the lysate was clarified by centrifugation at 40,000 x g. The supernatant was applied to Ni- NTA Superflow resin (QIAGEN), and the Casl2f protein was eluted with buffer B (20 mM Tris-HCl, pH 8.0, 300 mM imidazole, 300 mM NaCl, 3 mM 2-mercaptoethanol, 10% glycerol). The eluate was treated with Ulpl peptidase at 4°C overnight and then loaded onto a HiTrap SP column (GE Healthcare), equilibrated with buffer C (20 mM HEPES, pH 7.5, 300 mM NaCl, 2 mM DTT). The protein was eluted with a linear gradient of 0.3-2 M NaCl. The Casl2f protein was further purified on a Superdex 200 Increase 10 / 300 column (GE Healthcare), equilibrated with buffer D (20 mM HEPES, pH 7.5, 1 M NaCl, 2 mM DTT). The peak fractions were collected and stored at -80°C until use. The sgRNAs were transcribed in vitro with T7 RNA polymerase, using PCR-amplified DNA templates, and were purified by 10% denaturing (7 M urea) polyacrylamide gel electrophoresis.Example 2. In vitro DNA cleavage assay
[0047] The purified PsaCasl2f was diluted to 2 pM (20 mM HEPES, pH 7.5, 600 mM NaCl, 2 mM DTT) and mixed with an equal volume of the sgRNA (2 pM) at 37°C for 2 min. The pre-assembled PsaCasl2f-sgRNA complex (5 pL, 1 pM) was then mixed with the linearized plasmid target containing the target sequence and the TTA protospacer adjacent motif (PAM) (2.5 pl, 100 ng / pL) and buffer F (2.5 pL, 60 mM HEPES, pH 7.5, 40 mM MgCh, 2 mMDTT). The 10 pL reaction solution (500 nM PsaCasl2f-sgRNA, 250 ng target DNA, 20 mM HEPES, pH 7.5, 150 mM NaCl, 10 mM MgCh, 1 mM DTT) was incubated at 37°C. Aliquots (2 pL) were taken at 15, 30, and 60 min, and mixed with 6 pL of quench solution (20 mM HEPES, pH 7.5, 150 mM NaCl, 2 mM DTT, 3.5 pg Proteinase K, 17.5 mM EDTA). The reaction products were incubated at 95 °C for 2 min and then analyzed using a MultiNA microchip electrophoresis system (SHIMADZU).Example 3. Measurement of luciferase activity
[0048] Media containing secreted or intracellular luciferase was harvested 48 hours after transfection unless otherwise noted. 20pL of media is used to measure secreted luciferase activity using Targeting Systems Cypridinia and Targeting systems Gaussia luciferase assay kits (Targeting Systems) on a Biotek Synergy 4 plate reader with an injection protocol. All replicates were performed as biological replicates. Intracellular Nanoluc and firefly luciferase were measured by lysing the cell in the luciferase assay mix (Promega) according to the manufacturer’s protocol. 5 minutes after lysis at room temperature, the signal is read out using a Biotek Synergy 4 plate reader.Example 4. Quantification of protein expression
[0049] Two days after the transfection of HEK293FT or BJ Fibroblast cells, the Nano-Gio HiBiT Lytic Detection System (Promega) was used for the quantification of the HiBiT tags, in cell lysates. For the preparation of the Nano-Gio HiBiT Lytic Reagent, the Nano-Gio HiBit Lytic Buffer (Promega) was mixed with Nano-Gio HiBiT Lytic Substrate (Promega) and the LgBiT Protein (Promega) according to the manufacturer’s protocol. The volume of Nano-Gio HiBiT Lytic Reagent added was equal to the culture medium present in each well, and the samples were placed on an orbital shaker at 600 rpm for 3 minutes. After incubation of 10 minutes at room temperature, the readout took place with 125 gain and 2 seconds integration time using a plate reader (Biotek Synergy Neo 2). The control background was subtracted from the final measurements.Example 5. Harvest of total RNA and quantitative PCR
[0050] For gene expression experiments in mammalian cells, cell harvesting and reverse transcription for cDNA generation were performed using a previously described modification of the commercial Cells-to-Ct kit (Thermo Fisher Scientific) 48 h after transfection.Transcript expression was then quantified with qPCR using Fast Advanced Master Mix (Thermo Fisher Scientific) and TaqMan qPCR probes (Thermo Fisher Scientific) with GAPDH control probes (Thermo Fisher Scientific). All qPCR reactions were performed in 10-pl reactions with two technical replicates in a 384-well format and read out using a LightCycler 480 Instrument II (Roche). For multiplexed targeting reactions, readout of different targets was performed in separate wells. Expression levels were calculated by subtracting housekeeping control (GAPDH) cycle threshold (Ct) values from target Ct values to normalize for total input, resulting in ACt levels. Relative transcript abundance was computed as 2-ACt. All replicates were performed as biological replicates.Example 5. AAV production and purification
[0051] Recombinant AAV2 / 8 was produced by transient HEK293 cell transfection and CsCl sedimentation by the University of Massachusetts Medical School Viral Vector Core, as generally known. Vector preparations were monitored by ddPCR, and purity was assessed by 4%— 12% SDS-acrylamide gel electrophoresis and silver staining (Invitrogen).Example 6. Animal AAV injection and processing
[0052] For in vivo testing of enPsaCasl2, an AAV was prepared at a titer of 1.5 x 1013gc / mL in sterile phosphate-buffered saline PBS (University of Massachusetts Medical School Viral Vector Core). Animals were randomly assigned to experimental or control groups. The investigator was not blinded to assignments. AAV was delivered to 3-month-old male C57 / BL6 mice via retro-orbital injection at a dose of 1.5 x 1012gc adjusted to 100 pL with PBS, pH 7.4 (Gibco), before the injection. In control animals, 100 pL of sterile PBS was administered via retro-orbital injection. To monitor serum levels of PCSK9 and total cholesterol, blood was routinely drawn from mice pre- and post-injection. Mice were fasted for 12 hours overnight prior to the blood draw. Blood collection was performed by saphenous vein sampling, with no more than 1% of the blood volume collected over a 24-hours period.To collect serum, whole blood was incubated at room temperature to allow clotting for 1 h, followed by centrifugation at 10,000 x g for 10 min. Serum samples were used immediately for testing, with the remaining samples stored at -20°C for any subsequent analysis. Animal studies were performed in accordance with the recommendations in the Guide for the Care and Use of Laboratory Animals of the National Institutes of Health. The protocols were approved by the Institutional Animal Care and Use Committee at the Massachusetts Institute of Technology.Example 7. In vivo Flue mRNA delivery and comparative in vivo bioluminescence
[0053] Prior to bioluminescence imaging, 8 to 10-week-old Albino B6 were anesthetized with 3% isoflurane and injected with 5 pg of synthesized mRNA via retro-orbital injection using homemade lipid nanoparticles. At the indicated time points post-injection, the mice were anesthetized again with 3% isoflurane and immediately administered 200 pl of 15 mg / mL D-luciferin (PerkinElmer) for imaging. Ventral bioluminescence images were acquired using an IVIS Spectrum In Vivo Imaging System (PerkinElmer). The following conditions were used for image acquisition: exposure time = 60 sec, binning = medium: 4, field of view = 15 x 15 cm, and f / stop = 1. Bioluminescent images were analyzed using Living Image 4.3 software (PerkinElmer) and normalized radiance (photons / s) was reported.Example 8. Serum and tissue analysis
[0054] To quantify serum PCSK9, the Mouse Proprotein Convertase 9 / PCSK9 Quantikine ELISA Kit (R&D Systems) was used with fresh serum samples, according to the manufacturer’s protocol. Total cholesterol levels were measured with theCholesterol / Cholesterol Ester-Glo™ Assay (Promega), according to the manufacturer’s protocol. All assays were performed using a Biotek Synergy Neo2 plate reader. To assess genome editing in livers, mice were euthanized by carbon dioxide inhalation. Livers were extracted and placed in ice-cold, sterile PBS. The liver portions that were not used for immediate analysis were snap-frozen and stored at -80°C. To isolate genomic DNA, liver pieces were processed using the DNeasy Blood & Tissue Kit (Qiagen), according to the manufacturer’s protocols. The PCSK9 region of interest was amplified from purified genomic DNA and sequenced as described above.Example 9. High throughput Cloning of mutants
[0055] Expression constructs for PsaCasl2f nucleases were cloned for mammalian expression via Gibson cloning using Hifi Assembly mix (NEB) according to the manufacturer’s instructions. Overlapping reverse and forward primer- carrying mutations for desired amino acids are used to amplify the plasmid around the globe with 18bp of homology. Then DPNI is used to clean up the plasmid from PCR reactions followed by column cleanup. 50ng of the cleaned-up PCR product is then used to perform Gibson reactions according to the manufacture’s protocol. For all Gibson clonings, 2 pl of assembled reactions were transformed into 20 pl of competent Stbl3 cells generated by Mix and Go! competency kit (Zymo) and plated on agar plates supplemented with appropriate antibiotics. After growth overnight at 37 °C, colonies were picked into TB medium (Thermo Fisher Scientific) and incubated with shaking at 37 °C for 24 h. Cultures were collected using a QIAprep Spin Miniprep kit (Qiagen) according to the manufacturer’s instructions.Example 10. PsaCasl2f mutations
[0056] To determine whether the model began to focus on specific locations of the protein, we further analyzed the pairwise mutational distance between the 12 mutations in each round and observed a shift to bimodality in the distribution of distances as the model entered the later rounds, with a bimodality coefficient of 0.72 in Round 4, suggesting that the model learned specific regions. We interpreted the mutations using a biophysical model to predict the free energy change of the protein based on the nominated mutations by our model. We found that the mutations nominated by the model in later rounds did not lead to protein stabilization (Fig. 2H). Instead, there was a large decrease in the coefficient of variation (COR) in the distribution of free energy changes, with only 471% COR in Round 4 relative to the 30,323% in Round 1 (Fig. 2H). This selection runs contrary to traditional rational engineering approaches that try to maximize stability, and likely reflects the model’s gain of understanding in the relationship between the fitness / stability and functional activity landscapes over iterative rounds, allowing the selection of non-intuitive residues. Protein function is not necessarily correlated with stability, and thus the combination of the base LLM latent space model and the top layer “domain expert” model employed here represents an important step toward the efficient in silico evolution of higher activity protein variants.References:1. Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, N. Smetanin, R. Verkuil, O. Kabeli, Y. Shmueli, A. Dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido, A. Rives, Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1 123-1130 (2023).2. M. Heinzinger, K. Weissenow, J. G. Sanchez, A. Henkel, M. Mirdita, M. Steinegger, B. Rost, Bilingual Language Model for Protein Sequence and Structure, bioRxiv (2024)p. 2023.07.23.550085.3. A. Elnaggar, H. Essam, W. Salah-Eldin, W. Moustafa, M. Elkerdawy, C. Rochereau, B. Rost, Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, arXiv [cs.LG] (2023). http: / / arxiv.org / abs / 2301.06568.4. N. Brandes, D. Ofer, Y. Peleg, N. Rappoport, M. Linial, ProteinBERT: a universal deeplearning model of protein sequence and function. Bioinformatics 38, 2102-2110 (2022).5. Y. He, X. Zhou, C. Chang, G. Chen, W. Liu, G. Li, X. Fan, M. Sun, C. Miao, Q. Huang, Y. Ma, F. Yuan, X. Chang, Protein language models-assisted optimization of an uracil- N-glycosylase variant enables programmable T-to-G and T-to-C base editing. Mol. Cell 84, 1257-1270. e6 (2024).6. T. Hayes, R. Rao, H. Akin, N. J. Sofroniew, D. Oktay, Z. Lin, R. Verkuil, V. Q. Tran, J. Deaton, M. Wiggert, R. Badkundri, I. Shafkat, J. Gong, A. Derry, R. S. Molina, N. Thomas, Y. A. Khan, C. Mishra, C. Kim, L. J. Bartie, M. Nemeth, P. D. Hsu, T. Sercu, S. Candido, A. Rives, Simulating 500 million years of evolution with a language model, bioRxiv (2024)p. 2024.07.01.600583.7. N. Ferruz, S. Schmidt, B. Hocker, ProtGPT2 is a deep unsupervised language model for protein design. Nat. Commun. 13, 4348 (2022).8. A. Madani, B. Krause, E. R. Greene, S. Subramanian, B. P. Mohr, J. M. Holton, J. L. Olmos Jr, C. Xiong, Z. Z. Sun, R. Socher, J. S. Fraser, N. Naik, Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 41, 1099— 1106 (2023).9. J. A. Ruffolo, S. Nayfach, J. Gallagher, A. Bhatnagar, J. Beazer, R. Hussain, J. Russ, J. Yip, E. Hill, M. Pacesa, A. J. Meeske, P. Cameron, A. Madani, Design of highly functional genome editors by modeling the universe of CRISPR-Cas sequences, bioRxiv (2024)p. 2024.04.22.590591.10. K. K. Yang, Z. Wu, F. H. Arnold, Machine-leaming-guided directed evolution for protein engineering. Nat. Methods 16, 687-694 (2019).11. H. Lu, D. J. Diaz, N. J. Czarnecki, C. Zhu, W. Kim, R. Shroff, D. J. Acosta, B. R. Alexander, H. O. Cole, Y. Zhang, N. A. Lynd, A. D. Ellington, H. S. Alper, Machine learning-aided engineering of hydrolases for PET depolymerization. Nature 604, 662- 667 (2022).12. N. Thomas, D. Belanger, C. Xu, H. Lee, K. Hirano, K. Iwai, V. Polic, K. D. Nyberg, K. G. Hoff, L. Frenz, C. A. Emrich, J. W. Kim, M. Chavarha, A. Ramanan, J. J. Agresti, L. J. Colwell, Engineering of highly active and diverse nuclease enzymes by combining machine learning and ultra-high-throughput screening, bioRxiv (2024)p.2024.03.21.585615.13. Z. Wu, S. B. J. Kan, R. D. Lewis, B. J. Wittmann, F. H. Arnold, Machine learning- assisted directed protein evolution with combinatorial libraries. Proc. Natl. Acad. Sci. U. S A. 116, 8852-8858 (2019).14. B. J. Wittmann, Y. Yue, F. H. Arnold, Informed training set design enables efficient machine learning-assisted directed protein evolution. Cell Syst 12, 1026-1045. e7 (2021).15. S. Biswas, G. Khimulya, E. C. Alley, K. M. Esvelt, G. M. Church, Low-N protein engineering with data-efficient deep learning. Nat. Methods 18, 389-396 (2021).16. L. Brenan, A. Andreev, O. Cohen, S. Pantel, A. Kamburov, D. Cacchiarelli, N. S. Persky, C. Zhu, M. Bagul, E. M. Goetz, A. B. Burgin, L. A. Garraway, G. Getz, T. S. Mikkelsen, F. Piccioni, D. E. Root, C. M. Johannessen, Phenotypic Characterization of a Comprehensive Set of MAPK1 / ERK2 Missense Mutants. Cell Rep. 17, 1171-1183 (2016).17. P. Notin, A. W. Kollasch, D. Ritter, L. vanNiekerk, S. Paul, H. Spinner, N. Rollins, A. Shaw, R. Weitzman, J. Frazer, M. Dias, D. Franceschi, R. Orenbuch, Y. Gal, D. S. Marks, ProteinGym: Large-Scale Benchmarks for Protein Design and Fitness Prediction. bioRxiv, doi: 10.1101 / 2023.12.07.570727 (2023).18. T. Hino, S. N. Omura, R. Nakagawa, T. Togashi, S. N. Takeda, T. Hiramoto, S. Tasaka, H. Hirano, T. Tokuyama, H. Uosaki, S. Ishiguro, M. Kagieva, H. Yamano, Y. Ozaki, D. Motooka, H. Mori, Y. Kirita, Y. Kise, Y. Itoh, S. Matoba, H. Aburatani, N. Yachie, T. Karvelis, V. Siksnys, T. Ohmori, A. Hoshino, O. Nureki, An AsCasl2f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell 186, 4920-4935. e23 (2023).19. H. K. Haddox, A. S. Dingens, J. D. Bloom, Experimental Estimation of the Effects of All Amino-Acid Mutations to HIV’s Envelope Protein on Viral Replication in Cell Culture. PLoS Pathog. 12, el006114 (2016).20. E. D. Kelsic, H. Chung, N. Cohen, J. Park, H. H. Wang, R. Kishony, RNA Structural Determinants of Optimal Codons Revealed by MAGE-Seq. Cell Syst 3, 563-57 l.e6 (2016).21. M. A. Stiffler, D. R. Hekstra, R. Ranganathan, Evolvability as a function of purifying selection in TEM- 1 p-lactamase. Cell 160, 882-892 (2015).22. C. J. Markin, D. A. Mokhtari, F. Sunden, M. J. Appel, E. Akiva, S. A. Longwell, C. Sabatti, D. Herschlag, P. M. Fordyce, Revealing enzyme functional architecture via high-throughput micro fluidic enzyme kinetics. Science 373 (2021).23. A. O. Giacomelli, X. Yang, R. E. Lintner, J. M. McFarland, M. Duby, J. Kim, T. P. Howard, D. Y. Takeda, S. H. Ly, E. Kim, H. S. Gannon, B. Hurhula, T. Sharpe, A. Goodale, B. Fritchman, S. Steelman, F. Vazquez, A. Tshemiak, A. J. Aguirre, J. G. Doench, F. Piccioni, C. W. M. Roberts, M. Meyerson, G. Getz, C. M. Johannessen, D. E. Root, W. C. Hahn, Mutational processes shape the landscape of TP53 mutations in human cancer. Nat. Genet. 50, 1381-1387 (2018).24. E. M. Jones, N. B. Lubock, A. J. Venkatakrishnan, J. Wang, A. M. Tseng, J. M. Paggi, N. R. Latorraca, D. Cancilla, M. Satyadi, J. E. Davis, M. M. Babu, R. O. Dror, S. Kosuri, Structural and functional characterization of G protein-coupled receptors with deep mutational scanning. Elife 9 (2020).25. M. B. Doud, J. D. Bloom, Accurate Measurement of the Effects of All Amino-Acid Mutations on Influenza Hemagglutinin. Viruses 8 (2016).26. J. M. Lee, J. Huddleston, M. B. Doud, K. A. Hooper, N. C. Wu, T. Bedford, J. D. Bloom, Deep mutational scanning of hemagglutinin helps predict evolutionary fates of human H3N2 influenza variants. Proc. Natl. Acad. Sci. U. S. A. 115, E8276-E8285 (2018).27. A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma, R. Fergus, Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl. Acad. Sci. U. S. A. 118 (2021).28. E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, G. M. Church, Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods 16, 1315-1322 (2019).29. A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y. Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger, D. Bhowmik, B. Rost, ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning. IEEE Trans. Pattern Anal. Mach. Intell. 44, 7112-7127 (2022).30. B. L. Hie, V. R. Shanker, D. Xu, T. U. J. Bruun, P. A. Weidenbacher, S. Tang, W. Wu, J. E. Pak, P. S. Kim, Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol. 42, 275-283 (2024).31. A. Baum, D. Ajithdoss, R. Copin, A. Zhou, K. Lanza, N. Negron, M. Ni, Y. Wei, K. Mohammadi, B. Musser, G. S. Atwal, A. Oyejide, Y. Goez-Gazi, J. Dutton, E. Clemmons, H. M. Staples, C. Bartley, B. Klaffke, K. Alfson, M. Gazi, O. Gonzalez, E. Dick Jr, R. Carrion Jr, L. Pessaint, M. Porto, A. Cook, R. Brown, V. Ali, J. Greenhouse, T. Taylor, H. Andersen, M. G. Lewis, N. Stahl, A. J. Murphy, G. D. Yancopoulos, C. A. Kyratsous, REGN-COV2 antibodies prevent and treat SARS-CoV-2 infection in rhesus macaques and hamsters. Science 370, 1110-1115 (2020).32. C.-L. Hsieh, J. A. Goldsmith, J. M. Schaub, A. M. DiVenere, H.-C. Kuo, K. Javanmardi, K. C. Le, D. Wrapp, A. G. Lee, Y. Liu, C.-W. Chou, P. O. Byrne, C. K. Hjorth, N. V. Johnson, J. Ludes-Meyers, A. W. Nguyen, J. Park, N. Wang, D. Amengor, J. J. Lavinder, G. C. Ippolito, J. A. Maynard, I. J. Finkelstein, J. S. McLellan, Structure-based design of prefiision-stabilized SARS-CoV-2 spikes. Science 369, 1501-1505 (2020).33. C. Xin, J. Yin, S. Yuan, L. Ou, M. Liu, W. Zhang, J. Hu, Comprehensive assessment of miniature CRISPR-Casl2f nucleases for gene disruption. Nat. Commun. 13, 5623 (2022).34. Z. Wu, Y. Zhang, H. Yu, D. Pan, Y. Wang, Y. Wang, F. Li, C. Liu, H. Nan, W. Chen, Q. Ji, Programmed genome editing by a miniature CRISPR-Casl2f nuclease. Nat. Chem. Biol. 17, 1132-1138 (2021).35. X. Xu, A. Chemparathy, L. Zeng, H. R. Kempton, S. Shang, M. Nakamura, L. S. Qi, Engineered miniature CRISPR-Cas system for mammalian genome regulation and editing. Mol. Cell 81, 4333-4345.e4 (2021).36. B. P. Kleinstiver, A. A. Sousa, R. T. Walton, Y. E. Tak, J. Y. Hsu, K. Clement, M. M. Welch, J. E. Homg, J. Malagon-Lopez, I. Scarfo, M. V. Maus, L. Pinello, M. J. Aryee, J. K. Joung, Engineered CRISPR-Cas 12a variants with increased activities and improved targeting ranges for gene, epigenetic and base editing. Nat. Biotechnol. 37, 276-282 (2019).37. X. Kong, H. Zhang, G. Li, Z. Wang, X. Kong, L. Wang, M. Xue, W. Zhang, Y. Wang, J. Lin, J. Zhou, X. Shen, Y. Wei, N. Zhong, W. Bai, Y. Yuan, L. Shi, Y. Zhou, H. Yang, Engineered CRISPR-OsCasl2fl and RhCasl2fl with robust activities and expanded target range for genome editing. Nat. Commun. 14, 2046 (2023).38. L. Zhang, J. A. Zuris, R. Viswanathan, J. N. Edelstein, R. Turk, B. Thommandru, H. T. Rube, S. E. Glenn, M. A. Collingwood, N. M. Bode, S. F. Beaudoin, S. Lele, S. N. Scott, K. M. Wasko, S. Sexton, C. M. Borges, M. S. Schubert, G. L. Kurgan, M. S. McNeill, C. A. Fernandez, V. E. Myer, R. A. Morgan, M. A. Behlke, C. A. Vakulskas, AsCasl2a ultra nuclease facilitates the rapid generation of therapeutic cell medicines. Nat. Commun. 12, 3908 (2021).39. D. Y. Kim, J. M. Lee, S. B. Moon, H. J. Chin, S. Park, Y. Lim, D. Kim, T. Koo, J.-H. Ko, Y.-S. Kim, Efficient CRISPR editing with a hypercompact Casl2fl and engineered guide RNAs delivered by adeno-associated virus. Nat. Biotechnol. 40, 94-102 (2022).40. P. Pausch, B. Al-Shayeb, E. Bisom-Rapp, C. A. Tsuchida, Z. Li, B. F. Cress, G. J. Knott, S. E. Jacobsen, J. F. Banfield, J. A. Doudna, CRISPR-CasO from huge phages is a hypercompact genome editor. Science 369, 333-337 (2020).41. J. L. Doman, S. Pandey, M. E. Neugebauer, M. An, J. R. Davis, P. B. Randolph, A. McElroy, X. D. Gao, A. Raguram, M. F. Richter, K. A. Everette, S. Banskota, K. Tian, Y. A. Tao, J. Tolar, M. J. Osborn, D. R. Liu, Phage-assisted evolution and protein engineering yield compact, efficient prime editors. Cell 186, 3983-4002. e26 (2023).42. M. T. N. Yamall, E. I. loannidi, C. Schmitt-Ulms, R. N. Krajeski, J. Lim, L. Villiger, W. Zhou, K. Jiang, S. K. Garushyants, N. Roberts, L. Zhang, C. A. Vakulskas, J. A. Walker, A. P. Kadina, A. E. Zepeda, K. Holden, H. Ma, J. Xie, G. Gao, L. Foquet, G. Bial, S. K. Donnelly, Y. Miyata, D. R. Radiloff, J. M. Henderson, A. Ujita, O. O. Abudayyeh, J. S. Gootenberg, Drag-and-drop genome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases. Nat. Biotechnol., 1-13 (2022).43. J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, A. Rives, Language models enable zeroshot prediction of the effects of mutations on protein function, bioRxiv (202 l)p. 2021.07.09.450648.44. A. Dousis, K. Ravichandran, E. M. Hobert, M. J. Moore, A. E. Rabideau, An engineered T7 RNA polymerase that produces mRNA free of immunostimulatory byproducts. Nat. Biotechnol. 41, 560-568 (2023).45. Z. J. Kartje, H. I. Janis, S. Mukhopadhyay, K. T. Gagnon, Revisiting T7 RNA polymerase transcription in vitro with the Broccoli RNA aptamer as a simplified real-time fluorescent reporter. J. Biol. Chem. 296, 100175 (2021).46. R. Chen, S. K. Wang, J. A. Belk, L. Amaya, Z. Li, A. Cardenas, B. T. Abe, C.-K. Chen, P. A. Wender, H. Y. Chang, Author Correction: Engineering circular RNA for enhanced protein production. Nat. Biotechnol. 41, 293 (2023).47. S. R. Johnson, X. Fu, S. Viknander, C. Goldin, S. Monaco, A. Zelezniak, K. K. Yang, Computational scoring and experimental evaluation of enzymes generated by neural networks. Nat. Biotechnol., doi: 10.1038 / s41587-024-02214-2 (2024).48. V. R. Shanker, T. U. J. Bruun, B. L. Hie, P. S. Kim, Unsupervised evolution of protein and antibody complexes with a structure-informed language model. Science 385, 46-53 (2024).49. Y. Serrano, A. Ciudad, A. Molina, Are Protein Language Models Compute Optimal?, arXiv [q-bio.BM] (2024). http: / / arxiv.org / abs / 2406.07249.50. X. Cheng, B. Chen, P. Li, J. Gong, J. Tang, L. Song, Training Compute-Optimal Protein Language Models, bioRxiv (2024)p. 2024.06.06.597716.51. B. Chen, X. Cheng, P. Li, Y.-A. Geng, J. Gong, S. Li, Z. Bei, X. Tan, B. Wang, X. Zeng, C. Liu, A. Zeng, Y. Dong, J. Tang, L. Song, xTrimoPGLM: Unified lOOB-Scale Pretrained Transformer for Deciphering the Language of Protein, bioRxiv (2024)p. 2023.07.05.547496.52. C. N. Bedbrook, K. K. Yang, J. E. Robinson, E. D. Mackey, V. Gradinaru, F. H. Arnold, Machine learning-guided channelrhodopsin engineering enables minimally invasive optogenetics. Nat. Methods 16, 1176-1184 (2019).53. J. W. Thornton, Resurrecting ancient genes: experimental analysis of extinct molecules. Nat. Rev. Genet. 5, 366-375 (2004).54. D. Ghosh, J. Cabrera, Enriched Random Forest for High Dimensional Genomic Data. IEEE / ACM Trans. Comput. Biol. Bioinform. 19, 2817-2828 (2022).55. A. Kirjner, J. Yim, R. Samusevich, S. Bracha, T. S. Jaakkola, R. Barzilay, I. R. Fiete, Improving protein optimization with smoothed fitness landscapes (2023). https: / / openreview.net / pdf / idmxlF2Zv8xO.56. K. Huang, R. Lopez, J.-C. Hutter, T. Kudo, A. Rios, A. Regev, Sequential Optimal Experimental Design of Perturbation Screens Guided by Multi-modal Priors, bioRxiv (2023)p. 2023.12.12.571389.57. P. M. Groth, M. H. Kerrn, L. Olsen, J. Salomon, W. Boomsma, Protein property prediction with uncertainties, arXiv [q-bio.BM] (2024). http: / / arxiv.org / abs / 2407.00002.58. J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, “ImageNet: A large-scale hierarchical image database” in 2009 IEEE Conference on Computer Vision and Pattern Recognition (IEEE, 2009), pp. 248-255.59. J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figumov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Zidek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T.Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P.Kohli, D. Hassabis, Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021).60. M. Sourisseau, D. J. P. Lawrence, M. C. Schwarz, C. H. Storrs, E. C. Veit, J. D. Bloom, M. J. Evans, Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. J. Virol. 93 (2019).61. B. L. Hie, V. R. Shanker, D. Xu, T. U. J. Bruun, P. A. Weidenbacher, S. Tang, W. Wu, J. E. Pak, P. S. Kim, Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol. 42, 275-283 (2024).62. T. Hino, S. N. Omura, R. Nakagawa, T. Togashi, S. N. Takeda, T. Hiramoto, S. Tasaka, H. Hirano, T. Tokuyama, H. Uosaki, S. Ishiguro, M. Kagieva, H. Yamano, Y. Ozaki, D. Motooka, H. Mori, Y. Kirita, Y. Kise, Y. Itoh, S. Matoba, H. Aburatani, N. Yachie, T. Karvelis, V. Siksnys, T. Ohmori, A. Hoshino, O. Nureki, An AsCasl2f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell 186, 4920-4935. e23 (2023).63. A. J. Greaney, T. N. Starr, C. O. Barnes, Y. Weisblum, F. Schmidt, M. Caskey, C. Gaebler, A. Cho, M. Agudelo, S. Finkin, Z. Wang, D. Poston, F. Muecksch, T. Hatziioannou, P. D. Bieniasz, D. F. Robbiani, M. C. Nussenzweig, P. J. Bjorkman, J. D. Bloom, Mapping mutations to the SARS-CoV-2 RBD that escape binding by different classes of antibodies. Nat. Commun. 12, 4196 (2021).64. M. Sourisseau, D. J. P. Lawrence, M. C. Schwarz, C. H. Storrs, E. C. Veit, J. D. Bloom, M. J. Evans, Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. J. Virol. 93 (2019).65. A. Elnaggar, H. Essam, W. Salah-Eldin, W. Moustafa, M. Elkerdawy, C. Rochereau, B. Rost, Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, arXiv [cs.LG] (2023). http: / / arxiv.org / abs / 2301.06568.66. E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, G. M. Church, Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods 16, 1315-1322 (2019).67. J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, A. Rives, Language models enable zero-shot prediction of the effects of mutations on protein function, bioRxiv (2021) p.2021.07.09.450648.68. L. Gieselmann, C. Kreer, M. S. Ercanoglu, N. Lehnen, M. Zehner, P. Schommers, J. Potthoff, H. Gruell, F. Klein, Effective high-throughput isolation of fully human antibodies targeting infectious pathogens. Nat. Protoc. 16, 3639-3671 (2021).69. X. Wang, S. Liu, Y. Sun, X. Yu, S. M. Lee, Q. Cheng, T. Wei, J. Gong, J. Robinson, D. Zhang, X. Lian, P. Basak, D. J. Siegwart, Preparation of selective organ-targeting (SORT) lipid nanoparticles (LNPs) using multiple technical methods for tissue-specific mRNA delivery. Nat. Protoc. 18, 265-291 (2023). 70. K. Clement, H. Rees, M. C. Canver, J. M. Gehrke, R. Farouni, J. Y. Hsu, M. A. Cole, D.R. Liu, J. K. Joung, D. E. Bauer, L. Pinello, CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol. 37, 224-226 (2019).71. J. Schymkowitz, J. Borg, F. Stricher, R. Nys, F. Rousseau, L. Serrano, The FoldX web server: an online force field. Nucleic Acids Res. 33, W382-8 (2005). 72. S. Bae, J. Park, J.-S. Kim, Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473— 1475 (2014).
Claims
What is claimed:
1. A composition comprising:(a) a target specific miniature nuclease comprising an amino acid sequence at least 70% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs. 2-60; and(b) a guide RNA (gRNA), wherein the target specific miniature nuclease and the gRNA form a complex capable of binding to, nicking, unwinding, and / or cleaving a DNA target.
2. The composition of claim 1, wherein the target specific miniature nuclease comprises one or more point mutations, relative to wild-type sequence SEQ ID NO:1.
3. The composition of claim 1, wherein the target specific miniature nuclease comprises a point mutation at position 333, relative to wild-type sequence SEQ ID NO: 1.
4. The composition of claim 3, wherein the target specific miniature nuclease comprises a point mutation at position 333, encoding a valine relative to wild-type sequence SEQ IDNO:1.
5. The composition of claim 1, wherein the target specific miniature nuclease comprises a point mutation at position 178, relative to wild-type sequence SEQ ID NO: 1.
6. The composition of claim 5, wherein the target specific miniature nuclease comprises a point mutation at position 178 encoding an arginine, relative to wild-type sequence SEQ IDNO:1.
7. The composition of claim 1, wherein the target specific miniature nuclease comprises a point mutation at position 454, relative to wild-type sequence SEQ ID NO: 1.
8. The composition of claim 7, wherein the target specific miniature nuclease comprises a point mutation at position 454 encoding a proline, relative to wild-type sequence SEQ ID9. The composition of claim 1, wherein the target specific miniature nuclease comprises a point mutation at positions 333, 178, and 454, relative to wild-type sequence SEQ ID NO: 1.
10. The composition of claim 1, wherein the target specific miniature nuclease comprises a point mutation at positions 333 encoding a valine, 178 encoding an arginine, and 454 encoding a proline, relative to wild-type sequence SEQ ID NO: 1.
11. A nucleic acid molecule encoding the target specific miniature nuclease of claim 1.
12. A vector comprising the nucleic acid molecule of claim 11.
13. A cell comprising the vector of claim 12.
14. The composition of claim 1, wherein the target specific miniature nuclease and the gRNA are packaged into one or more AAV vector(s) for in vitro or in vivo delivery.
15. The composition of claim 1, wherein the gRNA comprises a PCSK9 targeting guide.
16. A method of gene editing comprising adding the composition of claim 1 to a DNA target, wherein the DNA target is bound to, nicked, unwound, and or cleaved.
17. A method of gene editing comprising:(a) adding to a DNA target, a composition comprising a miniature CRISPR nuclease and a guide RNA; wherein the miniature CRISPR nuclease and the guide RNA bind to the DNA target; and(b) cleaving the DNA target, wherein one or more mutations of the miniature CRISPR nuclease amino acid sequence is located in a WED domain of the miniature CRISPR nuclease amino acid sequence, or in an a-helix of a REC domain of the miniature CRISPR nuclease amino acid sequence, or in a C-terminus of a RuvC domain of the miniature CRISPR nuclease amino acid sequence.
18. The method of claim 17, wherein the mutation located in the WED domain is at position 33,3 relative to wild-type sequence SEQ ID NO: 1.
19. The method of claim 17, wherein the mutation located in the a-helix of a REC domain is at position 178, relative to wild-type sequence SEQ ID NO: 1.
20. The method of claim 17, wherein the mutation located at the C-terminus of a RUVC domain is at position 454, relative to wild-type sequence SEQ ID NO:
1.
21. A non-natural miniature CRISPR nuclease comprising an amino acid sequence with one or more mutations relative to a wild-type miniature nuclease sequence of SEQ ID NO: 1, wherein the one or more mutations are located at amino acid positions 20, 40, 44, 45, 54, 78, 88, 140, 147, 155, 159, 163, 178, 186, 188, 189, 190, 191, 192, 194, 197, 333, 365, 425, 452, 454, and combinations thereof.
Citation Information
Patent Citations
Systems, methods, and compositions comprising miniature crispr nucleases for gene editing and programmable gene activation and inhibition
WO2022266298A1