An optimized mmeFz2 protein and its gene editing applications
By mutating the amino acid sequence of the MmeFz2 protein and fusing it with the HMG-D protein, combined with the ωRNA system, the gene editing efficiency of the MmeFz2 protein was optimized, solving the problem of low delivery efficiency of CRISPR-Cas nucleases in mammalian cells and achieving highly efficient gene editing results.
Patent Information
- Application Number
- CN202510950520.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Existing CRISPR-Cas nucleases, such as Cas9 and Cas12a, are large in size, resulting in low delivery efficiency during in vivo gene therapy with adeno-associated virus (AAV), especially in gene editing efficiency in mammalian cells.
Enhanced variants enMmeFz2 and evoMmeFz2 were designed by merging the MmeFz2 protein with amino acid sequence mutations and HMG-D protein fusion, and combining it with the ωRNA system. The structural modification and evolutionary optimization were carried out using AlphaFold3 and EVOLVEpro, which improved the editing activity of the protein-ωRNA-DNA interface.
It significantly improved gene editing efficiency in mammalian cells, achieving an average activity increase of 30.1-fold and 32.3-fold, and demonstrated the recovery effect of dystrophin in a humanized Duchenne muscular dystrophy mouse model, showcasing the potential of Fanzor2 as a programmable genome editing tool.
Smart Images

Figure CN120718887B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of gene editing, and particularly relates to an optimized MmeFz2 protein and gene editing application thereof. BACKGROUND
[0002] The advent of programmable genome editing technologies, especially the CRISPR-Cas system, has revolutionized modern biotechnology and medicine. CRISPR effectors, such as Cas9 and Cas12 nucleases, enable precise DNA manipulation across species, becoming a powerful tool for biological research, gene therapy, and agricultural breeding. However, widely used Cas nucleases (such as Cas9 and Cas12a) are typically over 1,000 amino acids in length, posing significant challenges for efficient delivery, especially for in vivo gene therapy via adeno-associated virus (AAV).
[0003] Recently, researchers have discovered and characterized compact CRISPR nucleases and their ancestral proteins from prokaryotes, including mini-Cas12 effectors (Cas12f, Cas12j, and Cas12n, ranging in length from 400 to 800 amino acids), as well as their ancestral proteins TnpB and IscB (~400 amino acids). In addition, Fanzor (Fz), an omega RNA-guided endonuclease in eukaryotes, is widely present in fungi, algae, protozoa, metazoans, acellular organisms, and certain large double-stranded DNA viruses, representing a unique class of RNA-programmed genome editing enzymes with significant evolutionary differences compared to prokaryotic systems. Notably, through phylogenetic and structural studies, the newly discovered prokaryotic obligate mobile element-guided activity (OMEGA) protein TnpB is considered an evolutionary precursor of eukaryotic Fz proteins and prokaryotic CRISPR-Cas12 nucleases. Fanzor proteins are mainly divided into two major categories: Fz1 and Fz2. Fz1 proteins range in length from 600 to 900 amino acids, while Fz2 proteins are more compact (~480 amino acids) and more structurally similar to TnpB. The compact structure of Fz2 makes it an ideal candidate for viral delivery of therapeutic genome editing. However, the natural Fz2 system has very low activity in mammalian cells (<1% editing efficiency). SUMMARY
[0004] Based on the above technical problems, the application provides an optimized MmeFz2 protein and gene editing application thereof.
[0005] The application specifically adopts the following technical solutions:
[0006] An optimized MmeFz2 protein is obtained by mutating a wild-type MmeFz2 protein with an amino acid sequence as shown in SEQ ID NO. 1 as follows:
[0007] 1) mutating C at position 69, E at position 305 and E at position 326 of the wild-type MmeFz2 protein into K, N and Q respectively, and naming it as enMmeFz2.
[0008] 2) mutating E at position 178, E at position 305 and E at position 418 of the wild-type MmeFz2 protein into H, S and R respectively, and naming it as evoMmeFz2.
[0009] The present application develops enhanced variants (enMmeFz2 and evoMmeFz2) of MmeFz2, which achieve 30.1-fold and 32.3-fold average activity improvement on 38 human genome targets compared with the wild-type system.
[0010] Further, the C-terminal of the optimized MmeFz2 protein is further fused with an HMG-D protein.
[0011] Based on the same inventive concept, the present application further provides an optimized MmeFz2-omegaRNA system comprising an omegaRNA and the optimized MmeFz2 protein, wherein the omegaRNA is used to guide the optimized MmeFz2 protein to recognize a target site.
[0012] Further, the sequence of the omegaRNA is shown in SEQ ID NO. 3.
[0013] Based on the same inventive concept, the present application further provides a polynucleotide encoding the optimized MmeFz2-omegaRNA system.
[0014] Based on the same inventive concept, the present application further provides a recombinant vector comprising the polynucleotide.
[0015] Based on the same inventive concept, the present application further provides a cell comprising the recombinant vector.
[0016] Based on the same inventive concept, the present application further provides applications of the optimized MmeFz2-omegaRNA system, the polynucleotide, the recombinant vector or the cell in gene editing.
[0017] Based on the same inventive concept, the present application further provides applications of the optimized MmeFz2-omegaRNA system, the polynucleotide, the recombinant vector or the cell in preparing a preparation for gene editing.
[0018] Compared with the prior art, the present application has the following beneficial effects:
[0019] This invention designs MmeFz2 as a highly efficient genome editor by optimizing its omega RNA scaffold and protein sequence. Using AlphaFold3, we first identified structural defects in the wild-type omega RNA (WT-omega RNA) and rationally redesigned a truncated variant with 30% reduced length, significantly improving editing activity. At the same time, we performed structure-guided mutations on the protein-omega RNA-DNA interface guided by AlphaFold3 information, and the validation results were integrated into a PLM-guided iterative evolution pipeline, using EVOLVEpro to predict new functional mutations. Two evolved MmeFz2 variants, enMmeFz2 from structure-guided design and evoMmeFz2 from PLM-guided evolution, showed convergent improvements in editing efficiency. Further fusion of a single-stranded DNA binding domain (HMG-D) to evoMmeFz2 enhanced editing efficiency, exceeding the performance of the engineered TnpB system. With its compact size, we packaged the optimized evoMmeFz2 system in a single AAV vector and demonstrated strong dystrophin restoration in a humanized Duchenne muscular dystrophy (DMD) mouse model, achieving therapeutic levels of in vivo editing. Our work establishes the position of Fanzor2 as a programmable genome editing tool and proposes a combination engineering strategy for RNA-guided nucleases, highlighting the transformative potential of combining artificial intelligence with structural biology to advance precision medicine. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1Validation of the effect of engineering MmeFz2 proteins with AlphaFold3 or EVOLVEpro to enhance their activity in mammalian cells. a: Schematic of the evolutionary engineering of MmeFz2 proteins using AlphaFold3 or EVOLVEpro. b: Comparison of the gene editing efficiency mediated by MmeFz2 variants at the B2M site in HEK293T cells. 141 mutations in MmeFz2 residues that could potentially enhance the interaction of MmeFz2-omegaRNA with the target DNA were predicted (predicted by AlphaFold3). The top 15 mutations (>1.2-fold: C69K, C69R, Q158K, Q158R, S185N, E305N, E309Q, E309R, Y316R, E326N, E326Q, E326K, L356Q, S377N, S377Q) were labeled with red squares for further validation. c: Comparison of the average gene editing efficiency of the 15 high-efficiency mutants at nine endogenous sites in HEK293T cells. Each data point represents the average gene editing efficiency for each target site. d: Comparison of the gene editing efficiency mediated by C69K or C69R in combination with S185N, E305N, E309Q, E309R, Y316R, E326Q, L356Q, S377N, and S377Q at the B2M site in HEK293T cells. The top eight mutations (C69K, S185N, E305N, E309R, Y316R, E326Q, L356Q, S377Q) were selected for further optimization. e: Comparison of the gene editing efficiency mediated by C69K in combination with S185N, E305N, E309R, Y316R, L356Q, S377Q at the B2M site in HEK293T cells. The triple mutant C69K+E305N+E326Q (named en-Pro) was selected for further study due to its highest editing activity, which is marked with a red triangle. f: MmeFz2 was engineered through three rounds of EVOLVEpro. The high-efficiency mutants selected in each round (E178G, Y316T, E326A, E178S, E178H, E178Q, E178N, E305S, E305D, E418R) are marked with red squares and further validated. A total of 60 mutants were evaluated for their fold improvement in genome editing efficiency at the B2M site in HEK293T cells from the first and third rounds. g: Comparison of the average gene editing efficiency of the ten high-efficiency mutants at nine endogenous sites in HEK293T cells. Each data point represents the average gene editing efficiency for each target site. Two mutations in the first round (E178S and E178H) and three mutations in the third round (E305S, E305D, and E418R) are marked with red triangles for further validation.h: The gene editing efficiency of E178S, E178H, E305S, E305D, and E418R combinations at the B2M locus in HEK293T cells was compared. The triple mutant E178H+E305S+E418R (named evo-Pro) was selected for further study due to its highest editing activity and was labeled with a red triangle. i: The combination of en-omega RNA with two engineered protein variants (en-Pro and evo-Pro) further improved the gene editing efficiency of the B2M locus in HEK293T cells. j: Structural basis of activity-enhancing mutations C69K, E326Q, and E178H. Data are shown as mean ± s.e.m (n = 3). Fold change represents the ratio of editing efficiency of the protein variant to that of WT-MmeFz2.
[0021] Figure 2 To further optimize MmeFz2 architecture by ssDBD fusion and verify its gene editing efficiency at endogenous loci in mammalian cells. a: The gene editing efficiency of evoMmeFz2 fused with five ssDBDs and three exonucleases at their N- or C-termini at five genomic loci in HEK293T cells. b: The average gene editing efficiency of 16 variants mediated by evoMmeFz2 at five endogenous loci in HEK293T cells was compared. evoMmeFz2-HMG-D variant was selected for further study due to its highest editing efficiency and was labeled with a red triangle. Each data point represents the average gene editing efficiency at each target locus. c: The gene editing efficiency of WT-MmeFz2, enMmeFz2, enMmeFz2-HMG-D, evoMmeFz2, and evoMmeFz2-HMG-D at 38 endogenous loci in HEK293T cells was compared. d: The average gene editing efficiency of WT-MmeFz2, enMmeFz2, enMmeFz2-HMG-D, evoMmeFz2, and evoMmeFz2-HMG-D at 38 genomic loci was compared. Each data point represents the average indel frequency at each target locus, calculated based on three independent experiments. Error bars and p values were from these 38 data points. p values were determined by unpaired two-tailed Student's t test. Data are shown as mean ± s.e.m (n = 3).
[0022] Figure 3To assess the specificity of evoMmeFz2-mediated genome editing in mammalian cells, evoMmeFz2, evoMmeFz2-HMG-D, and WT-MmeFz2 were analyzed for on- and off-target activity at six genomic loci (KRAS-guide1, CXCR4-guide2, DYRK1A-guide1, B2M-guide1, B2M-guide2, and B2M-guide5). Off-target sites identified by Cas-OFFinder contained 2 to 5 mismatches. Values are expressed as the mean of three independent biological replicates (n = 3).
[0023] Figure 4 evoMmeFz2 and evoMmeFz2-HMG-D restored dystrophin expression in DMDAmE5051, KIhE50 / Y mice after a single AAV injection. a: Schematic of in vivo injection of a single AAV9-evoMmeFz2 or evoMmeFz2-HMG-D construct into the tibialis anterior (TA) muscle of the right leg of 3-week-old DMDAmE5051, KIhE50 / Y mice. b: Schematic of the exon skipping strategy used by evoMmeFz2 or evoMmeFz2-HMG-D to restore the DMD transcript open reading frame (ORF). c: RT-PCR products from muscle tissue of DMDAmE5051, KIhE50 / Y mice were analyzed by gel electrophoresis. d: RNA deletion insertion editing events were analyzed by targeted amplicon sequencing three weeks after intramuscular injection. e: Dystrophin expression restoration three weeks after TA injection of evoMmeFz2 or evoMmeFz2-HMG-D was detected by dystrophin immunofluorescence staining. Dystrophin and actin staining are shown in green and purple, respectively. Scale bar: 100 pm. f: Quantitative analysis of Dys+ fibers in TA muscle cross-sections. g: Dystrophin and paxillin expression in TA muscle three weeks after AAV9-evoMmeFz2, AAV9-evoMmeFz2-HMG-D, or saline injection was evaluated by Western blot analysis. Paxillin levels were used as an internal loading control. h: The percentage of restored dystrophin was quantified by gray intensity analysis. Values and error bars are expressed as the mean ± standard error (n = 3 independent biological replicates). p values were determined by unpaired two-tailed t test. Each dot in panels d, f, and h represents one mouse.
[0024] Figure 5To engineer the MmeFz2 protein using AlphaFold3 or EVOLVEpro. A: The gene editing efficiency of 15 high-performance MmeFz2 mutants (C69K, C69R, Q158K, Q158R, S185N, E305N, E309Q, E309R, Y316R, E326N, E326Q, E326K, L356Q, S377N, S377Q) selected through AlphaFold3 was evaluated in HEK293T cells at nine endogenous gene loci (B2M-guide4, CXCR4-guide1, CXCR4-guide2, DYRK1A-guide5, KRAS-guide1, VEGFA-guide1, EMX1-guide2, EMX1-guide6, DYRK1A-guide1). Data are presented as mean ± standard error (n = 3). B: Comparison of gene editing efficiencies at nine endogenous gene loci (KRAS-guide1, CXCR4-guide1 / 2, EMX1-guide2, DYRK1A-guide5, B2M-guide4, VEGFA-guide1, EMX1-guide6, DYRK1A-guide1) in HEK293T cells using 10 different mutants (E178G, Y316T, E326A, E178S, E178H, E178Q, E178N, E305S, E305D, E418R) selected via EVOLVEpro. Data are presented as mean ± standard error (n = 3).
[0025] Figure 6To compare the gene editing efficiencies of WT-MmeFz2, enMmeFz2, enMmeFz2-HMG-D, evoMmeFz2, evoMmeFz2-HMG-D, and the IS200 / IS605 transposon-derived TnpB gene editing system in mammalian cells. A: In HEK293T cells, targeted editing was performed on eight endogenous gene loci (CXCR4-guide1, B2M-guide1, EMX1-guide5, and DYRK1A-guide1 to guide5), and the gene editing efficiencies of WT-MmeFz2, enMmeFz2, enMmeFz2-HMG-D, evoMmeFz2, evoMmeFz2-HMG-D, and the IS200 / IS605 transposon-derived IsTfu1 TnpB system were compared. Data are presented as mean ± standard error (n = 3). B: In HEK293T cells, targeted editing was performed on eight additional endogenous gene sites (KRAS-guide1, CXCR4-guide1, B2M-guide6, DYRK1A-guide2, guide4, guide7, guide8, and DMD-guide5). The gene editing efficiency of the five MmeFz2 variants and the IsDge10 TnpB system derived from the IS200 / IS605 transposon was compared. Data are presented as mean ± standard error (n = 3).
[0026] Figure 7 The structure of the MmeFz2-ωRNA-dsDNA ternary complex predicted by AlphaFold3 is shown below. A: The structure of the complex formed by the binding of the MmeFz2 ribonucleoprotein (RNP) to the B2M target site double-stranded DNA, as predicted by AlphaFold3. B: The overall structure of the MmeFz2 nuclease. C: A schematic diagram of the domain composition of the MmeFz2 nuclease. Detailed Implementation
[0027] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments, but this should not be construed as limiting the invention. Unless otherwise specified, the technical means used in the following embodiments are conventional means well known to those skilled in the art, and the materials, reagents, etc. used in the following embodiments are commercially available unless otherwise specified.
[0028] The methods and materials involved in the following embodiments of the present invention:
[0029] 1. Structural prediction using AlphaFold3
[0030] The amino acid sequence of the wild-type MmeFz2 protein (SEQ ID NO.1), its corresponding / engineered full-length ωRNA (containing a 20-nucleotide B2M guide sequence; the wild-type coding sequence is shown in SEQ ID NO.2, and the optimized coding sequence is shown in SEQ ID NO.3), and a 40 bp endogenous B2M target DNA sequence were submitted to the AlphaFold3 online server (https: / / golgi.sandbox.google.com / ) to predict the structure of the ternary complex. The obtained structure was then refined using COOT. Molecular visualization images were generated using CueMol software (http: / / www.cuemol.org).
[0031] SEQ ID NO.1: MKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAK TNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLREGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPL SHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKEIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKL CSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCEVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAA.
[0032] SEQ ID NO. 2: TTCGGGTTCGATTCTATCCCCAGGGCTCGAATGCATTTTTGTCACAGATTTTGCCAATGCAAGATCTGGGGGCAAGAATGTCTCCGGGTGAAAAGAGTCAG.
[0033] SEQ ID NO. 3: TTCGGGTTCGATTCTATCCCCAGGGCTCGAATGCAGAATCAATATTCTGTCTCCGGGTGAAAAGAGTCAG.
[0034] 2. Enzyme activity enhancement strategies based on EVOLVEpro
[0035] This invention utilizes EVOLVEpro (https: / / github.com / mat10d / EvolvePro) for protein engineering optimization. This platform is based on a few-shot active learning framework, combining structural information provided by AlphaFold3 structure prediction results for modeling and optimization. First, based on the AlphaFold3-predicted MmeFz2–ωRNA–DNA ternary complex structure model, 141 single-point mutants were rationally designed at their interaction interfaces as the initial training set for the EVOLVEpro regression model. This regression model uses sequence embedding features extracted by the Protein Language Model (PLM; ESM-2 15B) as input and a random forest model as the top-level regressor to predict the relative activity of the mutants. The top 20 mutants with the best predicted activity in each round were experimentally validated at B2M gene sites, and their indel efficiency was used to update the model, thereby further optimizing the prediction of enzyme activity trends.
[0036] 3. Construction of plasmid vectors
[0037] Plasmid cloning employed standard molecular cloning techniques. Wild-type MmeFz2 with optimized human codons and the ωRNA scaffold were synthesized by Huajin Biotechnology Co., Ltd. PCR amplification was performed using Phanta Max Super-Fidelity DNA polymerase (Vazyme) during the construction of the MmeFz2-ωRNA plasmid, followed by fragment assembly using the Basic Seamless cloning and Assembly Kit (TransGen). Each plasmid contains a CBh promoter, a 3×FLAG tag, an SV40 nuclear localization signal, MmeFz2 protein, ribosomal NLS, bGH poly(A) signal, a U6 promoter, and ωRNA. Target oligonucleotides for the ωRNA were ordered from Qingke Biotechnology Co., Ltd., and ligated to a BsaI-digested backbone vector using T4 DNA ligase (Thermo) after annealing. The backbone vector was synthesized using the Cas12i editing system plasmid (from the literature "An engineeredxCas12i with high activity, high specificity, and broad PAM range"; this plasmid contains the CBh promoter to initiate Cas12i protein expression, the U6 promoter to initiate sgRNA expression, and the CMV promoter to initiate red fluorescent protein expression, thus enabling fluorescence flow cytometry sorting). The nucleotide sequences encoding the wild-type MmeFz2 protein and the wild-type ωRNA scaffold were obtained by seamless cloning. The nucleotide sequences encoding the Cas12i protein (only the Cas12i protein in the expression cassette was replaced, retaining the NLS sequences directly connected upstream and downstream: the 5' nuclear localization signal sequence SV40NLS and the 3' nuclear localization signal sequence nucleoplasmin NLS) and the nucleotide sequence encoding the sgRNA backbone were replaced within the Cas12i editing system plasmid to obtain the wild-type MmeFz2-ωRNA editing system backbone plasmid. The ωRNA spacer sequences (i.e., guide sequences) used in this invention are listed in Table 1.
[0038] Table 1: Target sites and related information
[0039]
[0040] The MmeFz2 mutants include: MmeFz2_C69K, MmeFz2_C69R, MmeFz2_Q158K, MmeFz2_Q158R, MmeFz2_S185N, MmeFz2_E305N, MmeFz2_E309Q, MmeFz2_E309R, MmeFz2_Y316R, MmeFz2_E326N, MmeFz2_E326Q, MmeFz2_E326K, MmeFz2_L356Q, MmeFz2_S377N, MmeFz2_S377Q, MmeFz2_E178G, MmeFz2_E178H, MmeFz2_E178S, MmeFz2_E178Q, MmeFz2_E178N, MmeFz2_E305S, MmeFz2_E305D, MmeFz2_Y316T, MmeFz2_E326A, MmeFz2_E418R, MmeFz2_C69K+S185N, MmeFz2_C69K+E305N, MmeFz2_C69K+E309Q, MmeFz2_C69K+E309R, MmeFz2_C69K+Y316R, MmeFz2_C69K+E326N, MmeFz2_C69K+E326Q, MmeFz2_C69K+E326K, MmeFz2_C69K+L356Q, MmeFz2_C69K+S377N, MmeFz2_C69K+S377Q, MmeFz2_C69R+S185N, MmeFz2_C69R+E305N, MmeFz2_C69R+E309Q, MmeFz2_C69R+E309R, MmeFz2_C69R+Y316R, MmeFz2_C69R+E326N, MmeFz2_C69R+E326Q, MmeFz2_C69R+E326K, MmeFz2_C69R+L356Q, MmeFz2_C69R+S377N, MmeFz2_C69R+S377Q, MmeFz2_C69K+S185N+E305N, MmeFz2_C69K+E305N+E309R, MmeFz2_C69K+E305N+Y316R, MmeFz2_C69K+E305N+E326Q (enMmeFz2)、MmeFz2_C69K+E305N+L356Q、MmeFz2_C69K+E305N+S377Q、MmeFz2_C69K+S185N+E309R、MmeFz2_C69K+E309R+Y316R、MmeFz2_C69K+E309R+E326Q、MmeFz2_C69K+E309R+L356Q、MmeFz2_C69K+E309R+S377Q、MmeFz2_C69K+S185N+Y316R、MmeFz2_C69K+S185N+E326Q、MmeFz2_C69K+S185N+L356Q、MmeFz2_C69K+S185N+S377Q、MmeFz2_C69K+Y316R+E326Q、MmeFz2_C69K+Y316R+L356Q、MmeFz2_C69K+Y316R+S377Q、MmeFz2_C69K+E326Q+L356Q、MmeFz2_C69K+E326Q+S377Q、MmeFz2_C69K+L356Q+S377Q、MmeFz2_E178H+E305D、MmeFz2_E178H+E305S、MmeFz2_E178H+E418R、MmeFz2_E178S+E305D、MmeFz2_E178S+E305S、MmeFz2_E178S+E418R、MmeFz2_E305D+E418R、MmeFz2_E305S+E418R、MmeFz2_E178H+E305D+E418R、MmeFz2_E178H+E305S+E418R(evoMmeFz2)、MmeFz2_E178S+E305D+E418R、MmeFz2_E178S+E305S+E418R、HMG-D-evoMmeFz2、HMGN1-evoMmeFz2、HMGB1-evoMmeFz2、H1G-evoMmeFz2、Sso7d-evoMmeFz2、T5-evoMmeFz2、TREX1-evoMmeFz2、TREX2-evoMmeFz2、evoMmeFz2-HMG-D、evoMmeFz2-HMGN1、evoMmeFz2-HMGB1、evoMmeFz2-H1G、evoMmeFz2-Sso7d、evoMmeFz2-T5、evoMmeFz2-TREX1、evoMmeFz2-TREX2、enMmeFz2-HMG-D。
[0041] In this invention, amino acid residues are represented by single letters. For example, MmeFz2_C69K (or simply C69K) is formed by replacing the amino acid C at position 69 of wild-type MmeFz2 with K; MmeFz2_C69K+E305N+E326Q is formed by replacing the amino acid C at position 69 of wild-type MmeFz2 with K, replacing the E at position 305 with N, and replacing the E at position 326 with Q.
[0042] The amino acid sequence of HMG-D-evoMmeFz2 is:
[0043] The amino acid sequence of HMGN1-evoMmeFz2 is as follows: MPKRKVSSAEGAAKEEPKRRSARLSAKPPAKVEAKPKKAAAKDKSSDKKVQTKGKRGAKGKQAEVANQETKEDLPAENGETKTEESPASDEAGEKEAKSDSGGSSGGSSGSETPGTSESATPESSGGSSGGSMKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAA。
[0044] The amino acid sequence of HMGB1-evoMmeFz2 is as follows: GKGDPKKPRGKMSSYAFFVQTCREEHKKKHPDASVNFSEFSKKCSERWKTMSAKEKGKFEDMAKADKARYEREMKTYIPPKGESGGSSGGSSGSETPGTSESATPESSGGSSGGSMKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAA。
[0045] The amino acid sequence of H1G-evoMmeFz2 is:。
[0046] The amino acid sequence of Sso7d-evoMmeFz2 is:。
[0047] The amino acid sequence of T5-evoMmeFz2 is:。
[0048] The amino acid sequence of TREX1-evoMmeFz2 is: MQTLIFFDMEATGLPFSQPKVTELCLLAVHRCALESPPTSQGPPPTVPPPPRVVDKLSLCVAPGKACSPAASEITGLSTAVLAAHGRQCFDDNLANLLLAFLRRQPQPWCLVAHNGDRYDFPLLQAELAMLGLTSALDGAFCVDSITALKALERASSPSEHGPRKSYSLGSIYTRLYGQSPPDSHTAEGDVLALLSICQWRPQALLRWVDAHARPFGTIRPMYGVTASARTKPRPSAVTTTASGGSSGGSSGSETPGTSESATPESSGGSSGGSMKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAA。
[0049] The amino acid sequence of TREX2-evoMmeFz2 is: MSEAPRAETFVFLDLEATGLPSVEPEIAELSLFAVHRSSLENPEHDESGALVLPRVLDKLTLCMCPERPFTAKASEITGLSSEGLARCRKAGFDGAVVRTLQAFLSRQAGPICLVAHNGFDYDFPLLCAELRRLGARLPRDTVCLDTLPALRGLDRAHSHGTAAAGAQGYSLGSLFHRYFRAEPSAAHSAEGDVHTLLLIFLHRAAELLAWADEQARGWAHIEPMYLPPDDPSLEASGGSSGGSSGSETPGTSESATPESSGGSSGGSMKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAA。
[0050] The amino acid sequence of evoMmeFz2-HMG-D is as follows: MKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAASGGSSGGSSGSETPGTSESATPESSGGSSGGSMSDKPKRPLSAYMLWLNSARESIKRENPGIKVTEVAKRGGELWRAMKDKSEWEAKAAKAKDDYDRAVKEFEANGGSSAANGGGAKKRAKPAKKVAKKSKKEESDEDDDDESE。
[0051] The amino acid sequence of evoMmeFz2-HMGN1 is as follows: MKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAASGGSSGGSSGSETPGTSESATPESSGGSSGGSMPKRKVSSAEGAAKEEPKRRSARLSAKPPAKVEAKPKKAAAKDKSSDKKVQTKGKRGAKGKQAEVANQETKEDLPAENGETKTEESPASDEAGEKEAKSD。
[0052] The amino acid sequence of evoMmeFz2-HMGB1 is as follows: MKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAASGGSSGGSSGSETPGTSESATPESSGGSSGGSGKGDPKKPRGKMSSYAFFVQTCREEHKKKHPDASVNFSEFSKKCSERWKTMSAKEKGKFEDMAKADKARYEREMKTYIPPKGE。
[0053] The amino acid sequence of evoMmeFz2-H1G is:。
[0054] evoMmeFz2-Sso7d's amino acid sequence is:。
[0055] evoMmeFz2-T5's amino acid sequence is:。
[0056] The amino acid sequence of evoMmeFz2-TREX1 is: MKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAASGGSSGGSSGSETPGTSESATPESSGGSSGGSMQTLIFFDMEATGLPFSQPKVTELCLLAVHRCALESPPTSQGPPPTVPPPPRVVDKLSLCVAPGKACSPAASEITGLSTAVLAAHGRQCFDDNLANLLLAFLRRQPQPWCLVAHNGDRYDFPLLQAELAMLGLTSALDGAFCVDSITALKALERASSPSEHGPRKSYSLGSIYTRLYGQSPPDSHTAEGDVLALLSICQWRPQALLRWVDAHARPFGTIRPMYGVTASARTKPRPSAVTTTA。
[0057] The amino acid sequence of evoMmeFz2-TREX2 is: MKRKREQMTLWKAAFVNGQETFKSWIDKARMLELNCDVSSASSTHYSDLNLKTKCAKTDDKFMCNYSVCIRPTSKQKRTLNQMLKVSNYAYNWCNYLVKEKDFKPKQFDLQRIVAKTNSTDVPAEYRLPGDDWFFDNKMSSIKLTACKNFCTMYKSTQTNQKKTKVDLRNKDIVQLRHGSFEVQSKYVRLLTEKDIPGERIRQSRIALMPDSFSKSKKDWKERFLRLSKNVSKIPPLSHDMKVCKRPNGKFILQISCDPICTRQIQVQTSDSICSIDPGGRTFATCYDPSNIKTFQIGPEADKKSIIHEFHNKIDYVHRLLSHAQEKKQTQAVQDRIGQLKKLHLKLKTYVDDVHLKLCSYLVKNYKLVVLGKISVSSIVRKDRPNHLAKSANRDLLCWQHYRFRQRLLHRVRGTDCRVIIQDERYTSKTCGNCGEKNNKLGGKETFTCESCNYKTHRDVNGARNILCKYLGLFPFAASGGSSGGSSGSETPGTSESATPESSGGSSGGSMSEAPRAETFVFLDLEATGLPSVEPEIAELSLFAVHRSSLENPEHDESGALVLPRVLDKLTLCMCPERPFTAKASEITGLSSEGLARCRKAGFDGAVVRTLQAFLSRQAGPICLVAHNGFDYDFPLLCAELRRLGARLPRDTVCLDTLPALRGLDRAHSHGTAAAGAQGYSLGSLFHRYFRAEPSAAHSAEGDVHTLLLIFLHRAAELLAWADEQARGWAHIEPMYLPPDDPSLEA。
[0058] enMmeFz2-HMG-D:
[0059] 4. Cell culture, transfection, and flow cytometry analysis
[0060] Human HEK293T cells were cultured in DMEM (Gibco) medium containing 10% fetal bovine serum, 1% non-essential amino acids, and 1% penicillin-streptomycin-glutamine. All cell types were cultured at 37°C and 5% CO2 and passaged every 2 days until 80% confluence was reached. To screen for protein and ωRNA variants at endogenous sites, 2 × 10⁶ cells were cultured. 5HEK293T cells were seeded into 24-well plates. At approximately 80% confluence, 1500 ng of the MmeFz2 expression plasmid was added to each well for transfection with 1500 ng of plasmid and 3 μL of polyethyleneimine (PEI) at a 1:2 ratio of DNA (µg) to PEI (µL). After 60–72 hours, the transfected cells were digested with 0.05% trypsin (Gibco) for fluorescence-activated cell sorting (FACS), and mCherry-positive cells were used for genomic DNA extraction.
[0061] 5. DNA extraction and insertion / deletion efficiency analysis
[0062] Approximately 10,000 flow-cytometry-sorted cells were lysed with 20 μL of lysis buffer (formulation: 10 mM Tris-HCl, pH 8.0; 0.05% SDS; 20 μg / ml proteinase K). After incubation at 55°C for 30 min, the proteinase was inactivated by heating at 95°C for 5 min. Then, 1 μL of the lysis product was used as a template for PCR amplification.
[0063] To evaluate the gene editing efficiency of different MmeFz2 variants in vivo, genomic DNA was extracted from the muscle tissue of gene-edited mice that were successfully born after treatment with AAV9-evoMmeFz2-DMD ωRNA. DNA extraction was performed using the TIANamp Genomic DNA Extraction Kit (TIANGEN).
[0064] For targeted amplicon sequencing, nested PCR amplification was performed using Phanta Max high-fidelity DNA polymerase (Vazyme, P505), with amplified fragment lengths of 200–250 bp, using barcode-tagged primers. After mixing, the PCR products were purified using a gel extraction kit (Omega). The amplicon library was constructed using the VAHTS Universal DNA Library Prep Kit (Vazyme), subsequently purified, and then subjected to 150 bp paired-end sequencing on an Illumina NovaSeq 6000 platform.
[0065] Sequencing data were first demultiplexed using Cutadapt (v2.8), and then analyzed using CRISPResso2 software to quantify indel (insertion / deletion) efficiency. Target site sequences and primer information are detailed in Table 1.
[0066] 6. Animals
[0067] All mouse experiments were approved by the Biomedical Research Ethics Committee of Shanghai Huida Gene Co., Ltd., China. Mice were housed in a controlled barrier environment with a 12-hour light / dark cycle, maintained at a temperature of 18°C to 23°C and a relative humidity of 40% to 60%, and had free access to food and water. DMDΔmE5051 and KIhE50 / Y mice were constructed using the CRISPR / Cas9 system on a C57BL / 6J background. Given that Duchenne muscular dystrophy (DMD) is the most common sex-linked lethal genetic disease in humans, male mice were used in this invention.
[0068] 7. Off-target site analysis predicted by Cas-OFFinder
[0069] To assess the specificity of the MmeFz2 system, we predicted potential off-target sites using the CRISPR RGEN tool (Cas-OFFinder, http: / / www.rgenome.net / cas-offinder / ). Since no TAM option for the MmeFz2 protein was provided, a 23-nucleotide sequence containing a 20-nucleotide target sequence and a 3-nucleotide TAM (5′-TAG) was input into the tool. Mismatches were limited to five nucleotides, and the PAM sequence was set to 5′-NNN, consistent with SpRY Cas9. Potential off-target sites with one or more mismatches were selected for primer design using the online Primer-BLAST tool (https: / / www.ncbi.nlm.nih.gov / tools / primer-blast). The first 10 predicted potential off-target sites were amplified by PCR and sequenced to assess the genome editing specificity of WT-MmeFz2, evoMmeFz2, and evoMmeFz2-HMG-D. All predicted off-target site sequences and their corresponding primers are listed in Table 2.
[0070] Table 2: Off-target detection related site information
[0071]
[0072] 8. Intramuscular injection
[0073] This invention uses adeno-associated virus type 9 (AAV9) as the delivery vector. After sequencing verification, the evoMmeFz2 and evoMmeFz2-HMG-D plasmids and their associated ωRNAs were co-transfected with pHelper, pRepCap, and the target gene expression plasmid (GOI) into HEK293T cells to prepare AAV9 viral particles. After three days of culture following transfection, AAV was collected and purified using iodixanol density gradient centrifugation.
[0074] In in vivo experiments, 3-week-old DMDΔmE50-51 and KI^hE50 / Y mice were anesthetized and injected with 50 μL of AAV9 formulation (1 × 10¹² vg) or an equal volume of physiological saline as a control into the tibialis anterior (TA) muscle. Three weeks after injection, the mice were sacrificed under anesthesia, and the tibialis anterior muscle was sampled in sections for targeted analysis: distal tissue was used to detect gene editing and exon skipping efficiency, mid-section tissue was used for immunoblotting to detect dystrophin protein expression, and proximal tissue was used to assess the distribution and expression level of dystrophin using immunofluorescence.
[0075] 9. Western blot analysis
[0076] Muscle tissue samples were lysed with RIPA lysis buffer, and the supernatant was collected for protein quantification. Protein concentration was determined using the Pierce BCA Protein Assay Kit (Thermo Fisher Scientific), and all samples were diluted to the same concentration with ultrapure water. 10 μg of total protein was loaded into each well and separated by SDS-polyacrylamide gel electrophoresis (SDS-PAGE). After electrophoresis, the protein was transferred to a PVDF membrane and wet-transferred at 350 mA for 3.5 hours. The membrane was then blocked with TBST buffer containing 5% skim milk at room temperature for 1 hour. The membrane strips were incubated overnight at 4°C with primary antibody in TBST buffer containing 0.05% BSA. The primary antibodies used included anti-dystrophin antibody (Sigma, D8168) and anti-vinculin antibody (CST, 13901S). The membrane strips were washed three times with TBST buffer for 5 minutes each time on a shaker. HRP-labeled secondary antibody was then added, and the membrane was incubated at room temperature for 1 hour. Finally, the target protein bands were detected using chemiluminescence reagent (Invitrogen).
[0077] 10. Histology and immunofluorescence
[0078] In histological analysis, paraffin-embedded tissue sections were first dewaxed in xylene, then sequentially dehydrated with 100%, 95%, 75%, and 50% ethanol, followed by washing with distilled water. The samples were then stained with hematoxylin and eosin (H&E) and 0.1% picrosirius red for routine histological observation of tissue structure and collagen deposition. During picrosirius red staining, sections were stained in 0.1% picrosirius red for 1 hour, then rinsed twice with acidified water and vigorously shaken to remove excess water. The sections were then dehydrated three times with anhydrous ethanol, cleared in xylene, and mounted in neutral resin. In immunofluorescence experiments, tissue samples were embedded in an OCT complex and rapidly frozen in liquid nitrogen. Serial frozen sections (10 μm thick) were fixed at 37°C for 2 hours, permeabilized with 0.4% Triton-X / PBS solution for 30 minutes, and then blocked with 10% goat serum at room temperature for 1 hour. Sections were then incubated overnight at 4°C with primary antibodies including anti-dystrophin antibody (Abcam, ab15277) and anti-spectrin antibody (Millipore, MAB1622). After thorough washing with PBS, compatible fluorescently labeled secondary antibodies (Jackson ImmunoResearch: donkey anti-rabbit IgG labeled with Alexa Fluor 488, or donkey anti-mouse IgG labeled with Alexa Fluor 647) and DAPI nuclear staining were added, and the sections were incubated at room temperature for 3 hours. After washing with PBS for 15 minutes, the sections were mounted with Fluoromount-G mounting medium. All images were acquired using a Nikon C2 fluorescence microscope. The proportion of dystrophin-positive (Dys⁺) myofibrils was calculated by counting the number of Dys⁺ fibers among all spectrin-positive myofibrils. The following specific examples illustrate the optimized MmeFz2 protein of the present invention and its applications.
[0079] Example 1: Structural-guided engineering of MmeFz2 using AlphaFold3 or EVOLVEpro
[0080] Using the AlphaFold3-predicted MmeFz2-ωRNA-DNA ternary complex structural model, we propose that MmeFz2 consists of four distinct regions: a RuvC nuclease domain (residues 1-61, 265-429, and 457-478) containing a zinc finger motif (ZF, 429-457); a recognition domain (REC, 72-177); and a wedge-shaped domain (WED, 61-72 and 177-265). Figure 7We designed 141 single-point mutants located at the protein-nucleic acid interface to enhance the interaction between MmeFz2 and nucleic acids or promote conformational flexibility. Figure 1 b). In the initial screening of the B2M site, we identified 15 single-point mutants (C69K, C69R, Q158K, Q158R, S185N, E305N, E309Q, E309R, Y316R, E326N, E326Q, E326K, L356Q, S377N, S377Q), whose insertion / deletion efficiency was more than 1.2 times higher than that of wild-type MmeFz2. Figure 1 b). In validation at nine genomic loci, C69K and C69R outperformed wild-type MmeFz2 and other single-point mutants ( Figure 1 c, Figure 5 A). Subsequently, we combined C69K or C69R with secondary mutations (S185N, E305N, E309Q, E309R, Y316R, E326N, E326Q, E326K, L356Q, S377N, S377Q) to evaluate their synergistic effects. Notably, in the double-mutation combination, C69K showed superior synergistic potential compared to C69R. Figure 1 d). The third round of engineering produced a triple mutant C69K+E305N+E326Q (named en-Pro), whose insertion / deletion efficiency was 2.1 times higher than that of wild-type MmeFz2. Figure 1 e).
[0081] To overcome the limitations of rational design, we utilized insertion and deletion efficiency data from the initial 141 rationally designed single-point mutants. Figure 1 a, b), and developed the active learning framework EVOLVEpro. Through three rounds of iterative prediction and experimental validation, we identified ten variants (E178G, Y316T, E326A, E178H, E178S, E178Q, E178N, E305S, E305D, E418R), which improved the insertion / deletion efficiency at the B2M site by ≥1.1 times. Figure 1 f). Multi-site validation showed that efficiency gradually improved with increasing iterative evolution rounds, with E178H, E178S, E305D, E305S, and E418R showing continuous improvement. Figure 1 g, Figure 5 B). By combining these mutations, we generated a triple mutant E178H+E305S+E418R (named evo-Pro), which showed a 2.0-fold increase in insertion / deletion efficiency compared to the wild-type MmeFz2. Figure 1 h).
[0082] The optimized en-ωRNA, when combined with engineered en-Pro and evo-Pro, showed improvements of 76.0-fold and 66.1-fold, respectively, compared to the original MmeFz2-ωRNA system. Figure 1 i). Structural predictions of AlphaFold3 suggest that the beneficial mutations derived from both methods may stabilize the key interaction between MmeFz2 and ωRNA. Figure 1 Specifically, the C69K mutation introduces a novel interaction between the WED domain and the phosphate backbone of the PK region in ωRNA (j). Figure 1 j). The E326Q and E178H mutations were predicted to enhance the interaction between the catalytic RuvC domain and the spacer sugar phosphate backbone. Figure 1 These results collectively demonstrate that the combination of AlphaFold3-guided rational design and EVOLVEpro-driven small-scale learning of protein evolution synergistically enhances the gene editing efficiency of the MmeFz2-ωRNA system. In conclusion, the EVOLVEpro-based laboratory model method for proteins exhibits competitive performance with methods utilizing AlphaFold3 in protein engineering.
[0083] Example 2: Efficient genome editing in human cells using an engineered MmeFz2-ωRNA system.
[0084] Previous studies have shown that fusing non-sequence-specific DNA-binding domains (ssDBDs) or exonuclease modules with RNA-guided nucleases can significantly improve genome editing efficiency. To further enhance the genome editing capabilities of the MmeFz2-ωRNA system, we systematically evaluated five ssDBDs (HMG-D, HMGN1, HMGB1, H1G, and Sso7d) and three exonucleases (TREX1, TREX2, and T5 exonucleases), which were fused to the N-terminus or C-terminus of evoMmeFz2. The study found that fusion of ssDBDs with evoMmeFz2 generally improved genome editing efficiency, while fusion of all exonucleases reduced editing activity. Figure 2 a, b). Among all the ssDBDs tested, the 112-amino acid HMG-D domain from the Drosophila high-mobility family of chromosomal proteins performed best. Figure 2 a, b). The C-terminal fusion of evoMmeFz2-HMG-D achieved a maximum insertion / deletion efficiency of 81.5% at five endogenous sites, an average improvement of 1.2 times compared to evoMmeFz2. Figure 2 a, b).
[0085] Next, we comprehensively evaluated 38 endogenous sites covering 9 genes to assess the genome editing efficiency of engineered MmeFz2-ωRNA systems. The results showed that all engineered MmeFz2-ωRNA systems exhibited significantly improved genome editing efficiency compared to the wild-type system. Figure 2 (c, d). Specifically, the editing efficiency of enMmeFz2-HMG-D, evoMmeFz2-HMG-D, enMmeFz2, and evoMmeFz2 is 40.2 times, 37.1 times, 30.1 times, and 32.3 times higher than that of WT-MmeFz2, respectively. Figure 2 c, d). Notably, the variants containing HMG-D (enMmeFz2-HMG-D and evoMmeFz2-HMG-D) showed slightly higher average insertion / deletion efficiencies at all test sites than the non-fusion enzyme. Figure 2 d). To compare performance with known compact editors, we compared these MmeFz2 variants with two previously characterized IS200 / IS605 transposon-encoded TnpB nucleases, IsTfu1 and IsDge10. Although all MmeFz2 variants recognize 5′-TAG target adjacent motifs (TAMs), IsTfu1 and IsDge10 recognize different TAM sequences, namely 5′-TGAT and 5′-TTAT, respectively. For direct comparison, we selected endogenous sites that are compatible with all systems and have overlapping TAM sequences. The results showed that all engineered MmeFz2 variants had significantly higher genome editing efficiency than IsTfu1 TnpB, and comparable efficiency to IsDge10 TnpB. Figure 6 (A, B). In summary, these findings establish the engineered MmeFz2-ωRNA system as an effective mammalian genome editing tool, offering broad target compatibility and significantly improved efficiency compared to WT-MmeFz2 and the established TnpB system.
[0086] Example 3: Evaluation of the genome editing specificity of the engineered MmeFz2-ωRNA system
[0087] To comprehensively evaluate the specificity profile of evoMmeFz2 and its HMG-D fusion variants in human cells, we first used the Cas-OFFinder algorithm to identify potential ωRNA-dependent off-target sites at six endogenous sites (including KRAS, CXCR4, DYRK1A, and B2M). Targeted amplicon sequencing analysis revealed that evoMmeFz2-HMG-D exhibited higher targeting efficiency than evoMmeFz2 at almost all six sites. Figure 3The engineered variants evoMmeFz2 and evoMmeFz2-HMG-D showed fewer but still detectable off-target edits, with evoMmeFz2 showing off-target edit rates of 6.15% and 7.15% at the DYRK1A-guide1 OT4 and B2M-guide2 OT7 sites, respectively, while evoMmeFz2-HMG-D showed off-target edit rates of 5.81% and 6.04% at the same sites, respectively. Figure 3 These results indicate that the engineered MmeFz2 variant maintained reasonable genome editing specificity in human cells, and that HMG-D fusion did not exacerbate nonspecific editing.
[0088] Example 4: Engineered MmeFz2 variant restored dystrophin expression in humanized Duchenne mouse model.
[0089] Duchenne muscular dystrophy (DMD) is a fatal muscle disease caused by a deficiency of dystrophin, affecting 1 in 3,500 to 5,000 male newborns. It is caused by multiple pathogenic mutations in the human X-chromosome-linked DMD gene. Most clinically documented DMD pathogenic mutations occur in the “hotspot” region of the DMD gene, which covers exons 45–55 and encodes the central rod domain of the protein. Mutations in the DMD gene typically involve the deletion of one or more exons, disrupting the open reading frame (ORF) and resulting in nonfunctional truncated dystrophin, leading to severe muscle degeneration. ORF restoration can be achieved by introducing small insertions or deletions at the splice acceptor site (SAS) or splice donor site (SDS), i.e., removing an exon by exon skipping. Due to its compact structure, the engineered MmeFz2-ωRNA system can be packaged into a single rAAV vector, making it a promising gene-editing tool for correcting DMD in vivo.
[0090] Previously, we generated and validated a genetically humanized DMD mouse model (named DMDΔmE5051, KIhE50 / Y) by knocking in the human exon 50 sequence to replace exons 50 and 51 of the mouse DMD gene. This model has human-specific exon deletion mutations. Figure 4 a). To evaluate the in vivo activities of evoMmeFz2 and evoMmeFz2-HMG-D, we first screened for effective target sequences in HEK293T cells and found that DMD-guide3 exhibited the highest insertion / deletion efficiency at the SDS site in exon 50 (a). Figure 2c). Subsequently, we injected AAV9 particles carrying the engineered MmeFz2-ωRNA system expression element into the tibialis anterior (TA) muscle of 3-week-old male DMDΔmE5051,KIhE50 / Y mice. Figure 4 a). Three weeks after injection, we collected TA muscle samples for further analysis. Figure 4 a). The study found that by disrupting the SDS of exon 50, the open reading frame of the DMD can be restored, thereby enabling correct splicing of exons 49 to 52 or exons 50 to 52 even in cases of exon 50 skipping or rearranging. Figure 4 b). PCR-based detection confirmed successful skipping of exon 50 of human DMD following evoMmeFz2 or evoMmeFz2-HMG-D-mediated insertion / deletion formation, a fact verified by gel electrophoresis. Figure 4 c). RT-PCR analysis of mRNA extracted from whole muscle showed that in the evoMmeFz2-HMG-D group, the misalignment efficiency was 9.38±0.82%, the in-frame efficiency was higher at 7.73±0.46%, and the skipping efficiency was also improved at 9.62±0.43%, compared with the evoMmeFz2 group ( Figure 4 d). Western blotting and immunostaining results further confirmed that evoMmeFz2 or evoMmeFz2-HMG-D can effectively restore the expression of dystrophin. Figure 4 e, g). Furthermore, compared to the evoMmeFz2 group, evoMmeFz2-HMG-D treatment significantly increased the number of dystrophin-positive myofibrils and improved protein expression levels (e, g). Figure 4 (f, g, h). In summary, our results demonstrate that the engineered MmeFz2-ωRNA system not only exhibits effectiveness in mammalian cells but also possesses the ability to effectively restore disease symptoms through in vivo single AAV delivery.
[0091] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.
Claims
1. An optimized MmeFz2 protein, characterized in that, It is obtained by performing any of the following mutations on the wild-type MmeFz2 protein with the amino acid sequence shown in SEQ ID NO.1: 1) Mutate the C at position 69 of the wild-type MmeFz2 protein to K, the E at position 305 to N, and the E at position 326 to Q; or 2) Mutate the E at position 178 of the wild-type MmeFz2 protein to H, the E at position 305 to S, and the E at position 418 to R.
2. The optimized MmeFz2 protein according to claim 1, characterized in that, The optimized MmeFz2 protein also has an HMG-D protein fused to its C-terminus.
3. An optimized MmeFz2-ωRNA system, characterized in that, The protein comprises ωRNA and the optimized MmeFz2 protein according to any one of claims 1 to 2, wherein the ωRNA is used to guide the optimized MmeFz2 protein to recognize a target site.
4. The optimized MmeFz2-ωRNA system according to claim 3, characterized in that, The sequence of the ωRNA is shown in SEQ ID NO.
3.
5. A polynucleotide encoding the optimized MmeFz2-ωRNA system of claim 4.
6. A recombinant vector comprising the polynucleotide of claim 5.
7. Cells comprising the vector of claim 6.
8. The use of the optimized MmeFz2-ωRNA system of any one of claims 3 to 4, the polynucleotide of claim 5, the recombinant vector of claim 6, or the cell of claim 7 in gene editing.
9. The use of the optimized MmeFz2-ωRNA system of any one of claims 3 to 4, the polynucleotide of claim 5, the recombinant vector of claim 6, or the cell of claim 7 in the preparation of gene-editing formulations.