COME nuclease variant and application thereof
By integrating the oligomerization domain in CRISPR nucleases and regulating the three-dimensional spatial arrangement of proteins, the trade-off between efficiency and specificity of CRISPR nucleases is solved, achieving high efficiency and high specificity while improving, and expanding the application scope of the CRISPR system.
Patent Information
- Application Number
- CN202510452836.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
The existing CRISPR nuclease technology has a trade-off between improving efficiency and specificity, and lacks universality and combinatorial compatibility across systems, making it difficult to further develop in precise gene editing.
By integrating the oligomerization domain, especially the gp41 trimerization and GCN4 dimerization domains in CRISPR nucleases, the three-dimensional spatial arrangement of proteins is regulated, and the structure of the CRISPR protein complex is optimized, so as to achieve high efficiency and high specificity while improving.
The targeting activity and specificity of CRISPR nucleases have been significantly improved, and they have shown enhancing effects across multiple CRISPR systems, and synergistically integrate with existing methods to achieve super-additive effects.
Smart Images

Figure CN120366270A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene editing, and particularly relates to a COME nuclease variant and its application. Background Art
[0002] The CRISPR-Cas system, as a programmable genome editing tool, has revolutionized the landscape of biological research, making it possible to precisely manipulate genetic material in various organisms. These systems are derived from bacterial adaptive immune systems, and nucleases such as SpCas9, FrCas9, and Cas12a orthologs have been widely used in various applications ranging from functional genomics to gene therapy.
[0003] Currently, the main technical routes to improve the performance of CRISPR nucleases include: structure-guided engineering: reducing non-specific DNA binding by modifying key residues to improve specificity, such as SpCas9-HF1 (Kleinstiver et al., 2016, Nature) and eSpCas9 (Slaymaker et al., 2016, Science), etc.; directed evolution: screening for variants with higher specificity through protein directed evolution, such as Sniper-Cas9 (Lee et al., 2018, Nature Communications) and xCas9 (Hu et al., 2018, Nature); rational design: rational design based on structural and functional analysis, such as SuperFiCas9 (Bravo et al., 2022, Nature).
[0004] However, these existing methods have several key drawbacks: efficiency-specificity trade-off dilemma: existing high-specificity CRISPR variants usually come at the cost of sacrificing 70%-95% of the targeting activity, while high-efficiency variants tend to produce unwanted off-target effects; system-specific improvement: existing methods usually only optimize for specific CRISPR systems (such as SpCas9), lacking universality among evolutionarily different CRISPR systems; mechanism limitations: existing methods mainly focus on modifying the DNA binding interface or changing catalytic properties, while ignoring the impact of protein spatial conformation on target recognition; poor combinatorial compatibility: the combined effects of existing high-fidelity variants are limited and often cannot achieve superadditive effects. The above problems make it difficult for CRISPR tools to simultaneously achieve high efficiency and high specificity, severely restricting the further development of CRISPR technology in precise gene editing, especially in medical applications. Summary of the Invention
[0005] The object of the present invention is to overcome the deficiencies of the prior art and provide a Controlled OligoMerization Engineering (COME) nuclease variant and its application. By manipulating the three-dimensional spatial arrangement of the CRISPR protein complex rather than only modifying its intrinsic catalytic properties, new optimization paths can be created; it can break through the traditional trade-off between efficiency and specificity and achieve a significant improvement in both simultaneously; it can be generally applied to a general enhancement platform for a variety of evolutionarily different CRISPR systems; and it can achieve synergistic effects with existing optimization methods (such as directed evolution), generating a comprehensive optimization effect that exceeds the sum of the effects of individual methods.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] In the first aspect, the present invention provides a method of controlled oligomerization engineering, which integrates an oligomerization domain into a CRISPR nuclease.
[0008] As a preferred embodiment of the method of the present invention, the oligomerization domain includes an oligomerization domain that is temperature-sensitive, small molecule-regulated, or tissue-specifically expressed.
[0009] As a preferred embodiment of the method of the present invention, the oligomerization domain includes a gp41 trimerization domain and / or a GCN4 dimerization domain.
[0010] As a further preferred embodiment of the method of the present invention, the amino acid sequence of the gp41 trimerization domain is as shown in SEQ ID NO:1; the amino acid sequence of the GCN4 dimerization domain is as shown in SEQ ID NO:2.
[0011] As a preferred embodiment of the method of the present invention, the CRISPR nuclease is any one of Cas7, Cas8, Cas9, Cas10, Cas11, Cas12, Cas14, IscB, and TnpB.
[0012] As a further preferred embodiment of the method of the present invention, the CRISPR nuclease is any one of SpCas9, FrCas9, LbCas12a, LtCas12a, Cas12f, and Cas12k.
[0013] In the second aspect, the present invention provides a SpCas9 variant, in which an oligomerization domain is integrated into the amino acid sequence of the SpCas9 protein.
[0014] As a preferred embodiment of the SpCas9 variant of the present invention, the integration site is behind any one of the amino acids at positions 1, 25, 54, 62, 170, 203, 213, 231, 249, 257, 308, 355, 365, 400, 532, 573, 584, 674, 719, 765, 768, 776, 782, 790, 808, 819, 826, 831, 846, 868, 890, 910, 924, 945, 975, 1010, 1020, 1033, 1050, 1051, 1055, 1059, 1068, 1072, 1102, 1110, 1120, 1130, 1227, 1246, 1248, 1252, 1260, 1276, 1290, 1300, 1302, 1327, 1332, 1340, 1368 of the SpCas9 protein, or the amino acids at positions 1048-1063 of the SpCas9 protein are integrally replaced.
[0015] In a third aspect, the present invention provides a FrCas9 variant, and an oligomerization domain is integrated into the amino acid sequence of the FrCas9 protein.
[0016] As a preferred embodiment of the FrCas9 variant of the present invention, the integration site is behind any one of the amino acids at positions 1, 40, 51, 55, 108, 112, 120, 125, 170, 192, 200, 215, 234, 250, 265, 270, 275, 290, 295, 315, 320, 325, 363, 379, 428, 539, 566, 610, 618, 690, 702, 707, 727, 755, 761, 800, 830, 882, 980, 1002, 1043, 1055, 1095, 1111, 1117, 1134, 1184, 1188, 1224, 1248, 1321, 1329, 1333, 1347, 1372 of the FrCas9 or eFrCas9 protein.
[0017] As a further preferred embodiment of the FrCas9 variant of the present invention, compared with FrCas9, the eFrCas9 protein has different amino acid sequences only at positions 1103 and 732, and the amino acid sequences at positions 1103 and 732 are as shown in SEQ ID NO:22.
[0018] In a fourth aspect, the present invention provides an LbCas12a variant, and an oligomerization domain is integrated into the amino acid sequence of the LbCas12a protein.
[0019] As a preferred embodiment of the LbCas12a variant of the present invention, the integration site is behind any one of the amino acids at positions 1, 11, 85, 171, 269, 270, 335, 354, 371, 406, 478, 572, 625, 653, 731, 805, 810, 825, 836, 865, 928, 965, 967, 972, 982, 1040, 1074, 1087, 1109, 1117, 1119, 1120, 1121, 1138, 1142, 1143, 1148, 1156, 1158, 1168, 1171, 1175, 1211, 1228 of the LbCas12a protein.
[0020] In a fifth aspect, the present invention provides an LtCas12a variant, in which an oligomerization domain is integrated into the amino acid sequence of the LtCas12a protein.
[0021] As a preferred embodiment of the LtCas12a variant of the present invention, the integration site is behind any one of the amino acids at positions 12, 85, 330, 354, 402, 572, 652, 687, 895, 992, 1029, 1039, 1049, 1108, 1141, 1150, 1173, 1181, 1186, 1208, 1220, 1222, 1232, 1235 of the LtCas12a protein.
[0022] As the SpCas9 variant of the present invention 、 The FrCas9 variant 、 The FrCas9 variant 、 The LbCas12a variant 、 In a further preferred embodiment of the LtCas12a variant, the oligomerization domain includes an oligomerization domain with temperature sensitivity, small molecule regulation, or tissue-specific expression.
[0023] As the SpCas9 variant of the present invention 、 The FrCas9 variant 、 The FrCas9 variant 、 The LbCas12a variant 、 In a further preferred embodiment of the LtCas12a variant, the oligomerization domain includes a gp41 trimerization domain and / or a GCN4 dimerization domain.
[0024] As the SpCas9 variant of the present invention 、 The FrCas9 variant 、 The FrCas9 variant 、 The LbCas12a variant、 A further preferred embodiment of the LtCas12a variant, wherein the amino acid sequence of the gp41 trimerization domain is as shown in SEQ ID NO: 1; the amino acid sequence of the GCN4 dimerization domain is as shown in SEQ ID NO: 2.
[0025] In a sixth aspect, the present invention provides a method for constructing a COME variant, comprising the following steps:
[0026] (1) Select an oligomerization domain insertion site based on structural analysis, optimize the oligomerization domain coding sequence, avoid potential RNA secondary structures, and design PCR primers containing restriction sites;
[0027] (2) PCR amplify the target region, integrate the oligomerization domain by digestion with restriction enzymes or Gibson assembly, transform Escherichia coli, and verify positive clones of the COME variant by colony sequencing;
[0028] (3) Clone the COME variant into a mammalian expression vector containing a CMV promoter, add a tag sequence, a nuclear localization signal, FLAG or HA, and purify the plasmid;
[0029] (4) Confirm purity and integrity by Nanodrop and electrophoresis, confirm the vector structure using a restriction map, verify the full-length COME variant sequence by sequencing, and confirm expression by protein electrophoresis.
[0030] In a seventh aspect, the present invention provides a method for evaluating the performance of a COME variant, comprising the following steps:
[0031] (1) Transfect the COME variant and the sgRNA plasmid into HEK293 cells, use amplicon deep sequencing analysis to evaluate the editing efficiency, verify at multiple genomic loci, and screen the best-performing variants for in-depth characterization;
[0032] (2) Perform genome-wide off-target analysis using GUIDE-seq, select multiple representative genomic loci, process the data analysis using a standard pipeline, and calculate the specificity ratio;
[0033] (3) Verify oligomerization using the Split-GFP complementation assay, quantitatively characterize protein interactions using FRET analysis, compare with a flexible linker control group, and perform structural simulation analysis to study the effect of oligomerization on conformation.
[0034] In an eighth aspect, the present invention relates to the SpCas9 variant 、 the FrCas9 variant 、 the FrCas9 variant 、 the LbCas12a variant、 The described LtCas12a variants and their derivatives are used in any of the following fields:
[0035] i. Genome editing, transcriptional activation or inhibition, base editing or primer editing;
[0036] ii. Preparing drugs for gene therapy, prevention and / or diagnosis;
[0037] iii. NNTA PAM recognition;
[0038] v. Genome imaging, epigenetic modification or nucleic acid detection;
[0039] vi. Biosensors.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] 1. The present invention significantly improves both the editing efficiency and specificity simultaneously: The SpCas9-308gp41 variant shows a 2.1-fold increase in targeting activity while eliminating detectable off-target effects; the SpCas9-584gp41 achieves a 41.8-fold increase in specificity while maintaining strong targeting activity; the LtCas12a-85gp41 exhibits up to a 4.5-fold enhancement in targeting activity with no significant increase in off-targets.
[0042] 2. The present invention realizes the universality across multiple CRISPR systems: It shows significant enhancement in four phylogenetically different CRISPR systems, namely SpCas9, FrCas9, LbCas12a, and LtCas12a; a conserved enhancement site (position 85) is found in Cas12a homologs despite only 33% sequence identity; the method of the present invention can be widely applied to other CRISPR systems, showing its potential as a general optimization platform.
[0043] 3. The present invention has significant advantages compared with existing methods: Compared with existing high-fidelity variants such as SpCas9-HF1, eSpCas9, SpRYCas9, SuperFiCas9, Sniper-Cas9, and xCas9, the COME-SpCas9 variant shows an advantage of 1.3 - 242 times in terms of targeting activity; existing high-fidelity variants usually sacrifice 70% - 95% of the targeting activity to obtain a moderate increase in specificity, while the COME variant maintains or improves the targeting activity while significantly enhancing the specificity; the COME technology does not require precise atomic-level knowledge of the structural information and can be more easily extended to newly discovered CRISPR systems.
[0044] 4. The present invention has a super synergy effect: Combining COME with directed evolution generates UltraFrCas9, achieving a up to 138-fold enhancement in targeting activity and a specificity ratio of 30,046; The combined effect far exceeds the additive effect of the two individual methods, demonstrating the orthogonality of the two optimization strategies; This synergy effect provides a new paradigm for future CRISPR system optimization.
[0045] 5. The present invention provides a new mechanism for regulating the function of CRISPR nucleases: Proven by FRET and Split-GFP complementation experiments, the enhancement stems from specific oligomerization-mediated spatial arrangement changes; Compared with the control using only a flexible linker of the same length, the oligomerization variants exhibit significantly higher editing efficiency. Brief Description of the Drawings
[0046] Figure 1 It is a schematic diagram of the SpCas9 protein structure and the potential insertion sites of 62 oligomerization domains; It shows the distribution of the SpCas9 domains and potential insertion sites on the secondary structures of α-helices and β-sheets; It shows the RuvC domain (blue), REC domain (green), HNH domain (orange), and PAM interaction domain (pink).
[0047] Figure 2 It is a comparison of the editing efficiencies of the COME-enhanced SpCas9 variants at each target site; In the figure, A: indel frequency analysis of the gp41-SpCas9 variant at the RNF2 and HEK293 site 2; B: indel frequency analysis of the GCN4-SpCas9 variant at the RNF2 and HEK293 site 2; C: editing efficiency of the best variant at 12 different genomic sites; D: statistical analysis of the specificity of the best variant at 12 different genomic sites.
[0048] Figure 3 It is the comparison of off-target of the gp41-COME enhanced variant with the wild type shown by GUIDE-seq analysis.
[0049] Figure 4 It is the comparison of off-target of the GCN4-COME enhanced variant with the wild type shown by GUIDE-seq analysis.
[0050] Figure 5Characterization of the molecular mechanism for enhanced oligomerization-mediated; in the figure, A: Schematic diagram of the SpCas9 construct, showing the insertion positions of gp41 or GS-linker; B: Comparison of TIDE analysis at the RNF2 and HEK293 site 2 loci; C: Design of the Split-GFP complementation experiment; D: Flow cytometry quantification showing the percentage of fluorescent cells of each protein combination; E-J: FRET analysis diagrams, including donor, acceptor concentration-dependent emission spectra and FRET signals of the mixed proteins.
[0051] Figure 6 Extension of the oligomerization strategy to FrCas9 and its synergy with directed evolution; in the figure, A: Relative editing efficiency of the GCN4-FrCas9 variant; B: Relative editing efficiency of the gp41-FrCas9 variant; C: GUIDE-seq analysis at representative genomic targets; D: Optimization of on-target read counts of the FrCas9 variant relative to wild-type FrCas9; E: Average targeting enhancement of UltraFrCas9 at multiple genomic loci; F: Specificity ratio of each FrCas9 variant.
[0052] Figure 7 Enhancement of Cas12a homologs by oligomerization domains; in the figure, A: Domain architecture of LbCas12a and potential insertion sites; B: Activity analysis of the gp41-modified LbCas12a variant; C: Amplicon sequencing results at the CDKN2A, DYRK1A, and RUNX1 loci; D: GUIDE-seq analysis results; E: Relative targeting reads of the optimized LbCas12a variant; F: Domain architecture of LtCas12a and potential insertion sites; G: Activity analysis of the gp41-modified LtCas12a variant; H: Amplicon sequencing results at the CDKN2A, DYRK1A, and RUNX1 loci; I: GUIDE-seq analysis showing the targeting and off-target distribution of the LtCas12a variant.
[0053] Figure 8 Comprehensive performance comparison of the COME technology with existing high-fidelity SpCas9 variants; in the figure, A: GUIDE-seq analysis at the RNF2 locus; B-E: Targeted GUIDE-seq read counts at the RNF2, CIITA, FANCF, and PCSK9 loci; F-I: Specificity ratios at the above loci. Detailed implementation manners
[0054] To better illustrate the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments. Those skilled in the art should understand that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0055] Unless otherwise specified, the test methods used in the examples are conventional methods; the materials, reagents, etc. used, unless otherwise specified, can be obtained from commercial sources.
[0056] The gene locus information for SpCas9 editing in the examples is shown in Table 1:
[0057] Table 1 Gene Locus Information
[0058]
[0059]
[0060] The gene locus information for FrCas9 editing is shown in Table 2:
[0061] Table 2 Gene Locus Information
[0062]
[0063] The gene locus information for LbCas12a and LtCas12a editing is shown in Table 3:
[0064] Table 3 Gene Locus Information
[0065]
[0066] The sgRNA sequences for GUIDE-seq analysis in the examples are shown in Tables 4 and 5:
[0067] Table 4 Gene Locus Information of SpCas9 Protein Series
[0068]
[0069]
[0070] Table 5 Gene Locus Information of FrCas9 Protein Series
[0071] Gene Name sgRNA Sequence DYRK1A-2 TGTTACCACTGGATTGAGGTCTGGTA EMX1-2 GGTCATTGTCATGTCCAGTTGTGGTA FANCF-3 CGGGGTCCCAGGTGCTGACGTAGGTA GRIN2B-3 TGGCCATGCGACCCTCTTCATAGGTA GRIN2B-9 AAAAAGGAAGGAGTTCTTTGTAGGTA HEK293-site2-2 TGAGCTAACTGTGACAGCATGTGGTA RNF2-1 ATGAGTTACAACGAACACCTCAGGTA RUNX1-3 AACCTCACCCTTGGAGAGTTCAGGTA RUNX1-5 TTAGTATAAGCGCTCAATCAATGGTA
[0072] The methods for synthetic fragments and insertion design in the examples are:
[0073] Precisely insert the oligomerization domain into specific sites of the CRISPR nuclease. This process requires designing PCR primers, which contain the oligomerization domain coding sequence and the sequences on both sides of the insertion site of the target enzyme. The specific design is as follows: 5'-[pre-sequence of SpCas9]-[oligomerization domain coding sequence]-[post-sequence of SpCas9]-3'. Using plasmid pST1374 (product number 13426, brand addgene) as the backbone, construct the Cas protein variant in the example into the plasmid, so that the final expression vector of the nuclease variant contains: high-efficiency expression controlled by the CMV promoter, nuclear localization signal (NLS) sequence to ensure nuclear localization, stop codon and poly(A) signal, and appropriate selection markers (such as antibiotic resistance).
[0074] Example 1: Use the gp41 trimerization domain to enhance the specificity and activity of SpCas9
[0075] 1. Selection of oligomerization domain
[0076] Select two oligomerization domains: the trimerization domain of HIV-1 glycoprotein 41 (gp41) (HIV-1 gp41 residues 628-683, sequence is KNEQELLELDKWASL, see SEQ ID NO:1); the dimerization domain of yeast GCN4 (GCN4 residues 249-281, sequence is EELLSKNYHLENEVARLKK, see SEQ ID NO:2).
[0077] The gp41 trimerization domain forms a stable triple-helix coiled-coil structure, which can induce protein assembly without affecting its original function; GCN4 has a well-studied dimer leucine zipper structure, which can promote the formation of protein dimers.
[0078] 2. Determination of insertion site
[0079] Through a comprehensive analysis of the domain architecture of SpCas9 (amino acid sequence see SEQ ID NO:3), 62 potential insertion sites were determined for inserting GCN4 or gp41, see Figure 1 .
[0080] Consider three key criteria: surface accessibility, structural flexibility, and minimal interference with key protein-protein or protein-DNA interactions. These sites are mainly distributed in the loop regions between α-helices and β-sheets, spanning the RuvC domain, REC domain, HNH domain, and PAM interaction domain. The insertion sites are at positions 1, 25, 54, 62, 170, 203, 213, 231, 249, 257, 308, 355, 365, 400, 532, 573, 584, 674, 719, 765, 768, 776, 782, 790, 808, 819, 826, 831, 846, 868, 890, 910, 924, 945, 975, 1010, 1020, 1033, 1050, 1051, 1055, 1059, 1068, 1072, 1102, 1110, 1120, 1130, 1227, 1246, 1248, 1252, 1260, 1276, 1290, 1300, 1302, 1327, 1332, 1340, 1368 of the SpCas9 protein, or replace amino acids 1048 - 1063 of SpCas9.
[0081] 3. Construct gp41-SpCas9 variants
[0082] Use the PCR method to insert the gp41 trimerization domain into the selected sites to construct SpCas9-Xgp41 series variants (X represents the insertion position, i.e., insert gp41 after the Xth amino acid). The specific construction steps are as follows:
[0083] Design a pair of primers with the gp41 coding sequence at the 5' end and sequences complementary to both sides of the specific site of SpCas9 at the 3' end; perform PCR amplification using a high-fidelity DNA polymerase; clone the PCR product into a mammalian expression vector containing the CMV promoter; verify the construction correctness using Sanger sequencing; purify the plasmid using the EndoFree Plasmid Maxi Kit (Qiagen).
[0084] 4. Preliminary screening of sites for improving editing efficiency
[0085] Perform a preliminary screening in HEK293 cells and analyze the editing efficiency at two gene loci, RNF2 and HEK293 site2, using amplicon deep sequencing. Specifically as follows:
[0086] HEK293 cells were seeded at a density of 1×10^5 cells per well in a 24-well plate; 500 ng of Cas variant plasmid and 200 ng of sgRNA plasmid were transfected using Lipofectamine 3000 24 hours after transfection; the culture medium was changed 6 hours after transfection; cells were harvested 72 hours after transfection for genomic DNA extraction; sequencing libraries were constructed using a two-step PCR method: the first round of PCR (25 cycles) used gene-specific primers with Illumina-compatible adapters, and the second round of PCR (11 cycles) added unique dual indexes; amplicons were purified using DNA purification beads and equimolarly mixed after quantification; sequencing was performed on the MGI2000 platform; sequencing results were analyzed using CRISPResso2 software to calculate insertion and deletion (indel) frequencies, see Figure 2 .
[0087] The screening results showed that for gp41-SpCas9 variants, variants at 12 insertion sites exhibited enhanced editing efficiency, among which variants at insertion sites 203, 213, and 584 showed the highest improvement (1.22-1.24-fold compared to the wild type).
[0088] 5. Genome-wide off-target analysis (GUIDE-seq)
[0089] To rigorously assess the targeting specificity of these engineered variants, a comprehensive analysis was performed using GUIDE-seq (genome-wide, unbiased identification of DSBs enabled by sequencing):
[0090] HEK293 cells were transfected with Cas variants, sgRNA, and double-stranded oligonucleotides (dsODNs); cells were harvested to extract genomic DNA; DNA was fragmented to an average size of 500 bp using Bioruptor pico; end repair, A-tailing, and adapter ligation were performed using the KAPA HTP library preparation kit; two rounds of nested PCR were performed using dsODN-specific primers; sequencing was performed on the MGI2000 platform; data were analyzed using the open-source GuideSeq software, see Figure 3 and 4 .
[0091] GUIDE-seq analysis revealed two optimal variants, SpCas9-308gp41 (amino acid sequence see SEQ ID NO: 4) and SpCas9-584gp41 (amino acid sequence see SEQ ID NO: 5):
[0092] At the CIITA locus, wild-type SpCas9 showed 44,682 on-target reads and 15 off-target reads, while SpCas9-308gp41 exhibited superior on-target activity (93,887 reads, 2.1-fold increase) and no detectable off-target effects.
[0093] At the PCSK9 locus, SpCas9-584gp41 maintained robust on-target activity (13,246 reads) while significantly reducing off-target events (12 off-target reads compared to 501 for wild-type), representing the superior specificity of SpCas9-584gp41 in reducing off-target effects.
[0094] In addition to the above two variants, variants such as SpCas9-1248gp41 and SpCas9-1368gp41 all showed superior on-target activity and low off-target effects.
[0095] Example 2: Characterization of the molecular mechanism underlying enhanced oligomerization regulation
[0096] 1. Comparison between oligomerization domain and flexible linker
[0097] To elucidate the molecular mechanism underlying the enhanced performance of engineered nucleases, it was investigated whether the gp41 and GCN4 domains function through specific oligomerization rather than simply providing flexible linker effects, as shown in Figure 5 . Specifically as follows:
[0098] The gp41 sequences in the SpCas9-1gp41 and SpCas9-231gp41 variants were replaced with an equal-length glycine-serine (GS) flexible linker (GGGGS)3; TIDE analysis was used to determine the editing efficiency at the RNF2 and HEK293 site 2 loci.
[0099] The results showed that at the RNF2 locus, the SpCas9-1gp41 and SpCas9-231gp41 variants showed enhanced editing efficiency (59.5% and 55.25% respectively), compared to wild-type SpCas9 (53.0%), while the corresponding GS linker counterparts showed significantly reduced activity (SpCas9-1GS: 39.2%, SpCas9-231GS: 40.75%). Similar functional differences were also observed at the HEK293 site 2 locus, where the gp41-integrated variants maintained superior performance while the GS linker variants performed poorly. These results clearly indicate that the improvement brought about by the oligomerization domain cannot be attributed to simple linker flexibility, but involves specific structural and functional contributions.
[0100] As shown in Figure 4As shown, both SpCas9-1GCN4 and SpCas9-308GCN4 can exhibit superior on-target activity and low off-target effects at 11 genomic target sites. This indicates that the incorporation of GCN4 can also significantly enhance the activity of SpCas9. The principle of its enhancing SpCas9 activity was further confirmed by FRET experiments ( Figure 5 H, I, and J).
[0101] 2. Verification of oligomerization ability by Split-GFP complementation assay
[0102] To directly verify the oligomerization ability of the two domains, the Split-GFP complementation assay was adopted. The specific steps are as follows:
[0103] Superfolder GFP was split between β-strands 10 and 11 (GFP1-10: residues 1-214; GFP11: residues 215-230); fusion proteins were systematically designed with gp41 or GCN4 located at the C-terminus of GFP1-10 or the N-terminus of GFP11; equimolar amounts of the Split-GFP construct pairs were transfected into HEK293 cells; the amino acid sequence of GFP1-10-gp41 is shown in SEQ ID NO:10, the amino acid sequence of gp41-GFP11 is shown in SEQ ID NO:11, the amino acid sequence of GFP1-10-GCN4 is shown in SEQ ID NO:12, and the amino acid sequence of GFP11-GCN4 is shown in SEQ ID NO:13.
[0104] The GFP fluorescence intensity was analyzed by flow cytometry 48 hours after transfection; unfused GFP fragments were used as negative controls, and full-length sfGFP was used as a positive control; co-transfection of mCherry was used to normalize the transfection efficiency.
[0105] The results of Split-GFP analysis showed that unfused GFP fragments produced the lowest fluorescence (0.8% GFP-positive cells), while the gp41-mediated combination exhibited significantly higher fluorescence signals, with the GFP1-10-gp41 and gp41-GFP11 pairs showing the highest complementation efficiency (12.8% GFP-positive cells). Similar results were obtained for GCN4-mediated dimerization, with GFP1-10-GCN4 and GFP11-GCN4 generating 8% GFP-positive cells compared to 0.5% for the negative control. These data directly demonstrated the ability of the gp41 and GCN4 domains to effectively mediate protein oligomerization in a protein environment.
[0106] 3. Resonance energy transfer (FRET) analysis
[0107] For more quantitative biophysical characterization, FRET was used to analyze SpCas9 variants fused to either mTurquoise2 (donor, λex = 434 nm, λem = 474 nm) or mVenus (acceptor, λex = 515 nm, λem = 527 nm) fluorescent proteins. Specifically as follows:
[0108] The C-termini of gp41-SpCas9 and GCN4-SpCas9 variants were fused to mTurquoise2 or mVenus, linked by an optimized (GGGGS)2 linker; proteins were purified using standard affinity chromatography; measurements were performed using a Hitachi F-4700 spectrofluorometer equipped with temperature control; measurements were carried out in reaction buffer (20 mM HEPES pH 7.5, 150 mM KCl, 1 mM DTT, 5% glycerol, supplemented with 0.05% Tween-20 to prevent non-specific protein adsorption).
[0109] The amino acid sequence of 1GCN4-SpCa9-mTurquoise2 is shown in SEQ ID NO:14, the amino acid sequence of 1GCN4-SpCa9-mVenus is shown in SEQ ID NO:15, the amino acid sequence of 308gp41-SpCa9-mTurquoise2 is shown in SEQ ID NO:16, and the amino acid sequence of 308gp41-SpCa9-mVenus is shown in SEQ ID NO:17.
[0110] For 308gp41-SpCas9, concentration-dependent fluorescence emission of the donor (mTurquoise2-labeled variant) and acceptor (mVenus-labeled variant) alone was observed. When mixed at an optimal donor ratio of 1:4 (determined by systematic titration experiments), a strong FRET signal was detected, indicating that the gp41 domain mediated successful protein-protein interactions. Importantly, the normalized FRET signal remained consistent at various protein concentrations, confirming the specificity and stability of gp41-mediated oligomerization.
[0111] Parallel experiments with 1GCN4-SpCas9 used the same donor ratio (1:4) established during gp41 optimization. When donor and acceptor molecules were combined, these experiments produced distinct FRET signals, and the normalized emission spectra showed comparable consistent efficiency to the gp41 system.
[0112] The FRET efficiency was calculated using the equation E = 1 - (FDA / FD), where FDA and FD represent the donor fluorescence intensities in the presence and absence of the receptor, respectively. Direct receptor excitation and spectral spillover were corrected using receptor-alone and donor-alone controls. The energy transfer data were fitted to an appropriate binding model using GraphPad Prism 9.0. All measurements were performed in technical triplicates and verified in multiple independent protein preparations.
[0113] These complementary methods provide compelling evidence that the enhanced performance of engineered nucleases directly stems from controlled oligomerization, which creates optimal spatial constraints that favor precise target recognition while disfavoring off-target binding. This mechanism explains how the methods of the present invention uniquely address the traditional efficiency-specificity trade-off by modulating the three-dimensional organization of the nuclease complex rather than altering its intrinsic catalytic properties.
[0114] Example 3: Extending the Oligomerization Strategy to FrCas9 and Co-Integrating with Directed Evolution
[0115] 1. Construction and Screening of FrCas9 Variants
[0116] To explore the generality of the oligomerization enhancement strategy beyond SpCas9, the method of Example 1 was applied to FrCas9 (amino acid sequence shown in SEQ ID NO: 18), a phylogenetically distinct Cas9 ortholog discovered and characterized from Faecalibaculum rodentium. FrCas9 has a unique 5'-NNTA-3' PAM, with superior specificity and efficiency, making it highly suitable for therapeutic applications.
[0117] Two series of FrCas9 variants were systematically engineered by integrating the GCN4 dimerization domain or the gp41 trimerization domain at 56 different positions in the protein distribution, as follows:
[0118] The positions where GCN4 or gp41 was inserted were behind the 1st, 40th, 51st, 55th, 108th, 112th, 120th, 125th, 170th, 192nd, 200th, 215th, 234th, 250th, 265th, 270th, 275th, 290th, 295th, 315th, 320th, 325th, 363rd, 379th, 428th, 539th, 566th, 610th, 618th, 690th, 702nd, 707th, 727th, 755th, 761st, 800th, 830th, 882nd, 980th, 1002nd, 1043rd, 1055th, 1095th, 1111th, 1117th, 1134th, 1184th, 1188th, 1224th, 1248th, 1321st, 1329th, 1333rd, 1347th, 1372nd amino acids of the FrCas9 protein, respectively.
[0119] Screening of GCN4-integrated FrCas9 variants revealed 10 positions that significantly enhanced editing efficiency (1.13 - 1.24-fold improvement), among which positions 215, 234, 707, and 830 showed the strongest improvement at the RNF2 and HEK293 site 2 loci; for gp41-integrated variants, positions 1, 265, 270, and 295 showed a consistent 1.04 - 1.08-fold enhancement in editing efficiency, see Figure 6 A and B in
[0120] 2. GUIDE-seq analysis of FrCas9-COME variants
[0121] GUIDE-seq analysis verified these improvements, see Figure 6 C and D in
[0122] At the RNF2 locus, FrCas9-1372GCN4 (amino acid sequence shown in SEQ ID NO:20) and FrCas9-265gp41 (amino acid sequence shown in SEQ ID NO:21) reached 145,479 and 152,965 targeted reads respectively, compared to 93,360 reads of wild-type FrCas9, representing an approximately 1.6-fold enhancement without affecting specificity.
[0123] Similar improvements were observed at the HEK293 site 2 locus, where FrCas9-265gp41 showed 121,545 reads, compared to 77,985 reads of the wild-type (a 1.57-fold increase).
[0124] At the GRIN2B-3 locus, both FrCas9-1372GCN4 and FrCas9-265gp41 maintained high targeting activity while minimizing off-target events.
[0125] Quantitative analysis at all three loci confirmed these enhancements, and FrCas9-265gp41 consistently showed a 1.6-fold increase in targeting activity relative to wild-type FrCas9.
[0126] 3. Synergistic effect of oligomerization and directed evolution
[0127] To investigate the potential synergistic effect between the oligomerization method and traditional protein engineering methods, an oligomerization domain was integrated into eFrCas9. The amino acid sequences at positions 1103 and 732 of eFrCas9 are shown in SEQ ID NO:22 compared to wild-type FrCas9.
[0128] The amino acid sequence of eFrCas9-1gp41 is shown in SEQ ID NO:23, and the amino acid sequence of eFrCas9-1372GCN4 is shown in SEQ ID NO:24.
[0129] GUIDE-seq analysis revealed a significant synergistic effect:
[0130] eFrCas9 alone showed a 6.26-fold enhancement relative to wild-type FrCas9; FrCas9-265gp41 demonstrated a 7.35-fold improvement; the combined eFrCas9-265gp41 (amino acid sequence shown in SEQ ID NO:25) variant (termed UltraFrCas9) achieved an extraordinary 29.22-fold enhancement relative to wild-type FrCas9, far exceeding the additive effects of the two methods.
[0131] This synergistic effect was consistently observed at different genomic targets:
[0132] At the RUNX1 locus, UltraFrCas9 reached 16,159 targeted reads (see Figure 6 in E), compared to 117 reads for wild-type FrCas9 (a 138-fold enhancement), significantly outperforming eFrCas9 (1,728 reads) and FrCas9-265gp41 (2,662 reads); a similar pattern was observed at the EMX1 locus (22,085 reads vs. 466 reads, representing a 47.4-fold improvement); at the DYRK1A locus (28,459 reads vs. 2,036 reads, a 14-fold enhancement); most importantly, these significant efficiency improvements were accompanied by exceptional specificity (see Figure 6 in F), with UltraFrCas9 achieving a specificity ratio as high as 30,046 at the HEK293 site2 target site. This unprecedented combination of enhanced activity and specificity indicates that the oligomerization strategy of the present invention operates through a mechanism orthogonal to traditional protein engineering methods, allowing combinable complementary optimization pathways to achieve maximum performance enhancement.
[0133] Example 4: Integrating oligomerization domains into Cas12a homologs
[0134] 1. Oligomerization engineering of LbCas12a
[0135] To expand the applicability of the oligomerization platform beyond the Cas9 nuclease, this example systematically applied the method to Cas12a homologs. Forty-four potential insertion sites were identified in the domains of LbCas12a (amino acid sequence shown in SEQ ID NO: 26) for inserting GCN4 or gp41, and the insertion sites are as follows: 1, 11, 85, 171, 269, 270, 335, 354, 371, 406, 478, 572, 625, 653, 731, 805, 810, 825, 836, 865, 928, 965, 967, 972, 982, 1040, 1074, 1087, 1109, 1117, 1119, 1120, 1121, 1138, 1142, 1143, 1148, 1156, 1158, 1168, 1171, 1175, 1211, 1228. See Figure 7 A in
[0136] The gp41 and GCN4 domains were evaluated using SSA (single-strand annealing) reporter assays and amplicon sequencing:
[0137] For gp41 integration in LbCas12a, positions 85, 406, 572, and C-terminal insertion showed the highest enhancements, reaching indel frequencies of 66.67%, 60.56%, 67.32%, and 65.43% respectively, compared to 50.34% in the wild type; for GCN4 integration, position 1117 was outstanding, achieving an editing efficiency of 80.91% at the CDKN2A locus, compared to 61.85% in the wild type. See Figure 7 B in
[0138] GUIDE-seq analysis verified these enhancements:
[0139] LbCas12a-1117GCN4 (amino acid sequence shown in SEQ ID NO: 27) showed a 1.29-fold increase in targeting activity while reducing off-target events by 62%; LbCas12a-406gp41 (amino acid sequence shown in SEQ ID NO: 28) showed excellent specificity, almost no off-target activity while maintaining robust targeting efficiency; LbCas12a-85gp41 (amino acid sequence shown in SEQ ID NO: 29) showed consistent enhancements at all sites, with a particularly strong improvement (1.98-fold increase) at the RUNX1 locus. See Figure 7 D and E in
[0140] 2. Oligomerization engineering of LtCas12a
[0141] This method was extended to LtCas12a (amino acid sequence shown in SEQ ID NO: 30), and 24 potential insertion sites were identified for the insertion of GCN4 or gp41. The insertion sites are as follows: 12, 85, 330, 354, 402, 572, 652, 687, 895, 992, 1029, 1039, 1049, 1108, 1141, 1150, 1173, 1181, 1186, 1208, 1220, 1222, 1232, 1235, see Figure 7 in F.
[0142] Using the same evaluation methods as for LbCas12a, namely the SSA reporter system and amplicon sequencing, the performance of each insertion site was evaluated.
[0143] Notably, despite only 33% sequence identity between LbCas12a and LtCas12a, position 85 emerged as the best site for both oligomerization domains in these two homologs. This conserved enhancer site strongly supports the existence of an evolutionarily conserved structural node that is particularly suitable for oligomerization optimization.
[0144] GUIDE-seq analysis confirmed a significant enhancement of the LtCas12a variant:
[0145] see Figure 7 and 8 , at the DYRK1A locus, LtCas12a-85GCN4 (amino acid sequence shown in SEQ ID NO: 31) achieved a 1.7-fold increase in targeting activity, while LtCas12a-85gp41 (amino acid sequence shown in SEQ ID NO: 32) demonstrated a 4.5-fold significant enhancement; at the CDKN2A locus, LtCas12a-85gp41 showed a 3.4-fold increase in targeting activity along with an improvement in specificity; the identification of position 85 as a conserved enhancer site between Cas12a homologs indicates the existence of evolutionarily conserved structural features suitable for oligomerization optimization. This conservation, combined with the successful enhancement of the Cas9 and Cas12a systems, establishes the method of the present invention as a universal optimization platform across multiple CRISPR nucleases in prokaryotic adaptive immune systems.
[0146] The performance of the best variants in Examples 1-4 above is shown in Table 6:
[0147] Table 6 Summary of the performance of the best COME variants
[0148]
[0149]
[0150] Example 5: Application and optimization expansion of controlled oligomerization engineering
[0151] 1. Applications of Controlled Oligomerization Engineering Technology
[0152] CRISPR variants optimized by controlled oligomerization engineering technology have broad application prospects, including:
[0153] Basic research applications: high-precision genome editing and functional research, genome screening at single-base resolution, precise regulation of multiplex gene editing systems; Therapeutic applications: precise gene editing for genetic diseases, improvement of the safety of somatic gene therapy, tumor-targeted gene editing; Diagnostic applications: highly sensitive nucleic acid detection, single-base mutation identification, rapid pathogen detection system; Biotechnological applications: improvement of industrial microbial genome engineering, precise modification of crop genomes, precise tools for synthetic biology.
[0154] COME uses a multimerized Cas protein and fuses a fluorescent protein with the Cas protein for expression, which can greatly increase the local concentration of the fluorescent protein, thus greatly enhancing the fluorescence intensity at the target site. This can be used for genome imaging and nuclear foci detection, etc.
[0155] Furthermore, the COME technology has broad application prospects in the industrial field:
[0156] Biopharmaceuticals: used for developing high-precision gene therapies, improving gene editing tools for CAR-T and other cell therapies, and simplifying gene therapy strategies for complex diseases; Agricultural biotechnology: developing high-precision crop gene editing tools, improving agricultural traits while minimizing off-target effects, and accelerating the plant breeding process; Industrial biotechnology: precisely engineering microorganisms for biomanufacturing, optimizing the production of biofuels and biomaterials, and creating customized modifications of industry-related enzymes; Diagnostic technology: developing ultrasensitive CRISPR diagnostic systems, improving point-of-care detection devices, and supporting precision medicine and personalized treatment.
[0157] 2. Optimization and Expansion of Controlled Oligomerization Engineering Technology
[0158] (1) Optimization of the Oligomerization Domain
[0159] To further optimize the performance of the oligomerization domain, the domain length and amino acid composition can be adjusted to increase stability; site-directed mutations can be introduced to change the oligomerization kinetics; conditional oligomerization domains that respond to external stimuli (such as light, small molecules, temperature) can be designed; and novel artificially designed oligomerization modules with more efficient and controllable oligomerization characteristics can be developed.
[0160] (2) Systematic Optimization of the Insertion Site
[0161] For each CRISPR nuclease, the insertion sites can be systematically optimized through the following strategies: virtual screening based on protein structure analysis; developing a high-throughput evaluation platform to test hundreds of insertion variants simultaneously; using machine learning to predict the optimal insertion sites based on existing experimental datasets; combining multiple beneficial insertion sites to explore synergistic effects.
[0162] (3) Extended applications of the COME technology
[0163] The COME technology can be extended to the following applications: improving the performance of base editors (BEs) and prime editors (PEs); enhancing the efficiency and specificity of transcriptional activation (CRISPRa) and inhibition (CRISPRi) systems; optimizing emerging CRISPR systems such as Cas12f, Cas12k, Cas14, Cas7-11, etc.; developing enhanced nucleic acid detection systems for disease diagnosis and environmental monitoring.
[0164] (4) Combinations with other engineering strategies
[0165] COME can be combined with the following strategies to further enhance the performance of the CRISPR system: combination with PAM-dependent modification to expand the targeting range; combination with rational protein design to optimize the catalytic center; combination with mRNA delivery systems to improve in vivo transfection; combination with nanomaterials to develop novel delivery systems.
[0166] Example 6: Standard operating procedure (SOP) of COME
[0167] 1. Standard operating procedure for constructing COME variants
[0168] (1) Design stage:
[0169] Select the insertion sites of the oligomerization domain based on structural analysis, optimize the coding sequence of the oligomerization domain, avoid potential RNA secondary structures, and design PCR primers containing appropriate restriction enzyme sites.
[0170] (2) Molecular cloning stage:
[0171] Amplify the target region using high-fidelity PCR, integrate the oligomerization domain using restriction enzyme digestion or Gibson assembly, transform the Escherichia coli DH5α strain, and verify by colony PCR and Sanger sequencing.
[0172] (3) Construction of expression vectors:
[0173] Clone the verified COME variants into a mammalian expression vector containing the CMV promoter, add tag sequences such as nuclear localization signal (NLS) and FLAG / HA, and purify the plasmid in large quantities using the EndoFree Plasmid Maxi Kit.
[0174] (4) Quality control standards:
[0175] Confirm purity and integrity by Nanodrop and agarose gel electrophoresis, confirm the vector structure using restriction mapping, verify the full-length COME variant sequence by sequencing, and confirm expression by Western blot.
[0176] 2. Standard Operating Procedure for the Performance Evaluation of COME Variants
[0177] (1) Preliminary activity screening:
[0178] Transfect COME variants and sgRNA plasmids into HEK293 cells, evaluate the editing efficiency using amplicon deep sequencing analysis, verify at multiple genomic loci, and screen the best-performing variants for in-depth characterization.
[0179] (2) Off-target analysis:
[0180] Perform genome-wide off-target analysis using GUIDE-seq, select multiple representative genomic loci, process the data analysis using a standard pipeline, and calculate the specificity ratio (on-target / off-target read ratio).
[0181] (3) Mechanism study:
[0182] Verify oligomerization using Split-GFP complementation experiments, quantitatively characterize protein interactions using FRET analysis, compare with a flexible linker control group, and perform structural simulation analysis to study the effect of oligomerization on conformation.
[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for controlling oligomerization engineering, characterized in that, Integrating an oligomerization domain into a CRISPR nuclease.
2. The method according to claim 1, wherein The oligomerization domain includes an oligomerization domain with temperature sensitivity, small molecule regulation, or tissue-specific expression.
3. The method according to claim 1, wherein The oligomerization domain includes a gp41 trimerization domain and / or a GCN4 dimerization domain.
4. The method according to claim 3, characterized in that, The amino acid sequence of the gp41 trimerization domain is as shown in SEQ ID NO:1; the amino acid sequence of the GCN4 dimerization domain is as shown in SEQ ID NO:
2.
5. The method according to claim 1, wherein The CRISPR nuclease is any one of Cas7, Cas8, Cas9, Cas10, Cas11, Cas12, Cas14, IscB, and TnpB.
6. The method according to claim 5, wherein The CRISPR nuclease is any one of SpCas9, FrCas9, LbCas12a, LtCas12a, Cas12f, and Cas12k.
7. A SpCas9 variant, characterized in that, An oligomerization domain is integrated into the amino acid sequence of the SpCas9 protein.
8. The SpCas9 variant according to claim 7, wherein The integration site is behind any one of the amino acids at positions 1, 25, 54, 62, 170, 203, 213, 231, 249, 257, 308, 355, 365, 400, 532, 573, 584, 674, 719, 765, 768, 776, 782, 790, 808, 819, 826, 831, 846, 868, 890, 910, 924, 945, 975, 1010, 1020, 1033, 1050, 1051, 1055, 1059, 1068, 1072, 1102, 1110, 1120, 1130, 1227, 1246, 1248, 1252, 1260, 1276, 1290, 1300, 1302, 1327, 1332, 1340, 1368 of the SpCas9 protein, or the amino acids at positions 1048 - 1063 of the SpCas9 protein are integrally replaced.
9. A FrCas9 variant, characterized in that, An oligomerization domain is integrated into the amino acid sequence of the FrCas9 protein.
10. The FrCas9 variant according to claim 9, wherein, The integration site is behind any one of the amino acids at positions 1, 40, 51, 55, 108, 112, 120, 125, 170, 192, 200, 215, 234, 250, 265, 270, 275, 290, 295, 315, 320, 325, 363, 379, 428, 539, 566, 610, 618, 690, 702, 707, 727, 755, 761, 800, 830, 882, 980, 1002, 1043, 1055, 1095, 1111, 1117, 1134, 1184, 1188, 1224, 1248, 1321, 1329, 1333, 1347, 1372 of the FrCas9 or eFrCas9 protein.
11. The FrCas9 variant according to claim 10, wherein, Compared with FrCas9, the eFrCas9 protein only has different amino acid sequences at positions 1103 and 732, and the amino acid sequences at positions 1103 and 732 are as shown in SEQ ID NO:
22.
12. An LbCas12a variant, characterized in that, An oligomerization domain is integrated into the amino acid sequence of the LbCas12a protein.
13. The LbCas12a variant according to claim 12, wherein The integration site is behind any one of the amino acids at positions 1, 11, 85, 171, 269, 270, 335, 354, 371, 406, 478, 572, 625, 653, 731, 805, 810, 825, 836, 865, 928, 965, 967, 972, 982, 1040, 1074, 1087, 1109, 1117, 1119, 1120, 1121, 1138, 1142, 1143, 1148, 1156, 1158, 1168, 1171, 1175, 1211, 1228 of the LbCas12a protein.
14. An LtCas12a variant, characterized in that, An oligomerization domain is integrated into the amino acid sequence of the LtCas12a protein.
15. The LtCas12a variant according to claim 14, wherein, The integration site is behind any one of the amino acids at positions 12, 85, 330, 354, 402, 572, 652, 687, 895, 992, 1029, 1039, 1049, 1108, 1141, 1150, 1173, 1181, 1186, 1208, 1220, 1222, 1232, 1235 of the LtCas12a protein.
16. The variant according to claims 7, 9, 12, and 14, characterized in that, The oligomerization domain includes an oligomerization domain that is temperature-sensitive, small molecule-regulated, or tissue-specifically expressed.
17. The variant according to claim 16, characterized in that, The oligomerization domain includes a gp41 trimerization domain and / or a GCN4 dimerization domain.
18. The variant according to claim 17, characterized in that, The amino acid sequence of the gp41 trimerization domain is as shown in SEQ ID NO:1; the amino acid sequence of the GCN4 dimerization domain is as shown in SEQ ID NO:
2.
19. A method for constructing a COME variant, characterized in that, Including the following steps: (1) Select the oligomerization domain insertion site based on structural analysis, optimize the oligomerization domain coding sequence, avoid potential RNA secondary structures, and design PCR primers containing restriction enzyme sites; (2) PCR amplify the target region, digest with restriction enzymes or perform Gibson assembly to integrate the oligomerization domain, transform Escherichia coli, and verify the COME variant positive clones by colony sequencing; (3) Clone the COME variant into a mammalian expression vector containing a CMV promoter, add tag sequences nuclear localization signal, FLAG or HA, and purify the plasmid; (4) Confirm the purity and integrity by Nanodrop and electrophoresis, confirm the vector structure using restriction enzyme digestion maps, verify the full-length COME variant sequence by sequencing, and confirm the expression by protein electrophoresis.
20. A method for evaluating the performance of a COME variant, characterized in that, Including the following steps: (1) Transfect the COME variant and sgRNA plasmid into HEK293 cells, use amplicon deep sequencing analysis to evaluate the editing efficiency, verify at multiple genomic sites, and screen the best-performing variants for in-depth characterization; (2) Perform genome-wide off-target analysis using GUIDE-seq, select multiple representative genomic sites, process the data analysis using a standard pipeline, and calculate the specificity ratio; (3) Use Split-GFP complementation experiments to verify oligomerization, use FRET analysis to quantitatively characterize protein interactions, compare with the flexible linker control group, and perform structural simulation to analyze the effect of oligomerization on conformation.
21. Use of the variant according to any one of claims 7-18 and its derivatives in any of the following fields: i. Genome editing, transcriptional activation or inhibition, base editing or primer editing; ii. Preparation of drugs for gene therapy, prevention and / or diagnosis; iii. NNTA PAM recognition; v. Genome imaging, epigenetic modification or nucleic acid detection; vi. Biosensors.