A cg12n-v4.6 nuclease, vector and ecg12n gene editing system and application

By modifying the CgCas12n nuclease and simplifying the sgRNA backbone, a highly efficient and compact eCg12n gene editing system was developed, solving the problems of low delivery efficiency and insufficient cleavage activity of the CRISPR/Cas system, and achieving efficient and accurate gene editing.

CN120866276BActive Publication Date: 2026-03-03AGRICULTURAL GENOMICS INSTITUTE AT SHENZHEN CHINESE ACADEMY OF AGRICULTURAL SCIENCES (SHENZHEN BRANCH GUANGDONG LABORATORY FOR LINGNAN MODERN AGRICULTURE)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511366364.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-03-03
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

In existing CRISPR/Cas systems, the large size of Cas9 and Cas12a proteins limits the delivery efficiency of gene editing systems, and the reported low cleavage activity of Cas12n nuclease limits its application in gene editing.

Method used

By modifying the wild-type CgCas12n nuclease, a Cg12n-v4.6 nuclease mutant was developed. Combined with a simplified sgRNA backbone, a highly efficient and compact eCg12n gene editing system was constructed, including a Cg12n-v4.6 nuclease expression vector and an sgRNA expression vector, thereby improving cleavage activity and gene editing efficiency.

Benefits of technology

It achieves high efficiency and accuracy in gene editing, with an editing efficiency of over 70%, and has a smaller system size, making it suitable for widespread application in gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120866276B_ABST
    Figure CN120866276B_ABST
Patent Text Reader

Abstract

This invention discloses a Cg12n-v4.6 nuclease, the amino acid sequence of which is shown in SEQ ID No. 12. This invention develops a Cg12n-v4.6 nuclease mutant with increased activity and improved editing efficiency, achieving an editing efficiency of over 70%. Simultaneously, the corresponding sgRNA of Cas12n is engineered to simplify the sgRNA backbone size. Combining the engineered Cg12n-v4.6 nuclease mutant and the simplified sgRNA backbone, an enhanced Cas12n gene editing tool, namely the eCg12n gene editing system, is developed. Furthermore, this invention constructs a compact cytosine base editor and a compact adenine base editor by fusing deaminases, and applies the cytosine base editor to induce early stop codon generation in disease genes, achieving safe gene knockout. The Cg12n-v4.6 nuclease mutant and its derived compact eCg12n gene editing system provided by this invention not only improve editing efficiency compared to the original wild-type editing system, but also further simplify the system size, making it more suitable for gene editing applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene editing technology, and in particular to a Cg12n-v4.6 nuclease, vector, eCg12n gene editing system and its applications. Background Technology

[0002] Clustered regularly interspaced short palindromic repeats (CRISPR / CRISPR associated) and their associated proteins constitute a natural "immune system" widely found in bacteria and archaea. This system defends against viral invasion by recognizing and cleaving exogenous nucleic acids. The development of CRISPR / Cas-based gene editing technology has had a tremendous impact on the field of biotechnology, greatly promoting its development and providing a powerful tool for efficient genome editing. Given its simplicity, efficiency, and low cost, CRISPR / Cas gene editing technology is widely used in research fields such as biomedicine, agricultural breeding, environmental monitoring, and disease diagnosis. Based on the composition of effector proteins, the currently discovered CRISPR / Cas systems are mainly divided into two categories. One category consists of multiple Cas proteins combined to form an effector protein complex, which works with crRNA to recognize and bind to target DNA, such as type I, III, and IV CRISPR / Cas systems. The other category consists of a single Cas protein with multiple functional domains acting as an effector protein, such as type II, V, and VI CRISPR / Cas systems. The first CRISPR / Cas family developed for in vitro and in vivo gene editing was the type II CRISPR / Cas9 family. With the discovery of more and more CRISPR / Cas systems, such as the type V CRISPR / Cas12 family and the type VI CRISPR / Cas13 family, the CRISPR / Cas-based gene editing toolkit has been continuously expanded and enriched. Among the reported Cas families, the Cas effector proteins of the type II CRISPR / Cas9 family and the type V CRISPR / Cas12a family are monomeric multifunctional domain nucleases, and due to their simplicity and high efficiency, they are widely used in gene editing research.

[0003] However, most Cas9 and Cas12a proteins contain approximately 950-1400 amino acids, and their large size limits the delivery efficiency of CRISPR / Cas editing systems, thus hindering their widespread application. Faced with this limitation, the development and exploration of smaller molecular weight CRISPR / Cas editing systems has become a research hotspot in recent years. Examples include the Cas12f, Cas12m, and Cas12n families of type V CRISPR systems, as well as the ancestor of type V CRISPR systems, TnpB, and its derived gene editing systems.

[0004] In reported miniature CRISPR / Cas systems, the Cas12n family of nucleases consists of 400-700 amino acids, approximately half the size of commonly used Cas nucleases, and belongs to the V-U4 CRISPR nuclease family. Unlike the Cas12f family of nucleases, which function as homodimers, the structural analysis of the Cas12n family of nucleases shows that they function as monomers. Due to their small size, gene editing systems based on this type of nuclease are easier to deliver and have greater application advantages. However, currently reported Cas12n nucleases have low (or even no) cleavage activity, limiting their application in gene editing. Therefore, this invention aims to develop a novel, efficient, and compact gene editing system by engineering the Cas12n nuclease and its corresponding sgRNA, providing technical support for gene editing applications. Summary of the Invention

[0005] The purpose of this invention is to provide a Cg12n-v4.6 nuclease, vector, eCg12n gene editing system, and its applications. By modifying the wild-type CgCas12n nuclease, a new Cg12n-v4.6 nuclease mutant with high gene editing efficiency is obtained. This Cg12n-v4.6 nuclease exhibits high cleavage activity and significantly improved gene editing efficiency compared to the wild-type. Then, its expression vector is constructed, and the sgRNA backbone is simplified to obtain a highly efficient and compact eCg12n gene editing system. This editing system has high gene editing accuracy and low off-target rate, providing better technical support for gene editing applications.

[0006] According to a first aspect of the present invention, a Cg12n-v4.6 nuclease is provided, the amino acid sequence of which is shown in SEQ ID No. 12. This Cg12n-v4.6 nuclease has been engineered to form a Cg12n-v4.6 nuclease mutant, exhibiting significantly enhanced cleavage activity and a gene editing efficiency exceeding 70%. This lays the foundation for the subsequent development of highly efficient and compact gene editing systems and tools.

[0007] According to a second aspect of the present invention, a vector for expressing the Cg12n-v4.6 nuclease is provided, the nucleotide sequence of which is shown in SEQ ID No. 6. This vector expresses an engineered Cg12n-v4.6 nuclease, and gene editing using this vector achieves a gene editing efficiency of over 70%.

[0008] According to a third aspect of the present invention, the application of Cg12n-v4.6 nuclease in gene editing for non-disease diagnostic and therapeutic purposes is provided. Therefore, due to its small size and engineered high gene editing efficiency, the application of Cg12n-v4.6 nuclease in the field of gene editing can significantly improve gene editing efficiency.

[0009] According to a fourth aspect of the present invention, a vector is provided for use in gene editing for non-disease diagnostic and therapeutic purposes, the nucleotide sequence of which is shown in SEQ ID No. 6. This vector expresses an engineered Cg12n-v4.6 nuclease, and because the vector is small, it is more suitable for gene editing. Gene editing using this vector achieves a gene editing efficiency of over 70%.

[0010] According to a fifth aspect of the present invention, an eCg12n gene editing system is provided, which expresses the Cg12n-v4.6 nuclease, the amino acid sequence of which is shown in SEQ ID No. 12. Because the expressed Cg12n-v4.6 nuclease in this gene editing system is small in size, it is more suitable for gene editing, and the modified Cg12n-v4.6 nuclease has high editing efficiency. Therefore, the eCg12n gene editing system has the characteristics of small size and high editing efficiency.

[0011] According to a sixth aspect of the present invention, an eCg12n gene editing system is provided, which contains a vector expressing a Cg12n-v4.6 nuclease. Because the Cg12n-v4.6 nuclease expressed by the vector in this gene editing system is small in size, it is more suitable for gene editing, and the modified Cg12n-v4.6 nuclease has high editing efficiency. Therefore, the eCg12n gene editing system has the characteristics of small size and high editing efficiency.

[0012] In some embodiments, the system contains a Cg12n-v4.6 expression vector, the nucleotide sequence of which is shown in SEQ ID No. 6.

[0013] In some embodiments, the system further includes an sgRNA expression vector constructed by inserting a target sequence into an sgRNA backbone, the nucleotide sequence of which is shown in SEQ ID No. 13. Because this sgRNA backbone is streamlined, it is more compact, more suitable for gene editing, and can improve gene editing efficiency.

[0014] According to a seventh aspect of the present invention, a cytosine base editor is provided, comprising an eCg12n-CBE vector mediating C-to-T base editing and an sgRNA expression vector. The eCg12n-CBE vector is constructed by fusing a vector with the cytosine deaminase ecDA1 gene, the nucleotide sequence of which is shown in SEQ ID No. 6. The nucleotide sequence of the eCg12n-CBE vector is shown in SEQ ID No. 10. The sgRNA expression vector is constructed by inserting a target sequence into the sgRNA backbone, the nucleotide sequence of which is shown in SEQ ID No. 13. Therefore, this cytosine base editor can be efficiently applied in the field of gene editing, exhibiting high editing efficiency, low off-target rate, and the ability to safely delete genes, demonstrating significant advantages over traditional tools.

[0015] According to an eighth aspect of the present invention, a cytosine base editor is provided, comprising an in-eCg12n-CBE vector and an sgRNA expression vector mediating C-to-T base editing. The in-eCg12n-CBE vector is constructed by embedding the cytosine deaminase ecDA1 into a Cg12n-v4.6 nuclease with an amino acid sequence as shown in SEQ ID No. 12. The nucleotide sequence of the in-eCg12n-CBE vector is shown in SEQ ID No. 11. The sgRNA expression vector is constructed by inserting a target sequence into an sgRNA backbone, the nucleotide sequence of which is shown in SEQ ID No. 13. Therefore, this cytosine base editor can be efficiently applied in the field of gene editing, exhibiting high editing efficiency, low off-target rate, and safe gene deletion, demonstrating significant advantages over traditional tools.

[0016] According to a ninth aspect of the present invention, an adenine base editor is provided, comprising an eCg12n-ABE vector mediating A-to-G base editing and an sgRNA expression vector. The eCg12n-ABE vector is constructed by fusing a vector with the adenine deaminase TAdA-8e-V106W gene, the nucleotide sequence of which is shown in SEQ ID No. 6. The nucleotide sequence of the eCg12n-ABE vector is shown in SEQ ID No. 15. The sgRNA expression vector is constructed by inserting a target sequence into the sgRNA backbone, the nucleotide sequence of which is shown in SEQ ID No. 13. Therefore, this adenine base editor can be efficiently applied in the field of gene editing, with high editing efficiency and a low off-target rate.

[0017] According to a tenth aspect of the present invention, the application of the described eCg12n gene editing system in gene editing for non-disease diagnosis and treatment purposes is provided. This improves gene editing efficiency, while offering good safety, high precision, and a low off-target rate.

[0018] According to the eleventh aspect of the present invention, the application of the cytosine base editor is provided in gene editing for purposes other than disease diagnosis and treatment. Thus, by using the cytosine base editor to edit disease genes and induce the generation of premature stop codons, safe gene knockout can be achieved.

[0019] According to a twelfth aspect of the present invention, the application of the adenine base editor is provided in gene editing for non-disease diagnostic and therapeutic purposes. Thus, the application of the adenine base editor enables highly efficient gene editing.

[0020] The beneficial effects of this invention are as follows: This invention develops a Cg12n-v4.6 nuclease mutant with increased activity and improved editing efficiency, and evaluates its gene editing characteristics, achieving an editing efficiency of over 70%. By engineering the sgRNA, the sgRNA backbone size is simplified, making the eCg12n gene editing system more compact. Combining the engineered Cg12n-v4.6 nuclease expression vector Cg12n-v4.6 and the simplified sgRNA backbone, an enhanced Cas12n editing tool, namely the eCg12n gene editing system, is developed. Simultaneously, by fusing deaminases, compact cytosine base editors and compact adenine base editors are constructed. The cytosine base editor is then used to induce premature stop codon generation in disease genes, achieving safe gene knockout. The Cg12n-v4.6 nuclease mutant and its derived compact eCg12n gene editing system provided by this invention not only improve editing efficiency compared to the original wild-type editing system, but also further simplify the system size, making it more suitable for gene editing applications. Attached Figure Description

[0021] Figure 1 A schematic diagram illustrating the principle of a fluorescence reporter system for detecting the cleavage activity of CgCas12n nuclease.

[0022] Figure 2 The graph shows the cleavage activity results of various engineered CgCas12n nuclease mutants in the fluorescent reporter system: where the horizontal axis represents the name of each CgCas12n nuclease variant, and the vertical axis represents the percentage of green fluorescent cells in the total number of cells.

[0023] Figure 3 The graph shows the gene editing efficiency of the Cg12n-v4.6 nuclease variant at multiple endogenous sites: the horizontal axis represents the number of multiple endogenous sites, and the vertical axis represents the insertion / deletion rate, which represents the editing efficiency of the Cg12n-v4.6 nuclease variant.

[0024] Figure 4 The following graph shows the gene editing efficiency of the Cg12n-v4.6 nuclease variant in different cell lines: The horizontal axis represents the name of the different cell lines, and the vertical axis represents the insertion / deletion rate, which represents the editing efficiency of the Cg12n-v4.6 nuclease variant. In the figure, WT represents the unmodified Cas12n nuclease, and v4.6 represents the modified Cg12n-v4.6 nuclease variant.

[0025] Figure 5 A schematic diagram illustrating the principle of engineered sgRNA backbone simplification;

[0026] Figure 6A comparison of the editing effects of different neck ring deletion sgRNA backbone versions in a fluorescence reporter system: where the horizontal axis represents the different sgRNA backbone names, and the vertical axis represents the ratio of green fluorescent cells to transfected positive cells, representing different sgRNA backbone editing efficiencies.

[0027] Figure 7 A comparison of the editing effects of different neck-loop deleted sgRNA backbone versions of Stem2 in a fluorescent reporter system: where the horizontal axis represents the different sgRNA backbone names, and the vertical axis represents the ratio of green fluorescent cells to transfected positive cells, representing different sgRNA backbone editing efficiencies.

[0028] Figure 8 A comparison of the editing effects of different versions of sgRNA backbones with stepwise deletion of bases in a fluorescent reporter system: where the horizontal axis represents different sgRNA backbone names, and the vertical axis represents the ratio of green fluorescent cells to transfected positive cells, representing different sgRNA backbone editing efficiencies.

[0029] Figure 9 A comparison of editing efficiency at endogenous sites for engineered sgRNA versions that delete D19 bases one by one by the Stem2 neck loop sgRNA backbone and Stem5: where the horizontal axis represents the names of different endogenous sites, the vertical axis represents the insertion / deletion rate, representing the editing efficiency of endogenous sites, and T6D19, T10D19, T11D19, and T12D19 are different sgRNA backbones;

[0030] Figure 10 The results show the gene editing efficiency of eCg12n-v4.6 nuclease at endogenous sites before and after sgRNA backbone modification: where the horizontal axis represents the different endogenous site numbers, and the vertical axis represents the insertion / deletion rate, representing the endogenous site editing efficiency. v4.6+WTsg represents the editing system composed of eCg12n-v4.6 nuclease and the unmodified sgRNA backbone, and v4.6+T6D19sg represents the editing system composed of eCg12n-v4.6 nuclease and the modified sgRNA backbone.

[0031] Figure 11 The graph shows the comparison of gene editing efficiency between the eCg12n gene editing system and the wild-type parent (unmodified CgCas12n gene editing system): where the horizontal axis represents different editing systems and the vertical axis represents the insertion / deletion rate, representing gene editing efficiency.

[0032] Figure 12The graph shows the results of the eCg12n gene editing system's editing efficiency at 5 endogenous sites and the off-target effects at potential sites: where the horizontal axis represents the endogenous sites and corresponding off-target site numbers where off-target effects were detected in this embodiment, and the vertical axis is the insertion / deletion rate, representing the gene editing efficiency;

[0033] Figure 13 The results of the accuracy test for the eCg12n gene editing system are shown in the figure: the left side shows the detection site sequence and the corresponding mismatched sgRNA sequence, the right side shows the editing efficiency of each sgRNA, and the horizontal axis represents the editing efficiency.

[0034] Figure 14 The graph shows a comparison of the editing efficiency of the eCg12n and SpCas9 gene editing systems at some endogenous sites: where the horizontal axis represents the endogenous site number detected in this embodiment, and the vertical axis represents the insertion / deletion rate, which represents the gene editing efficiency.

[0035] Figure 15 The graph shows a comparison of the editing efficiency of the eCg12n and AsCas12a gene editing systems at some endogenous sites: where the horizontal axis represents the endogenous site number detected in this embodiment, and the vertical axis represents the insertion / deletion rate, which represents the gene editing efficiency.

[0036] Figure 16 The graph shows the C-to-T editing efficiency of the eCg12n-CBE cytosine base editor at different endogenous sites: the horizontal axis represents the name of the endogenous site and the position of the corresponding cytosine base from the PAM sequence, and the vertical axis represents the C-to-T base editing efficiency. eCg12n-CBE represents the cytosine base editor derived from the eCg12n gene editing system, and WT-CBE represents the cytosine base editor derived from the unmodified wild-type CgCas12n gene editing system.

[0037] Figure 17 Characterization of the eCg12n-CBE cytosine base editor editing window: In the upper figure, the horizontal axis represents the different positions of cytosine base C in all target sequences, the vertical axis represents the C-to-T base editing efficiency, the middle figure shows the heatmap representation of the editing efficiency of C at different positions, and the lower figure shows the position representation of each sgRNA target site;

[0038] Figure 18 A comparison of the editing efficiency of the eCg12n-CBE and in-eCg12n-CBE cytosine base editors at the same site: where the horizontal axis represents the distance of the cytosine base C at the target site from PAM, and the vertical axis represents the C-to-T base editing efficiency;

[0039] Figure 19The graph shows the RNA off-target detection of the eCg12n-CBE and in-eCg12n-CBE cytosine base editors: where the vertical axis represents the number of detected RNA off-target events, and the horizontal axis represents the names of different cytosine base editors. WT-CBE represents the unmodified wild-type CgCas12n-derived cytosine base editor, in-eCg12n-CBE represents the embedded eCg12n-CBE, and eCg12n-CBE represents the cytosine base editor based on the eCg12n editing system.

[0040] Figure 20 The graph shows the editing efficiency of the eCg12n-ABE adenine base editor at different endogenous sites: where the horizontal axis represents the endogenous genes and sites detected in this embodiment, and the vertical axis represents the A-to-G base editing efficiency.

[0041] Figure 21 A schematic diagram illustrating the locations of multiple sgRNAs targeting the mouse Ldlr gene;

[0042] Figure 22 A comparison of editing efficiency of sgRNAs targeting the Ldlr gene at different sites: The horizontal axis represents the position of the cytosine base C in the 6 sgRNAs detected in this example, and the vertical axis represents the C-to-T base editing efficiency;

[0043] Figure 23 The peak diagram of the Sanger sequencing results shows that the appearance of overlapping peaks confirms that eCg12n-CBE can achieve C-to-T base editing at the Ldlr gene target site.

[0044] Figure 24 Location and sequence diagram of the sgRNA designed for this embodiment that can effectively target the mouse Dmd gene;

[0045] Figure 25 The graph shows the comparison of editing efficiency of multiple sgRNAs targeting the mouse Dmd gene: the horizontal axis represents the number of the sgRNA used, and the vertical axis represents the C-to-T base editing efficiency. Detailed Implementation

[0046] The invention will now be described in further detail with reference to the accompanying drawings.

[0047] Example 1. Engineering modification of CgCas12n nuclease to achieve efficient editing of CgCas12n nuclease in eukaryotic cells.

[0048] Based on homology alignment and arginine mutation strategies, this invention engineered the CgCas12n nuclease (NCBI Reference Sequence: WP_158407477.1). CgCas12n nucleases containing arginine mutations at specific sites were screened using a eukaryotic fluorescence reporter system, and beneficial mutations were enriched. These beneficial mutations were combined to obtain engineered CgCas12n mutants with significantly improved cleavage efficiency. Editing efficiency was compared across multiple endogenous sites in various cell lines. The specific steps include:

[0049] 1.1 Development of efficient CgCas12n nuclease mutants.

[0050] 1.1.1 Identification of high-efficiency CgCas12n nuclease mutant candidates and construction of expression vectors.

[0051] Candidate mutation sites were obtained through homology alignment and arginine substitution mutations were performed. In this invention, CgCas12n (amino acid sequence as shown in SEQ ID No. 1) was aligned with AcCas12n (amino acid sequence as shown in SEQ ID No. 2) and AsCas12a (amino acid sequence as shown in SEQ ID No. 3) to obtain candidate mutation sites. PCR primers were designed for each mutation site, and the PCR products were then inserted into the vector pGL3-U6-sgRNA-pCMV-BPNLS-CgCas12n-mCherry (sequence shown in SEQ ID No. 4) via homology recombination to construct the eukaryotic expression vector pCMV-CgCas12n-U6-sgRNA containing a single-mutant CgCas12n nuclease and the corresponding sgRNA. After all vectors were constructed, sequencing confirmed the correct sequences before use in downstream experiments.

[0052] 1.1.2 Screening for beneficial mutations in a eukaryotic fluorescence reporter system.

[0053] A eukaryotic fluorescent reporter system was constructed, which is a CMV-initiated EGFP expression vector. This system achieves EGFP expression inactivation by inserting a non-3-fold multiple sequence containing a CgCas12n PAM sequence and a target sequence after the EGFP start codon ATG, inducing a frameshift mutation. The constructed vector serves as the eukaryotic fluorescent reporter plasmid. When CgCas12n targets and cleaves this inserted sequence, DNA repair generates indels that restore the EGFP expression frame, allowing normal EGFP expression in the plasmid. The working principle of this reporter system is as follows: Figure 1 As shown.

[0054] After constructing the eukaryotic cell fluorescent reporter plasmid pCMV-△EGFP (this vector is artificially synthesized, and its sequence is shown in SEQ ID No. 5) according to the above method, HEK293T cells were seeded into 24-well plates. When the cell density in the wells reached 75%~85%, transient transfection was performed using EZ transfection reagent. The CgCas12n nuclease single mutant expression plasmid pCMV-CgCas12n-U6-sgRNA and the fluorescent reporter plasmid pCMV-△EGFP were co-transfected into HEK293T cells.

[0055] 72 hours after transfection, the culture medium was discarded, cells were washed with PBS, digested, and collected into centrifuge tubes. After centrifugation at 800 rpm for 3 min, the supernatant was discarded. Cells were then resuspended in PBS and filtered through a 100-mesh filter into flow cytometry tubes for fluorescence analysis. The cleavage activity of each CgCas12n mutant was determined based on the EGFP fluorescence recovery.

[0056] Based on the above experimental analysis, eleven beneficial mutations were selected from the CgCas12n single mutant constructed in this invention: E84R, I482R, Q170R, T184R, S374R, A56R, S218R, A89R, D157R, T423R, and A361R. Subsequently, two, three, and four mutation combinations were performed on these beneficial mutations to construct combined mutants, which significantly improved the cleavage efficiency of CgCas12n at the eukaryotic cell level (results are shown in Figure 1). Figure 2(As shown in the diagram). This invention selected four mutant combinations with the highest cleavage efficiency for subsequent experiments and named them Cg12n-v4.6 nuclease mutants, whose amino acid sequences are shown in SEQ ID No. 12.

[0057] 1.2 Validation of the editing efficiency of the Cg12n-v4.6 nuclease mutant at the endogenous site in eukaryotic cells.

[0058] 1.2.1 Construction of Cg12n-v4.6 expression vector and sgRNA vector targeting endogenous sites in eukaryotic cells.

[0059] Primers were designed targeting the Cg12n-v4.6 nuclease mutant sequence. Then, the nucleotide sequence expressing the Cg12n-v4.6 nuclease was amplified by PCR and inserted into the pCMV-EGFP expression vector by homologous recombination to construct the pCMV-Cg12n-v4.6-EGFP expression vector (hereinafter referred to as "Cg12n-v4.6 expression vector"). Its nucleotide sequence is shown in SEQ ID No. 6.

[0060] Twenty-one endogenous gene loci were selected, and targeting oligonucleotides were designed and synthesized based on their sequences. The sgRNA sequences used are shown in Table 1. A tccc sequence was added to the 5' end of the upstream sequence of each sgRNA, and an aaaa sequence was added to the 5' end of the downstream sequence. After synthesis, the upstream and downstream sequences were annealed using a pre-defined program. The annealed products were ligated overnight at 16°C into a linearized PGL3-U6-sgRNA-mCherry vector (this vector was artificially synthesized, and its sequence is shown in SEQ ID No. 7) to construct the sgRNA expression vector. The ligation product was transformed into *E. coli* and plated on ampicillin-resistant plates, then cultured overnight at 37°C. The next day, single clones were picked and sent for Sanger sequencing. After sequence alignment confirmed the correct clones, the culture was expanded, and plasmids were extracted. The concentration of the correctly sequenced sgRNA expression vector was determined for subsequent experiments.

[0061] Table 1. sgRNA sequence information targeting endogenous sites in eukaryotic cells

[0062] Target gene Forward oligonucleotide chain sequence Reverse oligonucleotide chain sequence VEGFA tcccCATGGAAACAGACCTGGCAG aaaaCTGCCAGGTCTGTTTCCATG VEGFA tcccCGACTCAACCTGGTAAACAT aaaaATGTTTACCAGGTTGAGTCG VEGFA tcccAATCATTTCCCCAAGAGGAA aaaaTTCCTCTTGGGGAAATGATT VEGFA tcccAGGCAAGCATGGAAACAGAC aaaaGTCTGTTTCCATGCTTGCCT AAVS1 tcccTGTAAGGAAGCTGCAGCACC aaaaGGTGCTGCAGCTTCCTTACA AAVS1 tcccGGTGCTGCAGCTTCCTTACA aaaaTGTAAGGAAGCTGCAGCACC AAVS1 tcccAGGAGAAGCAGTTTGGAAAA aaaaTTTTCCAAACTGCTTCTCCT AAVS1 tcccCAAACCTTAGAGGTTCTGGC aaaaGCCAGAACCTCTAAGGTTTG AAVS1 tcccGAATCTGCCTAACAGGAGGT aaaaACCTCCTGTTAGGCAGATTC AAVS1 tcccACCTCCTGTTAGGCAGATTC aaaaGAATCTGCCTAACAGGAGGT AAVS1 tcccAGGATGGAGAGGTGGCTAAA aaaaTTTAGCCACCTCTCCATCCT AAVS1 tcccAGCTAGCACAGACTAGAGAG aaaaCTCTCTAGTCTGTGCTAGCT AAVS1 tcccGCCATCCTAAGAAACGAGAG aaaaCTCTCGTTTCTTAGGATGGC PDCD1 tcccACAGTGGGGACTAGAGCTCA aaaaTGAGCTCTAGTCCCCACTGT PDCD1 tcccATGCTTCAGAGACGAGATGG aaaaCCATCTCGTCTCTGAAGCAT HEXA tcccGAGCTCTACACCACACCCAA aaaaTTGGGTGTGGTGTAGAGCTC HEXA tcccTTTAACTACTTACTGTTTGT aaaaACAAACAGTAAGTAGTTAAA HEXA tcccCTGTGGCATTTGCTTAAGCT aaaaAGCTTAAGCAAATGCCACAG HEXA tcccGACAAGGTTGAACAAACAGT aaaaACTGTTTGTTCAACCTTGTC EMX1 tcccGTGTCAGTTTAAGAATGGTG aaaaCACCATTCTTAAACTGACAC TP53 tcccGGTGCAGTTATGCCTCAGAT aaaaATCTGAGGCATAACTGCACC

[0063] 1.2.2 Transfection of HEK293T cells to detect the editing efficiency of Cg12n-v4.6 nuclease mutant.

[0064] One day before transfection, HEK293T cells were seeded in 24-well plates. When the cells reached a confluence density of 70%-80%, the Cg12n-v4.6 expression vector and the sgRNA expression vector were co-transfected into HEK293T cells using EZ transfection reagent. The cell culture medium was changed 8-10 hours after transfection. Cells were harvested 72 hours after transfection, and GFP-positive cells were sorted using a flow cytometry system. 5000-10000 cells (top 15% of fluorescence intensity) were collected from each sample. The sorted cells were centrifuged and lysed using cell lysis buffer to release DNA.

[0065] PCR primers targeting the desired sequence were designed and synthesized. The cell lysis buffer described above was used as a template for PCR amplification of DNA sequences containing the target site. The PCR products were identified and purified by gel electrophoresis, and amplicon sequencing was performed using an Illumina Novaseq PE150. Editing efficiency was analyzed using CRISPResso. Results are as follows: Figure 3 As shown, the Cg12n-v4.6 expression vector can perform gene editing at multiple endogenous sites, with a maximum editing efficiency of over 70%. Since the Cg12n-v4.6 expression vector is an expression vector for the modified CgCas12n nuclease, its high editing efficiency further demonstrates the successful modification of the CgCas12n nuclease and its high gene editing efficiency.

[0066] 1.2.3 Detection of the cleavage efficiency of Cg12n-v4.6 nuclease mutant in multiple eukaryotic cell lines.

[0067] Resuscitated frozen HEK293T, HeLa, and K562 cell lines were transferred to cell culture dishes and cultured at 37°C with 5% CO2. When the cells recovered and reached 90% confluence, they were seeded into 24-well plates and cultured overnight. The next day, when cell confluence reached 70-80%, the Cg12n-v4.6 expression vector and sgRNA expression vector were co-transfected into the target cell lines using EZ transfection reagent. The cell culture medium was changed 8-10 hours after transfection. 72 hours after transfection, the cells were digested, flow cytometry sorted, and enriched. The resulting cells were centrifuged, resuspended, and lysed. The lysate was amplified by PCR and sequenced to analyze the editing efficiency of the target sites. Results are as follows: Figure 4 As shown, the Cg12n-v4.6 expression vector can effectively edit endogenous sites in HEK293T, HeLa, and K562 cell lines. The editing efficiency of the Cg12n-v4.6 expression vector is comparable in HEK293T and HeLa cell lines, while the editing efficiency is lower in the K562 cell line. However, compared to the wild-type parent (unmodified CgCas12n nuclease), the Cg12n-v4.6 nuclease mutant shows a significantly improved editing efficiency, indicating that the modified CgCas12n nuclease gene editing efficiency is significantly enhanced.

[0068] Example 2: Engineered sgRNA to develop a compact eCg12n gene editing system.

[0069] Based on RNA structure prediction analysis, this invention simplifies the gene editing system by deleting its backbone nucleic acid sequence while maintaining the structural stability of sgRNA. The invention designs various deletion strategies, including neck loop deletion and base-by-base deletion (e.g., ...). ​ (As shown). Engineered sgRNAs were rapidly and efficiently screened in eukaryotic cells using a fluorescent reporter system to obtain the optimal sgRNA backbone (T6D19). Combined with an engineered Cg12n-v4.6 nuclease variant, a compact Cas12n gene editing system, eCg12n, was developed, and its editing efficiency was assessed at endogenous sites in eukaryotic cells.

[0070] 2.1 Engineered sgRNA to develop a compact Cas12n gene editing system.

[0071] Based on RNA structure analysis, sgRNA backbone modification schemes were designed, mainly including neck-loop deletion of Stem1-Stem4 and base-by-base deletion of Stem5. For each sgRNA backbone structure, primers were designed for PCR amplification to introduce mutations. The PCR amplification template was the pGL3-U6-sgRNA-pCMV-BPNLS-CgCas12n-mCherry vector (sequence shown in SEQ ID No. 4). Because the primers have homologous arms after amplification, they can recombine homologously with each other to construct expression vectors containing the modified sgRNA backbone. All vectors were confirmed by sequencing, and plasmids were extracted for later use.

[0072] Effective sgRNA backbone deletion versions were screened using a eukaryotic fluorescent reporter system: HEK293T cells were co-expressed with Cg12n-v4.6 expression vectors and co-expression vectors of different modified sgRNA backbones, along with fluorescent reporter plasmids. Effective sgRNA backbones were screened based on fluorescence retrieval rates. The first step was a deletion experiment of the neck loops Stem1-Stem4 that constitute the sgRNA backbone. The experimental results are as follows: ​ As shown, among the designed neck-loop deletion sgRNA versions of Stem1-Stem4, two versions (T4 and T6) showed improved editing efficiency despite the reduction in neck-loop deletion. Since both T4 and T6 are located in Stem2, more neck-loop deletion tests were performed on Stem2, with results as follows... ​ As shown, versions T4 and T6 with deletions exhibit the highest editing efficiency; excessive deletion of this neck loop reduces the system's editing efficiency. Secondly, among the sgRNA versions that delete each base of the Stem5 neck loop, two versions showed less reduction in editing efficiency (e.g., ​ As shown in d13 and d19, considering that the purpose of sgRNA engineering is to maintain efficiency while simplifying sgRNA, this invention selects the version with more base deletions (d19). Next, four sgRNA candidates (T6D19, T10D19, T11D19, T12D19) were constructed by combining the Stem2 neck-loop deleted sgRNA backbone (T11, T6, T10, T12) with the optimal Stem5 base deletion sgRNA backbone (d19), respectively. Editing efficiency was then tested at endogenous gene sites. The experimental results are as follows. ​ As shown, among the four candidate sgRNAs, T6D19 has the highest editing efficiency at multiple endogenous sites. Therefore, T6D19 was selected as the optimal sgRNA backbone, denoted as T6D19 sgRNA backbone. The nucleotide sequence of this sgRNA backbone is shown in SEQ ID No. 13.

[0073] The wild-type sgRNA expression vector (i.e., the unmodified sgRNA expression vector PGL3-U6-sgRNA-mCherry, nucleotide sequence shown in SEQ ID No. 7), the expression vector containing the modified T6D19 sgRNA backbone, and the Cg12n-v4.6 expression vector were co-transfected into HEK293T cells. Editing efficiency was assessed by targeting multiple endogenous sites. The experimental results are as follows: ​ As shown, compared to wild-type sgRNA, the modified sgRNA (T6D19) has higher editing efficiency at most endogenous sites, indicating that the modification of the sgRNA backbone increases the system's editing efficiency while reducing system size.

[0074] 2.2 The compact eCg12n gene editing system achieves efficient editing effect detection at endogenous gene sites in eukaryotic cells.

[0075] Based on the aforementioned engineering results of the Cas12n nuclease and its sgRNA backbone according to the present invention, a compact gene editing system was constructed by combining the optimal nuclease variant Cg12n-v4.6 with the optimal sgRNA backbone T6D19, named the eCg12n gene editing system. This system includes a Cg12n-v4.6 expression vector and an sgRNA expression vector, the nucleotide sequence of which is shown in SEQ ID No. 6. The sgRNA expression vector is constructed by inserting a target sequence into the sgRNA backbone, the nucleotide sequence of which is shown in SEQ ID No. 13. Based on the sgRNA backbone, sgRNA expression vectors targeting different endogenous sites were constructed. The eCg12n editing system (including the Cg12n-v4.6 expression vector and the sgRNA expression vector) was transfected into HEK293T cells using EZ transfection reagent. Cells were collected 72 hours after transfection, and the top 15% of cells with the highest fluorescence intensity were lysed by flow cytometry. The DNA sequence containing the target site was amplified by PCR, and its editing efficiency was analyzed by next-generation sequencing. The results are as follows: ​ As shown, compared to the original wild-type parent (i.e., the gene editing system composed of unmodified Cas12n nuclease and unmodified sgRNA backbone), the eCg12n gene editing system significantly improves the editing efficiency at endogenous sites, with an editing efficiency of up to 70% at some sites.

[0076] Example 3. Evaluation of the safety, accuracy, and editing efficiency of the eCg12n gene editing system.

[0077] This embodiment evaluates the safety, accuracy, and editing efficiency of the developed eCg12n gene editing system in eukaryotic cells.

[0078] 3.1 Off-target risk assessment of the eCg12n gene editing system.

[0079] This embodiment first assessed the off-target risks of the eCg12n gene editing system. For the highly efficient endogenous sites in Examples 1 and 2, such as VEGFA, AAVS1, PDCD1, and HEXA (5 sites in 4 genes), genomic alignment was performed, and sites with similar sequences were selected as potential off-target sites. HEK293T cells were transfected with the eCg12n editing system containing the target sites, and cells were sorted and lysed 72 hours after transfection. After obtaining the cell lysate, DNA fragments were amplified by PCR using primers containing the target sequence and primers containing potential off-target sites, respectively. The fragments were then purified and sequenced for analysis. The sequence information of the target sites and potential off-target sites, as well as the PCR amplification primers, are shown in Table 2 below. The experimental results are as follows: ​ As shown, no potential off-target editing was detected at the five high-efficiency editing sites examined in this embodiment. This indicates that the eCg12n gene editing system has a low off-target risk.

[0080] Table 2. Target sequences of multiple endogenous sites and primers for detecting potential off-target sites.

[0081] ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 43704153 43704172 ​ ​ - ​ ​ ​ L2 ​ 25693876 25693895 ​ - ​ ​ ​ L3 ​ 65692384 65692403 ​ - ​ ​ ​ ​ 55116092 55116111 ​ ​ + ​ ​ ​ L4 ​ 160532727 160532746 ​ + ​ ​ ​ L5 ​ 148424025 148424047 ​ + ​ ​ ​ L6 ​ 41378855 41378877 ​ - ​ ​ ​ L7 ​ 72679760 72679782 ​ - ​ ​ ​ L8 ​ 140635700 140635722 ​ + ​ ​ ​ L9 ​ 229763406 229763428 ​ - ​ ​ ​ L10 ​ 86841740 86841759 ​ - ​ gatggcttcagagtgtggct cgtttgaggggaatggggagt chr2 241853552 241853571 PDCD1 ACAGTGGGGACTAGAGCTCA + AAG ccccctttcaggacaagct gaggtccttgtcttgggagc L11 chr10 99897515 99897534 acagtggagactaggtgatg - AAG gcagaagtaacgataaaaagcgga agtagtaagacatccaattagacact L12 chr6 784679 784698 acagtggggcactagaggaag + AAG gctgcctgaagcttctggat tgtgaactggagatgcgcac L13 chr6 120550187 120550206 acagtggggattagagtcag - AAG agcagttcagagtctaggaaga gtcctcctgtcacacaaggg L14 chr3 130569329 130569348 acagtggggactaggttcct - AAG gcccttatatttggctgagct atcaaatgaggtgatgtttgcga chr15 72369140 72369159 HEXA-26 CTTAAGCAAATGCCACAGCT + AAG tctctgtattcctggctgtaatga gagtatacgcttccacagaaagg L16 chr18 46095192 46095211 cttaaaaaatgccagctta + AAG catggaagttccttgcagca aggagctgtataaaaggttttcaca L17 chr18 76364984 76365003 cttaagcaaatgagctgaat + AAG ttgggtaattcgccaagtct tggaaatcacgtaggttggt L18 chr9 22358170 22358189 cttaagcaaatgcttaagga - AAG tgactttgtccagcaaacatga tctccctccaaatcaaagcat chr19 55115141 55115160 AAVS1-11 AAGGATGGAGAAAGAGAAAG + AAG gactcaaacccagaagccca gtagccagccccgtcctg L19 chr5 132431819 132431838 tagtggaaaagggg - AAG acgcttctctggctcatcc tgatcatgtggctgggttcc L20 chr6 162866753 162866772 caggatggaaagaagaatt + AAG tgctgattccaattcccctgt tgcacagtggagcttgtcaa chr15 72368871 72368890 HEXA-27 TTTAACTACTTACTGTTTGT - AAG agcagccttccccttcttttt tgaacaggctgacacttcaa L21 chr18 26117138 26117157 ttttactacttactgcaata + AAG aagtgacccgcctgcctc aatttggtaagggggcaaatgg L22 chr2 41798969 41798988 tttaacaacttactttctca + AAG agctcgggatgcatcaactc tgggaggtgtttgggtcatg

[0082] 3.2 Evaluation of the editing accuracy of the eCg12n gene editing system.

[0083] This embodiment tested the sgRNA mismatch tolerance of the eCg12n gene editing system as an assessment of editing accuracy.

[0084] Targeting the AAVS1 site, this embodiment designed and constructed sgRNA expression vectors containing one and two base mismatches, respectively. Sequential mismatches at all 20 base positions were achieved through individual substitutions. The Cg12n-v4.6 expression vector was co-transfected with vectors expressing different mismatched sgRNAs into HEK293T cells. The 15% of cells with the strongest fluorescence were collected by flow cytometry, centrifuged, resuspended, and lysed. PCR amplification of the DNA fragment containing the target sequence was then performed for sequencing analysis. Experimental results are as follows: Figure 13As shown, when the sgRNA contains one mismatched base, the mismatches at positions 1, 2, 4-9, 11, 13, and 14 (proximal to the PAM) have the greatest impact on the eCg12n editing system; mismatches at positions 19 and 20 have the least impact. When the sgRNA contains two mismatched bases, mismatches at positions 17 and 18, and 19 and 20 have the least impact on the editing efficiency of the eCg12n editing system; mismatches at other positions result in the loss of cleavage activity of the eCg12n editing system. Combining the experimental results with one and two mismatched bases, the eCg12n editing system has low tolerance to base mismatches, especially to mismatches near the PAM, indicating that the eCg12n gene editing system has good editing precision.

[0085] 3.3 Evaluation of the editing efficiency of the eCg12n gene editing system.

[0086] This embodiment also evaluated the editing efficiency of the eCg12n gene editing system. SpCas9 (Uniprot Entry: Q99ZW2) and AsCas12a (Uniprot Entry: U2UMQ6) are the two most widely used CRISPR / Cas gene editing systems. Based on the PAM characteristics of SpCas9, AsCas12a, and CgCas12n, this embodiment designed and constructed targeting sgRNAs that simultaneously meet the PAM requirements of SpCas9 and CgCas12n, as well as AsCas12a and CgCas12n. The sgRNA sequences and primer information used for editing efficiency comparison with SpCas9 are shown in Table 3 below, and the sgRNA sequences and primer information used for editing efficiency comparison with AsCas12a are shown in Table 4 below. HEK293T cells were transfected with the eCg12n, SpCas9, and AsCas12a editing systems, targeting endogenous sites. The cells were then sorted by flow cytometry, lysed, subjected to PCR, sequencing, and analysis to determine the editing efficiency of the eCg12n, SpCas9, and AsCas12a editing systems at the same target sites. The experimental results are as follows: Figure 14 and 15 As shown, at the target sites tested in this embodiment, the eCg12n editing system exhibited comparable editing efficiency to the SpCas9 and AsCas12a editing systems. In fact, at some sites, eCg12n demonstrated superior editing efficiency, indicating that it is a highly efficient gene editing system. This demonstrates that the modified Cg12n-v4.6 nuclease possesses high gene editing activity, and the eCg12n gene editing system expressing this Cg12n-v4.6 nuclease exhibits good editing efficiency.

[0087] Table 3. sgRNA sequences and primer information used for editing efficacy comparison with SpCas9.

[0088] sg No. Gene Target PAM (CgCas12n) PAM (SpCas9) Primer-F Primer-R g1 HBB TCTGCCGTTACTGCCCTGTG AAG GGG ACATTTGCTTCTGACACAACTGT GTCAGTGCCTATCAGAAACCCA g2 HBB GGTAGACCACCAGCAGCCTA AAG AGG AGCCTTCACCTTAGGGTTGC TGGGCAGGTTGGTATCAAGG g3 IFNG TATGTAATGTAGAGAAATGC AAG TGG acagttgctaaagctatgaactca tgagcaagggacaatgagaga g4 IFNG GGAAACCAAGGTGTAAGTTT AAG TGG tcaaaatagcgtgggggcat actggctatggttatacatgacaca g5 LNX1 TAAAGGGATTTTAAAGGCAC AAG TGG caatgggtaacaaaagtcctggc gcagggggagcatcacatag g6 LNX1 GGAAGAGAGGATATAGCGGC AAG AGG ttgcattccagcagccaatg ggggaggtgaagaagagctt g7 KLHL29 TTGCCAGAGAAGTTAAAGGC AAG GGG tgtccgtgtacatgcgtgta aatgggaagcagggctgaaa g8 KLHL29 GGAAGACCTGGAAGTGTATG AAG AGG gcagtctacccatctcagcc aattgctggcaagtccaagc

[0089] Table 4. sgRNA sequences and primer information used for editing efficacy comparison with AsCas12a.

[0090] sgNo. Gene Target (CgCas12n) Target (AsCas12a) PAM(CgCas12n) PAM(AsCas12a) Primer-F Primer r-R g9 FGF18 AGGGTCTCCCGGTGGCCCTC GAGGGCCACCGGGAGACCCT AAG TTTA CCAGAATCCTGAGGTGCTGG GTCCCTTATTCCTGGAGCGG g10 FGF18 CAATAGCATCTTCCCAGATA TATCTGGGAAGATGCTATTG AAG TTTA CTCACTGGGCTGTGTGGTTT CCTGGCTCAGTAGGGGTAAC g11 HBB CTGTGATTCCAAATATTACG CGTAATATTTGGAATCACAG AAG TTTA CTTGGCCCCATACCATCAGT TCTGGAGACGCAGGAAGAGA g12 HBB ATGAGGAGGTCAAGAGATGA TCATCTCTTGACCTCCTCAT AAG TTTT GTAGGCCTGGGTGTCTGTTG TGGGACATGTGCCGAGCA g13 HBB GCTGAGAGATGCAGGATAAG GCTGAGAGATGCAGGATAAG AAG TTTG CCAGGCGAGGAGAAACCATC AGTAGAGGCTTGATTTGGAGGT g14 IFNG TATTTGGCACTTGGTGCATT AATGCACCAAGTGCCAAATA AAG TTTT TGGGATTGTGTATTGATAATCCAGAGA AGGCATTCAATTCAATAAGCACA g15 LNX1 AAGTCCAGCATCCTGGTCAA TTGACCAGGATGCTGGACTT AAG TTTG GTGGGACTTACCGGCTCG CTGCCTCACCAACTTCCTGG g16 KLHL29 TTCCAGTTCCACACCTTCTG CAGAAGGTGTGGAACTGGAA AAG TTTG AGTTGCTGCTGGAGTTTGTCT AGACCTCCCCGCTCAGTC g17 VEGFA AGTGACAGTATCCTCTGTAT ATACAGAGGATACTGTCACT AAG TTTA TGTGGTGCATTTGGAATCACC AAAGGGCCCATTGCTGTGA g18 AAVS1 GAAAGAAGGATGGAGAAAGA TCTTTCTCCATCCTTCTTTC AAG TTTC GCTTGCCAAGGACTCAAACC CGTGTGGAAAACTCCCTTTGT

[0091] Example 4: A compact base editor based on the eCg12n gene editing system can effectively mediate base editing in eukaryotic cells.

[0092] This embodiment constructs a compact base editor based on the eCg12n gene editing system by fusing the eCg12n gene editing system with cytosine base deaminase (ecDA1) and adenine base deaminase (TadA8eV106W), enabling efficient base editing in eukaryotic cells.

[0093] 4.1 Construction and evaluation of a compact cytosine base editor based on eCg12n.

[0094] 4.1.1 Construction of the compact cytosine base editor eCg-CBE based on eCg12n.

[0095] In this embodiment, an eCg12n-CBE cytosine base editor is constructed by linking a nuclease-inactivated Cg12n-v4.6 protein (dCg12n-v4.6, which is inactivated by a D297A amino acid mutation in the Cg12n-v4.6 nuclease) to the cytosine deaminase ecDA1 via the linker peptide XTEN. The amino acid sequence of the ecDA1 deaminase is shown in SEQ ID No. 8, and the amino acid sequence of the linker peptide XTEN is shown in SEQ ID No. 9. The vector expressing the eCg12n-CBE cytosine base editor and the sgRNA expression vector constitute the eCg12n-CBE cytosine base editing system. The nucleotide sequence of the sgRNA backbone in the sgRNA expression vector is shown in SEQ ID No. 13. The eCg12n-CBE cytosine base editor expression vector was constructed using PCR amplification and recombination. Nucleotide fragments expressing the linker peptide XTEN and ecDA1 deaminase were synthesized separately. Primers were designed, and the D297A mutation was introduced by PCR using the Cg12n-v4.6 expression vector (sequence shown in SEQ ID No. 6) as a template. The PCR product was then recombined with the synthesized nucleotide fragments to construct the vector. After confirmation of the vector's correctness by sequencing, the plasmid was extracted for later use. The sequence of the constructed eCg12n-CBE vector is shown in SEQ ID No. 10. Using the same method, the unmodified CgCas12n nuclease (amino acid sequence shown in SEQ ID No. 1) was fused with the linker peptide XTEN and ecDA1 deaminase to construct WT-CBE.

[0096] Simultaneously, an endogenous site sgRNA vector containing a C base in the target sequence was designed and constructed. The endogenous site target sequence used in this embodiment for eCg12n-CBE is shown in Table 5 below. An oligonucleotide chain containing the target sequence was synthesized and inserted into a vector expressing the T6D19 sgRNA backbone (sequence shown in SEQ ID No. 13) by enzyme digestion and ligation to construct an sgRNA expression vector. After the vector sequence was confirmed to be correct by sequencing, the plasmid was extracted.

[0097] Table 5. Endogenous site targeting sequences and amplification primers for evaluating eCg12n-CBE editing properties.

[0098] Locus Target sequence Forward primer Reverse primer VEGFA(1) ctccccgaaagattccaaag acagcaacatggaggaaggg agatgaggaaactgaggcaca VEGFA(2) ttgggccagatctggagcag acagcaacatggaggaaggg agatgaggaaactgaggcaca VEGFA(4) ccggggacccagagatgcat cccaacttccaccctctgtc cactcctctctggccagttg VEGFA(6) gatttggcccagatccctag cccaacttccaccctctgtc cactcctctctggccagttg PDCD1 (4) gacctcctgaaacatatgcc aaaatcaccccgttgcatgg gtaactggtctgggtgggtg HEXA (1) gagctctacaccacacccaa cacacactcagccctgatcc tccttccgtgttctacttccc HEXA (4) agagatgccttactaggtac cacacactcagccctgatcc tccttccgtgttctacttccc

[0099] 4.1.2 Evaluation of the editing performance of eCg12n-CBE in eukaryotic cells.

[0100] HEK293T cells were co-transfected with the constructed eCg12n-CBE and WT-CBE expression vectors, respectively, along with sgRNA vectors containing endogenous target sites. 72 hours post-transfection, cells were digested and sorted by flow cytometry to collect strongly fluorescent positive cells. After cell lysis, DNA fragments containing the target sequence were amplified by PCR and subjected to deep sequencing analysis. Experimental results are shown below. Figure 16 As shown, at the endogenous sites tested in this embodiment, eCg12n-CBE mediated efficient C-to-T base editing in all sites, with C-to-T base editing efficiency approaching 60% at some sites. Conversely, WT-CBE (unmodified wild-type CgCas12n-CBE) did not undergo C-to-T base editing at most sites. At all sites tested in this embodiment, the editing efficiency of eCg12n-CBE was higher than that of WT-CBE. This indicates that the engineered nuclease Cg12n-v4.6 and sgRNA backbone significantly improve the editing efficiency of the eCg12n-CBE developed in this invention.

[0101] This embodiment further analyzes the editing results of the test sites and finds that, taking the target base closest to PAM as the first position, eCg12n-CBE can edit cytosine between positions 2 and 17 of the target site (i.e., the editing window of eCg12n-CBE is C2-C17). In comparison, WT-CBE can mediate C-to-T base editing with target bases located between positions 2 and 13 (i.e., the editing window of WT-CBE is C2-C13). This indicates that the eCg12n-CBE developed in this invention has a wider editing window and can mediate a larger range of C-to-T base editing (results are shown in Figure 1). Figure 17 (As shown).

[0102] 4.1.3 Constructing an embedded eCg12n-CBE to reduce RNA off-target efficiency and improve the safety of base editing.

[0103] Comparing editing efficiency between in-eCg12n-CBE and eCg12n-CBE.

[0104] Considering the potential for off-target RNA editing in CBE, this embodiment constructs an embedded C-to-T base editor, in-eCg12n-CBE, which involves embedding the cytosine deaminase eCDA1 into the sequence of the Cg12n-v4.6 nuclease via the linker peptide XTEN. Specifically, the structure of the Cg12n-v4.6 nuclease protein was predicted using Alphafold3, revealing a loose, disordered region. This region was presumed to be non-essential for the function of the Cg12n-v4.6 protein, and therefore, the eCDA1 deaminase was considered for embedding into this region. Primers were designed for PCR amplification of the nucleotide fragment expressing the eCDA1 deaminase and the eCg12n-CBE vector backbone. The embedded cytosine base editor vector was constructed using homologous recombination. After sequencing confirmation, the plasmid was extracted, and the embedded vector was named in-eCg12n-CBE. The vector sequence is shown in SEQ ID No. 11. HEK293T cells were co-transfected with eCg12n-CBE or in-eCg12n-CBE and sgRNA expressing the target sequence. The editing efficiency of eCg12n-CBE and in-eCg12n-CBE at endogenous sites was compared. The results are as follows: Figure 18 As shown, the editing efficiency of eCg12n-CBE and in-eCg12n-CBE is comparable, indicating that inserting eCDA1 deaminase into the disordered region of the Cg12n-v4.6 protein does not affect the protein's function.

[0105] To investigate the RNA off-target effects of embedded cytosine base editors, HEK293T cells were co-transfected with WT-CBE, eCg12n-CBE, and in-eCg12n-CBE editors and sgRNA. RNA was extracted after flow cytometry sorting and then sequenced to determine the off-target efficiency. Results are as follows: Figure 19 As shown, compared to WT-CBE and eCg12n-CBE, the in-eCg12n-CBE editor has lower RNA off-target efficiency, indicating that the editor is safer.

[0106] 4.2 Construction and Efficiency Evaluation of a Compact Adenine Base Editor Based on the eCg12n Editing System

[0107] In this embodiment, the construction strategy for the compact adenine base editor based on the eCg12n editing system is the same as that described in Example 4.1.1. Specifically, the eCg12n-ABE adenine base editor is constructed by linking the dCg12n-v4.6 nuclease with the adenine deaminase variant TadA-8e-V106W via the linker peptide XTEN. The amino acid sequence of the TadA-8e-V106W adenine deaminase used in this embodiment is shown in SEQ ID No. 14, and the amino acid sequence of the linker peptide XTEN is shown in SEQ ID No. 9. The vector expressing eCg12n-ABE and the sgRNA expression vector constitute the eCg12n-ABE adenine base editing system, and the sgRNA backbone nucleotide sequence is shown in SEQ ID No. 13. As described in Example 4.1.1, the eCg12n-ABE adenine base editor expression vector was constructed using homologous recombination. After the obtained vector was confirmed to be correct by sequencing, the plasmid was extracted for later use. The sequence of the constructed eCg12n-ABE vector is shown in SEQ ID No. 15.

[0108] HEK293T cells were co-transfected with vectors expressing eCg12n-ABE or WT-ABE and sgRNA expression vectors targeting endogenous sites. After cell collection, the cells were lysed and DNA fragments containing the target sequence were amplified by PCR. Editing efficiency was evaluated by next-generation sequencing. The eCg12n-ABE target site sequence and PCR amplification primers used in this example are shown in Table 6 below. Experimental results are as follows: Figure 20 As shown, in this embodiment, eCg12n-ABE was able to mediate effective A-to-G base editing at all 12 endogenous sites tested.

[0109] Table 6 eCg12n-ABE target site sequence and PCR amplification primers

[0110] Locus Targeting sequence Forward primer Reverse primer VEGFA (1) AGTGACAGTATCCTCTGTAT acaggtttttgctctcaagacc tcaagccatgcctttttcgc VEGFA (2) ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ TTTAACTACTTACTGTTTGT agcagccttccccttctttt acaagagggacatggccttag HEXA (3) GACAAGGTTGAACAAACAGT gggtgtttgtcagctgtcatta actctgaatctcccccaacct

[0111] Example 5: Using the eCg12n-CBE editor to induce safe gene knockout through early termination.

[0112] Cytosine base editors (CBEs) can mediate C-to-T base editing, thus enabling the generation of premature stop codons through editing CAG, CGA, CAA, and TGG, thereby achieving gene knockout. Compared to cleavage knockout, CBE-mediated base editing does not produce DNA double-strand breaks, thus offering superior safety and controllability. This embodiment utilizes designed sgRNAs targeting the Ldlr and Dmd genes, respectively, to achieve site-specific knockout of these two genes.

[0113] Mutations in the Ldlr gene can lead to familial hypercholesterolemia. This embodiment designs seven sgRNA sequences (e.g., […]) targeting the Ldlr gene, capable of generating premature stop codons through C-to-T base editing. Figure 21 (As shown in Table 7). The sgRNA sequence targeting the Ldlr gene and the PCR amplification primer information are shown in Table 7. The sgRNA sequences in Table 7 were inserted into the optimal sgRNA backbone T6D19 to construct the corresponding target sgRNA expression vector. Editing was achieved by co-transfecting N2A cells with the eCg12n-CBE expression vector and the sgRNA expression vector. 72 hours after transfection, fluorescently positive cells were sorted by flow cytometry, lysed, and then PCR amplified the DNA fragment containing the target sequence. The amplified product was purified and recovered, and Sanger sequencing was performed to detect the editing of the target site. The experimental results are shown in Table 7. Figure 22 and 23 As shown, the eCg12n-CBE editor can mediate C-to-T editing of the Ldlr gene and successfully generate premature stop codons at some sites. This demonstrates that the eCg12n-CBE editor can mediate the safe knockout of the Ldlr gene.

[0114] The Dmd gene mutation causes Duchenne muscular dystrophy, a rare hereditary muscle atrophy disease. In this embodiment, 24 sgRNA vectors targeting the Dmd gene were designed and constructed. The targeted sgRNA sequences and PCR amplification primers are shown in Table 8. The sgRNA sequences in Table 8 were inserted into the optimal sgRNA backbone T6D19 to construct corresponding target sgRNA expression vectors. Editing was achieved by co-transfecting N2A cells with the eCg12n-CBE expression vector and the sgRNA expression vector, using the transfection system and steps described above. 72 hours after transfection, fluorescently positive cells were sorted by flow cytometry, lysed, and then PCR amplified DNA fragments containing the target sequence. The amplified products were purified and recovered, and Sanger sequencing was performed to detect the editing of the target site. Experimental results are as follows: Figure 24 and 25 As shown, eCg12n-CBE can mediate C-to-T editing of the Dmd gene, with sg2, sg10, sg16, sg19, sg24, and sg24-2 exhibiting editing efficiencies exceeding 10%, especially sg2, which has the highest editing efficiency, followed by sg24. Therefore, sg2 and sg24 sgRNA sites can be selected as preferred sites for disease model preparation. Furthermore, the eCg12n-CBE editor can mediate the safe knockout of the Dmd gene.

[0115] Table 7. sgRNA sequence and PCR amplification primers targeting the Ldlr gene.

[0116] sg Target sequence Forward primer Reverse primer sg1 ACGTGCTCCCAGGATGACTT CAGTTCGCAGCACAACCTAAA AGCAGGGGCTGCTAACGC sg2 TGCATCTCCCCGCAGTTTGT [[ID=]20]CAGTTCGCAGCACAACCTAAA AGCAGGGGCTGCTAACGC sg3 TGTGAGTGCCAGGCCGGCTT GGTGTCCCTTAGCCACTGGA CCAGGGAGGAGCTGCCCA sg4 ATGACCCTGGACCGCAGCGA CGGAGGTGTACAGTAGGGACCT GAGGCAGAAGAAAGGGGTCAC sg5 CCATTTTCAGTGCCAATCGA AATGGTTAGAATGGGAGGCTTTGT ACCCCTGGGAGGAAGAGAAG sg6 GTCACACAGCCTAGAGGTAA AATGGTTAGAATGGGAGGCTTTGT ACCCCTGGGAGGAAGAGAAG sg7 GCTGTCGTAAGTGTTGCTCA GGGAGATGGACCTGGTCTCAC AATGCTTACACACGTGCACG

[0117] Table 8. sgRNA sequence targeting the Dmd gene and PCR amplification primers

[0118] sg Target sequence Forward primer Reverse primer sg1 CCTCGATTCAAGAGTTATGC GTGAGTTTGACAGAAGTTCTTTCCA AGTCTTCAAACAAGTAAAAGAATGGT sg2 AACAGTTTCATGCTCATGAG TCAGCATTTGGAAGCTCCCA TGTTAACATCTCAAAATTAACCCACA sg3 ATCTAGAACAGGAGCAGGTC TGCCATTGTACTTGTAGTTCAAATGA CATGTATTCAGTTGCTTCAAAATTACA sg4 AACATTCAGACAAGTGGCTT TGGCAGCATTTTACTGAAGAACA CAAACAATTGCTTCCTGGTATTACT sg5 AAGAGGCAGATAACTGTGGA ACAAGCAGTGTCTGACCTCC TTGACATTACAAAGTACTCAAGCTT sg6 AGGCAGATAACTGTGGATTC ACAAGCAGTGTCTGACCTCC TTGACATTACAAAGTACTCAAGCTT sg7 CCAGTTAAAAATTTGTAAGg AGAACAACTGAACAGCCGGT TACAGAGCTCAATGCCTGGG sg8 AAAAGGGACAGGGGCCAATG TCACTGGCTATGAACAACTTTCT GGGGACTGGGAAAGGGGATA sg9 GGACAGGGGCCAATGTTTCT TCACTGGCTATGAACAACTTTCT GGGGACTGGGAAAGGGGATA sg10 AGAAAGAGCTACAGCAAGT GCTCTTCAGCCTCAAATTGAGC CAGGAGGGTAGGAGTGGGTGG sg11 AAACTTTCCTCCCAGTTGGT TGAGGCTCGCAAAGTTCTTTG TCAATATCTTTGAAGGACTCTGGGT sg12 GTGGGCAGAAGAATAAAGAGT GTCAAAATTCATGTCGAATGTCATGA ATTCCTATGTGCTAGAAGAAACCT sg13 TTGCTTGAACAGATCA GCTCGAAAGCTACGCAATCTG TGCCCATTCCCAGTTTGGAT sg14 TCCTTGCACTTAATTCAGGA GCTCGAAAGCTACGCAATCTG TGCCCATTCCCAGTTTGGAT sg15 CAGTTGGCAGCTTATATCAC AGTGTAGATTCTTCACGGACTGT ACCAACAGAATCCACTGAAGT sg16 AAACATAACCAGGGGAAGGA AAGGAATACTTGAGGCATTCTTTGC TTGTCTAATAATCAATCACCCTGGGGA sg17 AGTGTTGAACAGGAAGTAAT AGAAATGATAATTGACTGTTTGCATT ACAGTTGTGTAAAATACATTGTTACCA sg18 TAATTCAGTCACAACTAAGT AGAAATGATAATTGACTGTTTGCATT ACAGTTGTGTAAAATACATTGTTACCA sg19 CAGACAGAAAATCCCAAAGA TGCTCCCATATGTTGTAGAATATTTGTG AGAATCACAATAAGGGTTTCTGGT sg20 AGCTTGATGAACGAGTAACA TGCTCCATATGTTGTAGAATATTTGTG AGAATCACAATAAGGGTTTCTGGT sg21 AGATTGAGAAACAGAAGGCT TGACAACCAAACTCATAACCCT GACAAGAAGCTTCCAGCCTCT sg22 GAATTGGAGCAGTTTAACTC GTGTCCTCGGGTACTTTCCA TCCCAGTTCATTCACACTTTCA sg23 CGAGAGGAAATAAAGATAAA ACCAAAATGAAGACTGTACTTGTTGT ACTAAAGCACACATATTGAAATGGA sg24 ATAAAACAGCAGCTGTTACA ACCAAAATGAAGACTGTACTTGTTGT ACTAAAGCACACATATTGAAATGGA

[0119] The above descriptions are merely some embodiments of the present invention. For those skilled in the art, various modifications and improvements can be made without departing from the inventive concept of the present invention, and all such modifications and improvements fall within the scope of protection of the invention.

Claims

1. A Cg12n-v4.6 nuclease, wherein, The amino acid sequence of the Cg12n-v4.6 nuclease is shown in SEQ ID No.

12.

2. The use of the Cg12n-v4.6 nuclease as described in claim 1 in gene editing for purposes other than disease diagnosis and treatment.

3. An eCg12n gene editing system, wherein, The gene editing system contains the Cg12n-v4.6 nuclease as described in claim 1.

4. The eCg12n gene editing system according to claim 3, wherein, The system also includes an sgRNA expression vector, which is constructed by inserting a target sequence into the sgRNA backbone. The nucleotide sequence of the sgRNA backbone is shown in SEQ ID No.

13.

5. A cytosine base editor, wherein, The cytosine base editor includes an eCg12n-CBE vector and an sgRNA expression vector that can mediate C-to-T base editing. The nucleotide sequence of the eCg12n-CBE vector is shown in SEQ ID No.

10. The sgRNA expression vector is constructed by inserting a target sequence into the sgRNA backbone. The nucleotide sequence of the sgRNA backbone is shown in SEQ ID No.

13.

6. A cytosine base editor, wherein, The cytosine base editor includes an in-eCg12n-CBE vector and an sgRNA expression vector that can mediate C-to-T base editing. The nucleotide sequence of the in-eCg12n-CBE vector is shown in SEQ ID No.

11. The sgRNA expression vector is constructed by inserting a target sequence into the sgRNA backbone. The nucleotide sequence of the sgRNA backbone is shown in SEQ ID No.

13.

7. An adenine base editor, wherein, The adenine base editor includes an eCg12n-ABE vector and an sgRNA expression vector that can mediate A-to-G base editing. The nucleotide sequence of the eCg12n-ABE vector is shown in SEQ ID No.

15. The sgRNA expression vector is constructed by inserting a target sequence into the sgRNA backbone. The nucleotide sequence of the sgRNA backbone is shown in SEQ ID No.

13.

8. The use of the eCg12n gene editing system of claim 3 or 4, or the cytosine base editor of claim 5 or 6, or the adenine base editor of claim 7, in gene editing for purposes other than disease diagnosis and treatment.

Citation Information

Patent Citations

  • Base editing system and base editing method

    CN118207185A

  • KR20230007218A