Compact cas9 nuclease and its mediated gene editing system
By discovering and developing compact Cas9 nucleases SedCas9 and PhyCas9 and their sgRNA sequences, the limitations of traditional Cas9 nuclease size and PAM recognition have been overcome, enabling more efficient and wider gene editing applications.
Patent Information
- Application Number
- CN202510660577.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The large size of traditional Cas9 nucleases and their specific PAM recognition sequence limit their use in certain applications, necessitating the development of more compact and efficient gene editing tools.
By combining protein spatial clustering with in vitro DNA cutting experiments, two compact Cas9 nucleases, SedCas9 and PhyCas9, were discovered and identified. sgRNA sequences for their use were developed, and a gene editing system was constructed.
It provides miniaturized gene editing tools that can enter cells more efficiently, recognize specific PAM sequences, expand the application scope of gene editing, and enhance the diversity and flexibility of tools.
Smart Images

Figure BDA0005413668000000071 
Figure BDA0005413668000000072 
Figure BDA0005413668000000081
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gene editing technology, specifically, it relates to two compact Cas9 nucleases and the gene editing technology they mediate. Background Technology
[0002] Gene editing technology plays a crucial role in life science research, agriculture, and medical applications. The CRISPR / Cas system, as a revolutionary gene editing tool, is widely used in gene function research, gene therapy, and bio-breeding due to its high efficiency and precision. Among these, the Cas9 nuclease is one of the most commonly used nucleases in the CRISPR / Cas system; it achieves targeted cleavage of specific DNA sequences by binding to guide RNA (sgRNA).
[0003] However, traditional Cas9 nucleases have some limitations, such as their large protein size and specific PAM (prespacer adjacent motif) recognition sequences, which limit their use in certain applications. To overcome these limitations, researchers have been exploring and developing novel Cas9 homologs in order to obtain more compact and efficient gene editing tools.
[0004] This invention, through protein spatial clustering combined with in vitro DNA cleavage experiments, screened and found that SedCas9 and PhyCas9 exhibit good site-specific genome cleavage activity. The discovery of these novel, compact Cas9 nucleases provides new tools and strategies for the field of gene editing, and is expected to improve the delivery efficiency of gene editing complexes.
[0005] Although various Cas9 homologs have been reported in previous studies, compact Cas9 nucleases with better genome-directed cleavage activity and specific PAM recognition sequences still have significant research and application value. The SedCas9 and PhyCas9 nucleases provided in this invention not only have shorter amino acid lengths but also recognize specific PAM sequences, giving them unique advantages in gene editing applications.
[0006] In summary, the present invention aims to provide two novel compact Cas9 nucleases, SedCas9 and PhyCas9, and sgRNA sequences for use with them, to meet the needs of the gene editing field for efficient and miniaturized gene editing tools. Summary of the Invention
[0007] The present invention aims to provide two novel compact Cas9 nucleases, SedCas9 and PhyCas9, and gene editing systems mediated by them, to overcome the limitations of conventional Cas9 nucleases due to their large size in gene editing applications.
[0008] To achieve this objective, the present invention adopts the following technical solution:
[0009] Discovery of Compact Cas9 Nucleases: This invention identified two compact nucleases, SedCas9 and PhyCas9, homologous to type II-C Cas9, using protein spatial clustering combined with in vitro DNA cleavage experiments. Their amino acid sizes are 994 aa and 983 aa, respectively, and their amino acid sequences are shown in SEQ ID NO. 4 and 5. Both exhibited good genome-directed cleavage activity. The discovery of these two novel compact Cas9 nucleases provides new tools and strategies for the field of gene editing.
[0010] Development of sgRNA sequences: This invention provides sgRNA sequences for use with SedCas9 and PhyCas9. These sgRNAs can guide compact Cas9 nucleases to perform targeted cleavage, achieving efficient gene editing. Preferably, the scaffold sequence of the sgRNA of the SedCas9 nuclease is shown in SEQ ID NO. 8-13, and the scaffold sequence of the sgRNA of the PhyCas9 nuclease is shown in SEQ ID NO. 14-17.
[0011] Construction of a gene editing system: This invention constructs a gene editing system based on SedCas9 and PhyCas9, which includes binding sgRNA with a compact Cas9 nuclease to form a complex for genome editing.
[0012] Development of gene editing methods: This invention provides a method for gene editing using the above-mentioned gene editing system, which includes binding sgRNA with a compact Cas9 nuclease to form a complex, guiding the nuclease to cut the target DNA sequence, thereby achieving gene editing.
[0013] Kit Development: This invention provides a kit for gene editing using the above-mentioned compact Cas9 nuclease, comprising a compact Cas9 nuclease and sgRNA used in conjunction with it.
[0014] The technical solution of the present invention has the following beneficial effects:
[0015] 1. Miniaturized gene editing tools: SedCas9 and PhyCas9, as compact Cas9 nucleases, have smaller protein sizes, enabling them to enter cells more efficiently and improve the efficiency of gene editing.
[0016] 2. Expanding the scope of gene editing applications: SedCas9 and PhyCas9 recognize specific PAM sequences, which makes them more widely applicable in gene editing and can be used to edit more gene sites.
[0017] 3. Enhancing the diversity and flexibility of gene editing tools: The two compact Cas9 nucleases and their mediated gene editing systems provided by this invention offer researchers more options, enabling them to select appropriate tools according to specific research needs, thereby enhancing the diversity and flexibility of gene editing. Attached Figure Description
[0018] Figure 1 This study utilizes bioinformatics techniques to discover novel compact endonucleases, specifically Cas9. A: Phylogenetic analysis of candidate compact endonucleases PirCas9, AciCas9, Phy2Cas9, SedCas9, and PhyCas9 predicted using protein spatial clustering. B: Comparison analysis of amino acid sequence similarity between candidate Cas9 homologs and known CjCas9 homologs is presented.
[0019] Figure 2 Locus and sequence similarity analysis of candidate compact Cas9. A: Locus distribution of candidate compact Cas9 homologs and transcription direction pattern of guide RNA are shown. B: Three-dimensional spatial structure of candidate compact Cas9 predicted using AlphaFold3 combined with Pymol is presented, where the red part represents the Ruvc domain. C: DR sequence conservation analysis of candidate compact Cas9 is shown.
[0020] Figure 3 In bacteria, the characteristics of candidate compact Cas9 endonucleases recognizing PAM were identified by PAM library subtraction assays. The PAM motifs recognized by SedCas9 and PhyCas9 were NRRRRH and NNNTCC, respectively (where N represents A, T, C, G, R represents A or G, and H represents A, T, or C).
[0021] Figure 4 The targeting and cleavage capabilities of PirCas9, AciCas9, Phy2Cas9, SedCas9, and PhyCas9 endonucleases on target DNA were detected by gel electrophoresis. In the experiment, the targets were the PCR amplification products of the human FANCF gene, and the corresponding PAM sequences were CCN, CCNC, NGG, NRRRRH, and NNNTCC, respectively.
[0022] Figure 5Engineered SedCas9 and PhyCas9 sgRNA scaffolds. A: Shows the sequence alignment and in vitro DNA cleavage activity assay of the optimized SedCas9 sgRNA; B: Shows the sequence alignment and in vitro DNA cleavage activity assay of the optimized PhyCas9 sgRNA; PCR amplification products targeting human EXM1 and FANCF genes. Note: Spacer sequence and length are consistent; mutated tracrRNA sequence; sequence alignment diagrams were drawn by SnapGene. Detailed Implementation
[0023] Terminology Explanation
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0025] The protospacer adjacent motif (PAM) is a short DNA sequence (typically 2-6 base pairs in length). Generally, the PAM is essential for cleavage by Cas endonucleases and is usually located 3-4 nucleotides downstream of the cleavage site. Many different Cas endonucleases can be purified from different bacteria, and each enzyme may recognize a different PAM sequence.
[0026] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments that do not specify specific conditions are generally performed under conventional conditions.
[0027] Example 1: Mining Compact Cas9 Homologous Nucleases Based on Bioinformatics Strategies
[0028] In this embodiment, a bioinformatics workflow developed by the inventors was used to obtain the amino acid sequences and three-dimensional structural data of known Cas proteins from public databases (such as UniProt and NCBI), with a focus on protein family members containing the RuvC domain. The RuvC domain was annotated using tools (such as Pfam and InterPro) to determine its conserved sequence motifs (e.g., the DEDD motif of the catalytic triplet) and spatial topology. Based on the three-dimensional spatial coordinates of the RuvC domain, key feature parameters were extracted, including the spatial distance between domains (e.g., the interaction interface between RuvC and REC leaves).
[0029] In clustering algorithms, density-based spatial clustering algorithms (such as DBSCAN) are used to group RuvC domains with similar spatial characteristics. Specific parameters include: neighborhood radius (Eps), which defines the spatial distance threshold of key residues within the RuvC domain; and minimum density (MinPts), ensuring functional homology among members of each cluster. By comparing the functions of known Cas proteins (such as DNA cleavage and trans-cleavage activities), the potential functions of unknown proteins within the clusters can be predicted. For example, if a cluster contains the RuvC domain of Cas9, it is inferred that its members may have similar endonuclease activities.
[0030] Further evolutionary analysis involved constructing a phylogenetic tree to assess the evolutionary conservation and functional differentiation potential of the candidate proteins, ultimately identifying five novel, previously unknown bacterial proteins. Phylogenetic analysis revealed that these new bacterial proteins reside on different CRISPR-Cas9 phylogenetic branches. Figure 1 A) It is speculated that these may be novel RNA-guided Cas9 endonucleases. For ease of subsequent research, based on their bacterial species origin, the inventors named these new, unknown bacterial proteins Pirellulales (PirCas9), Acidobacteriota (AciCas9), Phycisphaerae (Phy2Cas9), Sedimentisphaerales (SedCas9), and Phycisphaerae (PhyCas9). The amino acid sizes encoded by these five Cas9 proteins are 989aa, 983aa, 990aa, 994aa, and 983aa, respectively, and their amino acid sequences are shown in SEQ ID NO. 1, 2, 3, 4, and 5, which are roughly equivalent to the known 984aa amino acid count of CjCas9.
[0031] Subsequently, the inventors used the Blast program to compare the sequence similarity of these five newly discovered bacterial proteins with CjCas9. The results showed that the amino acid sequence conservation of the five candidate proteins PirCas9, AciCas9, Phy2Cas9, SedCas9, and PhyCas9 with the known CjCas9 was 25.4%, 26.1%, 26.9%, 25.3%, and 24.1%, respectively. Figure 1 B). These results indicate that the candidate Cas9 protein is small in size and has relatively low sequence conservation, therefore, it is urgent to experimentally identify whether it has genome editing activity.
[0032] Example 2: Bioinformatics strategy analysis of loci, protein 3D structure and DR sequence of candidate compact Cas9
[0033] In this embodiment, the inventors analyzed the loci of these proteins using CRISPRCasFinder software. The results showed that the compact Cas9 homolog possessed a CRISPR array sequence containing multiple repeat sequences and spacer sequences, as well as Cas1 and Cas2 proteins. Next, the inventors used the RNAfold web server (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi) to predict candidate tracrRNA sequences found through the DR sequences of candidate compact Cas9 homologs, thus determining the sgRNA-tracrRNA transcription direction of the newly predicted bacterial protein. Figure 2 A).
[0034] To accurately resolve the spatial structures of these proteins, AlphaFold3 was used for structure prediction, and the prediction results were visualized using PyMol software. During the analysis, CjCas9 was used as a reference, and the RuvC domain of the compact Cas9 protein was highlighted. Figure 2 B). To further investigate the structure-function relationship of these proteins, the DR (Direct Repeat) sequences of known Cas9 nucleases were compared and analyzed. Figure 2 C). These results provide data support for evaluating the gene-editing activity of candidate compact Cas9 endonucleases.
[0035] Example 3: Detection of PAM sequences for SedCas9 and PhyCas9 protein recognition
[0036] In this embodiment, the PAM sequence recognized by the candidate compact Cas9 nuclease was identified by a bacterial PAM library reduction experiment.
[0037] The construction process for the randomized mixed PAM vector library is as follows: synthesize the DNA oligo sequence 5'-GGCCAGTGAATTCGAGCTCGGTACCCGGG ACTTTAAAAGTATTCGCCAT NNNNNNNAGCTTGGCGTAATCATGGTCATAGCTGTTT-3', where N is a random deoxyribonucleotide. Using Oligo-F: 5'-GGCCAGTGAATTCGAGCTCGG-3' and Oligo-R: 5'-AAACAGCTATGACCATGATTACGCCAA-3' as upstream and downstream primers, after PCR amplification, the ligation into the pUC19 vector via homologous recombination, followed by transformation into E. coli and plasmid extraction, yields a random mixed PAM library. The sgRNA sequences used for the candidate compact Cas9 are: PirCas9: 5'- GACTTTAAAAGTATTCGCCATGUUGCGGAUUGGUCGCAGGACGGGAUCGACUACACUGUUAGUGUAGUCGAUCUCGUUCUGCGAUGCUUUUCGUAACAAGACAAUCGUCUAACGACGAAACAUUCGCAGGGCAAAGCCCCACGGGGCUCCCG-3';AciCas9:5'- GACTTTAAAAGTATTCGCCAT GUUGUGAGUUGCGGCGAUUCUCUUAUCUGCUAAACUUGCGUUUAGUAGAUAAGAGAAUACGCCGCUUCUUAUAACAAGUUAGUUUUUCGAAGCUGACGUAGGAGAUCCGAAAGGAACCUAU-3'; PhyCas9: 5'- GACTTTAAAAGTATTCGCCAT GUUGUGGCUUGCACACAGCCGGGUCAGUUACAAUUCAAUUGUAACUGACCUGUGCUUUUGUGCUUGUCAUAACAAGUGCGAAAGCACGCGGACCACAGCCGGCGAAAGCCGGCUGUUCC-3';SedCas9: 5'- GACTTTAAAAGTATTCGCCAT GUUGUGACUUGCACUCCGACACGGAUCAGUUAUAUGAUCCGUGUCUGCGUGCUUGUCAUAACAAGUAAGAUUUCGCAAGAAAUCUCGCAGGCACUGCCCCAUUGGGCACUCCUACGGUGCUCAAUGGGGUAAACC-3'; PhyCas9: 5'- GACTT TAAAAGTATTCGCCAT GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACUAGACGUGUAACCAGCCUCGGUUCGAGUCUUUUCACAACAAGUGAACAAGAUUCCGCCAGGCAAUGCCUCACCUUGACGGGUGGGGCAUCCUUUUUU-3' (The underlined area is the target recognition sequence).
[0038] In the bacterial PAM library reduction experiment, the constructed vector pACYC-Duet-1-Cas9-sgRNA, co-expressing the candidate compact Cas9 protein and sgRNA, was first transformed into DE3(BL21) competent cells to prepare stably expressing bacterial strains. Simultaneously, a stably transformed bacterial strain without the sgRNA expression vector pACYC-Duet-1-Cas9 was constructed as a negative control. Next, 100 ng of the PAM library plasmid was electroporated into these stably expressing bacterial strains, and selection was performed using ampicillin and chloramphenicol-treated medium. Figure 3 A). After 16 hours, the colonies on the culture medium were scraped off and plasmids were extracted. Then, using 100 ng of the extracted plasmid as templates, PCR amplification was performed using library sequencing primers Seq-F: 5'-GGCCAGTGAATTCGAGCTCGG-3' and PAM-Seq-R: 5'-CAATTTCACACAGGAAACAGCTATGACC-3'. After product recovery, next-generation high-throughput sequencing was performed on the experimental and control groups, and the sequencing results were analyzed and displayed using WebLogo 3.0.
[0039] To identify the PAM sequence characteristics recognized by the compact Cas9 protein, 16,384 different types of PAM sequences in the starting vector library were statistically analyzed. The frequency of each sequence in the experimental and control groups during high-throughput sequencing was calculated, and the sequences were normalized using the total number of PAM sequences in each group. The change in PAM consumption was calculated as log2(normalized value of control group / normalized value of experimental group). A value greater than 3.5 was considered a significantly consumed PAM. Subsequently, WebLogo 3.0 was used to visualize the base frequencies at various positions in the significantly consumed PAM sequences.
[0040] Experimental results showed that the three nucleases, PirCas9, AciCas9, and Phy2Cas9, did not exhibit significant changes in their PAM sequence recognition preferences (e.g., Figure 3 As shown in BD), SedCas9 and PhyCas9 exhibit significant differences in PAM sequence recognition preference (e.g., As shown in EF, this indicates that SedCas9 and PhyCas9 have in vitro cleavage activity, and the identified PAM sequences are 3'-“NRRRRH”-5' and 3'-“NNNTCC”-5', respectively (N represents A, T, C or G, R represents A or G, and H represents A, T or C). This is inconsistent with the known PAM motif of compact CjCas9, thus expanding the toolbox of compact Cas9.
[0041] Example 4: Compact SedCas9 and PhyCas9 endonucleases exhibit in vitro DNA-targeting cleavage activity.
[0042] In this embodiment, the in vitro cleavage activities of PirCas9, AciCas9, Phy2Cas9, SedCas9, and PhyCas9 endonucleases on target DNA were tested using in vitro experiments. The compact Cas9 protein is guided by sgRNA paired with the target nucleic acid to recognize and bind to the target nucleic acid, thereby stimulating the protein's cleavage activity on the target nucleic acid. Figure 3 A). Next, agarose gel electrophoresis was performed to observe changes in the size of the target band in order to detect genome cleavage efficiency.
[0043] In this embodiment, PirCas9, AciCas9, and Phy2Cas9 selected human FANCF gene as the target double-stranded DNA (dsDNA), with PAM values of 3'-"CCN"-5', 3'-"CCNA"-5', and 3'-"NGG"-5', respectively, and the sequences are as follows: Bold text indicates PAM, and underlined text indicates the target sequence. The crRNA-CCG sequence corresponding to PirCas9 is: Figure 4 CUUUUGAC GUUGCGGAUUGGUCGCAGGACGGGAUCGACUACACU, crRNA-CCT is: UUUAGUGACUAG GGUCAACGUUUGCAC GUUGCGGAUUGGUCGCAGGACGGGAUCGACUACACU, crRNA-CCC is: UAUGA GUUGCGGAUUGGUCGCAGGACGGGAUCGACUACACU (underlined is the target sequence). The TracrRNA sequence corresponding to PirCas9 is: GUUAGUGUAGUCGAUCUCGUUCUGCGAUGCUUUUCGUAACAAGACAAUCGUCUAACGACGAAACAUUCGCAGGGCAAAGCCCCACGGGGCUCCCG. The crRNA-CCTC sequence corresponding to AciCas9 is: CAUUUGGGUUGGAACUGAGU GUUGUGAGUUGCGGCGAUUCUCUUAUCUGCUAAACU, crRNA-CCAC is UUGCACAAUAGGUUUCAAAGGUUGUGAGUUGCGGCGAUUCUCUUAUCUGCUAAACU, crRNA-CCTC-2 is AGGAAGUGAUUGGAAGUACU GUUGUGAGUUGCGGCGAUUCUCUUAUCUGCUAAACU (underlined is the target sequence). The TracrRNA sequence corresponding to AciCas9 is: UGCGUUUAGUAGAUAAGAGAAUACGCCGCUUCUUAUAACAAGUUAGUUUUUCGAAGCUGACGUAGGAGAUCCGAAAGGAACCUAU. The crRNA-AGG sequence corresponding to Phy2Cas9 is: GGUCUUAGCAUCUGGACAGA GUUGUGGCUUGCACACAGGCACGGGGUCAGUUACAAU, crRNA-TGG is: CGCAGGCCTCAGTTCTGTAT GUUGUGGCUUGCACACAGGCACGGGGUCAGUUACAAU, crRNA-AGG-2 is: AGAUAAAGUUCUAACUGCCC GUUGUGGCUUGCACACAGGCACGGGUCAGUUACAAU. The corresponding TracrRNA sequence for Phy2Cas9 is: UCAAUUGUAACUGACCUGUGCUUUUGUGCUUGUCAUAACAAGUGCGAAAGCACGCGGACCACAGCCGGCGAAAGCCGGCUGUUCC.
[0044] In this embodiment, the target double-stranded DNA (dsDNA) selected by SedCas9 is the human FANCF gene, with PAM 3'-"NGGGGA"-5', and its sequence is as follows: The bolded markings are PAM. The crRNA-GGAGAC sequence corresponding to SedCas9 is: AUAGUUUGCUUUUUAAUAGA TUAGCAGACCCAG GUUGUGACUUGCACUCCGACACGGAUCAGUUAUA, crRNA-AGAAGC sequence is AUAGACA CAAAGACUUCCGAAUU GUUGUGACUUGCACUCCGACACGGAUCAGUUAUA, crRNA-GGAGGA sequence is CCCC GGUUGUGACUUGCACUCCGACACGGAUCAGUUAUA (underlined is the target sequence); the TracrRNA sequence corresponding to SedCas9 is UGAUCCGUGUCUGCGUGCUUGUCAUAACAAGUAAGAUUUCGCAAGAAAUCUCGCAGGCACUGCCCCAUUGGGCACUCCUACGGUGCUCAAUGGGGUAAACC.
[0045] In this embodiment, the target double-stranded DNA (dsDNA) selected by PhyCas9 is the human FANCF gene, with PAM 3'-"NNNTCC"-5', and its sequence is as follows: Bold text indicates PAM, and underlined text indicates the target sequence. The corresponding crRNA-GGAACC sequence for PhyCas9 is: GAGUCCCAAGAUGUGCCCU The crRNA-GCCACC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. CGUCGGCCCCAAGAAGAGUU The crRNA-TGTACC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. UUUGACUUUAGUGACUAGCC The crRNA-CCCTCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. GUAAGAAUGUUGAAAAUAUG The crRNA-TACTCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. CUUUGCUGCCUUUUGUCGCG The crRNA-GAGTCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. AUAGAGGAAGUGAUUGGAAG The crRNA-CTGCCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. AAAGGCAUUUGGGUUGGAACU The crRNA-AGACCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. UCUGAAAGAUAAAGUUCUAA The crRNA-CCTCCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. CAGAACUGAGGCCUGCGCUG The crRNA-TATGCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. UUGAAACCUAUUGUGCAACUThe crRNA-CATGCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. UACAAUGUUCUCACCAAAUA The crRNA-AAGGCC sequence is GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU. CCAGAAAAUCCGUGACACUA GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU (underlined is the target sequence). The corresponding TracrRNA sequence for PhyCas9 is: ACGUGUAACCAGCCUCGGUUCGAGUCUUUUCACAACAAGUGAACAAGAUUCCGCCAGGCAAUGCCUCACCUUGACGGGUGGGGCAUCCUUUUUU.
[0046] Using HEK293 cell genome as a template, PCR amplification was performed using primers FANCF-F: 5'-CGCTTGCCTCAGAACAACTT-3' and FANCF-R: 5'-CAAACTCCAGATAGGCCAACAG-3' to obtain FANCF gene double-stranded DNA. Next, DNA sequences encoding PirCas9, AciCas9, Phy2Cas9, SedCas9, or PhyCas9 endonucleases were synthesized after E. coli codon optimization, and NLS nuclear localization signals were added to their C-termini. These sequences were then ligated into the pET-28a prokaryotic expression vector, transformed into E. coli BL21 strain, and after identifying positive clones, IPTG-induced expression was performed. The target protein was then purified by affinity chromatography.
[0047] The Cas9 in vitro cleavage reaction system was as follows: 2 μL of 10×CutSmart Buffer, 1 mM DTT, 500 ng of predicted Cas9-NLS-His protein, 500 ng of sgRNA, and 2 μL of FANCF target amplification product. The reaction was incubated at 37℃ for 30 min. After the reaction, 1 μL of proteinase K was added, and the reaction was terminated by incubation at 60℃ for 10 min. The experimental group received both sgRNA and target nucleic acid, while the control group received no sgRNA. The results were detected by 1.5% agarose gel electrophoresis.
[0048] The results are as follows AAUAGUUUGCUUUUUAAUAG As shown in Figure B, the newly discovered PirCas9, AciCas9, and Phy2Cas9 lack double-strand cleavage activity, while the newly discovered SedCas9 and PhyCas9 exhibit in vitro cleavage activity. Figure 4(CD). In comparison, the DNA-targeting cleavage activities of SedCas9 and PhyCas9 endonucleases are comparable, both exhibiting two distinct cleavage target bands. This confirms that both of these newly discovered compact Cas9 endonucleases possess DNA-targeting cleavage activity.
[0049] Example 5: Engineered sgRNA enhances SedCas9 and PhyCas9 gene editing activity
[0050] In this embodiment, to improve the gene editing efficiency mediated by SedCas9 and PhyCas9, the paired sgRNAs were engineered. On one hand, the structure of the sgRNAs was optimized by shortening their length; on the other hand, the aim was to enhance gene editing activity. After truncating the sgRNAs, their folded secondary structures were characterized using the RNAfold web server (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi), and sgRNA mutants that met the typical stem-loop structure characteristics of compact Cas9 proteins were screened. Subsequently, these mutants were purified by in vitro transcription and then synergized with SedCas9 and PhyCas9 to recognize and bind to target nucleic acids. Gene editing activity was evaluated using in vitro DNA cleavage experiments.
[0051] In this embodiment, the target double-stranded DNA (dsDNA) used is a partial sequence of the human EXM1 and FANCF genes. The target double-stranded DNA (dsDNA) selected by SedCas9 is the human EXM1 gene, with PAM of 3'-"NGGGGA"-5', and its sequence is as follows: Bold text indicates PAM, and underlined text indicates the target sequence. The crRNA sequence corresponding to SedCas9 is... Figure 4 GUUGUGACUUGCACUCCGACACGGAUCAGUUAUA (underlined is the target sequence); the tracRNA sequence corresponding to SedCas9 is UGAUCCGUGUCUGCGUGCUUGUCAUAACAAGUAAGAUUUCGCAAGAAAUCUCGCAGGCACUGCCCCAUUGGGCACUCCUACGGUGCUCAAUGGGGUAAACC.
[0052] The sgRNA sequences engineered for the SedCas9 endonuclease are as follows:
[0053] sgRNA WT: CCCAUAGGGAAGGGGGACAC GUUGUGACUUGCACUCCGACACGGAUCAGUUAUAUGAUCCGUGUCUGCGUGCUUGUCAUAACAAGUAAGAUUUCGCAAGAAAUCUCGCAGGCACUGCCCCAUUGGGCACUCCUACGGUGCUCAAUGGGGUAAACC (The underlined area is the target area).
[0054] sgRNA MT1: CCCAUAGGGAAGGGGGACAC GUUGUGACUUGCACUCCGACUCAGUUAUAUGAGUCUGCGUGCUUGUCAUAACAAGUAAGAUUUCGCAAGAAAUCUCGCAGGCACUGCCCCAUUGGGCACUCCUACGGUGCUCAAUGGGGUAAACC (The underlined area is the target area).
[0055] sgRNA MT2: CCCAUAGGGAAGGGGGACAC GUUGUGACUUGCACUCCGACACGGAUCAGUUAUAUGAUCCGUGUCUGCGUGCUUGUCAUAACAAGUAAGCGCAAGCUCGCAGGCACUGCCCCAUUGGGCACUCCUACGGUGCUCAAUGGGGUAAACC (The underlined area is the target area).
[0056] sgRNA MT3: CCCAUAGGGAAGGGGGACAC GUUGUGACUUGCACUCCGACACGGAUCAGUUAUAUGAUCCGUGUCUGCGUGCUUGUCAUAACAAGUAAGAUUUCGCAAGAAAUCUCGCAGGCACUGCACUCCUACGGUGUAAACC (The underlined area is the target area).
[0057] sgRNA MT4: CCCAUAGGGAAGGGGGACAC GUUGUGACUUGCACUCCGACUCAGUUAUAUGAGUCUGCGUGCUUGUCAUAACAAGUAAGCUCGCAGGCACUGCCCCAUUGGGCACUCCUACGGUGCUCAAUGGGGUAAACC (The underlined area is the target area).
[0058] sgRNA MT5: CCCAUAGGGAAGGGGGACACGUUGUGACUUGCACUCCGACUCAGUUAUAUGAGUCUGCGUGCUUGUCAUAACAAGUAAGCGCAAGCUCGCAGGCACUGCACUCCUACGGUGUAAACC (The underlined area is the target area).
[0059] PhyCas9 selected the human FANCF gene as its target double-stranded DNA (dsDNA), with PAM 3'-"NNNTCC"-5', and its sequence is as follows: Bold markings indicate PAM, and underlined markings indicate the target sequence.
[0060] The crRNA sequence corresponding to PhyCas9 is: CCCAUAGGGAAGGGGGACAC GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACU (underlined is the target sequence), the tracrRNA sequence corresponding to PhyCas9 is: ACGUGUAACCAGCCUCGGUUCGAGUCUUUUCACAACAAGUGAACAAGAUUCCGCCAGGCAAUGCCUCACCUUGACGGGUGGGGCAUCCUUUUUU.
[0061] The sgRNA sequences engineered for the PhyCas9 endonuclease are as follows:
[0062] sgRNAWT: AUAGAGGAAGUGAUUGGAAG GCUGUGGAUUGAUCUCGGGCCGGGGCUGGUUACACUAGACGUGUAACCAGCCUCGGUUCGAGUCUUUUCACAACAAGUGAACAAGAUUCCGCCAGGCAAUGCCUCACCUUGACGGGUGGGGCAUCCUUUUUU (The underlined area is the target area).
[0063] sgRNAMT1: AUAGAGGAAGUGAUUGGAAG GCUGUGGAUUGAUCUCGGGCCGGGGCGAAAGCCUCGGUUCGAGUCUUUUCACAACAAGUGAACAAGAUUCCGCCAGGCAAUGCCUCACCUUGACGGGUGGGGCAUCCUUUUUU (The underlined area is the target area).
[0064] sgRNAMT2:AUAGAGGAAGUGAUUGGAAG GCCGCGGAUUGAUCUCGGAAACGAGUCUUUUCACAACAAGUGAACAAGAUUCCGCCAGGCAAUGCCUCACCUUGACGGGUGGGGCAUCCUUUUUU (The underlined area is the target area);
[0065] sgRNAMT3: AUAGAGGAAGUGAUUGGAAG GCCGCGGAUUGAUCUCGGAAACGAGUCUUUUCACAACAAGUGAACAAGAUUCCGCCAGGCAAUGCCUCACCCUGACGGGUGGGGCAUCCUUUUUU (The underlined area is the target area);
[0066] sgRNAMT4: AUAGAGGAAGUGAUUGGAAG GCUGUGGAUUGAUCUCGGAAACGAGUCUUCAACAAGCAAGAUUCCGCCAGGCAAUGCCUCACCUUGACGGGUGGGGCAUCCUUUUUU (The underlined area is the target area.)
[0067] Using HEK293 cell genome as a template, PCR amplification was performed using primers EXM1-F: 5'-CATTGACAGAGGGACAAGCAATG-3', EXM1-R: 5'-CTGTCTCTGCAGTCTCAACC-3', FANCF-F: 5'-CGCTTGCCTCAGAACAACTT-3', and FANCF-R: 5'-CAAACTCCAGATAGGCCAACAG-3' to obtain double-stranded DNA of the EXM1 and FANCF genes. Next, DNA sequences encoding SedCas9 and PhyCas9 were synthesized after codon optimization for *E. coli*, and NLS nuclear localization signals were added to their C-termini, respectively. The DNA sequences are shown in SEQ ID NO. 18 and 19. These sequences were then ligated into the pET-28a prokaryotic expression vector, transformed into *E. coli* strain BL21, and after identification of positive clones, IPTG-induced expression was performed. The target proteins were then purified by affinity chromatography.
[0068] The in vitro cleavage reaction system was as follows: 2 μL of 10×CutSmart Buffer, 1 mM DTT, 500 ng of the predicted Cas9-NLS tag protein, 500 ng of engineered sgRNA, and 2 μL of FANCF or EXM1 target amplification product. The reaction was incubated at 37℃ for 30 min. After the reaction, 1 μL of proteinase K was added, and the reaction was terminated by incubation at 60℃ for 10 min. The experimental groups received the corresponding sgRNA and target nucleic acid, while the control group received no sgRNA. The results were analyzed by 1.5% agarose gel electrophoresis, and the changes in bands in each group were detected using a UV irradiation system.
[0069] The results are as follows AUAGAGGAAGUGAUUGGAAG As shown in Figure A, the MTsgRNAs engineered to target SedCas9 sgRNA all exhibited target cleavage activity and were shorter in length. For example... Figure 5 Figure 5 As shown in Figure B, the engineered sgRNAs of PhyCas9 exhibit target cleavage activity, with sgRNA MT1, sgRNA MT2, and sgRNA MT3 showing no gene editing activity, while sgRNA MT4 showed no detected gene editing activity. This demonstrates that the modified truncated sgRNAs can be used in conjunction with SedCas9 or PhyCas9, expanding the application scope of the compact CRISPR-Cas9 system.
[0070] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-described technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A compact Cas9 nuclease, characterized in that, The nuclease is either the SedCas9 nuclease with the amino acid sequence shown in SEQ ID NO.4 or the PhyCas9 nuclease with the amino acid sequence shown in SEQ ID NO.
5.
2. A polynucleotide, characterized in that, The polynucleotide is the polynucleotide encoding the compact Cas9 nuclease of claim 1.
3. The polynucleotide according to claim 2, characterized in that, The polynucleotide sequence encoding the SedCas9 nuclease is shown in SEQ ID NO.6, and the polynucleotide sequence encoding the PhyCas9 nuclease is shown in SEQ ID NO.
7.
4. A carrier, characterized in that, The vector comprises the polynucleotide of claim 2.
5. The application of the compact Cas9 nuclease of claim 1, the polynucleotide of claim 2, or the vector of claim 4 in gene editing.
6. A gene editing system or kit, characterized in that, It includes the PhyCas9 nuclease or SedCas9 nuclease as described in claim 1, and the sgRNA bound to the nuclease.
7. The gene editing system or kit according to claim 6, characterized in that, The scaffold sequence in the sgRNA of the SedCas9 nuclease is shown in SEQ ID NO. 8-13, and the scaffold sequence in the sgRNA of the PhyCas9 nuclease is shown in SEQ ID NO. 14-17.
Citation Information
Patent Citations
Cas9 protein, gene editing system containing Cas9 protein and application
CN113652411A
II-B type endonuclease mediated gene editing system and application
CN119040297A