VpCas9 protein, single site mutant of VpCas9 protein and application of VpCas9 protein in gene editing
Through deep learning models, the Cas9 protein was modified into VpCas9 and designed unit point mutants, which solved the limitations of Cas9 protein in terms of gene editing efficiency and PAM compatibility, and achieved more efficient genome editing effects.
Patent Information
- Application Number
- CN202510725509.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing Cas9 protein has limitations in gene editing efficiency, PAM compatibility and specificity, affecting its effectiveness and flexibility in a wider range of genome editing applications.
Cas9 protein was mined and engineered through deep learning models to obtain VpCas9 protein and its unit point mutants, and optimize their amino acid sequences to improve editing efficiency and PAM compatibility.
The VpCas9 protein and its unit point mutants show better editing effects in gene editing, extending its application prospects in plant, animal and microbial genome editing.
Smart Images

Figure CN120249251A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to homologs of Cas9 protein and their applications, in particular to the VpCas9 protein, a homolog of Cas9, its single-site mutants, and their applications in genome editing, belonging to the field of Cas9 homologs and their applications. Background Art
[0002] Gene editing refers to operations such as deletion, replacement, and insertion of target genes to obtain new functions or phenotypes. As an important research means with rapid development in life sciences, gene editing technologies represented by CRISPR-Cas9 are one of the current technical means with the highest efficiency, simplest operation, lowest cost, and easiest to master.
[0003] However, existing Cas9 proteins have certain limitations to varying degrees in terms of editing efficiency, PAM (protospacer adjacent motif) compatibility, and specificity, which affect their effectiveness and flexibility in broader genome editing applications and urgently need to be improved. Summary of the Invention
[0004] One object of the present invention is to provide a new homolog of Cas9, the VpCas9 protein, and its coding gene; Another object of the present invention is to provide single-site mutants of the VpCas9 protein, a homolog of Cas9, and their coding genes; A third object of the present invention is to provide a vector and a host cell containing the said coding gene; A fourth object of the present invention is to provide a CRISPR-Cas system, which includes the VpCas9 protein, a homolog of Cas9, or its mutant; A fourth object of the present invention is to apply the said homolog of the VpCas9 protein or its mutant, the vector containing the said coding gene, or the CRISPR-Cas system including the VpCas9 protein, a homolog of Cas9, or its mutant to gene editing, editing of target nucleic acids, gene cleavage, or preparation of drugs for detection of broken target nucleic acids or targeted gene therapy, etc.
[0005] The above objects of the present invention are achieved by the following technical solutions: One aspect of the present invention is to provide a homolog of Cas9, the VpCas9 protein, whose amino acid sequence is shown as SEQ ID No.5.
[0006] The present invention obtains the homolog of Cas9, the VpCas9 protein, through mining and modification by a deep learning model. Compared with SpCas9, it has better editing effects in terms of gene editing efficiency, PAM compatibility, or specificity.
[0007] The second aspect of the present invention is to provide a single-site mutant of the VpCas9 protein, a homolog of Cas9; wherein the single-site mutant is obtained by subjecting the amino acid sequence shown in SEQ ID No.5 to any one of V623I, E372K, T437P, I526V, F573L, D1080G, S438F, I439L or A110D amino acid single-site mutation; preferably, the amino acid sequence shown in SEQ ID No.5 is subjected to V623I amino acid single-site mutation to obtain a single-site mutant, and the gene editing cleavage activity of the single-site mutant is most significantly improved compared with the wild-type protein.
[0008] In the present invention, the single-site mutant "V623I" means that the 623rd amino acid of the amino acid sequence shown in SEQ ID NO.5 is mutated from valine (Val, V) to isoleucine (Ile, I); the expressions of the remaining single-site mutants of the present invention are similar.
[0009] Another aspect of the present invention is to provide a gene encoding a Cas9 homolog VpCas9 protein or a gene encoding a mutant of a Cas9 homolog VpCas9 protein.
[0010] Another aspect of the present invention is to provide a vector, which contains the coding gene and a regulatory element operably connected to the coding gene; wherein the vector can be an expression vector, a cloning vector or a shuttle vector, etc.
[0011] In a preferred embodiment, the regulatory element is selected from one or more of a promoter, a terminator, an enhancer, a transposon, a leader sequence or a marker gene.
[0012] Another aspect of the present invention is to provide a CRISPR-Cas system, which includes a Cas protein and at least one sgRNA; the Cas protein can bind to the sgRNA, and the sgRNA includes a direct repeat sequence and a spacer sequence that can hybridize with a target nucleic acid, wherein the Cas protein is the above-mentioned Cas9 homolog VpCas9 protein or a mutant thereof.
[0013] Another aspect of the present invention provides a kit for gene editing or gene cutting, which includes the above-mentioned Cas9 homolog VpCas9 protein or its mutant, or a polynucleotide encoding Cas9 homolog VpCas9 protein or its mutant, or a vector containing the polynucleotide sequence, or a CRISPR-Cas system containing the above-mentioned Cas9 homolog VpCas9 protein or its mutant.
[0014] Another aspect of the present invention is to apply the homolog VpCas9 protein of Cas9 or its mutant, or the polynucleotide encoding the homolog VpCas9 protein of Cas9 or its mutant, or the vector containing the polynucleotide sequence, or the CRISPR-Cas system containing the above-mentioned homolog VpCas9 protein of Cas9 or its mutant, or the kit for gene editing or gene cleavage to aspects such as gene editing, gene targeting, gene cleavage, or preparation of target nucleic acid detection or preparation of targeted gene therapy drugs.
[0015] In a specific embodiment of the present invention, the gene editing, gene targeting or gene cleavage is carried out intracellularly and / or extracellularly; the gene editing or editing of the target nucleic acid includes modifying genes, knocking out genes, mutating genes or changing the expression level of gene products, etc.
[0016] In a specific embodiment of the present invention, the corresponding operations of the gene editing, gene targeting or gene cleavage can be carried out in prokaryotic cells or eukaryotic cells.
[0017] In the present invention, the mining and modification of Cas9 are carried out through a deep learning model, and finally the homolog VpCas9 protein of Cas9 is screened. The VpCas9 protein has a better editing effect in terms of gene editing efficiency, PAM compatibility or specificity, etc.; in the present invention, through sequence screening and optimization design, single-site mutants with higher editing efficiency and extended PAM compatibility are obtained. Compared with the SpCas9 protein or VpCas9 protein, these single-site mutants have better gene editing effects in genome editing; the VpCas9 protein and its single-site mutants provided by the present invention have application prospects in genome editing of plants, animals or microorganisms.
[0018] Term definitions involved in the present invention Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods, devices, and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods, devices, and materials are now described.
[0019] The terms "polynucleotide" or "nucleotide" mean deoxyribonucleotides, deoxyribonucleosides, ribonucleosides or ribonucleotides in single-stranded or double-stranded form, and polymers thereof. Unless otherwise specified, the term encompasses nucleic acids containing known analogs of natural nucleotides, which have binding properties similar to those of the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise specifically limited, the term also means oligonucleotide analogs, which include PNA (peptide nucleic acid), DNA analogs (such as phosphorothioates, phosphoroamidates, etc.) used in antisense technology. Unless otherwise specified, a particular nucleic acid sequence also implicitly encompasses its conservatively modified variants (including, but not limited to, degenerate codon substitutions) and complementary sequences, as well as the explicitly specified sequences. Specifically, degenerate codon substitutions can be achieved by generating a sequence in which the third position of one or more selected (or all) codons is replaced by a mixed base and / or deoxyinosine residue ( Mol Cell. Probes 8:91-98 (1994)).
[0020] The terms "polypeptide", "peptide" and "protein" are used interchangeably herein to mean a polymer of amino acid residues. That is, a description of a polypeptide applies equally to a description of a peptide and a description of a protein, and vice versa. The term applies to both naturally occurring amino acid polymers and amino acid polymers in which one or more amino acid residues are non-naturally encoded amino acids. As used herein, the term encompasses amino acid chains of any length, including full-length proteins (i.e., antigens), in which the amino acid residues are joined by covalent peptide bonds.
[0021] The terms "mutation" and "mutant" have their ordinary meanings herein, referring to a genetic, naturally occurring or introduced change in a nucleic acid or polypeptide sequence, the meaning of which is the same as that commonly known to those skilled in the art.
[0022] The term "recombinant host cell line" or "host cell" means a cell containing the polynucleotide of the present invention, regardless of the method used for insertion to produce the recombinant host cell, such as direct uptake, transduction, f-mating or other methods known in the art. The exogenous polynucleotide can be maintained as a non-integrating vector such as a plasmid or can be integrated into the host genome. The host cell can be a prokaryotic cell or a eukaryotic cell.
[0023] The term "operably linked" refers to a functional linkage between two or more elements, and the elements that are operably linked can be adjacent or non-adjacent.
[0024] The term "target sequence" refers to the polynucleotide targeted by the guide sequence in the gRNA, such as a sequence complementary to the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of the CRISPR / Cas complex (including the Cas protein and the gRNA). BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Results of protein purification of SpCas9 and VpCas9; "Sp" represents SpCas9, and "Vp" represents VpCas9.
[0026] Figure 2 Results of detection of PAM and cleavage pattern; A) PAM preference of SpCas9; B) Cleavage pattern of SpCas9; C) PAM preference of SpCas9; D) Cleavage pattern of SpCas9.
[0027] Figure 3 Results of detection of in vitro nucleic acid fragment cleavage efficiency; A) Gel images of detection of nucleic acid fragment cleavage efficiency of SpCas9 and VpCas9; B) Grayscale analysis of the generated small fragments by imageJ.
[0028] Figure 4 Results of design of VpCas9 mutants and qPCR activity detection; A) 12 single-site mutants were designed based on the CasMiner characteristic matrix of VpCas9 and the conserved matrix of PSAP; B) Process of detecting Cas9 editing efficiency by qPCR; C) Editing activities of 12 single-site mutants; D) Editing activities of double-site mutants; In subfigure D), t-test was used to verify whether there was a significant difference between the wild type and the mutants, ns, P>0.05; *, P<0.05. DETAILED DESCRIPTION OF THE INVENTION
[0029] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description progresses. However, it should be understood that the described embodiments are exemplary only and do not constitute any limitation to the scope of the present invention. Those skilled in the art should understand that modifications or substitutions can be made to the details and forms of the technical solutions of the present invention without departing from the spirit and scope of the present invention, but such modifications or substitutions all fall within the protection scope of the present invention.
[0030] Experimental Materials and Methods 1. Model Prediction The collected protein sequence data was predicted by CasMiner (registration number: 2023SR0464752) software, and potential new Cas9 sequences were screened out for further analysis through the model scores of the sequences.
[0031] 2. Protein Expression and Purification The recombinant plasmid containing the target gene and His tag was heat-shock transformed into the competent cells of Escherichia coli BL21(DE3), and positive monoclonal colonies were obtained through positive identification. Then, the positive monoclonal colonies were placed in 50 mL of LB (lysogeny broth) medium containing kanamycin (50 μg / mL) and cultured in a constant temperature shaker at 37 °C and 200 rpm until the OD600 value was between 0.6 and 0.8. Then, 20 μL of 1 mol / L IPTG was added. Subsequently, the bacterial liquid system was placed under low-temperature induction conditions of 16 °C and 200 rmp for 18 hours. The cells were collected at 8000 rpm for 10 minutes and resuspended in 8 mL of 20 mmol / L phosphate (PB) buffer (pH = 7.0). After that, the cells were lysed by ultrasonic waves with a power of 35 W, lysed for 4 seconds, paused for 4 seconds, and lysed on an ice-water mixture for a total of 10 minutes. The supernatant after lysis was centrifuged at 8000 rpm and 4 °C for 30 minutes. Subsequently, after eluting with NTA40 buffer (containing 40 mmol / L imidazole), the eluate of NTA200 buffer (containing 200 mmol / L imidazole) was directly collected. The collected eluate was dialyzed overnight and concentrated with polyethylene glycol 8000 (PEG800), and finally the target protein could be obtained.
[0032] 3. PAM preference determination The PAM library and sgRNA were diluted to 200 ng / μL and 2 pmol / μL, respectively. Then, the purified Cas9 protein was quantified using a BCA protein quantification kit and diluted to 100 ng / μL. In the reaction solution containing Mg 2+ , 8 pmol (4 μL) of sgRNA and 300 ng of Cas9 protein were added, and the reaction was carried out at 37 °C for 10 minutes to form an RNP complex. Then, 200 ng of the PAM library fragment was added, and the reaction was carried out at 37 °C for 20 minutes. Next, a terminator containing 0.5% SDS, 150 mmol / L EDTA, and pH = 8.0 was added to terminate the reaction. Finally, the reaction system was subjected to next-generation sequencing.
[0033] 4. In vitro cleavage of the target fragment In the reaction solution containing Mg 2+ , 8 pmol (4 μL) of sgRNA and 300 ng of Cas9 protein were added, and the reaction was carried out at 37 °C for 10 minutes to form an RNP complex. Then, 200 ng of the nucleic acid fragment to be cleaved was added to the RNP system, and the reaction was carried out at 37 °C for 20 minutes. Finally, nucleic acid electrophoresis was performed, and the cleavage result of Cas9 was analyzed using imageJ.
[0034] 5. qPCR detection of cleavage efficiency Add 2500 ng of the genome of engineered Escherichia coli BMLacZ containing cleavage fragments to the RNP system, then add sterile water to 20 μL and digest at 37 °C for 20 minutes. After diluting the digestion system 125-fold, add 1 μL of the diluted solution to the qPCR reaction system. Finally, perform the reaction and fluorescence detection in a real-time fluorescence quantitative PCR instrument. By calculating 2 −ΔΔCt to reflect the gene editing efficiency of Cas9.
[0035] Experimental Example 1 Gene Mining and Protein Expression and Purification of VpCas9 Protein 1. Gene Mining of VpCas9 Protein This experiment was based on predictive analysis of the "representative genomes" (https: / / progenomes.embl.de / data / repGenomes / progenomes3.proteins.representatives.fasta.bz2) from the EMBL database. Potential Cas9 was screened by setting the sequence length range to 1301 - 1400 amino acid residues, and the number of Repeats in the genome corresponding to Cas9 was analyzed using the CRISPR Recognition Tool (CRT). Finally, a sequence with a probability of 99.9703%, 56 Repeats, and derived from Vagococcus penaei strain CD276T, 633807.SAMN06245872.BW732_04735, was selected and named VpCas9. The predictive analysis data are shown in Table 1.
[0036] Table 1 Predictive Analysis Results of CasMiner
[0037] Among them, the amino acid sequence of the VpCas9 protein is shown in SEQ ID No.5:
[0038] Among them, the amino acid sequence of the SpCas9 protein is shown in SEQ ID No. 6:
[0039] 2. Expression and purification of VpCas9 protein The protein sequence of VpCas9 was codon-optimized for Escherichia coli strains and the sequence was inserted between Nde I and Xho I of pET-28a, and a TAA stop codon was added to the 3' end of the VpCas9 coding sequence. The designed vector retained a His tag His-His-His-His-His-His (HHHHHH) upstream of the protein sequence for subsequent purification tags.
[0040] The results of expression and purification are as Figure 1 shown, demonstrating that the proteins of SpCas9 and VpCas9 were successfully expressed and purified in this experiment.
[0041] Experimental Example 2 Verification experiment of PAM and cleavage mode of VpCas9 A 150 bp PAM fragment with 5'-NNNNN-3' and the corresponding sgRNA were synthesized. The PAM library sequence is shown in the following SEQ ID No.7: GGTGAAGAGAACAACATGGCTATTATTAAGGAGTTCATGCGTTTTAAGGTCCACATGGAGGGTTCCGTTAACGGTCATGAATTTgaaattgagggtgagggtgaNNNNAGACCATACGAAGCTTTTCAAACTGCTAAGTTGAAGGTCACC (SEQ ID No.7).
[0042] The sgRNA sequence of SpCas9 is shown in the following SEQ ID No.8: gaaattgagggtgagggtgaGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCtttttt (SEQ ID No. 8).
[0043] The sgRNA sequence of VpCas9 is shown in the following SEQ ID No.9: gaaattgagggtgagggtgaGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTTAGTCCGTAAGCAACTATTCTAGTGGCACTGTCTCGGTGCtttttt (SEQ ID No. 9).
[0044] Subsequently, the results of PAM and cleavage methods were analyzed.
[0045] The test results are as Figure 2 shown. It was found that VpCas9 has an NGGV PAM (N: A / T / C / G, V: G / T). The first three bases of NGG in this PAM are the same as those of SpCas9, while VpCas9 shows a more obvious preference for guanine (G) and thymine (T) at the fourth position of the PAM ( Figure 2 A, 2C), thus introducing additional base restrictions. By analyzing the sequencing fragments, it was found that, similar to SpCas9, VpCas9 produced significant double-strand cleavage activity at -3 to -4 nt upstream of the PAM, generating cleavage gaps of 75.52% and 76.91% in the target sequence (TS) and non-target sequence (NTS), respectively ( Figure 2 B, 2D).
[0046] Test Example 3 Detection Test of the Basic Cleavage Ability of VpCas9 To further detect the basic cleavage abilities of VpCas9 and SpCas9, a nucleic acid fragment with a length of 1,657 bp was designed. This nucleic acid fragment can be cleaved into 600 bp and 1,057 bp and can be used to detect the cleavage activities of VpCas9 and SpCas9 in vitro. The nucleotide sequence of this nucleic acid fragment is shown as SEQ ID No.10 below:
[0047] The detection results are as follows Figure 3 shown: The results of the small fragment in vitro cleavage experiment illustrate that VpCas9 and SpCas9 can utilize each other's sgRNAs to cleave nucleic acid fragments ( Figure 3 A), and VpCas9&Vp-sgRNA exhibits the largest grayscale value ( Figure 3 B), which also indicates that VpCas9 has comparable in vitro cleavage efficiency to SpCas9.
[0048] Test Example 4 Design of VpCas9 Mutants and Detection of Mutant Cleavage Efficiency To further improve the activity of VpCas9, mutant design of VpCas9 was carried out in this experiment. First, the protein sequence of VpCas9 was submitted using jackhmmer, and homologous sequences of VpCas9 were retrieved from the UniRef90 database. Secondly, the features of the homologous sequences were extracted using Grad-CAM of CasMiner, and the sum of the feature matrices was divided by the number of homologous sequences as the feature matrix to assist in the mutation of VpCas9. At the same time, the conserved (functional) matrix of VpCas9 was obtained through position-specific amino acid probability (PSAP). The score differences (Diff) between the optimal mutants and the wild type at each site in the feature matrix and the conserved (functional) matrix were calculated respectively. Under the condition of ensuring the consistency of the optimal mutants at the same site, the top 12 single-site mutation points (total score less than 30) were selected by sorting and adding the two Diffs, and the process is as shown in Figure 4 A; To more accurately detect the editing efficiency of VpCas9 mutants, the qPCR method was used to detect the editing efficiency of Cas9, and the principle is as shown in Figure 4 B.
[0049] Analysis of the results found that among the 12 designed mutants, the editing activities of 9 mutants (VpCas9-V623I, VpCas9-E372K, VpCas9-T437P, VpCas9-I526V, VpCas9-F573L, VpCas9-D1080G, VpCas9-S438F, VpCas9-I439L, VpCas9-A110D) were improved compared with the wild type. Among them, the mutant VpCas9-V623I had the best cleavage activity ( Figure 4 C).
[0050] Based on the best single-point mutant VpCas9-V623I, on this basis, mutant sites with better efficiency than the wild type were further superimposed, and 8 double-point mutants were designed and their activities were verified. Finally, it was found that the editing activities of three double-point mutants were significantly improved compared with VpCas9-V623I ( Figure 4 D), and these three double-point mutants were VpCas9-V623I-I439L (VPM2-1), VpCas9-V623I-E372K (VPM2-2), and VpCas9-V623I-I526V (VPM2-3), respectively.
[0051] Among them, the amino acid sequence of VpCas9-V623I-I439L (VPM2-1) is shown in SEQ ID No. 11:
[0052] The amino acid sequence of VpCas9-V623I-E372K (VPM2-2) is shown in SEQ ID No. 12:
[0053] The amino acid sequence of VpCas9-V623I-I526V (VPM2-3) is shown in SEQ ID No. 13:
Claims
1. The VpCas9 protein for gene editing, characterized in that, Its amino acid sequence is shown in SEQ ID No.
5.
2. The single-site mutant of the VpCas9 protein according to claim 1, characterized in that, The single-site mutant is a single-site mutant obtained by performing any one of the amino acid single-site mutations of V623I, E372K, T437P, I526V, F573L, D1080G, S438F, I439L or A110D on the amino acid sequence shown in SEQ ID No.
5.
3. The coding gene of the VpCas9 protein according to claim 1.
4. The coding gene of the single-site mutant according to claim 2.
5. A carrier, characterized in that, The vector contains the coding gene according to claim 3 or 4 and regulatory elements operably linked to the coding gene.
6. The carrier according to claim 5, wherein The vector is selected from an expression vector, a cloning vector or a shuttle vector.
7. A CRISPR-Cas system, said system comprising a Cas protein and at least one sgRNA; said Cas protein being capable of binding to said sgRNA, said sgRNA comprising direct repeat sequences and a spacer sequence capable of hybridizing with a target nucleic acid, characterized in that, The Cas protein is the VpCas9 protein according to claim 1 or the single-site mutant according to claim 2.
8. A kit for gene editing or gene cleavage, characterized in that, The kit includes the VpCas9 protein according to claim 1, the single-site mutant according to claim 2, the coding gene according to claim 3 or 4, the vector according to claim 5 or 6, or the CRISPR-Cas system according to claim 7.
9. Use of the VpCas9 protein according to claim 1, the single-site mutant according to claim 2, the coding gene according to claim 3 or 4, the vector according to claim 5 or 6, or the CRISPR-Cas system according to claim 7 in gene editing, editing of target nucleic acids, gene cleavage, or in the preparation of target nucleic acid detection or targeted gene therapy drugs.
10. The application according to claim 9, characterized in that, The gene editing or editing of target nucleic acids includes modifying genes, knocking out genes, mutating genes, or changing the expression level of gene products; the gene editing, editing of target nucleic acids, or gene cleavage is performed in prokaryotic cells or eukaryotic cells for corresponding operations.
Citation Information
Patent Citations
CRISPR SpCas9 (K510A) mutant and application thereof
CN112538471A
CRISPR-FrCas9 protein mutant and application thereof
CN117866926A
VpCas9 protein double-site mutant and application thereof in gene editing
CN120230738A
Crispr / cas9 gene editing system and application thereof
US20240175055A1
Cited By
Efficient sgRNA for improving gene editing efficiency and application thereof
CN120866325A
A Highly Efficient sgRNA for Improving Gene Editing Efficiency and Its Applications
CN120866325B