VpCas9 proteins, unit point mutants thereof and applications in gene editing

By modifying the Cas9 protein using a deep learning model, VpCas9 and its unit point mutants were developed, overcoming the limitations of the Cas9 protein in gene editing efficiency and PAM compatibility, achieving more efficient genome editing results, and making it suitable for gene editing applications in a variety of organisms.

CN120249251BActive Publication Date: 2025-10-17INSTITUTE OF ANIMAL SCIENCES OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510725509.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-17
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing Cas9 proteins have limitations in gene editing efficiency, PAM compatibility, and specificity, which affect their effectiveness and flexibility in a wider range of genome editing applications.

Method used

By mining and modifying deep learning models, the VpCas9 protein, a homolog of Cas9, and its unit point mutants were obtained. Their amino acid sequences were optimized to improve editing efficiency and PAM compatibility. Combined with sgRNA, a CRISPR-Cas system was formed.

Benefits of technology

VpCas9 protein and its mutants show better editing effects in gene editing, improve the efficiency and PAM compatibility of gene editing, and are suitable for genome editing of plants, animals and microorganisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120249251B_ABST
    Figure CN120249251B_ABST
Patent Text Reader

Abstract

The application discloses a VpCas9 protein and a unit point mutant thereof and application in gene editing. The application screens and modifies Cas9 through a deep learning model, and finally obtains a new Cas9 protein VpCas9 protein, and the amino acid sequence of the VpCas9 protein is shown in SEQ ID No. 5; compared with a SpCas9 protein, the VpCas9 protein has better editing effect in gene editing efficiency, PAM compatibility or specificity; the application further obtains a unit point mutant with higher editing efficiency through sequence screening and optimization design, and compared with the SpCas9 protein and the VpCas9 protein, the unit point mutant has better gene editing effect in genome editing; the VpCas9 protein and the unit point mutant thereof provided by the application have application prospects in genome editing or editing target nucleic acid of plants, animals or microorganisms.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a homolog of Cas9 protein and its application, in particular to a homolog of Cas9 protein VpCas9 protein and its unit point mutant and its application in genome editing, and belongs to the field of Cas9 homolog and its application. BACKGROUND

[0002] Gene editing refers to the operation of deleting, replacing and inserting target genes to obtain new functions or phenotypes, and is an important research means for the rapid development of life science, wherein the gene editing technology represented by CRISPR-Cas9 is one of the most efficient, simplest, lowest cost and easiest to use technology means.

[0003] However, the existing Cas9 protein has certain limitations in editing efficiency, PAM (protospacer adjacent motif) compatibility and specificity, which affects its effectiveness and flexibility in more extensive genome editing applications, and needs to be improved. SUMMARY

[0004] One of the purposes of the present application is to provide a new homolog of Cas9 protein VpCas9 protein and its coding gene.

[0005] The second purpose of the present application is to provide a unit point mutant of the homolog of Cas9 protein VpCas9 protein and its coding gene.

[0006] The third purpose of the present application is to provide a vector containing the coding gene and a host cell.

[0007] The fourth purpose of the present application is to provide a CRISPR-Cas system, which comprises a homolog of Cas9 protein VpCas9 protein or its mutant.

[0008] The fourth purpose of the present application is to apply the homolog of Cas9 protein VpCas9 protein or its mutant, the vector containing the coding gene or the CRISPR-Cas system comprising the homolog of Cas9 protein VpCas9 protein or its mutant to gene editing, editing target nucleic acid, gene cutting or preparing target nucleic acid detection or targeted gene therapy drug, etc.

[0009] The above purposes of the present application are achieved by the following technical solutions:

[0010] One aspect of the present application is to provide a homolog of Cas9 protein VpCas9 protein, and its amino acid sequence is shown in SEQ ID No. 5.

[0011] The application mines and modifies Cas9 through a deep learning model to obtain a homolog VpCas9 protein of Cas9, which has better editing effect in gene editing efficiency, PAM compatibility or specificity compared with SpCas9.

[0012] The second aspect of the application is to provide a unit point mutant of the homolog VpCas9 protein of Cas9; wherein the unit point mutant is obtained by performing any one of V623I, E372K, T437P, I526V, F573L, D1080G, S438F, I439L or A110D amino acid unit point mutation on the amino acid sequence shown in SEQ ID No. 5; preferably, the unit point mutant obtained by performing V623I amino acid unit point mutation on the amino acid sequence shown in SEQ ID No. 5, the cutting activity of the gene editing of the unit point mutant is most significantly improved compared with the wild type protein.

[0013] In the application, the unit point mutant "V623I" means that the 623th amino acid of the amino acid sequence shown in SEQ ID No. 5 is mutated from valine (Val, V) to isoleucine (Ile, I); the expressions of the remaining unit point mutants in the application are similar.

[0014] Another aspect of the application is to provide a coding gene of the homolog VpCas9 protein of Cas9 or a coding gene of the mutant of the homolog VpCas9 protein of Cas9.

[0015] Another aspect of the application is to provide a vector, wherein the vector comprises the coding gene and a regulatory element operably linked to the coding gene; wherein the vector can be an expression vector, a cloning vector or a shuttle vector, etc.

[0016] In a preferred embodiment, the regulatory element is selected from one or more of a promoter, a terminator, an enhancer, a transposon, a leader sequence or a marker gene.

[0017] Another aspect of the application is to provide a CRISPR-Cas system, wherein the system comprises a Cas protein and at least one sgRNA; the Cas protein can bind to the sgRNA, the sgRNA comprises a direct repeat sequence and a spacer sequence capable of hybridizing with a target nucleic acid, wherein the Cas protein is the homolog VpCas9 protein of Cas9 or a mutant thereof.

[0018] Another aspect of the present application provides a kit for gene editing or gene cutting, the kit comprising the homolog of Cas9, VpCas9 protein or its mutant, or a polynucleotide encoding the homolog of Cas9, VpCas9 protein or its mutant, or a vector containing the polynucleotide sequence, or a CRISPR-Cas system containing the homolog of Cas9, VpCas9 protein or its mutant.

[0019] Another aspect of the present application is the application of the homolog of Cas9, VpCas9 protein or its mutant, or a polynucleotide encoding the homolog of Cas9, VpCas9 protein or its mutant, or a vector containing the polynucleotide sequence, or a CRISPR-Cas system containing the homolog of Cas9, VpCas9 protein or its mutant, or a kit for gene editing or gene cutting in gene editing, gene targeting, gene cutting, or preparation of target nucleic acid detection or preparation of targeted gene therapy drugs, etc.

[0020] In a specific embodiment of the present application, the gene editing, gene targeting or gene cutting is carried out in cells and / or outside cells; the gene editing or editing target nucleic acid includes modifying genes, knocking out genes, mutating genes or changing the expression amount of gene products, etc.

[0021] In a specific embodiment of the present application, the gene editing, gene targeting or gene cutting can be carried out in prokaryotic cells or eukaryotic cells.

[0022] The present application screens and modifies the homolog of Cas9, VpCas9 protein through a deep learning model, and the VpCas9 protein has better editing effect in terms of gene editing efficiency, PAM compatibility or specificity; the present application further obtains unit point mutants with higher editing efficiency and PAM compatibility expansion through sequence screening and optimization design, and these unit point mutants have better gene editing effect in genome editing compared with SpCas9 protein or VpCas9 protein; the VpCas9 protein and its unit point mutants provided by the present application have application prospects in genome editing of plants, animals or microorganisms.

[0023] Definitions of terms involved in the present invention

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Although any methods, devices and materials similar or equivalent to those described herein can be used in the practice or testing of the present application, the preferred methods, devices and materials are now described.

[0025] The term "polynucleotide" or "nucleotide" means deoxyribonucleotides, deoxyribonucleosides, ribonucleosides or ribonucleotides and polymers thereof in single or double stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have binding properties similar to the reference nucleic acids and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise specifically limited, the term also means oligonucleotide analogs, which include PNA (peptide nucleic acid), DNA analogs used in antisense technology (phosphorothioates, phosphamidates, etc.). Unless otherwise specified, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (including but not limited to degenerate codon substitutions) and complementary sequences as well as explicitly specified sequences. In particular, degenerate codon substitutions ( Mol Cell. Probes 8:91-98 (1994)).

[0026] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. That is, a description directed to a polypeptide equally applies to describing a peptide and describing a protein, and vice versa. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers in which one or more amino acid residues is a non-naturally encoded amino acid. As used herein, the terms encompass amino acid chains of any length, including full-length proteins (i.e., antigens), in which the amino acid residues are linked via covalent peptide bonds.

[0027] The terms "mutation" and "mutant" have their ordinary meanings herein and refer to genetic, naturally occurring or introduced changes in nucleic acid or polypeptide sequences, and their meanings are the same as those generally understood by those skilled in the art.

[0028] The term "recombinant host cell strain" or "host cell" refers to a cell comprising a polynucleotide of the present invention, regardless of the method used for insertion to produce the recombinant host cell, such as direct uptake, transduction, f-mating, or other methods known in the art. The exogenous polynucleotide may be maintained as a non-integrating vector, such as a plasmid, or may be integrated into the host genome. The host cell may be a prokaryotic cell or a eukaryotic cell.

[0029] The term "operably linked" refers to a functional connection between two or more elements. Operably linked elements may be contiguous or non-contiguous.

[0030] The term "target sequence" refers to a polynucleotide targeted by a guide sequence in a gRNA, e.g., a sequence having complementarity to the guide sequence, wherein hybridization between the target sequence and the guide sequence will facilitate formation of a CRISPR / Cas complex, including a Cas protein and a gRNA. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 The results of protein purification of SpCas9 and VpCas9; "Sp" is SpCas9, and "Vp" is VpCas9.

[0032] Figure 2 The results of PAM and cleavage mode detection; A) PAM preference of SpCas9; B) cleavage mode of SpCas9; C) PAM preference of SpCas9; D) cleavage mode of SpCas9.

[0033] Figure 3 The results of in vitro nucleic acid fragment cleavage efficiency detection; A) nucleic acid fragment cleavage efficiency detection gel map of SpCas9 and VpCas9; B) imageJ gray scale analysis of the small fragments generated.

[0034] Figure 4 The results of VpCas9 mutant design and qPCR activity detection; A) 12 single-point mutants were designed according to the CasMiner characteristic matrix of VpCas9 and the conservative matrix of PSAP; B) the process of detecting Cas9 editing efficiency by qPCR; C) editing activity of 12 single-point mutants; D) editing activity of double-site mutants; In subgraph D), t-test was used to verify whether there was a significant difference between wild type and mutants, ns, P>0.05; *, P<0.05. DETAILED DESCRIPTION

[0035] The advantages and features of the present application will become more apparent with the description of specific embodiments. However, it should be understood that the described embodiments are only exemplary and do not constitute any limitation on the scope of the present application. Those skilled in the art should understand that modifications or substitutions can be made to the details and forms of the technical solutions of the present application without departing from the spirit and scope of the present application, and such modifications or substitutions all fall within the protection scope of the present application.

[0036] Test materials and test methods

[0037] 1. Model prediction

[0038] The collected protein sequence data were subjected to prediction by CasMiner (Accession No. 2023SR0464752) software, and potential new Cas9 sequences were screened out by model score of the sequence for one-step analysis.

[0039] 2. Protein expression and purification

[0040] The recombinant plasmid containing the target gene and His tag was heat shock transformed into the plasmid BL21 (DE3) E. coli competent cells, and positive single clones were obtained by positive identification. Then, the positive single clones were placed in 50 mL LB (lysozyme broth) medium containing kanamycin (50 μg / mL) and cultured in a constant temperature shaker at 37 °C and 200 rpm until the OD600 value was between 0.6 and 0.8. Then, 20 μL of 1 mol / L IPTG was added. Subsequently, the bacterial solution was placed in a low-temperature induction condition at 16 °C and 200 rpm for 18 hours. The cells were collected at 8000 rpm for 10 minutes, and the bacterial bodies were resuspended with 8 mL of 20 mmol / L phosphate (PB) buffer (pH=7.0). Then, the cells were broken by ultrasonic wave with a power of 35 W, broken for 4 seconds, paused for 4 seconds, and broken for a total of 10 minutes on an ice water mixture. The supernatant after breaking was centrifuged at 8000 rpm and 4 °C for 30 minutes. Then, after elution into NTA40 buffer (containing 40 mmol / L imidazole), the eluate of NTA200 buffer (containing 200 mmol / L imidazole) was directly collected. The collected eluate was subjected to overnight dialysis and polyethylene glycol 8000 (PEG800) concentration, and finally the target protein was obtained.

[0041] 3. PAM preference determination

[0042] The PAM library and sgRNA were diluted to 200 ng / μL and 2 pmol / uL, respectively. Then, the purified Cas9 protein was quantified by BCA protein quantification kit and diluted to 100 ng / μL. In the reaction solution containing Mg 2+ , 8 pmol (4 uL) of sgRNA and 300 ng of Cas9 protein were added, and the reaction was carried out at 37 °C for 10 minutes to form an RNP complex. Then, 200 ng of PAM library fragments were added, and the reaction was carried out at 37 °C for 20 minutes. Then, the reaction was terminated by adding a termination agent containing 0.5 % SDS, 150 mmol / L EDTA, pH=8.0. Finally, the reaction system was subjected to second-generation sequencing.

[0043] 4. In vitro cleavage of target fragments

[0044] In the reaction solution containing Mg 2+The reaction solution of 8 pmol (4 uL) sgRNA and 300 ng Cas9 protein was added to the RNP complex at 37 ℃ for 10 minutes. Then 200 ng of nucleic acid fragment to be cut was added to the RNP system and reacted at 37 ℃ for 20 minutes. Finally, nucleic acid electrophoresis was performed, and the cleavage results of Cas9 were analyzed by imageJ.

[0045] 5. qPCR detection of cleavage efficiency

[0046] In the RNP system, 2500 ng of BMLacZ engineered E. coli genome containing the cleavage fragment was added, and then 20 μL of sterile water was injected to digest for 20 minutes at 37 ℃. After diluting the digestion system by 125 times, 1 μL of dilution was added to the qPCR reaction system. Finally, the reaction and fluorescence detection were performed in the real-time fluorescent quantitative PCR instrument. The gene editing efficiency of Cas9 was reflected by calculating 2 −ΔΔCt .

[0047] Test Example 1 Gene mining and protein expression and purification test of VpCas9 protein

[0048] 1. Gene mining of VpCas9 protein

[0049] In this test, the "representative genome" (https: / / progenomes.embl.de / data / repGenomes / progenomes3.proteins.representatives.fasta.bz2) from EMBL database was analyzed. By setting the sequence length range to 1301-1400 amino acid residues, potential Cas9 was screened, and the number of repeats in the corresponding genome of Cas9 was analyzed by CRISPR Recognition Tool (CRT). Finally, 633807.SAMN06245872.BW732_04735 from strain CD276T with a probability of 99.9703% and 56 repeats was screened, and the sequence was named VpCas9. The prediction analysis data is shown in Table 1. Vagococcus penaei

[0050] Table 1 Prediction analysis results of CasMiner

[0051]

[0052] Among them, the amino acid sequence of VpCas9 protein is shown in SEQ ID No. 5:

[0053]

[0054] wherein the amino acid sequence of the SpCas9 protein is set forth in SEQ ID No. 6:

[0055]

[0056] 2. Expression and purification of VpCas9 protein

[0057] The protein sequence of VpCas9 was codon-optimized for E. coli strain and the sequence was inserted into pET-28a Nde I and Xho I, and a TAA stop codon was added at the 3' end of the VpCas9 coding sequence. The designed vector retained a His tag His-His-His-His-His-His (HHHHHH) upstream of the protein sequence for subsequent purification tag.

[0058] The results of expression and purification are shown in Figure 1 , which prove that the proteins of SpCas9 and VpCas9 are successfully expressed and purified in this experiment.

[0059] Example 2 Verification experiment of PAM and cutting mode of VpCas9

[0060] A 150 bp PAM fragment with 5'-NNNNN-3' and the corresponding sgRNA were synthesized, and the PAM library sequence is shown in SEQ ID No. 7 as follows:

[0061] GGTGAAGAGAACAACATGGCTATTATTAAGGAGTTCATGCGTTTTAAGGTCCACATGGAGGGTTCCGTTAACGGTCATGAATTTgaaattgagggtgagggtgaNNNNAGACCATACGAAGCTTTTCAAACTGCTAAGTTGAAGGTCACC (SEQ ID No. 7).

[0062] The sgRNA sequence of SpCas9 is shown in SEQ ID No. 8 as follows:

[0063] gaaattgagggtgagggtgaGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCtttttt (SEQ ID No. 8).

[0064] The sgRNA sequence of VpCas9 is shown in SEQ ID No. 9 as follows:

[0065] gaaattgagggtgagggtgaGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTTAGTCCGTAAGCAACTATTCTAGTGGCACTGTCTCGGTGCtttttt (SEQ ID No. 9).

[0066] The results of PAM and cleavage mode were analyzed subsequently.

[0067] The experimental results are shown in Figure 2 The experimental results found that VpCas9 has NGGV PAM (N: A / T / C / G, V: G / T). The first three bases of NGG in the PAM are the same as SpCas9, while VpCas9 shows more obvious preference for guanine (G) and thymine (T) at the fourth position of the PAM Figure 2 A, 2C), thereby introducing additional base restrictions. By analyzing the sequencing fragments, it was found that, similar to SpCas9, VpCas9 produced significant double-stranded cleavage activity at -3 to -4 nt upstream of the PAM, producing 75.52% and 76.91% cleavage gaps in the target sequence (TS) and non-target sequence (NTS), respectively Figure 2 B, 2D).

[0068] Experimental Example 3: Test for basic cleavage ability of VpCas9

[0069] In order to further detect the basic cleavage ability of VpCas9 and SpCas9, a nucleic acid fragment with a length of 1,657 bp was designed, which can be cleaved into 600 bp and 1,057 bp, and can be used to detect the cleavage activity of VpCas9 and SpCas9 in vitro. The nucleotide sequence of the nucleic acid fragment is shown in SEQ ID No. 10 as follows:

[0070]

[0071] The results are shown in Figure 3 Figure 2. The results of the small fragment in vitro cleavage experiment show that VpCas9 and SpCas9 can use each other's sgRNA to cleave nucleic acid fragments (A), and VpCas9 & Vp-sgRNA exhibits the largest gray value (B), which also shows that VpCas9 has the same in vitro cleavage efficiency as SpCas9. Figure 3 Figure 3 Figure 2. The results of the small fragment in vitro cleavage experiment show that VpCas9 and SpCas9 can use each other's sgRNA to cleave nucleic acid fragments (A), and VpCas9 & Vp-sgRNA exhibits the largest gray value (B), which also shows that VpCas9 has the same in vitro cleavage efficiency as SpCas9.

[0072] Test Example 4: Design of VpCas9 mutants and test of cleavage efficiency of the mutants

[0073] In order to further improve the activity of VpCas9, the VpCas9 mutants were designed in this test. First, the protein sequence of VpCas9 was submitted to jackhmmer, and the homologous sequences of VpCas9 were retrieved from UniRef90 database. Second, the features of the homologous sequences were extracted using Grad-CAM of CasMiner, and the feature matrix was summed and divided by the number of homologous sequences as a feature matrix to assist the mutation of VpCas9. At the same time, the conservation (function) matrix of VpCas9 was obtained by position-specific amino acid probability (PSAP). The score difference (Diff) between the optimal mutant and the wild type at each site in the feature matrix and the conservation (function) matrix was calculated, respectively. Under the condition of ensuring the consistency of the optimal mutant at the same point, the top 12 unit point mutation points (total score less than 30) were selected as candidate mutation points by sorting and adding the two Diffs, as shown in Figure 4 A. In order to more accurately detect the editing efficiency of VpCas9 mutants, the editing efficiency of Cas9 was detected by qPCR method, and the principle is shown in Figure 4 B.

[0074] The analysis results show that the editing activity of 9 mutants (VpCas9-V623I, VpCas9-E372K, VpCas9-T437P, VpCas9-I526V, VpCas9-F573L, VpCas9-D1080G, VpCas9-S438F, VpCas9-I439L, VpCas9-A110D) among the designed 12 mutants is improved compared with the wild type, among which the cleavage activity of mutant VpCas9-V623I is the best (C). Figure 4

[0075] ​​Based on the best single-point mutant VpCas9-V623I, eight double-site mutants were designed and their activities were verified. Finally, it was found that the editing activities of three double-site mutants were significantly improved compared with VpCas9-V623I (VPM2-1, VPM2-2 and VPM2-3). Figure 4 D), which are VpCas9-V623I-I439L (VPM2-1), VpCas9-V623I-E372K (VPM2-2) and VpCas9-V623I-I526V (VPM2-3), respectively.

[0076] The amino acid sequence of VpCas9-V623I-I439L (VPM2-1) is shown in SEQ ID No. 11:

[0077]

[0078] The amino acid sequence of VpCas9-V623I-E372K (VPM2-2) is set forth in SEQ ID No. 12:

[0079]

[0080] The amino acid sequence of VpCas9-V623I-I526V (VPM2-3) is set forth in SEQ ID No. 13:

[0081]

Claims

1. A single-site mutant of the VpCas9 protein, characterized in that The single-site mutant is a single-site mutant obtained by performing a V623I amino acid single-site mutation on the amino acid sequence shown in SEQ ID No.

5.

2. A gene encoding the single-site mutant according to claim 1.

3. A carrier, characterized in that The vector comprises the coding gene according to claim 2 and a regulatory element operably linked to the coding gene.

4. The carrier according to claim 3, characterized in that The vector is selected from an expression vector, a cloning vector or a shuttle vector.

5. A CRISPR-Cas system comprising a Cas protein and at least one sgRNA; the Cas protein is capable of binding to the sgRNA, the sgRNA comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid, characterized in that: The Cas protein is the single-site mutant described in claim 1.

6. A kit for gene editing or gene cutting, characterized in that: The kit comprises the single-site mutant according to claim 1, the encoding gene according to claim 2, the vector according to claim 3 or the CRISPR-Cas system according to claim 5.

7. Use of the single-site mutant according to claim 1, the encoding gene according to claim 2, the vector according to claim 3 or the CRISPR-Cas system according to claim 5 in the preparation of drugs for target nucleic acid detection or targeted gene therapy.

Citation Information

Patent Citations

  • VpCas9 protein double-site mutant and application thereof in gene editing

    CN120230738A