Cas mutant protein and application thereof

By specifically mutating the amino acids of the Cas12i.3 protein and constructing a fusion protein, the problem of low editing efficiency of Cas12i.3 was solved, and a more efficient gene editing effect was achieved.

CN120665841AActive Publication Date: 2025-09-19CHINA AGRI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510729815.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-19
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing Cas12i.3 nuclease has low editing efficiency, which limits its widespread application.

Method used

By mutating amino acids 168, 273, 332, 478, and 599 of the Cas12i.3 protein to arginine, the N168R+S273R+L332R+G478R+S599R mutant protein was formed, and then fused with a nuclear localization signal, a tag, and a T5 exonuclease to construct the fusion protein NLS-N168R+S273R+L332R+G478R+S599R-T5 to improve editing efficiency.

Benefits of technology

It significantly improves the efficiency of gene editing and is suitable for multi-gene editing applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005431585590000111
    Figure BDA0005431585590000111
  • Figure BDA0005431585590000131
    Figure BDA0005431585590000131
  • Figure BDA0005431585590000141
    Figure BDA0005431585590000141
Patent Text Reader

Abstract

The invention discloses a Cas mutant protein and application thereof, and belongs to the technical field of gene editing. The Cas mutant protein provided by the invention is a mutant protein obtained by mutating the 168th amino acid, the 273th amino acid, the 332th amino acid, the 478th amino acid and the 599th amino acid of SEQ ID NO: 1 into arginine and keeping other amino acid sequences unchanged. By mutating multiple amino acids of the wild type Cas12i protein, the editing efficiency is improved, and the Cas12i protein can be used for multi-gene editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of gene editing technology, and specifically relates to a Cas mutant protein and its application. Background Art

[0002] CRISPR-Cas is an adaptive immune system developed by prokaryotes to defend against viral infection or phage invasion. It uses RNA to recognize DNA targets and generate double-strand breaks, followed by site-specific gene editing through non-homologous end joining or homologous recombination.

[0003] There are six types of CRISPR-Cas systems (types I-VI), among which type II SpyCas9 and type V AsCas12a gene editing systems have been widely used in the fields of microorganisms, animals and plants due to their relatively simple composition, high editing efficiency and easy operation. Chinese invention patent CN111757889B discloses a type V Cas protein Cas12f.4, which is referred to as Cas12i.3 in this article. Compared with SpCas9 and AsCas12a, Cas12i.3 is a relatively small Cas nuclease (1,045 amino acids) with a 5'-TTN PAM motif and can autonomously process pre-crRNA. Although Cas12i.3 has the above advantages, its editing efficiency is still an important factor restricting its widespread application.

[0004] Therefore, it is necessary to provide a Cas nuclease with higher editing efficiency. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a Cas enzyme with higher editing efficiency. The technical problem to be solved is not limited to the technical subject described above, and those skilled in the art can clearly understand other technical subjects not mentioned herein through the following description.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] The present invention provides a Cas mutant protein, wherein the Cas mutant protein is a mutant protein in which the 168th amino acid, the 273rd amino acid, the 332nd amino acid, the 478th amino acid and the 599th amino acid of SEQ ID NO: 1 are all mutated to arginine, and the other amino acid sequences remain unchanged.

[0008] Specifically, the above-mentioned Cas mutant protein is a mutant protein obtained by mutating the 168th amino acid of SEQ ID NO: 1 from asparagine to arginine, the 273rd amino acid from serine to arginine, the 332nd amino acid from leucine to arginine, the 478th amino acid from glycine to arginine, and the 599th amino acid from serine to arginine, while keeping the other amino acid sequences unchanged.

[0009] Specifically, the amino acid sequence of the Cas mutant protein is SEQ ID NO: 6. The Cas mutant protein is hereinafter referred to as N168R+S273R+L332R+G478R+S599R or N168R+S273R+L332R+G478R+S599R mutant.

[0010] Furthermore, the coding sequence of the above-mentioned Cas mutant protein is obtained by replacing the codon encoding the 168th amino acid N of SEQ ID NO: 1 (nucleotides 502-504 of SEQ ID NO: 2 aac) with a codon encoding R (aga), the codon encoding the 273rd amino acid S of SEQ ID NO: 1 (nucleotides 817-819 of SEQ ID NO: 2 agc) with a codon encoding R (aga), the codon encoding the 332nd amino acid L of SEQ ID NO: 1 (nucleotides 994-996 of SEQ ID NO: 2 ctg) with a codon encoding R (aga), the codon encoding the 478th amino acid G of SEQ ID NO: 1 (nucleotides 1432-1434 of SEQ ID NO: 2 ggc) with a codon encoding R (aga), the codon encoding the 599th amino acid S of SEQ ID NO: 1 (nucleotides 817-819 of SEQ ID NO: 2 agc) with a codon encoding R (aga), and the codon encoding the 332nd amino acid L of SEQ ID NO: 1 (nucleotides 994-996 of SEQ ID NO: 2 ctg) with a codon encoding R (aga). The nucleotide sequence of SEQ ID NO: 2 at positions 1795-1797 (agc) is replaced with the codon encoding R (aga) while keeping the other nucleotide sequences of SEQ ID NO: 2 unchanged.

[0011] The present invention also provides a fusion protein comprising the aforementioned Cas mutant protein, hereinafter referred to as the fusion protein NLS-N168R+S273R+L332R+G478R+S599R-T5.

[0012] Furthermore, the modified portion includes, but is not limited to, a tag, a nuclear localization signal, a linker, a T5 exonuclease, and may also include other sequences that facilitate gene editing of the fusion protein. Specifically, the modified portion may be one or more of the foregoing.

[0013] Furthermore, the nuclear localization signal is located at, near or close to the end (e.g., the N-terminus, the C-terminus or both) of the Cas mutant protein.

[0014] Furthermore, the epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can select other suitable epitope tags (e.g., purification, detection or tracing).

[0015] Furthermore, the Cas mutant protein is optionally coupled, conjugated or fused to the modified portion via a linker.

[0016] Furthermore, the modified portion is directly connected to the N-terminus or C-terminus of the Cas mutant protein.

[0017] Further, the modified portion is connected to the N-terminus or C-terminus of the Cas protein of the present invention via a linker. Such linkers are well known in the art, and examples thereof include but are not limited to one or more (e.g., 1, 2, 3, 4 or 5) amino acids (e.g., Gly or Ser).

[0018] It is well known to those skilled in the art that the aforementioned Cas mutant proteins and fusion proteins are not limited by the method of their production. For example, they can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.

[0019] Furthermore, the fusion protein is obtained by linking the 22nd amino acid residue of the tag protein to the 1st amino acid residue of the first nuclear localization signal sequence through a peptide bond, the 31st amino acid residue of the first nuclear localization signal sequence to the 1st amino acid residue of the first linker through a peptide bond, the 8th amino acid residue of the first linker to the 1st amino acid residue of the Cas mutant protein through a peptide bond, the 1045th amino acid residue of the Cas mutant protein to the 1st amino acid residue of the second linker through a peptide bond, the 10th amino acid residue of the second linker to the 1st amino acid residue of the T5 exonuclease through a peptide bond, and the 291st amino acid residue of the T5 exonuclease to the 1st amino acid residue of the second nuclear localization sequence through a peptide bond, totaling 1423 amino acids.

[0020] The tag protein in the fusion protein comprises 22 amino acids, which are encoded by nucleotides 35-100 of SEQ ID NO: 3; the first nuclear localization signal sequence in the fusion protein comprises 31 amino acids, which are encoded by nucleotides 101-193 of SEQ ID NO: 3; the first linker connecting the nuclear localization signal sequence in the fusion protein and the Cas mutant protein (amino acid sequence of SEQ ID NO: 6) comprises 8 amino acids, whose amino acid sequence is GIHGVPAA, which is encoded by nucleotides 194-217 of SEQ ID NO: 3; the second linker connecting the fusion protein and the T5 exonuclease comprises 10 amino acids, whose amino acid sequence is SGGSGGSGGS, which is encoded by nucleotides 3353-3382 of SEQ ID NO: 3; the T5 exonuclease in the fusion protein comprises 291 amino acids, which are encoded by nucleotides 3383-4255 of SEQ ID NO: 3; the second nuclear localization sequence in the fusion protein comprises 16 amino acids, which is encoded by nucleotides 3383-4255 of SEQ ID NO: 3 NO:3 encodes nucleotides 4256-4303.

[0021] Furthermore, the recombinant vector expressing the aforementioned fusion protein (N168R+S273R+L332R+G478R+S599R mutant) expresses the fusion protein.

[0022] The present invention also provides a biomaterial related to a mutant protein, wherein the mutant protein is the aforementioned Cas mutant protein, and the biomaterial is any one of the following:

[0023] B1), a nucleic acid molecule encoding the aforementioned Cas mutant protein;

[0024] B2), an expression cassette containing the nucleic acid molecule described in B1);

[0025] B3), a recombinant vector containing the nucleic acid molecule described in B1), or a recombinant vector containing the expression cassette described in B2);

[0026] B4), a recombinant microorganism containing the nucleic acid molecule described in B1), or a recombinant microorganism containing the expression cassette described in B2), or a recombinant microorganism containing the recombinant vector described in B3).

[0027] The above-mentioned biological material related to the mutant protein, the nucleotide sequence of the nucleic acid molecule B1) is obtained by replacing the codon encoding the 168th amino acid N in SEQ ID NO: 1 (nucleotides 502-504 of SEQ ID NO: 2 aac) with a codon encoding R (aga), the codon encoding the 273rd amino acid S in SEQ ID NO: 1 (nucleotides 817-819 of SEQ ID NO: 2 agc) with a codon encoding R (aga), the codon encoding the 332nd amino acid L in SEQ ID NO: 1 (nucleotides 994-996 of SEQ ID NO: 2 ctg) with a codon encoding R (aga), the codon encoding the 478th amino acid G in SEQ ID NO: 1 (nucleotides 1432-1434 of SEQ ID NO: 2 ggc) with a codon encoding R (aga), the codon encoding the 599th amino acid S in SEQ ID NO: 1 (nucleotides 817-819 of SEQ ID NO: 2 agc) with a codon encoding R (aga), and the codon encoding the 332nd amino acid L in SEQ ID NO: 1 (nucleotides 994-996 of SEQ ID NO: 2 ctg) with a codon encoding R (aga). The nucleotide sequence of SEQ ID NO: 2 at positions 1795-1797 (agc) is replaced with the codon encoding R (aga) while keeping the other nucleotide sequences of SEQ ID NO: 2 unchanged.

[0028] The present invention also provides a biomaterial related to the fusion protein, wherein the fusion protein is the aforementioned fusion protein, and the biomaterial is any one of the following:

[0029] B1), a nucleic acid molecule encoding the aforementioned fusion protein;

[0030] B2), an expression cassette containing the nucleic acid molecule described in B1);

[0031] B3), a recombinant vector containing the nucleic acid molecule described in B1), or a recombinant vector containing the expression cassette described in B2);

[0032] B4), a recombinant microorganism containing the nucleic acid molecule described in B1), or a recombinant microorganism containing the expression cassette described in B2), or a recombinant microorganism containing the recombinant vector described in B3).

[0033] In the above-mentioned biological materials related to the fusion protein, the nucleotide sequence of the nucleic acid molecule described in B1) is obtained by replacing nucleotides 218 to 3352 of SEQ ID NO: 3 with nucleotides 1-3135 of the coding sequence of the aforementioned Cas mutant protein while keeping the other nucleotides of SEQ ID NO: 3 unchanged.

[0034] In the above-mentioned biological materials, the expression cassette containing a nucleic acid molecule described in B3) refers to a DNA capable of expressing the RNA molecule described above in a host cell. The expression cassette may also include a single-stranded or double-stranded nucleic acid molecule containing all regulatory sequences necessary to express the nucleic acid molecule of any of the above-mentioned proteins or the DNA of the RNA molecule. The regulatory sequences can, under their compatible conditions, direct the coding sequence to express any of the above-mentioned proteins or the DNA of the RNA molecule in a suitable host cell. The regulatory sequences include, but are not limited to, a leader sequence, a polyadenylation sequence, a propeptide sequence, a promoter, a signal sequence, and a transcription terminator. At a minimum, the regulatory sequence must include a promoter and termination signals for transcription and translation. In order to introduce specific restriction enzyme sites into the vector so as to connect the regulatory sequence to the coding region of the nucleic acid sequence encoding the protein or the DNA of the RNA molecule, a regulatory sequence with a linker can be provided. The regulatory sequence can be a suitable promoter sequence, i.e., a nucleic acid sequence that is recognized by the host cell expressing the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of the protein or the DNA of the RNA molecule. The promoter can be any nucleic acid sequence that is transcriptionally active in the selected host cell, including mutant, truncated, and hybrid promoters, and can be derived from genes encoding extracellular or intracellular proteins that are homologous or heterologous to the host cell. The regulatory sequence can also be a suitable transcription termination sequence, i.e., a sequence that is recognized by the host cell to terminate transcription. The termination sequence can be operably linked to the 3' end of the nucleic acid sequence encoding the protein or the DNA of the RNA molecule. Any terminator that is functional in the selected host cell can be used in the present invention. The regulatory sequence can also be a suitable leader sequence, i.e., an untranslated region of an mRNA that is important for translation in the host cell. The leader sequence can be operably linked to the 5' end of the nucleic acid sequence encoding the protein or the DNA of the RNA molecule. Any leader sequence that is functional in the selected host cell can be used in the present invention. The regulatory sequence can also be a signal peptide coding region, which encodes an amino acid sequence attached to the amino terminus of the protein that directs the DNA encoding the protein or the RNA molecule into the cellular secretory pathway. Any signal peptide coding region that directs the expressed protein or the DNA of the RNA molecule into the secretory pathway of the selected host cell can be used in the present invention. It may also be desirable to add regulatory sequences that regulate the expression of the protein or RNA molecule according to the growth conditions of the host cell. Examples of regulatory sequences are systems that can turn gene expression on or off in response to chemical or physical stimuli, including in the presence of regulatory compounds. Other examples of regulatory sequences are those that allow gene amplification.

[0035] In the above biological materials, the vector may be a plasmid, cosmid, phage or viral vector.

[0036] In the above-mentioned biological materials, the microorganisms may be yeast, bacteria, algae or fungi.

[0037] The present invention also provides a composition for gene editing, which comprises the aforementioned Cas mutant protein and at least one gRNA; the gRNA is capable of binding to the aforementioned Cas mutant protein.

[0038] In the above composition, the gRNA includes a first segment and a second segment; the first segment is also called a "skeleton region", "protein binding segment", "protein binding sequence", or "direct repeat sequence"; the second segment is also called a "targeting sequence for targeting nucleic acid" or "targeting segment for targeting nucleic acid", or "guide sequence for targeting target sequence".

[0039] The first segment of the gRNA is capable of interacting with the aforementioned Cas mutant protein or fusion protein, thereby forming a complex between the Cas mutant protein and the gRNA or forming a complex between the fusion protein and the gRNA.

[0040] The present invention also provides an application of a Cas protein in gene editing or in preparing a product for gene editing, wherein the Cas protein is the aforementioned Cas mutant protein.

[0041] The present invention also provides a use of a fusion protein in gene editing or in the preparation of a product for gene editing, wherein the fusion protein is the aforementioned fusion protein.

[0042] The product is a combination reagent or kit for gene editing.

[0043] The present invention also provides an application of a biomaterial, wherein the biomaterial is the aforementioned biomaterial associated with a mutant protein or a biomaterial associated with a fusion protein, and the application is the application of the biomaterial in gene editing or in the preparation of products for gene editing.

[0044] The present invention also provides a kit for gene editing, which includes the aforementioned Cas mutant protein, the aforementioned fusion protein or the aforementioned combination.

[0045] The above kits also include reagents necessary for gene editing, such as containers, reagents, culture media, cytokines, buffers, antibodies, etc.

[0046] The present invention improves editing efficiency by mutating multiple amino acids of the wild-type Cas12i protein and can be used for multi-gene editing. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1It is the editing efficiency of each combination after the efficient mutation in region 1 and region 3 based on the combination of 168 / 599 in region 2. Wherein WT represents the recombinant vector expressing Cas12i.3 mutant (S273R), 168 / 599 represents the recombinant vector expressing mutant protein 168 / 599, and mutant protein 168 / 599 is obtained by mutating the 168th amino acid of SEQ ID NO: 1 to arginine (R), the 273rd amino acid to arginine (R), the 599th amino acid to arginine (R), and keeping the other amino acids unchanged. 1-1 indicates the recombinant vector expressing the mutant protein N168R+L332R+S599R+S273R was transferred into, 1-2 indicates the recombinant vector expressing the mutant protein N168R+S599R+D851R+S273R was transferred into, 1-3 indicates the recombinant vector expressing the mutant protein N168R+G478R+S599R+S273R was transferred into, 1-4 indicates the recombinant vector expressing the mutant protein N168R+G478R+D551R+S599R+S273R was transferred into, 1-5 indicates the recombinant vector expressing the mutant protein N 1-6 indicates the recombinant vector expressing the mutant protein N168R+L332R+G478R+S599R+S273R, 1-7 indicates the recombinant vector expressing the mutant protein N168R+G478R+S599R+D851R+S273R, 1-8 indicates the recombinant vector expressing the mutant protein N168R+G478R+D551R+S599R+D851R+S273R.

[0048] Figure 2The figure shows the editing efficiency of the mutant N168R+S273R+L332R+G478R+S599R in ADRB2, CHRM4, FANCF, and CXCR4. The control group at ADRB2 represents the S273R gene editing vector targeting the ADRB2 gene, and 168 / 332 / 478 / 599 represents the N168R+S273R+L332R+G478R+S599R gene editing vector targeting the ADRB2 gene. The control group at CHRM4 represents the S273R gene editing vector targeting the CHRM4 gene, and 168 / 332 / 478 / 599 represents the N168R+S273R+L332R+G478R+S599R gene editing vector targeting the CHRM4 gene. The control group at FANCF indicated the transfer of the S273R gene editing vector targeting the FANCF gene, and 168 / 332 / 478 / 599 indicated the transfer of the N168R+S273R+L332R+G478R+S599R gene editing vector targeting the FANCF gene; the control group at CXCR4 indicated the transfer of the S273R gene editing vector targeting the CXCR4 gene, and 168 / 332 / 478 / 599 indicated the transfer of the N168R+S273R+L332R+G478R+S599R gene editing vector targeting the CXCR4 gene.

[0049] Figure 3 Schematic diagram of the construction of 20-target and 30-target vectors using the mutant N168R+S273R+L332R+G478R+S599R.

[0050] Figure 4 This is a gel image of T7E1 digestion of the mutant N168R+S273R+L332R+G478R+S599R at positions 3, 5, 7, 8, 9, and 12 of 20 target sites. Target 3 indicates GAPDH-1, 5 indicates LMNA-1, 7 indicates AR, 8 indicates ADRB2, 9 indicates CCR4, and 12 indicates CD2. Each lane in the figure has three bands marked with "*." The band with the highest molecular weight is the specific band that was not cleaved, while the remaining two bands with smaller molecular weights indicate cleaved specific bands.

[0051] Figure 5This is a gel image of T7E1 digestion of the mutant N168R+S273R+L332R+G478R+S599R at positions 3, 5, 7, 8, 9, 12, 21, 24, 28, and 30, out of 30 target sites. Target 3 indicates GAPDH-1, 5 indicates LMNA-1, 7 indicates AR, 8 indicates ADRB2, 9 indicates CCR4, 12 indicates CD2, 21 indicates DNMT1-1, 24 indicates EMX1-1, 28 indicates FANCF-3, and 30 indicates VEGFA. Each lane has three bands marked with "*." The band with the highest molecular weight is the specific band that was not cleaved, while the remaining two bands with smaller molecular weights indicate cleaved specific bands. DETAILED DESCRIPTION

[0052] definition:

[0053] In the present invention, amino acid residues can be represented by single letters or three letters, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamine (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine ​​(Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), arginine (Arg, R).

[0054] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.

[0055] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.

[0056] The data in the following examples were processed using GraphPad Prism 8 statistical software. The experimental results were expressed as mean ± standard deviation using an unpaired t-test. P < 0.05 (*) indicated a significant difference.

[0057] Example 1: Screening Cas12i efficient mutation combinations in three regions

[0058] 1. Obtaining mutant proteins

[0059] 1. Obtaining the S273R mutant

[0060] The inventors provided the amino acid sequence of wild-type Cas12i.3 (its amino acid sequence is SEQ ID NO: 1, and its CDS sequence is SEQ ID NO: 2) during the preliminary research process. Later, during the experiment, a mutation was accidentally introduced, in which the third C of the codon of the 273rd amino acid of Cas12i.3 was mutated to A (C→A), causing the amino acid to mutate from serine S to arginine R (S273R). The S273R variant of Cas12i.3 was subsequently referred to as S273R.

[0061] 2. Obtaining L332R and other mutants

[0062] The tag, nuclear localization signal, linker, T5 exonuclease and other sequences were synthesized with S273R and further constructed into a fusion protein of the Cas protein variant S273R (nucleotide sequence is SEQ ID NO: 3, a total of 4303 bp), wherein nucleotides 1-31 are vector homologous sequences (including gccaccatgg (kozak sequence)), nucleotides 32-34 are the start codon ATG, nucleotides 35-100 are 3×FLAG tag sequences (encoding 22 amino acids), nucleotides 101-193 are the first nuclear localization signal sequence (C-Myc NLS and BP NLS, encoding 31 amino acids), nucleotides 194-217 are linker (encoding amino acids GIHGVPAA, a total of 8 amino acids), nucleotides 218-3352 are the nucleotide sequence of S273R (SEQ ID NO: 3). The 273rd amino acid of NO:1 is mutated from serine S to arginine R while keeping the other amino acid sequences unchanged, encoding 1045 amino acids), nucleotides 3353-3382 are the nucleotide sequence of linker (encoding amino acids SGGSGGSGGS, a total of 10 amino acids), nucleotides 3383-4255 are the T5 exonuclease coding sequence (encoding 291 amino acids), and nucleotides 4256-4303 are the second nuclear localization signal sequence (nucleoplasmin NLS, encoding 16 amino acids). The nucleotide sequence is a DNA fragment of SEQ ID NO:3 (encoding a total of 1423 amino acids) synthesized by sequence.

[0063] The nucleotide sequence of S273R is obtained by replacing the codon encoding amino acid S at position 273 of SEQ ID NO: 1 (nucleotides 817-819 agc of SEQ ID NO: 2) with the codon encoding R (aga). That is, the nucleotide sequence of S273R differs from that of SEQ ID NO: 2 in that the nucleotide at position 819 of SEQ ID NO: 2 is replaced by adenine deoxyribonucleotide (A) instead of cytosine deoxyribonucleotide (C).

[0064] Based on the coding sequence SEQ ID NO of the fusion protein of the aforementioned synthetic Cas protein variant S273R:3, the S273R mutation is divided into three regions (region 1, region 2, region 3) according to spatial position, and efficient mutation combination is found in each region. The following are the mutations or mutation combination types distributed in 3 regions. It should be noted that each of the above-mentioned mutant proteins is mutated on the basis of S273R, so the S273R mutation site is not shown in the naming, such as L332R is actually the wild-type Cas12i.3 (amino acid sequence is SEQ ID NO:1) the 273rd amino acid is mutated to R (arginine) from S (serine), the 332nd amino acid is mutated to R (arginine) from L (leucine) and the mutant protein obtained by keeping other amino acids unchanged.

[0065] Various combinations of distribution area 1: L332R, T850R, D851R, L332R+T850R, L332R+D851R, L332R+T850R+D851R.

[0066] Various combinations distributed in area 2: N168R, D233R, T235R, D267R, N168R+D267R, N168R+D233R, N168R+T235R, D233R+D267R, T235R+D267R, N168R+T235R+D267R, S7R, T505R, S599R, S7R+N168R, N168R+T5 05R、N168R+S599R、S7R+N168R+T505R、N168R+T505R+S599R、S7R+N168R+D267R、N168R+D2 67R+T505R, N168R+D267R+S599R, S7R+N168R+D267R+T505R, N168R+D267R+T505R+S599R.

[0067] Distributed in area 3: S477R, G478R, D551R, L662R, D551R+L662R, S477R+L662R, G478R+L662R, S477R+D551R, G478R+D551R, S477R+D551R+L662R, G478R+D551R+L662R.

[0068] The mutant proteins corresponding to the combinations of the above-mentioned regions are named after the mutation positions. Taking T850R, L332R+T850R, S477R+D551R+L662R and N168R+D267R+T505R+S599R as examples, T850R is a mutant protein obtained by mutating the 273rd amino acid of wild-type Cas12i.3 (amino acid sequence is SEQ ID NO: 1) from S to R, the 850th amino acid from T to R, and keeping other amino acids unchanged. L332R+T850R is a mutant protein obtained by mutating the 273rd amino acid of wild-type Cas12i.3 (amino acid sequence is SEQ ID NO: 1) from S to R, the 332nd amino acid from L to R, the 850th amino acid from T to R, and keeping other amino acids unchanged. S477R+D551R+L662R is a mutant protein obtained by mutating the 273rd amino acid of the wild-type Cas12i.3 (amino acid sequence is SEQ ID NO: 1) from S to R, the 477th amino acid from S to R, the 551st amino acid from D to R, the 662nd amino acid from L to R, and keeping the other amino acids unchanged. The N168R+D267R+T505R+S599R mutant is a mutant protein obtained by mutating the 273rd amino acid of the wild-type Cas12i.3 (amino acid sequence is SEQ ID NO: 1) from S to R, the 168th amino acid from N to R, the 267th amino acid from D to R, the 505th amino acid from T to R, the 599th amino acid from S to R, and keeping the other amino acids unchanged. Subsequent reference is made to the corresponding mutant protein with each of the aforementioned combinations.

[0069] The preparation methods of the DNA fragments for each mutant protein in the above combination are similar. Taking L332R as an example, the DNA sequence encoding the aforementioned L332R is based on SEQ ID NO: 2, except that nucleotides 817-819 in SEQ ID NO: 2 (encoding amino acid S at position 273 in SEQ ID NO: 1) are changed to AGA (encoding arginine), and nucleotides 994-996 in SEQ ID NO: 2 (encoding amino acid L at position 332 in SEQ ID NO: 1) are changed to AGA (encoding arginine). With the aforementioned AGA as the center, the DNA sequence encoding the aforementioned L332R is divided into two parts, and two pairs of primers are designed for amplifying these two parts of the DNA sequence.

[0070] First pair of primers:

[0071] Primer-F: 5'-TCACTTTTTTTCAGGTTGGACCGGTGCC-3';

[0072] L332R-R: 5'-GTCGCCCTCGCTCAGTCTGCCtctCAGCTCCAC-3';

[0073] Second pair of primers:

[0074] L332R-F: 5'-GTGGAGCTGagaGGCAGACTGAGCGAGGGCGAC-3';

[0075] Primer-R: 5'-CTTTTTCTTTTTTGCCTGGCCGGCCT-3'.

[0076] L332R-R and L332R-F were designed to contain the codons for the amino acid mutation point at position 332 (lowercase letters in L332R-R and L332R-F), totaling 33 bp of sequence.

[0077] Using S273R as a template, Primer-F and L332R-R were used to amplify the first part of the nucleotide sequence of the L332R mutant protein, and L332R-F and Primer-R were used to amplify the second part of the nucleotide sequence of the L332R mutant protein to obtain two fragments containing overlapping sequences. Fragment amplification kit: PrimeSTAR Max DNA Polymerase (Baoriyi Biotechnology Co., Ltd., R045A), for details of the experimental process, please refer to the instructions. The two fragments were then connected using the seamless cloning kit (pEASY-Basic Seamless Cloning and Assembly Kit (Beijing Quanshijin Biotechnology Co., Ltd., CU201-02), for details of the experimental process, please refer to the instructions. A PCR product containing the nucleotide sequence encoding the L332R mutant protein was obtained.

[0078] Through the above methods and examples, target mutations can be accurately and efficiently introduced as points to obtain DNA fragments of mutant proteins represented by various combinations in distribution area 1, distribution area 2, and distribution area 3, and any combination thereof.

[0079] It should be noted that the DNA fragments of the above mutant proteins all contain sequences such as tags, nuclear localization signals, linkers, and T5 exonuclease.

[0080] 2. Construction of a gene editing plasmid targeting tdTomato and containing mutant proteins

[0081] 1. Construction of U6-CBh-Cas9-T2A-EGFP-bGH polyA vector

[0082] PX458: also known as U6-sgRNA-CBh-Cas9-T2A-EGFP-bGH polyA, purchased from Addgene vector library, catalog number 48138.

[0083] PX458 was double-digested with the restriction endonucleases BbsI (NEB) and XbaI (NEB) to remove the sgRNA scaffold sequence, generating a linearized PX458 vector. Primers TF and TR were synthesized and annealed to generate a DNA duplex complementary to the linearized PX458 vector after digestion. This annealing product was the annealing product. Annealing system: 2.5 μL of 100 μM TF, 2.5 μL of 100 μM TR, 1 μL of T4 ligase buffer, and ddH2O to 10 μL. Annealing procedure: Heat in a metal bath at 95°C for 5 min, then open the bath, close the bath, and cool to room temperature. The recovered linearized PX458 vector (double-digested with BbsI and XbaI) was ligated to the annealing product using a T4 ligase kit (Baori Biotechnology Co., Ltd.) to generate the vector U6-CBh-Cas9-T2A-EGFP-bGH polyA.

[0084] TF: 5′-CACCACTAGTT-3′;

[0085] TR: 5′-CTAGAACTAGT-3′.

[0086] The difference between U6-CBh-Cas9-T2A-EGFP-bGH polyA and PX458 is that U6-CBh-Cas9-T2A-EGFP-bGH polyA is a recombinant vector obtained by replacing the sgRNA scaffold sequence between the BbsI and XbaI restriction enzyme recognition sites of PX458 with a DNA fragment formed by annealing TF and TR, while keeping the other nucleotides of PX458 unchanged.

[0087] 2. Construction of a gene editing plasmid targeting tdTomato and containing mutant proteins

[0088] (1) Construction of plasmid targeting tdTomato

[0089] In this example, a gRNA targeting tdTomato was designed for the tdTomato coding sequence (GenBank Accession No. KT878736.1, positions 2529-3959 (Update Date 06-OCT-2015)). The target sequence was 5'-AAGACCAUCUACAUGGCCAAGAA-3' (targeting nucleotides 3063 to 3085 of GenBank Accession No. KT878736.1). The U6-CBh-Cas9-T2A-EGFP-bGH polyA vector was double-digested with KpnI (NEB) and SpeI (NEB), and the large fragment containing the U6 promoter (with SpeI and KpnI restriction sites after the U6 promoter) was recovered to obtain the U6-CBh-Cas9-T2A-EGFP-bGH polyA linear vector. The PCR product expressing the DNA fragment containing the gRNA sequence targeting tdTomato was amplified by primers tdTomato-F and tdTomato-R, and the PCR product was recovered and the concentration was determined using a product recovery kit (Guangzhou Meiji Biotechnology Co., Ltd., Cat. No. D2111-02). The gRNA fragment (DNA fragment containing the gRNA sequence targeting tdTomato) was homologously recombined with the U6-CBh-Cas9-T2A-EGFP-bGH polyA linear vector using a seamless cloning kit to form a U6-crRNA-CBh-Cas9-T2A-EGFP-bGH polyA vector.

[0090] The DNA fragment expressing the gRNA sequence targeting tdTomato is 5'-aaaggacgaaacaccGAGAGAATGTGcGCATAGTCgCACAAGACCATCTACATGGCCAAGAATTTTTTTgtacccgttacataa-3' (84 bp). Nucleotides 17 to 39 are direct repeats, and nucleotides 40 to 62 are the target sequence, targeting bases 3063 to 3085 of GenBank Accession No. KT878736.1. Nucleotides 63 to 69 are transcription termination signals. The nucleotides in lowercase letters are homologous to the U6-CBh-Cas9-T2A-EGFP-bGH polyA vector. The additional G at position 16 enhances transcription.

[0091] tdTomato-F: 5'-aaaggacgaaacaccGAGAGAATGTGCGCATAGTCGCAC-3';

[0092] tdTomato-R: 5'-ttatgtaacgggtacAAAAAAATTCTTGGCCATGTAGATGGTCTTGTGCGACTATGCGCA-3.

[0093] The difference between the U6-crRNA-CBh-Cas9-T2A-EGFP-bGH polyA vector and the U6-CBh-Cas9-T2A-EGFP-bGH polyA vector is that the U6-crRNA-CBh-Cas9-T2A-EGFP-bGH polyA vector is a recombinant vector obtained by replacing the DNA fragment between the KpnI and SpeI enzyme recognition sites of U6-CBh-Cas9-T2A-EGFP-bGH polyA with a DNA fragment expressing the gRNA sequence targeting tdTomato, while keeping the other nucleotide sequences of the U6-CBh-Cas9-T2A-EGFP-bGH polyA vector unchanged.

[0094] (2) Insertion of mutant proteins and plasmid construction

[0095] The plasmid U6-crRNA-CBh-Cas9-T2A-EGFP-bGH polyA was double-digested with restriction endonucleases AgeI (NEB) and FseI (NEB) to remove the Cas9 coding sequence, obtaining a linearized plasmid without the Cas9 coding sequence.

[0096] The DNA fragments of the mutant proteins in "I. Obtaining mutant proteins" above were constructed by seamless cloning on the aforementioned linearized plasmid without the Cas9 coding sequence to obtain U6-crRNA-CBh-T2A-EGFP-bGH polyA vectors expressing each mutant protein one by one.

[0097] For example, the DNA fragment of the L332R mutant protein was constructed on the U6-crRNA-CBh-T2A-EGFP-bGH polyA vector by seamless cloning to obtain a recombinant vector that can express the L332R mutant protein.

[0098] Recombinant vectors expressing the aforementioned other mutant proteins can be prepared by referring to the above methods.

[0099] The above-mentioned recombinant vector may be subsequently referred to as a gene editing vector containing a Cas protein variant.

[0100] 2. Cell culture and plasmid transfection

[0101] The tdTomato ovine fibroblast fluorescent reporter cell line (Wang Linli, 2024, Establishment and Application of a High-Efficiency Multiplex CRISPR / Cas12i.3 Gene Editing System, China Agricultural University, 2024[D]) was cultured in DMEM (Gibco) supplemented with 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Well-functioning cells were plated onto 10 cm culture dishes (Corning) and cultured until confluence reached approximately 80%. Cells were digested with trypsin-EDTA (0.25%, Gibco) and collected into EP tubes. The cells were suspended in 100 μL of electroporation buffer (Beijing Ingen Biotechnology Co., Ltd., Cat. No. 98668-20) and 7 μg of plasmid (the gene editing vector containing the Cas protein variant constructed above) was added and mixed thoroughly. The cells were placed in a Lonza Amaxa Nucleofector 2B Nucleofector, set to program A-033, and electroporated. After electroporation, 500 μL of DMEM high-glucose medium was immediately added and the cells were incubated at 37°C for 10 minutes. The cells were plated into 6-well plates with complete medium containing 20% ​​FBS. After 6 hours of transfection, the culture medium was changed to complete medium containing 15% FBS. After 48 hours of transfection, the cells were digested with trypsin-EDTA (0.25%, Gibco) and analyzed by flow cytometry to analyze the changes in tdTomato fluorescence intensity in EGFP-positive cells, specifically by calculating the quenching ratio of the average fluorescence intensity of tdTomato and the proportion of cells with weak fluorescence intensity.

[0102] 3. Summary of results

[0103] In region 1, the L332R and D851R mutants showed the highest mean fluorescence intensity quenching ratios. In region 2, the N168R and N168R+D267R mutants were first detected, followed by the introduction of the remaining mutations. Ultimately, the N168R+S599R and S7R+N168R+T505R mutants showed the highest mean fluorescence intensity quenching ratios. In region 3, the G478R and G478R+D551R mutants did not show the highest mean fluorescence intensity quenching ratios, but did have the highest proportion of weak fluorescence intensity.

[0104] Example 2: Further optimization of the combination of highly effective mutations in the three regions screened

[0105] Through Example 1, it was found that the L332R and D851R mutants in region 1, the N168R+S599R and S7R+N168R+T505R mutants in region 2, and the G478R and G478R+D551R in region 3 had a good effect of improving editing efficiency, and the N168R+S599R mutant in region 2 was a two-step combination with the best efficiency improvement effect. Therefore, based on the N168R+S599R in region 2, the efficient combinations in regions 1 and 3 were further combined to further improve the editing efficiency.

[0106] Based on the N168R+S599R mutant in region 2, a mutant was constructed by a method similar to that of introducing mutations in Example 1.

[0107] Table 1. Mutation combination types of N168R+S599R in region 2

[0108] name Amino acid mutation type 168 / 599 N168R+S599R+S273R 1-1 N168R+L332R+S599R+S273R 1-2 N168R+S599R+D851R+S273R 1-3 N168R+G478R+S599R+S273R 1-4 N168R+G478R+D551R+S599R+S273R 1-5 N168R+L332R+G478R+S599R+S273R 1-6 N168R+L332R+G478R+D551R+S599R+S273R 1-7 N168R+G478R+S599R+D851R+S273R 1-8 N168R+G478R+D551R+S599R+D851R+S273R

[0109] Referring to the experimental method in Example 1, a plasmid was constructed, and referring to the experimental method in the example, the aforementioned plasmid was used for subsequent cell culture and plasmid transfection.

[0110] In the above plasmids:

[0111] Recombinant vector expressing the N168R+S273R+L332R+G478R+S599R mutant (fusion protein NLS-N168R+S273R+L332R+G478R+S599R-T5): The only difference between this recombinant vector and U6-crRNA-CBh-Cas9-T2A-EGFP-bGH polyA is that the recombinant fragment M replaces the coding sequence of the Cas9 protein between the AgeI and FseI enzyme recognition sites on the U6-crRNA-CBh-Cas9-T2A-EGFP-bGH polyA vector, and the other nucleotide sequences are exactly the same. The recombinant fragment M is a recombinant fragment obtained by replacing nucleotides 218 to 3352 of SEQ ID NO: 3 with a DNA fragment of the N168R + S273R + L332R + G478R + S599R mutant (excluding the stop codon, a total of 3135 bp) while keeping the other nucleotides of SEQ ID NO: 3 unchanged. The nucleotide sequence of the DNA fragment of the N168R + S273R + L332R + G478R + S599R mutant (excluding the stop codon, a total of 3135 bp) is as follows: the codon encoding amino acid N at position 168 of SEQ ID NO: 1 (nucleotides 502-504 of SEQ ID NO: 2 aac) is replaced with a codon encoding R (aga), the codon encoding amino acid S at position 273 of SEQ ID NO: 1 (nucleotides 817-819 of SEQ ID NO: 2 agc ... The codon encoding the 332nd amino acid L in SEQ ID NO: 1 (nucleotides 994-996 ctg in SEQ ID NO: 2) is replaced with a codon encoding R (aga), the codon encoding the 478th amino acid G in SEQ ID NO: 1 (nucleotides 1432-1434 ggc in SEQ ID NO: 2) is replaced with a codon encoding R (aga), and the codon encoding the 599th amino acid S in SEQ ID NO: 1 (nucleotides 1795-1797 agc in SEQ ID NO: 2) is replaced with a codon encoding R (aga), while keeping the other nucleotide sequences of SEQ ID NO: 2 unchanged.

[0112] The recombinant vector expressing the N168R+S273R+L332R+G478R+S599R mutant expresses the fusion protein NLS-N168R+S273R+L332R+G478R+S599R-T5, wherein the 22nd amino acid residue of the tag protein is linked to the 1st amino acid residue of the first nuclear localization signal sequence through a peptide bond, the 31st amino acid residue of the first nuclear localization signal sequence is linked to the 1st amino acid residue of the first linker through a peptide bond, and the 8th amino acid residue of the first linker is linked to the 1st amino acid residue of the first linker through a peptide bond. The amino acid residue at position 1045 of the N168R+S273R+L332R+G478R+S599R mutant is linked to the amino acid residue at position 1 by a peptide bond, the amino acid residue at position 1045 of the N168R+S273R+L332R+G478R+S599R mutant is linked to the amino acid residue at position 10 of the second linker by a peptide bond, the amino acid residue at position 10 of the second linker is linked to the amino acid residue at position 10 of the T5 exonuclease by a peptide bond, and the amino acid residue at position 291 of the T5 exonuclease is linked to the amino acid residue at position 1 of the second nuclear localization sequence by a peptide bond, resulting in a total of 1423 amino acids.

[0113] The tag protein in the fusion protein NLS-N168R+S273R+L332R+G478R+S599R-T5 comprises 22 amino acids, which are encoded by nucleotides 35-100 of SEQ ID NO: 3; the first nuclear localization signal sequence in the fusion protein comprises 31 amino acids, which are encoded by nucleotides 101-193 of SEQ ID NO: 3; the first linker connecting the nuclear localization signal sequence in the fusion protein and the N168R+S273R+L332R+G478R+S599R mutant comprises 8 amino acids, whose amino acid sequence is GIHGVPAA, which is encoded by nucleotides 194-217 of SEQ ID NO: 3; the second linker connecting the fusion protein and the T5 exonuclease comprises 10 amino acids, whose amino acid sequence is SGGSGGSGGS, which is encoded by nucleotides 194-217 of SEQ ID NO: 3. NO: 3 is encoded by nucleotides 3353-3382, the T5 exonuclease in the fusion protein contains 291 amino acids and is encoded by nucleotides 3383-4255 in SEQ ID NO: 3, and the second nuclear localization sequence in the fusion protein contains 16 amino acids and is encoded by nucleotides 4256-4303 in SEQ ID NO: 3.

[0114] 2. Summary of results

[0115] The results are as follows Figure 1As shown in Figure 2, based on the N168R+S599R in region 2 with the highest efficiency improvement, the N168R+S273R+L332R+G478R+S599R mutant obtained had the most significant improvement in editing activity. Therefore, this N168R+S273R+L332R+G478R+S599R mutant based on S273R was named N168R+S273R+L332R+G478R+S599R (amino acid sequence is SEQ ID NO: 6).

[0116] Example 3, Verification of Editing Efficiency of Engineered Cas12i Nuclease N168R+S273R+L332R+G478R+S599R

[0117] The recombinant vector expressing the S273R mutant obtained in Example 1 and the recombinant vector expressing the N168R+S273R+L332R+G478R+S599R mutant obtained in Example 2 were double-digested with KpnI (NEB) and SpeI (NEB) and recovered (there are SpeI and KpnI enzyme recognition sites after the U6 promoter) to obtain the S273R recombinant vector after enzyme digestion and the N168R+S273R+L332R+G478R+S599R recombinant vector after enzyme digestion.

[0118] Targets were designed for ADRB2, CHRM4, FANCF, and CXCR4 in HEK293T cells. DNA fragments expressing gRNA sequences targeting ADRB2 were amplified using primers F and ADRB2-R, DNA fragments expressing gRNA sequences targeting CHRM4 were amplified using primers F and CHRM4-R, DNA fragments expressing gRNA sequences targeting FANCF were amplified using primers F and FANCF-R, and DNA fragments expressing gRNA sequences targeting CXCR4 were amplified using primers F and CXCR4-R. PCR products were recovered using product recovery kits (Guangzhou Meiji Biotechnology Co., Ltd., Cat. No. D2111-02) and their concentrations were determined. The above DNA fragments were homologously recombined with the above enzyme-digested S273R recombinant vector (or enzyme-digested N168R+S273R+L332R+G478R+S599R recombinant vector) using a seamless cloning kit to form S273R gene editing vectors targeting ADRB2, CHRM4, FANCF, or CXCR4 genes, and N168R+S273R+L332R+G478R+S599R gene editing vectors targeting ADRB2, CHRM4, FANCF, or CXCR4 genes (DNA fragments of the gRNA sequences targeting the four genes were homologously recombined with any combination of the two enzyme-digested recombinant vectors). HEK293T cells were transfected with Lipofectamine 3000, and 20,000 EFGP-positive cells were collected for each combination. The above primer information is as follows:

[0119] F: 5'-aaaggacgaaacaccGAGAGAATGTGCGCATAGTCGCAC-3';

[0120] ADRB2-R: 5'-taagttatgtaacggAAAAAAAACCGAGGCACGCACATACAGGCAGTGcGACTATGCgCA-3';

[0121] CHRM4-R: 5'-taagttatgtaacggAAAAAAACGTGTCTGGGGAGGAAGGGGAGAGTGcGACTATGCgCA-3';

[0122] FANCF-R: 5'-taagttatgtaacggAAAAAAAGTGCTAGTCCACTGGCTTCTGGGGTGcGACTATGCgCA-3';

[0123] CXCR4-R: 5'-taagttatgtaacggAAAAAAACTTCAGGCGCATCCCGCTTCCCTGTGcGACTATGCgCA-3'.

[0124] Table 2. DNA fragments expressing gRNA sequences targeting the above genes

[0125]

[0126] In the above sequence, nucleotides 1 to 15 are the vector left homologous sequence, G at position 16 is used to enhance RNA transcription, nucleotides 17 to 39 are direct repeat sequences, and nucleotides 40 to 62 are target sequences (nucleotides 40 to 62 in ADRB2-R target nucleotides 148826203 to 148826225 of NC_000005.10, and nucleotides 40 to 62 in CHRM4-R target nucleotides 4638 to 4643 of NC_000011.10). Nucleotides 6582 to 46386604, nucleotides 40-62 in FANCF-R target nucleotides 22625083 to 22625105 of NC_000011.10, nucleotides 40-62 in CXCR4-R target nucleotides 136118388 to 136118410 of NC_000002.12), nucleotides 63 to 69 are transcription termination signals, and nucleotides 70 to 84 are vector right homologous sequences.

[0127] DNA was extracted from the cells collected above, and PCR reactions were performed using the primers in the table below to amplify the target region. Amplicon sequencing was performed, and the editing efficiency of each target was calculated using CRISPResso2 software: Editing efficiency (%) = number of mutated reads / total number of reads × 100%.

[0128] Table 3. PCR reaction primers

[0129] Primer name Sequence (5'-3') HEK293T_ADRB2_seq_F GGAGGGTGTGTCTCAGTGTC HEK293T_ADRB2_seq_R GCTTTTGGCTCTTCTGTGGC HEK293T_CHRM4_seq_F ATTCTGCCAGAGAATGTCCCTC HEK293T_CHRM4_seq_R ATTTCCACCGTCTCATAGCGA HEK293T_CXCR4_seq_F CCTGGGCCTCAGTGTCTCTA HEK293T_CXCR4_seq_R CAGGGGACCCTGCTGTTTG HEK293T_VEGFA_seq_F GAGAAGGCCAGGGGTCACTC HEK293T_VEGFA_seq_R AACTCTGTCCAGAGACACGC

[0130] The results are as follows Figure 2 As shown, the editing efficiency of the Cas mutant protein N168R+S273R+L332R+G478R+S599R was significantly improved compared with S273R.

[0131] Example 4, Application of Engineered Cas12i Nucleic Acid (N168R+S273R+L332R+G478R+S599R) in Multi-Gene Editing

[0132] The recombinant vector expressing the N168R+S273R+L332R+G478R+S599R mutant obtained in Example 3 was digested with Sa cI (NEB) and recovered to obtain the digested N168R+S273R+L332R+G478R+S599R recombinant vector.

[0133] The triplex-crRNA array-20 target (T20) and triplex-crRNA array-30 target (T30) targeting 20 target sites (12 genes) and 30 target sites (17 genes) of genes such as ACTB, GAPDH, and LMNA were synthesized, and homologously recombined with the above-mentioned enzyme-digested N168R+S273R+L332R+G478R+S599R recombinant vector to form gene editing vectors T20-U6 and T30-U6. The above two gene editing vectors target 20 target sites (12 genes) and 30 target sites (17 genes) of genes such as ACTB, GAPDH, and LMNA, respectively (as shown in FIG). Figure 3 shown).

[0134] T20-U6 is a recombinant expression vector obtained by inserting a DNA fragment with the nucleotide sequence of SEQ ID NO: 4 (T20, the lowercase sequence agagaatgtgcgcatagtcgcac is a direct repeat sequence in crRNA, totaling 1128bp) into the SacI restriction site of the N168R+S273R+L332R+G478R+S599R recombinant vector while keeping other nucleotide sequences unchanged.

[0135] In SEQ ID NO:4, nucleotides 1-25 are vector homologous sequences used for seamless cloning with the vector backbone after enzyme digestion, nucleotides 26-135 are a triplex structure (triplex, which can stabilize mRNA in the absence of a poly (A) tail), nucleotides 136-185 are meaningless, nucleotides 186-1128 are lowercase direct repeat sequences, and the rest are target ACTB-1, target ACTB-2, target GAPDH-1, target GAPDH-2, target LMNA-1, target LMNA-2, target AR, target ADRB2, target CCR4, target CCR10-1, target CCR10-2, target CD2, target CHR M4-1, target CHRM4-2, target CXCR4-1, target CXCR4-2, target HBB-1, target HBB-2, target IL1RN-1, and target IL1RN-2.

[0136] T30-U6 is a recombinant expression vector obtained by inserting a DNA fragment with the nucleotide sequence of SEQ ID NO: 5 (T30, the lowercase sequence agagaatgtgcgcatagtcgcac is a direct repeat sequence in crRNA, totaling 1588bp) into the SacI restriction site of the N168R+S273R+L332R+G478R+S599R recombinant vector while keeping other nucleotide sequences unchanged.

[0137] In SEQ ID NO: 5, nucleotides 1-25 are vector homologous sequences, nucleotides 26-135 are the first triplex-crRNA array, nucleotides 136-185 are meaningless, nucleotides 186-1588 are lowercase letters in repeated sequences, and the rest are target ACTB-1, target ACTB-2, target GAPDH-1, target GAPDH-2, target LMNA-1, target LMNA-2, target AR, target ADRB2, target CCR4, target CCR10-1, target CCR 10-2, target CD2, target CHRM4-1, target CHRM4-2, target CXCR4-1, target CXCR4-2, target HBB-1, target HBB-2, target IL1RN-1, target IL1RN-2, target DNMT1-1, target DNMT1-2, target DNMT1-3, target EMX1-1, target EMX1-2, target FANCF-1, target FANCF-2, target FANCF-3, target GRIN2B, target VEGFA. The aforementioned repeat sequence is used to connect the target and the target. The Cas12i protein has RNase activity and can cut the expressed crRNA array at the 5' end of the repeat sequence to produce multiple crRNAs targeting different target genes.

[0138] HEK293T cells were transfected with Lipofectamine 3000, and 100,000 EGFP-positive cells were collected. T7E1 assay was performed on the 3rd, 5th, 7th, 8th, 9th, and 12th targets and the 3rd, 5th, 7th, 8th, 9th, 12th, 21st, 24th, 28th, and 30th targets, respectively, of the 20 and 30 targets. Gene editing efficiency (% indel) = 100 × (1-(1-Fraction Cl eaved) 1 / 2 ). Fraction Cleaved = the sum of the grayscale of the sheared band / the sum of the grayscale of the sheared band and the unsheared specific band. The target sequence information of the above crRNA is as follows, where targets 1-20 are the first 20 targets in Table 4, and targets 1-30 are all 30 targets in Table 4.

[0139] The T7E1 detection method is as follows: 48 hours after transfection, genomic DNA was extracted from cells using a genomic extraction kit (Guangzhou Meiji Biotechnology Co., Ltd., Cat. No. D3018-02). 100 ng of the extracted genomic DNA was used as a template for PCR amplification. Amplification reaction system and procedure: The total volume of the amplification reaction was 50 μL, and the primers listed in Table 5 were used for amplification. The components were: 100 ng DNA template, 1 μL each of 10 μmol / L upstream and downstream primers, 25 μL of PrimeS TAR (Baoriyi Biotechnology Co., Ltd.), and the total volume was made up to 50 μL with sterile deionized water. The PCR reaction procedure was as follows: initial denaturation at 98°C for 3 minutes; denaturation at 98°C for 10 seconds, annealing at 60°C for 15 seconds, and extension at 72°C for 30 seconds (33 cycles); and a final extension at 72°C for 5 minutes. After PCR, the PCR product was recovered using a product recovery kit and the concentration was determined.

[0140] Prepare the enzyme digestion system as follows: 500 ng of amplified product, 1.1 μL of cutsmart, and 10.5 μL of ddH2O. Mix thoroughly and follow the hybridization protocol: 95°C for 10 min; ramp down to 85°C at -2°C / s; then ramp down to 25°C at -0.1°C / s. Add 0.5 μL of T7E1 (NEB) and digest at 37°C for 15 min. Immediately add 2 μL of Loading Buffer and prepare a 2% agarose gel for electrophoresis analysis. Analyze the digestion results using a gel imaging system.

[0141] Table 4. Gene targets

[0142]

[0143]

[0144]

[0145] In Table 4, NC_000007.14 and the like are the NCBI Genebank accession numbers of the corresponding genes. The sequences corresponding to the aforementioned NCBI Genebank accession numbers are based on the update date closest to the filing date of this application.

[0146] Table 5. Amplification primers

[0147]

[0148] The results are as follows Figure 4 and Figure 5As shown, 3 represents the target GAPDH-1, 5 represents the target LMNA-1, 7 represents the target AR, 8 represents the target ADRB2, 9 represents the target CCR4, 12 represents the target CD2, 21 represents the target DNMT1-1, 24 represents the target EMX1-1, 28 represents the target FANCF-3, and 30 represents the target VEGFA. It can be seen that all selected sites showed effective cutting, indicating that the engineered Cas12i nucleic acid (N168R+S273R+L332R+G478R+S599R) can be used for multi-gene editing.

[0149] Table 6. Editing efficiency of each target in the 20-target array (Indel%)

[0150] Target name Position in the array Editing efficiency (Indel%) GAPDH-1 3 36.0 LMNA-1 5 * AR 7 26.4 ADRB2 8 30.8 CCR4 9 12.2 CD2 12 41.3

[0151] In the above table, “*” indicates that the target band was not detected and there was no editing efficiency.

[0152] Table 7. Editing efficiency of each target in the 30-target array (Indel%)

[0153]

[0154] SEQ ID NO: 1 (wild-type amino acid sequence)

[0155]

[0156] SEQ ID NO: 2 (wild-type nucleotide sequence)

[0157]

[0158] SEQ ID NO:3(3xFLAG-Myc NLS-BP NLS-Cas12i3(S273R)-linker(SGGSGGSGGS)-T5

[0159] exo-nucleoplasmin NLS (marked by a mutation in the DNA sequence) 4303 bp

[0160]

[0161] SEQ ID NO: 4 (T20, lowercase sequence marks the direct repeat sequence in crRNA)

[0162]

[0163] SEQ ID NO: 5 (T30, lowercase sequence marks the direct repeat sequence in crRNA)

[0164]

[0165] Sequence 6 engineered Cas12i nuclease N168R+S273R+L332R+G478R+S599R protein sequence (amino acid mutations are marked)

[0166]

[0167] The present invention has been described in detail above. For those skilled in the art, without departing from the purpose and scope of the present invention, and without the need to carry out unnecessary experimental conditions, the present invention can be implemented in a wide range under equivalent parameters, concentrations and conditions. Although the present invention provides specific embodiments, it should be understood that further improvements can be made to the present invention. In short, according to the principles of the present invention, this application is intended to include any changes, uses or improvements to the present invention, including changes that depart from the disclosed scope in this application and are made using conventional techniques known in the art.

Claims

1. A Cas mutant protein, characterized in that The Cas mutant protein is a mutant protein in which the 168th amino acid, the 273rd amino acid, the 332nd amino acid, the 478th amino acid and the 599th amino acid of SEQ ID NO: 1 are all mutated to arginine, and the other amino acid sequences remain unchanged.

2. A fusion protein, characterized in that The fusion protein is a protein comprising the Cas mutant protein according to claim 1.

3. A biomaterial associated with a mutant protein, characterized in that The biological material is any one of the following: B1), a nucleic acid molecule encoding the protein according to claim 1; B2), an expression cassette containing the nucleic acid molecule described in B1); B3), a recombinant vector containing the nucleic acid molecule described in B1), or a recombinant vector containing the expression cassette described in B2); B4), a recombinant microorganism containing the nucleic acid molecule described in B1), or a recombinant microorganism containing the expression cassette described in B2), or a recombinant microorganism containing the recombinant vector described in B3).

4. A biomaterial associated with a fusion protein, characterized in that The biological material is any one of the following: B1), a nucleic acid molecule encoding the protein according to claim 2; B2), an expression cassette containing the nucleic acid molecule described in B1); B3), a recombinant vector containing the nucleic acid molecule described in B1), or a recombinant vector containing the expression cassette described in B2); B4), a recombinant microorganism containing the nucleic acid molecule described in B1), or a recombinant microorganism containing the expression cassette described in B2), or a recombinant microorganism containing the recombinant vector described in B3).

5. A composition for gene editing, characterized in that The composition comprises the Cas mutant protein according to claim 1 and at least one gRNA; the gRNA is capable of binding to the Cas mutant protein according to claim 1.

6. A composition for gene editing, characterized in that The composition comprises the fusion protein of claim 2 and at least one gRNA; the gRNA is capable of binding to the fusion protein of claim 2.

7. Use of Cas protein in gene editing or in the preparation of products for gene editing, characterized in that: The Cas protein is the Cas mutant protein according to claim 1.

8. Use of a fusion protein in gene editing or in the preparation of a product for gene editing, characterized in that: The fusion protein is the fusion protein according to claim 2.

9. Use of the biomaterial according to claim 3 or 4 in gene editing or in the preparation of products for gene editing.

10. A kit for gene editing, characterized in that The kit comprises the Cas mutant protein according to claim 1, the fusion protein according to claim 2, or the composition according to claim 5 or 6.

Citation Information

Patent Citations

  • Novel crispr enzymes and systems

    CN113348245A

  • Mutated Cas protein and application thereof

    CN116286739A

  • Base editing tool and application thereof

    CN117327679A

  • Mutated Cas protein and application thereof

    CN118147110A

  • Mutant cas13b with improved efficiency

    WO2024155245A2