Cas mutant proteins and uses thereof

By performing site-directed mutagenesis and fusion protein design on the Cas-sf2044 protein, the problems of low editing efficiency and easy prediction of target sites in the CRISPR/Cas system were solved, achieving more efficient and specific gene editing.

CN119776319BActive Publication Date: 2025-10-10SHANDONG SHUNFENG BIOTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510045956.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-01-22
Filing Date
2025-01-13
Publication Date
2025-10-10
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

The existing CRISPR/Cas system has problems in gene editing, such as low editing efficiency, complex and diverse PAM sequences, and easy prediction of target sites leading to off-target effects. In addition, different Cas proteins have their own advantages and disadvantages, making it difficult to meet diverse editing needs.

Method used

By performing site-directed mutagenesis on the Cas-sf2044 protein, especially modifying specific amino acid sites, its editing activity is improved and its application range is expanded, including minor changes in the amino acid sequence and the design of fusion proteins, to optimize its expression and function in different cells.

Benefits of technology

It improves the editing activity and efficiency of Cas proteins, expands their application range, enhances their specificity for target sequences, reduces off-target effects, and adapts to diverse gene editing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119776319B_ABST
    Figure CN119776319B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of nucleic acid editing, in particular, the field of clustered regularly interspaced short palindromic repeats (CRISPR) technology. Specifically, the present application provides a Cas mutant protein with improved activity and increased editing efficiency. Compared with the wild-type parent Cas protein, the Cas mutant protein of the present application has significantly improved activity and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese patent application CN202410087586.3, filed on January 22, 2024. This application incorporates the entire text of the aforementioned Chinese patent application. Technical Field

[0002] The present invention relates to the field of gene editing, in particular to the field of clustered regularly interspaced short palindromic repeats (CRISPR). Specifically, the present invention relates to a Cas mutant protein with improved activity and editing efficiency and its application. Background Art

[0003] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to guide specific binding to target sequences on the genome and cut DNA to produce double-strand breaks, and uses biological non-homologous end joining or homologous recombination to perform site-directed gene editing.

[0004] The CRISPR / Cas9 system is the most commonly used Type II CRISPR system, which recognizes a 3'-NGG PAM motif and performs blunt-end cleavage on target sequences. A newly discovered class of CRISPR / Cas systems, Type V, utilizes a 5'-TTN motif and performs sticky-end cleavage on target sequences. Examples include Cpf1, C2c1, CasX, and CasY. However, the various CRISPR / Cas systems currently available have varying advantages and disadvantages. For example, Cas9, C2c1, and CasX all require two guide RNAs, while Cpf1 requires only a single guide RNA and can be used for multiplexed gene editing. CasX is 980 amino acids long, while the more common Cas9, C2c1, CasY, and Cpf1 are typically around 1300 amino acids. Furthermore, the PAM sequences of Cas9, Cpf1, CasX, and CasY are more complex and diverse, while C2c1 recognizes a strict 5'-TTN motif, making its target site more predictable than other systems, thereby reducing potential off-target effects.

[0005] Chinese invention patent CN117230042A discloses a Cas protein Cas-sf2044. In order to improve the editing efficiency of the protein, this application optimizes the protein and improves its editing efficiency. Summary of the Invention

[0006] After extensive experiments and repeated exploration, the inventors of this application improved the editing activity of the Cas-sf2044 protein and expanded its scope of application by site-directed mutagenesis of the protein.

[0007] Cas effector proteins

[0008] On the one hand, the present invention provides a Cas mutant protein with improved or enhanced activity, wherein the mutant protein has a mutation at any one or several of the following amino acid positions corresponding to the amino acid sequence shown in SEQ ID No. 1: position 160, position 163, position 674, position 694, position 697, position 335, position 602, position 644, position 171, position 254, position 337, position 482, position 514, position 517, position 528, position 538, position 700, position 866 or position 681 compared with the amino acid sequence of the parent Cas protein.

[0009] The above amino acid site refers to the site from the N-terminus of SEQ ID No. 1.

[0010] In one embodiment, the amino acid at position 160 or the amino acid at position 674 mutates to a non-N amino acid, for example, A, V, G, D, L, F, W, Y, Q, S, E, K, M, T, C, P, H, R, or I; preferably, the amino acid at position 160 or the amino acid at position 674 mutates to R.

[0011] In one embodiment, the amino acid at position 697 is mutated to a non-N amino acid, for example, A, V, G, D, L, F, W, Y, Q, S, E, K, M, T, C, P, H, R, I; preferably, it is mutated to G, Q, K or R.

[0012] In one embodiment, the amino acid at position 163 or the amino acid at position 602 mutates to a non-L amino acid, for example, A, V, G, N, Q, F, W, Y, D, S, E, K, M, T, C, P, H, R, or I; preferably, the amino acid at position 163 or the amino acid at position 602 mutates to R.

[0013] In one embodiment, the amino acid at position 694 or the amino acid at position 644 mutates to an amino acid other than A, for example, L, V, G, N, Q, F, W, Y, D, S, E, K, M, T, C, P, H, R, or I; preferably, the amino acid at position 694 or the amino acid at position 644 mutates to R.

[0014] In one embodiment, the amino acid at position 335 is mutated to a non-S amino acid, for example, A, V, G, L, Q, F, W, Y, D, N, E, K, M, T, C, P, H, R, I; preferably, it is mutated to R.

[0015] In one embodiment, the 171stamino acid or the 254thamino acid or the 337thamino acid or the 514thamino acid or the 517thamino acid or the 528thamino acid or the 538thamino acid or the 866thamino acid is mutated to an amino acid other than E, for example, A, V, G, L, Q, F, W, Y, D, S, K, N, M, T, C, P, H, R, I; preferably, the 171stamino acid or the 254thamino acid or the 337thamino acid or the 514thamino acid or the 517thamino acid or the 528thamino acid or the 538thamino acid or the 866thamino acid is mutated to R.

[0016] In one embodiment, the 482ndamino acid is mutated to an amino acid other than E, for example, A, V, G, L, Q, F, W, Y, D, S, K, N, M, T, C, P, H, R, I; preferably, to G, A, V, L, I, M, F, Y, W, S, T, C, P, N, D, K, R, or H.

[0017] In one embodiment, the 700thamino acid is mutated to an amino acid other than E, for example, A, V, G, L, Q, F, W, Y, D, S, K, N, M, T, C, P, H, R, I; preferably, to G, L, I, M, Y, S, K, R, or H.

[0018] In one embodiment, the 681stamino acid is mutated to an amino acid other than D, for example, A, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, to R.

[0019] In one embodiment, the amino acid sequence of the parent Cas protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to SEQ ID No. 1.

[0020] In one embodiment, the parent Cas protein is Cas-sf2044.

[0021] Those skilled in the art will appreciate that protein structure can be altered without adversely affecting its activity and functionality. For example, one or more conservative amino acid substitutions can be introduced into a protein's amino acid sequence without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples and implementations of conservative amino acid substitutions will be apparent to those skilled in the art. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the substituted residue, i.e., a non-polar amino acid residue can be substituted for another non-polar amino acid residue, a polar uncharged amino acid residue can be substituted for another polar uncharged amino acid residue, a basic amino acid residue can be substituted for another basic amino acid residue, and an acidic amino acid residue can be substituted for another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, where one amino acid is replaced with another amino acid belonging to the same group, fall within the scope of the present invention, as long as the substitution does not inactivate the biological activity of the protein. Therefore, the proteins of the present invention may contain one or more conservative substitutions in their amino acid sequences, preferably generated by substitutions according to Table 1. Furthermore, the present invention also encompasses proteins containing one or more other non-conservative substitutions, as long as such non-conservative substitutions do not significantly affect the desired function and biological activity of the proteins of the present invention.

[0022] Conservative amino acid replacement can be carried out at one or more predicted non-essential amino acid residues.A "non-essential" amino acid residue is an amino acid residue that can be changed (deleted, substituted or replaced) without changing biological activity, while an "essential" amino acid residue is required for biological activity.A "conservative amino acid replacement" is a replacement in which an amino acid residue is replaced by an amino acid residue with a similar side chain.Amino acid replacement can be carried out in the non-conserved region of the above-mentioned Cas mutant protein.In general, such replacement is not carried out for conserved amino acid residues, or is not carried out for amino acid residues located within a conserved motif, where such residues are required for protein activity.However, it will be appreciated by those skilled in the art that functional variants can have less conservative or non-conservative changes in conserved regions.

[0023] Table 1

[0024] Initial residue Representative replacement Preferred substitutions Ala(A) Val; Leu; Ile Val Arg(R) Lys; Gln; Asn Lys Asn(N) Gln; His; Lys; Arg Gln Asp(D) Glu Glu Cys(C) Ser Ser Gln(Q) Asn Asn Glu(E) Asp Asp Gly(G) Pro; Ala Ala His(H) Asn; Gln; Lys; Arg Arg Ile(I) Leu; Val; Met; Ala; Phe Leu Leu(L) Ile; Val; Met; Ala; Phe Ile Lys(K) Arg; Gln; Asn Arg Met(M) Leu; Phe; Ile Leu Phe(F) Leu; Val; Ile; Ala; Tyr Leu Pro(P) Ala Ala Ser(S) Thr Thr Thr(T) Ser Ser Trp(W) Tyr; Phe Tyr Tyr(Y) Trp; Phe; Thr; Ser Phe Val(V) Ile;Leu;Met;Phe;Ala Leu

[0025] It is well known in the art that one or more amino acid residues can be changed (replaced, deleted, truncated or inserted) from the N and / or C terminus of a protein while still retaining its functional activity. Therefore, proteins in which one or more amino acid residues are changed from the N and / or C terminus of a Cas protein while retaining its desired functional activity are also within the scope of the present invention. These changes may include changes introduced by modern molecular methods such as PCR, which includes PCR amplification of a protein coding sequence by means of including an amino acid coding sequence among the oligonucleotides used in the PCR amplification.

[0026] It will be appreciated that proteins can be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such manipulations are generally known in the art. For example, amino acid sequence variants of the above-described proteins can be prepared by mutations in the DNA. Other forms of mutagenesis and / or directed evolution can also be accomplished, for example, using known mutagenesis, recombination, and / or shuffling methods, in combination with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.

[0027] It will be appreciated by those skilled in the art that these minor amino acid changes in the Cas proteins of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be changed, but the polypeptide may retain its activity. If the mutations present are not close to the catalytic domain, active site, or other functional domains, a smaller effect can be expected.

[0028] Those skilled in the art can identify the essential amino acids of the Cas mutant protein of the present invention according to methods known in the art, such as site-directed mutagenesis or protein evolution or analysis of bioinformatics systems. The catalytic domain, active site or other functional domain of the protein can also be determined by physical analysis of the structure, such as by the following techniques: such as nuclear magnetic resonance, crystallography, electron diffraction or photoaffinity labeling, combined with mutations of amino acids at putative key sites.

[0029] In the present invention, amino acid residues can be represented by single letters or three letters, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamine (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine ​​(Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), arginine (Arg, R).

[0030] The term "AxxB" indicates that the amino acid A at position xx is changed to amino acid B. For example, N160R indicates that the N at position 160 is mutated to R. When multiple amino acid positions are mutated simultaneously, a similar format like N160R-L163R can be used. For example, N160R-L163R indicates that the N at position 160 is mutated to R and the L at position 163 is mutated to R.

[0031] Specific amino acid positions (numbers) within the proteins of the present invention are determined by aligning the amino acid sequence of the target protein with SEQ ID No. 1 using standard sequence alignment tools, such as the Smith-Waterman algorithm or the CLUSTALW2 algorithm, wherein the sequences are considered aligned when the alignment score is the highest. The alignment score can be calculated according to the method described in Wilbur, WJ and Lipman, DJ (1983) Rapid similarity searches of nucleic acid and protein data banks. Proc. Natl. Acad. Sci. USA, 80: 726-730. The default parameters for the ClustalW2 (1.82) algorithm are preferably used: protein gap open penalty = 10.0; protein gap extension penalty = 0.2; protein matrix = Gonnet; protein / DNA end gap = -1; protein / DNAGAPDIST = 4. Preferably, the AlignX program (part of the vectorNTI group) is used to determine the position of specific amino acids in the protein of the present invention by aligning the amino acid sequence of the protein with SEQ ID No. 1 using default parameters suitable for multiple alignment (gap opening penalty: 10 log gap extension penalty 0.05).

[0032] In one embodiment, the amino acid sequence of the parent Cas protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to SEQ ID No. 1.

[0033] In some embodiments, the parent Cas protein is a natural wild-type Cas protein; in other embodiments, the parent Cas protein is an engineered Cas protein.

[0034] In one embodiment, the parent Cas protein is Cas-sf2044.

[0035] Cas proteins from a variety of organisms can be used as parent Cas proteins, and in some embodiments, the parent Cas proteins have nuclease activity. In some embodiments, the parent Cas protein is a nuclease, i.e., it cuts two chains of a target double-helical nucleic acid (e.g., double-helical DNA). In some embodiments, the parent Cas protein is a nickase, i.e., it cuts a single strand of a target double-helical nucleic acid (e.g., double-helical DNA).

[0036] In one embodiment, the Cas mutant protein is selected from any one of the following groups I-III:

[0037] I. A Cas mutant protein obtained by generating a mutation at any one or more of the following amino acid positions in the amino acid sequence shown in SEQ ID No. 1: position 160, position 163, position 674, position 694, position 697, position 335, position 602, position 644, position 171, position 254, position 337, position 482, position 514, position 517, position 528, position 538, position 700, position 866 or position 681;

[0038] II. Compared with the Cas mutant protein described in I, it has the mutation site described in I; and, compared with the Cas mutant protein described in I, it has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity;

[0039] III. Compared with the Cas mutant protein described in I, it has the mutation site described in I; and, compared with the Cas mutant protein described in I, it has a sequence of one or more amino acid substitutions, deletions or additions; the one or more amino acids include 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions.

[0040] The biological functions of the Cas protein include, but are not limited to, the activity of binding to the guide RNA, the endonuclease activity, the activity of binding to and cutting a specific site of the target sequence under the guidance of the guide RNA, including but not limited to Cis cleavage activity and Trans cleavage activity.

[0041] In the present invention, "Cas mutant protein" can also be referred to as a mutated Cas protein, or a Cas protein variant.

[0042] The present invention also provides a fusion protein, which includes the Cas mutant protein as described above and other modified parts.

[0043] In one embodiment, the modifying moiety is selected from another protein or polypeptide, a detectable label, or any combination thereof.

[0044] In one embodiment, the modifying portion is selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting portion, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), and a domain having an activity selected from the following: nucleotide deaminase, methylase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity and nucleic acid binding activity; and any combination thereof. The NLS sequence is well known to those skilled in the art, and examples thereof include, but are not limited to, the SV40 large T antigen, EGL-13, c-Myc and TUS protein.

[0045] In one embodiment, the NLS sequence is located at, near, or proximal to a terminus (e.g., the N-terminus, the C-terminus, or both) of the Cas protein of the invention.

[0046] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can select other suitable epitope tags (for example, purification, detection or tracing).

[0047] The reporter gene sequence is well known to those skilled in the art, and examples thereof include but are not limited to GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.

[0048] In one embodiment, the fusion protein of the present invention comprises a domain capable of binding to a DNA molecule or an intracellular molecule, such as maltose binding protein (MBP), the DNA binding domain (DBD) of Lex A, the DBD of GAL4, and the like.

[0049] In one embodiment, the fusion protein of the invention comprises a detectable label, such as a fluorescent dye, eg, FITC or DAPI.

[0050] In one embodiment, the Cas protein of the present invention is optionally coupled, conjugated or fused to the modifying portion via a linker.

[0051] In one embodiment, the modification portion is directly linked to the N-terminus or C-terminus of the Cas protein of the present invention.

[0052] In one embodiment, the modified portion is connected to the N-terminus or C-terminus of the Cas protein of the present invention via a linker. Such linkers are well known in the art, and examples thereof include but are not limited to linkers comprising one or more (e.g., 1, 2, 3, 4 or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA or Ava), or PEG, etc.

[0053] The Cas protein, protein derivative or fusion protein of the present invention is not limited by the method of its production. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.

[0054] Cas protein nucleic acid

[0055] In another aspect, the present invention provides an isolated polynucleotide comprising:

[0056] (a) a polynucleotide sequence encoding a Cas mutant protein or fusion protein of the present invention;

[0057] Alternatively, a polynucleotide complementary to the polynucleotide described in (a).

[0058] In one embodiment, the nucleotide sequence is codon optimized for expression in prokaryotes. In one embodiment, the nucleotide sequence is codon optimized for expression in eukaryotic cells.

[0059] In one embodiment, the cell is an animal cell, eg, a mammalian cell.

[0060] In one embodiment, the cell is a human cell.

[0061] In one embodiment, the cell is a plant cell, such as a cell from a cultivated plant (such as cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.

[0062] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.

[0063] Guide RNA (gRNA)

[0064] On the other hand, the present invention provides a gRNA, which includes a first segment and a second segment; the first segment is also called a "skeleton region", "protein binding segment", "protein binding sequence", or "direct repeat (DirectRepeat) sequence"; the second segment is also called a "targeting sequence for targeting nucleic acid" or "targeting segment for targeting nucleic acid", or "guide sequence for targeting target sequence".

[0065] The first segment of the gRNA is capable of interacting with the Cas protein of the present invention, thereby forming a complex between the Cas protein and the gRNA.

[0066] In a preferred embodiment, the first segment is a direct repeat sequence as described above.

[0067] In a preferred embodiment, the direct repeat sequence of the Cas protein is AGUGCAAUAGUUACAGAAUAGUAAUUAUAUUCGCA.

[0068] The targeting sequence of the targeting nucleic acid of the present invention or the targeting section of the targeting nucleic acid comprises a nucleotide sequence complementary to the sequence in the target nucleic acid. In other words, the targeting sequence of the targeting nucleic acid of the present invention or the targeting section of the targeting nucleic acid interacts with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the targeting sequence of the targeting nucleic acid or the targeting section of the targeting nucleic acid can be changed, or can be modified to hybridize any desired sequence in the target nucleic acid. The nucleic acid is selected from DNA or RNA.

[0069] The percent complementarity between the targeting sequence of a targeting nucleic acid or the targeting segment of a targeting nucleic acid and the target sequence of a target nucleic acid can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).

[0070] The "backbone region", "protein binding segment", "protein binding sequence", or "direct repeat sequence" of the gRNA of the present invention can interact with the CRISPR protein (or Cas protein). The gRNA of the present invention guides the interacting Cas protein to a specific nucleotide sequence within the target nucleic acid through the action of the targeting sequence of the target nucleic acid.

[0071] Preferably, the guide RNA comprises a first segment and a second segment from the 5' to the 3' direction.

[0072] In the present invention, the second segment can also be understood as a guide sequence that hybridizes with the target sequence.

[0073] The gRNA of the present invention is capable of forming a complex with the Cas protein.

[0074] carrier

[0075] The present invention also provides a vector comprising the Cas mutant protein, isolated nucleic acid molecule or polynucleotide as described above; preferably, it further comprises a regulatory element operably linked thereto.

[0076] In one embodiment, the regulatory element is selected from one or more of the following groups: enhancer, transposon, promoter, terminator, leader sequence, polyadenylation sequence, marker gene.

[0077] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.

[0078] In some embodiments, the vector included in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a phage, etc., which are well known to those skilled in the art.

[0079] CRISPR system

[0080] The present invention provides an engineered non-naturally occurring vector system, or a CRISPR-Cas system, which includes a Cas mutant protein or a nucleic acid sequence encoding the Cas mutant protein and a nucleic acid encoding one or more guide RNAs.

[0081] In one embodiment, the nucleic acid sequence encoding the Cas mutant protein and the nucleic acid encoding one or more guide RNAs are artificially synthesized.

[0082] In one embodiment, the nucleic acid sequence encoding the Cas mutant protein and the nucleic acid encoding one or more guide RNAs do not naturally co-occur.

[0083] The one or more guide RNAs target one or more target sequences in the cell. The one or more target sequences hybridize to the genomic loci of the DNA molecules encoding the one or more gene products, and guide the Cas protein to the genomic loci of the DNA molecules encoding the one or more gene products. After the Cas protein reaches the target sequence position, it modifies, edits or cuts the target sequence, thereby changing or modifying the expression of the one or more gene products.

[0084] The cells of the present invention include one or more of animals, plants, or microorganisms.

[0085] In some embodiments, the Cas protein is codon-optimized for expression in a cell.

[0086] In some embodiments, the Cas protein directs cleavage of one or both strands at the location of the target sequence.

[0087] The present invention also provides an engineered non-naturally occurring vector system, which may include one or more vectors, wherein the one or more vectors include:

[0088] a) a first regulatory element, which is operably linked to the gRNA,

[0089] b) a second regulatory element, which is operably linked to the Cas protein;

[0090] Components (a) and (b) are located on the same or different carriers of the system.

[0091] The first and second regulatory elements include a promoter (e.g., a constitutive promoter or an inducible promoter), an enhancer (e.g., a 35S promoter or a 35S enhanced promoter), an internal ribosome entry site (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences).

[0092] In some embodiments, the vector in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a phage, etc., which are well known to those skilled in the art.

[0093] In some embodiments, the systems provided herein are in a delivery system. In some embodiments, the delivery system is a nanoparticle, a liposome, an exosome, a microbubble, and a gene gun.

[0094] In one embodiment, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell. In one embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0095] In one embodiment, the target sequence is present in a cell. In one embodiment, the target sequence is present in the nucleus or in the cytoplasm (e.g., an organelle). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.

[0096] In one embodiment, the Cas protein is connected to one or more NLS sequences. In one embodiment, the fusion protein comprises one or more NLS sequences. In one embodiment, the NLS sequence is connected to the N-terminus or C-terminus of the protein. In one embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the protein.

[0097] On the other hand, the present invention relates to an engineered CRISPR system, comprising the above-mentioned Cas protein and one or more guide RNAs, wherein the guide RNA includes a direct repeat sequence and a spacer sequence capable of hybridizing with a target nucleic acid, and the Cas protein is capable of binding to the guide RNA and targeting a target nucleic acid sequence complementary to the spacer sequence.

[0098] Protein-nucleic acid complexes / compositions

[0099] In another aspect, the present invention provides a compound or composition comprising:

[0100] (i) a protein component selected from the group consisting of: the above-mentioned Cas proteins, derivatized proteins or fusion proteins, and any combination thereof; and

[0101] (ii) a nucleic acid component comprising (a) a guide sequence capable of hybridizing to a target sequence; and (b) a direct repeat sequence capable of binding to a Cas protein of the present invention.

[0102] The protein component and the nucleic acid component combine with each other to form a complex.

[0103] In one embodiment, the nucleic acid component is a guide RNA in a CRISPR-Cas system.

[0104] In one embodiment, the complex or composition is non-naturally occurring or modified. In one embodiment, at least one component of the complex or composition is non-naturally occurring or modified. In one embodiment, the first component is non-naturally occurring or modified; and / or the second component is non-naturally occurring or modified.

[0105] Activated CRISPR complex

[0106] On the other hand, the present invention also provides an activated CRISPR complex, the activated CRISPR complex comprising: (1) a protein component selected from: a Cas protein, a derivatized protein or a fusion protein of the present invention, and any combination thereof; (2) a gRNA comprising (a) a guide sequence capable of hybridizing with a target sequence; and (b) a direct repeat sequence capable of binding to the Cas protein of the present invention; and (3) a target sequence bound to the gRNA. Preferably, the binding is carried out by binding of the targeting sequence of the targeting nucleic acid on the gRNA to the target nucleic acid.

[0107] As used herein, the terms "activated CRISPR complex," "activated complex," or "ternary complex" refer to the complex formed after the Cas protein, gRNA, and target nucleic acid in the CRISPR system are bound or modified.

[0108] The Cas protein and gRNA of the present invention can form a binary complex that is activated when bound to a nucleic acid substrate to form an activated CRISPR complex. The nucleic acid substrate is complementary to the spacer sequence in the gRNA (or referred to as a guide sequence that hybridizes with the target nucleic acid). In some embodiments, the spacer sequence of the gRNA fully matches the target substrate. In other embodiments, the spacer sequence of the gRNA matches a portion (continuous or discontinuous) of the target substrate.

[0109] In a preferred embodiment, the activated CRISPR complex can exhibit collateral nuclease cleavage activity, which refers to the non-specific cleavage activity or random cleavage activity of the activated CRISPR complex on single-stranded nucleic acids, also known as trans cleavage activity in the art.

[0110] Delivery and delivery compositions

[0111] The Cas proteins, gRNAs, fusion proteins, nucleic acid molecules, vectors, systems, complexes and compositions of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipofection, nucleofection, microinjection, sonoporation, gene guns, calcium phosphate-mediated transfection, cationic transfection, lipofection, dendritic transfection, heat shock transfection, nucleofection, magnetofection, lipofection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial virions, etc.

[0112] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: the Cas protein, fusion protein, nucleic acid molecule, vector, system, complex and composition of the present invention.

[0113] In one embodiment, the delivery vehicle is a particle.

[0114] In one embodiment, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).

[0115] host cells

[0116] The present invention also relates to an in vitro, ex vivo or in vivo cell or cell line or their progeny, wherein the cell or cell line or their progeny comprises: the Cas protein, fusion protein, nucleic acid molecule, protein-nucleic acid complex, activated CRISPR complex, vector, and delivery composition of the present invention.

[0117] In certain embodiments, the cell is a prokaryotic cell.

[0118] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, such as a cell of a non-human primate, cattle, sheep, pig, dog, monkey, rabbit, rodent (such as rat or mouse). In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as a cell of poultry (such as chicken), fish or crustacean (such as clams, shrimp). In certain embodiments, the cell is a plant cell, such as a cell or cultivated plant or food crop such as cassava, corn, sorghum, soybean, wheat, oat or rice that a monocot or dicot has, such as an algae, tree or production plant, fruit or vegetable (for example, trees such as citrus trees, nut trees; Solanaceae, cotton, tobacco, tomato, grape, coffee, cocoa, etc.).

[0119] In certain embodiments, the cell is a stem cell or a stem cell line.

[0120] In certain cases, the host cells of the invention comprise genetic or genomic modifications that are not present in their wild-type form.

[0121] Gene Editing Methods and Applications

[0122] The Cas mutant protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition, activated CRISPR complex, or host cell of the present invention can be used for any one or more of the following purposes: targeting and / or editing target nucleic acid; cutting double-stranded DNA, single-stranded DNA, or single-stranded RNA; non-specific cutting and / or degradation of side branch nucleic acid; non-specific cutting of single-stranded nucleic acid; nucleic acid detection; detection of nucleic acid in target sample; specific editing of double-stranded nucleic acid; base editing of double-stranded nucleic acid; base editing of single-stranded nucleic acid. In other embodiments, it can also be used to prepare reagents or kits for any one or more of the above purposes.

[0123] The present invention also provides the use of the above-mentioned Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition or activated CRISPR complex in gene editing, gene targeting or gene cleavage; or, use in the preparation of reagents or kits for gene editing, gene targeting or gene cleavage.

[0124] In one embodiment, the gene editing, gene targeting or gene cleavage is performed inside and / or outside the cell.

[0125] The present invention also provides a method for editing a target nucleic acid, targeting a target nucleic acid, or cutting a target nucleic acid, the method comprising contacting the target nucleic acid with the above-mentioned Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition, or activated CRISPR complex. In one embodiment, the method is to edit, target, or cut a target nucleic acid inside or outside a cell.

[0126] The gene editing or editing of target nucleic acid includes modifying genes, knocking out genes, changing the expression of gene products, repairing mutations, and / or inserting polynucleotides, gene mutations.

[0127] The editing can be performed in prokaryotic cells and / or eukaryotic cells.

[0128] On the other hand, the present invention also provides the use of the above-mentioned Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition or activated CRISPR complex in nucleic acid detection, or in the preparation of reagents or kits for nucleic acid detection.

[0129] On the other hand, the present invention also provides a method for cutting single-stranded nucleic acids, the method comprising contacting a nucleic acid population with the above-mentioned Cas protein and gRNA, wherein the nucleic acid population comprises a target nucleic acid and a plurality of non-target single-stranded nucleic acids, and the Cas protein cuts the plurality of non-target single-stranded nucleic acids.

[0130] The gRNA is capable of binding to the Cas protein.

[0131] The gRNA is capable of targeting the target nucleic acid.

[0132] The contacting can be inside the cell in vitro, ex vivo or in vivo.

[0133] Preferably, the cleavage of the single-stranded nucleic acid is non-specific cleavage.

[0134] On the other hand, the present invention also provides the use of the above-mentioned Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition or activated CRISPR complex in non-specific cleavage of single-stranded nucleic acid, or in the preparation of a reagent or kit for non-specific cleavage of single-stranded nucleic acid.

[0135] On the other hand, the present invention also provides a kit for gene editing, gene targeting or gene cleavage, which comprises the above-mentioned Cas protein, gRNA, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition, the above-mentioned activated CRISPR complex or the above-mentioned host cell.

[0136] On the other hand, the present invention also provides a kit for detecting a target nucleic acid in a sample, the kit comprising: (a) a Cas protein, or a nucleic acid encoding the Cas protein; (b) a guide RNA, or a nucleic acid encoding the guide RNA, or a precursor RNA comprising the guide RNA, or a nucleic acid encoding the precursor RNA; and (c) a single-stranded nucleic acid detector that is single-stranded and does not hybridize with the guide RNA.

[0137] It is known in the art that precursor RNA can be cleaved or processed into the mature guide RNA described above.

[0138] In another aspect, the invention provides the use of the above-mentioned Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition, activated CRISPR complex or host cell in preparing a preparation or kit, wherein the preparation or kit is used for:

[0139] (i) gene or genome editing;

[0140] (ii) target nucleic acid detection and / or diagnosis;

[0141] (iii) editing a target sequence in a target locus to modify an organism or non-human organism;

[0142] (iv) treatment of disease;

[0143] (iv) Targeting target genes.

[0144] Preferably, the above-mentioned gene or genome editing is performed inside or outside the cell.

[0145] Preferably, the target nucleic acid detection and / or diagnosis is performed in vitro.

[0146] Preferably, the treatment of the disease is the treatment of a condition caused by a defect in the target sequence in the target locus.

[0147] In another aspect, the present invention provides a method for detecting a target nucleic acid in a sample, the method comprising contacting the sample with the Cas protein, gRNA (guide RNA) and a single-stranded nucleic acid detector, the gRNA comprising a region that binds to the Cas protein and a guide sequence that hybridizes with the target nucleic acid; detecting a detectable signal generated by the Cas protein cleaving the single-stranded nucleic acid detector, thereby detecting the target nucleic acid; the single-stranded nucleic acid detector does not hybridize with the gRNA.

[0148] Method for specifically modifying target nucleic acid

[0149] On the other hand, the present invention also provides a method for specifically modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition or the above-mentioned activated CRISPR complex.

[0150] The specific modification can occur in vivo or in vitro.

[0151] The specific modification can occur inside or outside the cell.

[0152] In some cases, the cell is selected from a prokaryotic cell or a eukaryotic cell, eg, an animal cell, a plant cell, or a microbial cell.

[0153] In one embodiment, the modification refers to a break in the target sequence, such as a single-strand / double-strand break in DNA, or a single-strand break in RNA.

[0154] In some cases, the method further comprises contacting the target nucleic acid with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target nucleic acid.

[0155] In one embodiment, the modification further comprises inserting an editing template (eg, an exogenous nucleic acid) into the break.

[0156] In one embodiment, the method further comprises: contacting the target nucleic acid with an editing template, or delivering the editing template to a cell containing the target nucleic acid. In this embodiment, the method repairs the broken target gene by homologous recombination with an exogenous template polynucleotide; in some embodiments, the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides in the target gene; in other embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene containing the target sequence.

[0157] Detection (non-specific cleavage)

[0158] On the other hand, the present invention provides a method for detecting a target nucleic acid in a sample, the method comprising contacting the sample with the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition or the above-mentioned activated CRISPR complex and a single-stranded nucleic acid detector; detecting a detectable signal generated by the Cas protein cleaving the single-stranded nucleic acid detector, thereby detecting the target nucleic acid.

[0159] In the present invention, the target nucleic acid includes ribonucleotides or deoxyribonucleotides; including single-stranded nucleic acids and double-stranded nucleic acids, such as single-stranded DNA, double-stranded DNA, single-stranded RNA, and double-stranded RNA.

[0160] In one embodiment, the target nucleic acid is derived from a sample such as a virus, bacteria, microorganism, soil, water, human body, animal, plant, etc. Preferably, the target nucleic acid is a product enriched or amplified by methods such as PCR, NASBA, RPA, SDA, LAMP, HAD, NEAR, MDA, RCA, LCR, RAM, etc.

[0161] In one embodiment, the target nucleic acid is a viral nucleic acid, a bacterial nucleic acid, a specific nucleic acid associated with a disease, such as a specific mutation site or SNP site or a nucleic acid that differs from a control; preferably, the virus is a plant virus or an animal virus, for example, a papillomavirus, a hepadnavirus, a herpes virus, adenovirus, a poxvirus, a parvovirus, a coronavirus; preferably, the virus is a coronavirus, preferably, SARS, SARS-CoV2 (COVID-19), HCoV-229E, HCoV-OC43, HCoV-NL63, HCoV-HKU1, Mers-Cov.

[0162] In the present invention, the gRNA has at least 50% matching with the target sequence on the target nucleic acid, preferably at least 60%, preferably at least 70%, preferably at least 80%, preferably at least 90%.

[0163] In one embodiment, when the target sequence contains one or more characteristic sites (such as a specific mutation site or SNP), the characteristic sites are completely matched with the gRNA.

[0164] In one embodiment, the detection method may comprise one or more gRNAs with different guide sequences, which target different target sequences.

[0165] In the present invention, the single-stranded nucleic acid detector includes but is not limited to single-stranded DNA, single-stranded RNA, DNA-RNA hybrids, nucleic acid analogs, base modifiers, and single-stranded nucleic acid detectors containing a base-free spacer, etc.; "nucleic acid analogs" include but are not limited to: locked nucleic acid, bridge nucleic acid, morpholino nucleic acid, ethylene glycol nucleic acid, hexitol nucleic acid, threose nucleic acid, arabinose nucleic acid, 2'oxymethyl RNA, 2'methoxyacetyl RNA, 2'fluoro RNA, 2'amino RNA, 4'thio RNA and combinations thereof, including optional ribonucleotides or deoxyribonucleotide residues.

[0166] In the present invention, the detectable signal is achieved by the following means: vision-based detection, sensor-based detection, color detection, fluorescence signal-based detection, gold nanoparticle-based detection, fluorescence polarization, colloidal phase transition / dispersion, electrochemical detection and semiconductor-based detection.

[0167] In the present invention, preferably, a fluorescent group and a quenching group are provided at both ends of the single-stranded nucleic acid detector, respectively, so that when the single-stranded nucleic acid detector is cleaved, a detectable fluorescent signal can be exhibited. The fluorescent group is selected from one or any combination of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red, or LC RED460; and the quenching group is selected from one or any combination of BHQ1, BHQ2, BHQ3, Dabcy1, or Tamra.

[0168] In other embodiments, different labeling molecules are respectively set at the 5' end and the 3' end of the single-stranded nucleic acid detector, and the colloidal gold test results of the single-stranded nucleic acid detector before and after being cut by the Cas protein are detected by colloidal gold detection; the single-stranded nucleic acid detector will show different color development results on the colloidal gold detection line and the quality control line before and after being cut by the Cas protein.

[0169] In some embodiments, the method of detecting a target nucleic acid may further include comparing the level of the detectable signal to a reference signal level, and determining the amount of the target nucleic acid in the sample based on the level of the detectable signal.

[0170] In some embodiments, the method of detecting a target nucleic acid can also include using RNA reporter nucleic acid and DNA reporter nucleic acid on different channels (e.g., fluorescent colors), and determining the level of detectable signal by measuring the signal levels of the RNA and DNA reporter molecules, and by measuring the amount of target nucleic acid in the RNA and DNA reporter molecules, and sampling based on the level of the combined (e.g., using a minimum or product) detectable signal.

[0171] In one embodiment, the target gene is present in a cell.

[0172] In one embodiment, the cell is a prokaryotic cell.

[0173] In one embodiment, the cell is a eukaryotic cell.

[0174] In one embodiment, the cell is an animal cell.

[0175] In one embodiment, the cell is a human cell.

[0176] In one embodiment, the cell is a plant cell, such as a cell from a cultivated plant (such as cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.

[0177] In one embodiment, the target gene is present in a nucleic acid molecule (eg, a plasmid) in vitro.

[0178] In one embodiment, the target gene is present in a plasmid.

[0179] Definition of terms

[0180] Unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, procedures in molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA used herein are conventional procedures widely used in the relevant fields. To facilitate a better understanding of the present invention, definitions and explanations of relevant terms are provided below.

[0181] Nucleic acid cleavage or nucleic acid cleavage herein includes: DNA or RNA breakage (Cis cleavage) in the target nucleic acid produced by the Cas enzyme described herein, DNA or RNA breakage in the side branch nucleic acid substrate (single-stranded nucleic acid substrate) (i.e., non-specific or non-targeted, Trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.

[0182] CRISPR system

[0183] As used herein, the terms "clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements associated with the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of directing the activity of the Cas genes.

[0184] CRISPR / Cas complex

[0185] As used herein, the term "CRISPR / Cas complex" refers to a complex formed by the binding of guide RNA or mature crRNA to the Cas protein, which comprises a direct repeat sequence that hybridizes to the guide sequence of the target sequence and binds to the Cas protein, and the complex is capable of recognizing and cleaving a polynucleotide that can hybridize to the guide RNA or mature crRNA.

[0186] Guide RNA (gRNA)

[0187] As used herein, the terms "guide RNA (gRNA)", "mature crRNA", and "guide sequence" are used interchangeably and have meanings generally understood by those skilled in the art. Generally speaking, a guide RNA can comprise a direct repeat sequence and a guide sequence, or consist essentially of or consist of a direct repeat sequence and a guide sequence.

[0188] In some cases, a guide sequence is any polynucleotide sequence that has sufficient complementarity to a target sequence to hybridize with the target sequence and guide specific binding of the CRISPR / Cas complex to the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between a guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining optimal alignment is within the capabilities of one of ordinary skill in the art. For example, there are publicly available and commercially available alignment algorithms and programs such as, but not limited to, ClustalW, Smith-Waterman in matlab, Bowtie, Geneious, Biopython, and SeqMan.

[0189] Target sequence

[0190] "Target sequence" refers to a polynucleotide targeted by a guide sequence in a gRNA, such as a sequence having complementarity with the guide sequence, wherein hybridization between the target sequence and the guide sequence promotes the formation of a CRISPR / Cas complex (including Cas protein and gRNA). Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex.

[0191] The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the nucleus or cytoplasm of the cell. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast. A sequence or template that can be used to recombine into a target locus comprising the target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence". In one embodiment, the editing template is an exogenous nucleic acid. In one embodiment, the recombination is homologous recombination.

[0192] In the present invention, a "target sequence" or "target polynucleotide" or "target nucleic acid" can be any polynucleotide that is endogenous or exogenous to a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide that is present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).

[0193] Single-stranded nucleic acid detector

[0194] The single-stranded nucleic acid detector of the present invention refers to a sequence containing 2-200 nucleotides, preferably 2-150 nucleotides, preferably 3-100 nucleotides, preferably 3-30 nucleotides, preferably 4-20 nucleotides, and more preferably 5-15 nucleotides. It is preferably a single-stranded DNA molecule, a single-stranded RNA molecule, or a single-stranded DNA-RNA hybrid.

[0195] The single-stranded nucleic acid detector includes different reporter groups or labeling molecules at both ends. When it is in the initial state (i.e., uncleaved state), it does not present a reporter signal. When the single-stranded nucleic acid detector is cleaved, it presents a detectable signal, i.e., a detectable difference is shown after cleavage and before cleavage.

[0196] In one embodiment, the reporter group or labeling molecule includes a fluorescent group and a quencher group, wherein the fluorescent group is selected from one or any combination of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red or LC RED460; and the quencher group is selected from one or any combination of BHQ1, BHQ2, BHQ3, Dabcy1 or Tamra.

[0197] In one embodiment, the single-stranded nucleic acid detector has a first molecule (such as FAM or FITC) connected to the 5' end and a second molecule (such as biotin) connected to the 3' end. The reaction system containing the single-stranded nucleic acid detector is used in conjunction with a flow strip to detect target nucleic acid (preferably, colloidal gold detection method). The flow strip is designed to have two capture lines, with an antibody that binds to the first molecule (i.e., the first molecule antibody) at the sample contact end (colloidal gold), an antibody that binds to the first molecule antibody at the first line (control line), and an antibody that binds to the second molecule (i.e., the second molecule antibody, such as avidin) at the second line (test line). When the reaction flows along the strip, the first molecule antibody binds to the first molecule and carries the cut or uncut oligonucleotide to the capture line. The cut reporter will bind to the antibody of the first molecule antibody at the first capture line, and the uncut reporter will bind to the second molecule antibody at the second capture line. The binding of the reporter group to each line will result in a strong readout / signal (e.g., color). As more reporters are cut, more signals will accumulate at the first capture line, and less signals will appear at the second line. In certain aspects, the present invention relates to the use of a flow strip as described herein for detecting nucleic acids. In certain aspects, the present invention relates to a method for detecting nucleic acids using a flow strip as defined herein, such as a (lateral) flow test or a (lateral) flow immunochromatographic assay. In certain aspects, the molecules in the single-stranded nucleic acid detector may be interchangeable or their positions may be altered, and any such modifications are encompassed by the present invention as long as the reporting principle is the same or similar to that of the present invention.

[0198] The detection method of the present invention can be used for quantitative detection of target nucleic acids. The quantitative detection index can be quantified based on the signal strength of the reporter group, such as the luminescence intensity of the fluorescent group, or the width of the color band.

[0199] wild type

[0200] As used herein, the term "wild type" has the meaning generally understood by those skilled in the art to refer to the typical form of an organism, strain, gene, or characteristic as it exists in nature, as distinguished from mutant or variant forms, which can be isolated from a source in nature and has not been intentionally modified by man.

[0201] Derivatization

[0202] As used herein, the term "derivatization" refers to the chemical modification of an amino acid, polypeptide, or protein wherein one or more substituents have been covalently attached to the amino acid, polypeptide, or protein. The substituents may also be referred to as side chains.

[0203] A derivatized protein is a derivative of the protein. Generally, the derivatization of the protein does not adversely affect the desired activity of the protein (e.g., the activity of binding to the guide RNA, the endonuclease activity, the activity of binding to and cutting a specific site of the target sequence under the guidance of the guide RNA), that is, the derivative of the protein has the same activity as the protein.

[0204] Derivatized proteins

[0205] Also known as "protein derivatives", refers to modified forms of proteins, for example, wherein one or more amino acids of the protein may be deleted, inserted, modified and / or substituted.

[0206] Non-naturally occurring

[0207] As used herein, the terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of human effort. When these terms are used to describe a nucleic acid molecule or polypeptide, they indicate that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is associated in nature or as found in nature.

[0208] Orthologue (ortholog)

[0209] As used herein, the term "orthologue" has the meaning generally understood by those skilled in the art. As a further guide, an "orthologue" of a protein as described herein refers to a protein belonging to a different species that performs the same or a similar function as the protein of which it is an orthologue.

[0210] Identity

[0211] As used herein, the term "identity" is used to refer to the match of sequence between two polypeptides or between two nucleic acids. When a position in each of two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percentage of identity" between two sequences is a function of the number of matching positions shared by the sequences divided by the number of positions compared x 100. For example, if 6 of 10 positions in two sequences are matched then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (3 of 6 positions are matched). Typically, the comparison is made over the length of the sequences being compared, after aligning the two sequences to produce maximum identity. Such alignment can be achieved by methods such as that of Needleman et al. (1970) J. Mol. Biol. 48:443-453, which can be conveniently performed by computer programs such as the Align program (DNAstar, Inc.). Percentage identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl. Biosci., 4:11-17 (1988)) as integrated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. In addition, percentage identity between two amino acid sequences can be determined using the algorithm of Needleman and Wunsch (J MoI Biol. 48:444-453 (1970)) as implemented in the GAP program, using either a Blossum 62 matrix or a PAM250 matrix, and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6.

[0212] carrier

[0213] The term "vector" refers to a nucleic acid molecule that is capable of transporting another nucleic acid molecule to which it is attached. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends, or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other various polynucleotides known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection so that the genetic material elements it carries are expressed in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. In addition, the vector may also contain a replication initiation site.

[0214] One type of vector is a "plasmid," which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, eg, by standard molecular cloning techniques.

[0215] Another type of vector is a viral vector, in which a virally derived DNA or RNA sequence is present in a vector for packaging a virus (e.g., a retrovirus, a replication-defective retrovirus, adenovirus, a replication-defective adenovirus, and adeno-associated virus). The viral vector also comprises a polynucleotide carried by a virus for transfection into a host cell. Some vectors (e.g., bacterial vectors and episomal mammalian vectors with a bacterial origin of replication) can replicate autonomously in the host cell into which they are introduced.

[0216] Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell upon introduction into the host cell and are thereby replicated along with the host genome. Furthermore, some vectors are capable of directing the expression of genes to which they are operably linked. Such vectors are referred to herein as "expression vectors."

[0217] host cells

[0218] As used herein, the term "host cell" refers to cells that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.

[0219] Those skilled in the art will appreciate that the design of the expression vector may depend on factors such as the choice of the host cell to be transformed, the level of expression desired, and the like.

[0220] Regulatory elements

[0221] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, CA (1990). In some cases, regulatory elements include those that direct the constitutive expression of a nucleotide sequence in many types of host cells and those that direct the expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or special cell types (e.g., lymphocytes). In some cases, regulatory elements can also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell-type-specific. In some cases, the term "regulatory element" encompasses enhancer elements such as WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I ((Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).

[0222] promoter

[0223] As used herein, the term "promoter" has a meaning well known to those skilled in the art and refers to a non-coding nucleotide sequence located upstream of a gene that can initiate expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to the promoter.

[0224] NLS

[0225] A "nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" a protein for import into the cell nucleus via nuclear transport, i.e., proteins with an NLS are transported to the cell nucleus. Typically, an NLS comprises a positively charged Lys or Arg residue exposed on the surface of the protein. Exemplary nuclear localization sequences include, but are not limited to, NLSs from the following: SV40 large T antigen, EGL-13, c-Myc, and TUS protein. In some embodiments, the NLS comprises the PKKKRKV sequence. In some embodiments, the NLS comprises the AVKRPAATKKAGQAKKKKLD sequence. In some embodiments, the NLS comprises the PAAKRVKLD sequence. In some embodiments, the NLS comprises the MSRRRKANPTKLSENAKKLAKEVEN sequence. In some embodiments, the NLS comprises the KLKIKRPVK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the sequence KIPIK in the yeast transcription repressor Matα2, and PY-NLS.

[0226] operably connected

[0227] As used herein, the term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to the one or more regulatory elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0228] Complementarity

[0229] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence by means of traditional Watson-Crick or other non-traditional types. Percent complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues of a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, "substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.

[0230] Strict conditions

[0231] As used herein, "stringent conditions" for hybridization refer to conditions under which a nucleic acid having complementarity with a target sequence predominantly hybridizes to the target sequence and does not substantially hybridize to non-target sequences. Stringent conditions are generally sequence-dependent and vary depending on many factors. Generally speaking, the longer the sequence, the higher the temperature at which the sequence specifically hybridizes to its target sequence.

[0232] hybridization

[0233] The terms "hybridize" or "complementary" or "substantially complementary" refer to a nucleic acid (e.g., RNA, DNA) comprising a nucleotide sequence that enables it to non-covalently bind, i.e., form base pairs and / or G / U base pairs, "anneal" or "hybridize" with another nucleic acid in a sequence-specific, antiparallel manner (i.e., a nucleic acid specifically binds to a complementary nucleic acid).

[0234] Hybridization requires that the two nucleic acids contain complementary sequences, although there may be mismatches between the bases. Suitable conditions for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, which are variables well known in the art. Typically, the length of a hybridizable nucleic acid is 8 nucleotides or more (e.g., 10 nucleotides or more, 12 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 22 nucleotides or more, 25 nucleotides or more, or 30 nucleotides or more).

[0235] It is understood that the sequence of a polynucleotide need not be 100% complementary to the sequence of its target nucleic acid to hybridize specifically. A polynucleotide may comprise 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% sequence complementarity to the target region in the target nucleic acid sequence with which it hybridizes.

[0236] The hybridization of the target sequence and the gRNA represents that at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or represents that at least 12, 15, 16, 17, 18, 19, 20, 21, 22 or more bases of the nucleic acid sequences of the target sequence and the gRNA can complement each other and hybridize to form a complex.

[0237] Express

[0238] As used herein, the term "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.

[0239] connector

[0240] As used herein, the term "linker" refers to a linear polypeptide formed by connecting multiple amino acid residues via peptide bonds. The linker of the present invention can be an artificially synthesized amino acid sequence, or a naturally occurring polypeptide sequence, such as a polypeptide having a hinge region function. Such linker polypeptides are well known in the art (see, for example, Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90: 6444-6448; Poljak, RJ et al. (1994) Structure 2: 1 121-1123).

[0241] treat

[0242] As used herein, the term "treat" refers to treating or curing a disorder, delaying the onset of symptoms of a disorder, and / or delaying the progression of a disorder.

[0243] Subjects

[0244] As used herein, the term "subject" includes, but is not limited to, various animals, plants, and microorganisms.

[0245] animal

[0246] For example, mammals, such as bovines, equines, ovines, porcines, canines, felines, lagomorphs, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., a human) has a disorder (e.g., a disorder caused by a disease-associated gene defect).

[0247] plant

[0248] The term "plant" is to be understood as any differentiated multicellular organism capable of performing photosynthesis, including crop plants at any stage of maturation or development, in particular monocotyledonous or dicotyledonous plants, vegetable crops, including artichokes, Brussels sprouts, cress, leeks, asparagus, lettuce (e.g. head lettuce, leaf lettuce, long-leaf lettuce), bok choy, yellow fleshed yam, melons (e.g. muskmelons, watermelons, crenshaws, cantaloupes, honeydews), oilseed crops (e.g. turnips, cabbages, cauliflowers, broccoli, collards, kale, Chinese cabbage, bok choy), artichokes, radishes, napa, okra, onions, celery, parsley, chickpeas, parsnips, chicory, peppers, potatoes, gourds (e.g. zucchini, cucumbers, crookneck squash, calabaza, pumpkin), radishes, dry bulb onions, rutabagas, eggplants (also known as aubergines), burdock, endive, green onions, endive, garlic, spinach, green onions, squash, greens, sugar beets (sugar beet and fodder beet), sweet potatoes, Swiss chard, wasabi, tomatoes, turnips, and spices; fruits and / or vine crops, such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, aronia, almonds, chestnuts, hazelnuts, pecans, pistachios, walnuts, citrus, blueberries, boysenberries, cranberries, currants, goji berries, raspberries, strawberries, blackberries, grapes, avocados, bananas, kiwis, persimmons, pomegranates, pineapples, tropical fruits, pomes, melons, mangoes, papayas, and lychees; field crops, such as clover, alfalfa, milkweed, meadow grass, corn / maize (fodder corn, sweet corn, popcorn), hops, jojoba, peanuts, rice, safflower, small grain crops (barley, oats, rye, wheat, etc.), sorghum, tobacco, kapok, legumes (beans, lentils, peas, soybeans), oil plants (rape, mustard, poppy, olive, sunflower, coconut, castor oil plant, cocoa, groundnuts), Arabidopsis, fiber plants (cotton, flax, jute), lauraceae (cinnamon, camphor), or a plant such as coffee, sugar cane, tea, and natural rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants, and trees such as forests (broad-leaved trees and evergreens, such as conifers), fruit trees, ornamental trees, and nut-bearing trees, and shrubs and other young plants.

[0249] Advantageous Effects of the Invention

[0250] The present application improves the activity of Cas-sf2044 protein by mutation, which has wide application prospects.

[0251] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples, but it will be understood by those skilled in the art that the following drawings and examples are intended only to illustrate the present invention and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art based on the following detailed description of the accompanying drawings and preferred embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0252] Figure 1 .Verification results of the editing efficiency of Cas proteins with single-site amino acid mutations.

[0253] Figure 2 .Verification results of different Cas protein editing efficiencies after amino acid mutation at position 482.

[0254] Figure 3 .Verification results of different Cas protein editing efficiencies after amino acid mutation at position 700.

[0255] Figure 4 .Verification results of different Cas protein editing efficiencies after amino acid mutation at position 697. DETAILED DESCRIPTION

[0256] The following examples are only used to describe the present invention, but are not intended to limit the present invention. Unless otherwise specified, the experiment and method described in the embodiment are carried out substantially according to the conventional methods well known in the art and described in various references. For example, the conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA used in the present invention can be found in Sambrook, Fritsch and Maniatis, " molecular cloning: laboratory manual " (MOLECULAR CLONING:ALABORATORY MANUAL), 2nd edition (1989); " current protocols in molecular biology " (FM Ausubel et al. edit, (1987)); " methods in enzymes " (METHODS IN ENZYMOLOGY) series (Academic Publishing Company): " PCR 2: A PRACTICAL METHOD " (PCR 2: A PRACTICAL METHOD) APPROACH) (MJ MacPherson, BD Hames and GR Taylor, eds. (1995)), ANTIBODIES: A LABORATORY MANUAL (Harlow and Lane, eds. (1988)), and ANIMAL CELL CULTURE (RI Freshney, ed. (1987)).

[0257] In addition, if specific conditions are not specified in the examples, the experiments were performed under conventional conditions or the conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments is not specified, they are all conventional products that can be obtained commercially. It is understood that the examples describe the present invention by way of example and are not intended to limit the scope of the present invention. All publications and other references mentioned herein are incorporated herein by reference in their entirety.

[0258] Example 1. Acquisition of Cas mutant proteins

[0259] For the known Cas protein (Cas-sf2044 in CN117230042A), the applicant predicted the key amino acid sites that may affect its biological function through bioinformatics, and mutated the amino acid sites to obtain a Cas mutant protein with improved editing activity. Specifically, the Cas-sf2044 coding sequence was codon-optimized and synthesized. The amino acid sequence of the wild-type Cas-sf2044 is shown in SEQ ID No. 1, and its nucleic acid sequence is shown in SEQ ID No. 2. The amino acids that bind to the potential Cas-sf2044 and the target sequence were subjected to site-directed mutagenesis using bioinformatics methods.

[0260] SEQ ID No. 1:

[0261] msnyknikfklvpfsqkdlinmqlnvnlhqqcyrefveqfcvlcnipfpglskdqieqkrkqlnlseddekdinyikdlvknknnign

[0262] siyafftgtkkempsrktdltplyrllkanilpfsllkgrenykksifqtvinqtlekfksyfkcnesvennfklslnkdsneeqvlnesemk

[0263] dlqnlfenlsknqsfsffnfnknwfskdkiktkllnnetnkikslsseeidlilsykdklysnefdlismfvefnlqkqkaeslksqadlnlf

[0264] knnnysfrigsnyenfnltqnnkdilleinssmgekitfkiiphkktqiwnleknnvkitsgenlgnyksvdvikmkrpadikakllkts

[0265] elnieiknnqiycnfiyeykcsdhgvyffhcsgnkkpdeknenilkerertfsfidlglfpmysistfkynnksndgeilvksgsgnekld

[0266] fgsafkihsiqigknstnlnkikqlleklkdlktylkfsksissfdensyqrqlktgveiselnslsfqkiseiksinlgfnesfnkeyflklien

[0267] qtftqkellllnckikdlfkilykeysniknsrifkfnkeddlicdgyywlqvideiinkksltyfnskpsekgnkskfiflkdfnyknnfa

[0268] nnyakiaasrlkkyclehkvdvcvfeknlnnflqskdndkktnktlinwanrnlfekiklaleehdicvsevdgkhssqldpqtmnwg

[0269] ardnlngngnkekifferngqiiqqnadlsasevlakrfftryedivhiyidqkikddktilklvkgkvrvesylkktinscyaivdengflkpiskkdynkfqelpskprtdiksnemyrhgskwyhfqqhrefqqdllargrelkki;

[0270] SEQ ID No.2:

[0271] atgtcaaattataaaaatataaaaaattaaaaacttgttccattttctcaaaaagatttaattaacatgcaattaatgttaatttacatcaacaatgttatcg

[0272] agaatttgttgaacaattttgtgttttatgcaacataccttttccaggattaagtaaagatcaaattgaacaaaaaagaaaacaacttaacctttctg

[0273] aagatgatgaaaaagatattatattaaagatttagttaaaaaaataaaataatataggaaatagtatatgcttttttaccggagaacaaaaaaa

[0274] gagatgccaagtcgcaaaacagatttaactcctttgtatagacttttaaaagcaaaatatactccattttctttgttaaaaggaagagaaaattataa

[0275] aaaatctatatttcaaacagtaattaatcaaactttagaaaaatttaaatcttattttaaatgtaatgaatcagttgaaaataatttcaaactttcattaa

[0276] ataaagattcaaatgaagaacaagttttaaacgaaagtgaaatgaaagatcttcaaaatctttttgaaaatttatctaaaaatcaatctttttcattttt

[0277] taactttaataaaaattggttttctaaagataaaattaaaacaaaactacttaataatgaaacaaataagattaaatcacttagtagcgaagaaatt

[0278] gattgattctatcttataaagataaattatattcaaatgaatttgatttaatttcaatgtttgttgaatttaaatttacaaaaacaaaaggctgaatcacta

[0279] aaatctcaagcagatttaaatttatttaaaaataattattcatttagaattgggtcaaattatgaaaattttaaatttaacacaaaataataaagatat

[0280] tctattagaaattaattcttctatgggagaaaaaattacttttaaaataattccacataaaaaaactcaaatttggaacttagaaaaaaataatgttaa

[0281] gataacatcaggtgaaaatttaggtaattataaatctgtagatgttataaaaatgaaacgacccgctgatattaaagctaaattacttaaaacatct

[0282] gaattaaatatagaaattaaaaataatcaaatatattgtaattttatatacgaatataaatgtagcgatcatggagtttatttttttcattgttctggaaat

[0283] aaaaaccagatgaaaaaaatgaaaatatcttaaaagaaagagaaagaactttttcttttatagatttgggattgtttccaatgtatagtatatcaa

[0284] catttaaatataataataaatctaatgatggagaaattttagttaaatctggatctggtaacgaaaaattagattttggaagtgcatttaaaattcatt

[0285] ctattcaaattggtaaaaattcaacaaatttaaacaaaattaaacaacttcttgaaaaactaaaagatttaaaaacatatttgaagttttcaaaatcta

[0286] tatcatcttttgatgaaaattcatatcaaaggcaattgaaaacaggtgtagaaatatctgaattaaactctctttcttttcaaaaaatatctgaaattaa

[0287] aagtattaatttaggatttaaatgaatcttttaataaagaatattttttgaaattaatagaaaatcaaacctttaacccaaaaaaattacttttattaaattg

[0288] taaaataaaagatttattcaaaattttatacaaagaatattccaatattaaaaattcaagaatatttaaatttaataaagaagatgatttaatttgtgatg

[0289] gatattattggctgcaagttattgatgaaataataaatattaaaaaatctttaacatatttcaattcaaaaccttcagaaaaaggaaataaatctaaat

[0290] ttattttcttaaaagattttaattataaaaataattttgctaataattatgctaaaatagcagcttcaagattgaaaaaatattgtctagaacataaagtt

[0291] gatgtttgtgtatttgaaaaaaatcttaataactttctacaaagtaaagataatgataaaaaaaaacaaataaaactttaattaattgggctaatagaaa

[0292] tttatttgaaaaaattaaattagctcttgaagaacatgatatatgtgtatctgaagtagatggtaaacattctagccaattagatcctcaaactatga

[0293] attggggagcaagagataatttaaatggaaatggaaataaagaaaaaatattctttgaaagaaatggacaaataatacagcaaaatgcagattt

[0294] gtctgcgtcagaagtattagcaaaaagattctttacaagatatgaagatattgttcatatttatattgatcaaaaaattaaagatgataaaactatatt

[0295] aaaacttgttaaaggaaaagtaagagttgaaagctatttgaaaaaaacaattaactcttgttatgctattgtagatgaaaatggatttctaaaacca

[0296] atatctaaaaaagattataataaatttcaagaattaccatcaaaaccaagaactgatattaaatcaaatgaaatgtatcgtcatggttcaaaatggtatcactttcaacaacatagagaatttcaacaagatctacttgcgagaggtagagaactaaaaaaaatagcttga;

[0297] Variants of the Cas protein were generated by PCR-based site-directed mutagenesis. The specific method is to divide the DNA sequence of the Cas-sf2044 protein into two parts with the mutation site as the center, design two pairs of primers to amplify the two parts of the DNA sequence respectively, and introduce the sequence to be mutated on the primers. Finally, the two fragments are loaded into the pcDNA3.3-eGFP vector by Gibson cloning. The combination of mutants is constructed by splitting the DNA of the Cas-sf2044 protein into multiple segments and using PCR and Gibson cloning. Fragment amplification kit: TransStart FastPfu DNAPolymerase (containing 2.5mM dNTPs), please refer to the instructions for the specific experimental process. Gel recovery kit: Gel DNA Extraction Mini Kit. Detailed experimental procedures are detailed in the instructions. Kit used for vector construction: pEASY-Basic Seamless Cloning and Assembly Kit (CU201-03). Detailed experimental procedures are detailed in the instructions.

[0298] The amino acid sites of the Cas-sf2044 mutation involved in this embodiment include N160R, L163R, N674R, A694R, N697R, S335R, L602R, A644R, E171R, E254R, E337R, E482R, E514R, E517R, E528R, E538R, E700R, E866R, D681R and other mutated amino acid sites; based on the above mutations, in SEQ ID Based on No. 1, multiple Cas mutant proteins with single site mutations were obtained, including (named after the mutation type): N160R, L163R, N674R, A694R, N697R, S335R, L602R, A644R, E171R, E254R, E337R, E482R, E514R, E517R, E528R, E538R, E700R, E866R, D681R; the above mutation sites are SEQ ID No.1 The 160th, 163rd, 674th, 694th, 697th, 335th, 602nd, 644th, 171st, 254th, 337th, 482nd, 514th, 517th, 528th, 538th, 700th, 866th or 681st amino acid from the N-terminus is mutated to R, respectively.

[0299] Example 2. Verification of the editing activity of Cas mutant proteins

[0300] A fluorescent reporter system suitable for verifying Cas-sf2044 cleavage was constructed with reference to the literature (Yang, Yi, et al. "Highly efficient and rapid detection of the cleavage activity of Cas9 / gRNA via a fluorescent reporter." Applied biochemistry and biotechnology 180.4(2016):655-667.). The fluorescent reporter vector contains RFP and non-luminescent GFFP fluorescent proteins, and the Cas vector contains CFP fluorescent protein. TTT The target site TTTAACAGTGGCCTTATTAA was tested. The underlined position is the PAM sequence. After 48 hours of transfection into CHO cells, the editing efficiency was determined by flow cytometry analysis of the ratio of GFP fluorescence in CFP and RFP. Figure 1 As shown, taking the editing efficiency of WT wild-type Cas-sf2044 (SEQ ID No. 1) as the benchmark (100%), the results showed that the editing efficiency of some Cas-sf2044 mutant proteins was lower than that of wild-type Cas-sf2044, and the editing efficiency of some Cas-sf2044 mutant proteins was higher than that of wild-type Cas-sf2044; Figure 1 Different dots represent different mutants. Among them, the Cas-sf2044 mutant proteins N160R, L163R, N674R, A694R, N697R, S335R, L602R, A644R, E171R, E254R, E337R, E482R, E514R, E517R, E528R, E538R, E700R, E866R, and D681R all showed significantly improved editing efficiencies compared to the wild-type Cas-sf2044, with increases exceeding 20%, as shown in the table below.

[0301] Mutant protein name Editing efficiency relative to wild-type Cas-sf2044 (%) N160R 133 L163R 122 N674R 126 A694R 125 N697R 148 S335R 130 L602R 124 A644R 128 E171R 120 E254R 133 E337R 155 E482R 156 E514R 121 E517R 133 E528R 135 E538R 121 E700R 151 E866R 122 D681R 126

[0302] The above results show that the 160th, 163rd, 674th, 694th, 697th, 335th, 602nd, 644th, 171st, 254th, 337th, 482nd, 514th, 517th, 528th, 538th, 700th, 866th or 681st amino acid of SEQ ID No. 1 from the N-terminus is the key site for the activity of Cas-sf2044. By mutating the above amino acid sites, the editing efficiency of Cas-sf2044 can be greatly improved.

[0303] Example 3. Editing activity of saturated mutants at amino acid positions 482, 700, and 697 of the Cas protein verify

[0304] From the results of Example 2, it can be seen that the editing activity of Cas-sf2044 (shown in SEQ ID No. 1) is greatly improved after mutation of the amino acid sites at positions 482, 700 and 697 from the N-terminus. In order to further verify the effect of mutation of the amino acid sites at positions 482, 700 and 697 to other forms of amino acids on the editing activity of the Cas-sf2044 protein, the applicant used the method of Example 1 to saturate mutate the amino acid site at position 482 to G, A, V, L, I, M, F, Y, W, S, T, C, P, N, Q, D, K , R or H, and obtained Cas proteins with single amino acid mutations E482G, E482A, E482V, E482L, E482I, E482M, E482F, E482Y, E482W, E482S, E482T, E482C, E482P, E482N, E482Q, E482D, E482K, E482R or E482H; the 700th amino acid site was saturated mutated to G, A, V, L, I, M, F, Y, W, S, T, C, P, N, Q, D, K, R or H, the Cas protein E700G, E700A, E700V, E700L, E700I, E700M, E700F, E700Y, E700W, E700S, E700T, E700C, E700P, E700N, E700Q, E700D, E700K, E700R or E700H with a single amino acid mutation was obtained; the amino acid site 697 was saturated mutated to The Cas proteins with single amino acid mutations such as G, A, V, L, I, M, F, Y, W, S, T, C, P, Q, D, E, K, R or H were obtained, and the Cas proteins with single amino acid mutations such as N697G, N697A, N697V, N697L, N697I, N697M, N697F, N697Y, N697W, N697S, N697T, N697C, N697P, E482Q, N697D, N697E, N697K, N697R or N697H were obtained.

[0305] Based on the above amino acid mutation sites, the wild-type protein (WT) of Cas-sf2044 and proteins with mutations at the above amino acid single sites (named after the mutation type) were obtained, respectively: E482G, E482A, E482V, E482L, E482I, E482M, E482F, E482Y, E482W, E482S, E482T, E482C, E482P, E482N, E482Q, E482D, E482K, E482R or E482H, which are relative to SEQ ID 1, wherein the amino acid at position 482 from the N-terminus is mutated to G, A, V, L, I, M, F, Y, W, S, T, C, P, N, Q, D, K, R or H; E700G, E700A, E700V, E700L, E700I, E700M, E700F, E700Y, E700W, E700S, E700T, E700C, E700P, E700N, E700Q, E700D, E700K, E700R or E700H, which is relative to SEQ ID In the sequence shown in SEQ ID No. 1, the amino acid at position 700 from the N-terminus is mutated to G, A, V, L, I, M, F, Y, W, S, T, C, P, N, Q, D, K, R or H, respectively; and N697G, N697A, N697V, N697L, N697I, N697M, N697F, N697Y, N697W, N697S, N697T, N697C, N697P, E482Q, N697D, N697E, N697K, N697R or N697H, which, relative to the sequence shown in SEQ ID No. 1, the amino acid at position 697 from the N-terminus is mutated to G, A, V, L, I, M, F, Y, W, S, T, C, P, Q, D, E, K, R or H, respectively.

[0306] The gene editing activity of the different Cas proteins obtained above was verified in animal cells, and the target was designed for the FUT8 gene of Chinese hamster ovary cells (CHO): TTC CCCCTTGGCTGTACCAGAAGACCTT, the italic part is the PAM sequence, and the underlined area is the targeting region. The vector pcDNA3.3 is modified to carry EGFP fluorescent protein and PuroR resistance gene. The SV40 NLS-Cas-sf2044 fusion protein is inserted through the enzyme cutting sites XbaI and PstI; the U6 promoter and gRNA sequence are inserted through the enzyme cutting site Mfe1. The CMV promoter drives the expression of the fusion protein SV40 NLS-Cas-sf2044-NLS-GFP. The protein Cas-sf2044-NLS is connected to the protein GFP with the connecting peptide T2A. The promoter EF-1α drives the expression of the puromycin resistance gene. Plating: CHO cells are plated when the confluence reaches 70-80%, and the number of cells seeded in a 12-well plate is 8*10^4 cells / well. Transfection: Transfection is performed 24 hours after plating, and 6.25μl Hieff Trans is added to 100μl opti-MEM. TM Liposome nucleic acid transfection reagent, mix well; add 2.5ug plasmid to 100μl opti-MEM and mix well. TM Mix the diluted plasmid with the liposome nucleic acid transfection reagent and incubate at room temperature for 20 minutes. Add the incubated mixture to the culture medium containing the cells for transfection. Add puromycin selection: Add puromycin 24 hours after transfection to a final concentration of 10 μg / ml. After 24 hours of puromycin treatment, replace the culture medium with normal medium and continue culturing for another 24 hours. 48 hours after transfection, digest the cells with trypsin-EDTA (0.05%) and sort the cells with GFP signal using flow cytometry (FACS).

[0307] Extract DNA, PCR amplify the area near the editing region, and send for hiTOM sequencing: The cells were collected after trypsin digestion, and genomic DNA was extracted using a cell / tissue genomic DNA extraction kit (Biotech). The genomic DNA was amplified near the target site. The PCR product was sequenced by hiTOM. Sequencing data analysis was performed to count the types and proportions of sequences within 15nt upstream and 10nt downstream of the target site. The sequences with SNV frequencies greater than / equal to 1% or non-SNV mutation frequencies greater than / equal to 0.06% in the sequences were counted to obtain the editing efficiency of the target site by the Cas-sf2044 protein. CHO cell FUT8 gene target sequence: TTC CCCCTTGGCTGTACCAGAAGACCTT The italic part is the PAM sequence, and the underlined area is the targeting region. The gRNA sequence is: AGUGCAAUAGUUACAGAAUAGUAAUUAU AUUCGCA CCCCUUGGCUGUACCAGAAGACCUU , the underlined region is the target region, and the other regions are DR (direct repeat sequence) regions.

[0308] The results of the E482G, E482A, E482V, E482L, E482I, E482M, E482F, E482Y, E482W, E482S, E482T, E482C, E482P, E482N, E482Q, E482D, E482K, E482R, or E482H mutant Cas proteins are shown in Figure 2 , Figure 2 WT represents the wild-type Cas-sf2044 protein. Most mutant proteins obtained by mutating amino acid position 482 of Cas-sf2044 to different amino acid residues significantly enhance the editing activity of the Cas-sf2044 protein. In particular, mutant Cas proteins containing E482G, E482A, E482V, E482L, E482I, E482M, E482F, E482Y, E482W, E482S, E482T, E482C, E482P, E482N, E482D, E482K, E482R, or E482H significantly enhance editing efficiency compared to the wild-type Cas-sf2044 protein.

[0309] The results for E700G, E700A, E700V, E700L, E700I, E700M, E700F, E700Y, E700W, E700S, E700T, E700C, E700P, E700N, E700Q, E700D, E700K, E700R, or E700H mutant Cas proteins are shown in the table. Figure 3 , Figure 3 WT in this data represents the wild-type Cas-sf2044 protein. Some mutant proteins obtained by mutating amino acid residues at position 700 of Cas-sf2044 to different residues significantly enhance the editing activity of the Cas protein. In particular, the E700G, E700L, E700I, E700M, E700Y, E700S, E700K, E700R, or E700H mutant Cas proteins significantly enhance editing efficiency compared to the wild-type Cas-sf2044 protein.

[0310] Results for N697G, N697A, N697V, N697L, N697I, N697M, N697F, N697Y, N697W, N697S, N697T, N697C, N697P, E482Q, N697D, N697E, N697K, N697R, or N697H mutant Cas proteins are shown in Figure 4 , Figure 4The WT in this figure represents the wild-type Cas-sf2044 protein. Mutants obtained by mutating amino acid position 697 of Cas-sf2044 to different residues significantly enhance the editing activity of the Cas protein. In particular, the N697G, E482Q, N697K, and N697R mutants significantly improve editing efficiency compared to the wild-type Cas-sf2044 protein.

[0311] Although the specific embodiments of the present invention have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details based on all the teachings published, and these changes are all within the scope of protection of the present invention. The entire invention is given by the appended claims and any equivalents thereof.

Claims

1. A Cas mutant protein, which has a mutation at any of the following amino acid positions corresponding to the amino acid sequence of SEQ ID No. 1 compared to the amino acid sequence of the parent Cas protein: 482, 337, 700, 697, 528, 160, 254, 517, 335, 644, 674, 681, 694, 602, 163, 866, 514, 53 8 or 171; the 482nd, 337th, 700th, 697th, 528th, 160th, 254th, 517th, 335th, 644th, 674th, 681st, 694th, 602nd, 163rd, 866th, 514th, 538th or 171st amino acid is mutated to R.

2. A fusion protein comprising the Cas mutant protein according to claim 1 and other modified parts.

3. An isolated polynucleotide, characterized in that The polynucleotide is a polynucleotide sequence encoding the Cas mutant protein according to claim 1, or a polynucleotide sequence encoding the fusion protein according to claim 2.

4. A carrier, characterized in that The vector comprises the polynucleotide according to claim 3 and a regulatory element operably linked thereto.

5. A CRISPR-Cas system, characterized in that The system comprises the Cas mutant protein according to claim 1 and gRNA, and the gRNA is capable of binding to the Cas mutant protein according to claim 1.

6. A composition, characterized in that The composition comprises: (i) a protein component selected from the group consisting of: the Cas mutant protein of claim 1 or the fusion protein of claim 2; (ii) a nucleic acid component selected from: gRNA, or a nucleic acid encoding a gRNA, or a precursor RNA of a gRNA, or a nucleic acid encoding a precursor RNA of a gRNA, wherein the gRNA is capable of binding to the Cas mutant protein according to claim 1; The protein component and the nucleic acid component combine with each other to form a complex.

7. An engineered host cell, characterized in that The host cell comprises the Cas mutant protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6.

8. Use of the Cas mutant protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7 in gene editing, wherein the use is for purposes other than disease diagnosis and treatment; Alternatively, use in the preparation of a preparation or a kit for gene editing.

9. A method for editing a target nucleic acid, the method being a method for non-disease diagnosis and treatment purposes, the method comprising contacting the target nucleic acid with the Cas mutant protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7.

10. A kit for gene editing, comprising the Cas mutant protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7.

Citation Information

Patent Citations

  • Cas protein with improved editing activity and application thereof

    CN116004573A

  • Novel CRISPR enzyme and system and application

    CN117230042A