Optimized cas proteins and uses thereof
By performing site-directed mutations on the Cas12f.4 protein, particularly by replacing amino acids at positions 876 and/or 890, the off-target effects in the CRISPR/Cas system were resolved, enabling more efficient and precise gene editing.
Patent Information
- Application Number
- CN202411088340.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-26
- Filing Date
- 2024-03-15
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2044-03-15
AI Technical Summary
Existing CRISPR/Cas systems suffer from off-target effects during gene editing, leading to difficulties in predicting target sites and potential non-specific cleavage problems.
By performing site-directed mutagenesis on the Cas12f.4 protein, particularly by replacing amino acids at positions 876 and/or 890, the Cas protein was optimized to reduce off-target effects and improve editing activity.
It significantly reduced off-target effects, expanded the application range of Cas proteins, and improved the precision and efficiency of gene editing.
Smart Images

Figure CN118834854B_ABST
Abstract
Description
[0001] The present application is a divisional application of Chinese Patent Application No. CN 202410296507.X, filed on March 15, 2024, entitled “An Optimized Cas Protein and Its Application”.
[0002] This application claims priority to Chinese Patent Application No. CN 202310491001.X, filed on April 26, 2023. This application incorporates the entirety of the aforementioned Chinese Patent Application. TECHNICAL FIELD
[0003] The present application relates to the field of gene editing, in particular, the field of Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) technology. Specifically, the present application relates to a Cas protein with reduced off-target efficiency and its application. BACKGROUND
[0004] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to guide specific binding to target sequences on the genome and cut DNA to produce double-strand breaks, using biological non-homologous end joining or homologous recombination for site-directed gene editing.
[0005] The CRISPR / Cas9 system is the most commonly used Type II CRISPR system, which recognizes the PAM motif of 3’-NGG and performs blunt-end cleavage on target sequences. The CRISPR / Cas Type V system is a newly discovered CRISPR system that has a motif of 5’-TTN and performs sticky-end cleavage on target sequences, such as Cpf1, C2c1, CasX, and CasY. However, the different CRISPR / Cas systems currently available each have different advantages and disadvantages. For example, Cas9, C2c1, and CasX all require two RNA guide RNAs, while Cpf1 only requires one guide RNA and can be used for multiplex gene editing. CasX has a size of 980 amino acids, while common Cas9, C2c1, CasY, and Cpf1 typically have a size of around 1300 amino acids. In addition, the PAM sequences of Cas9, Cpf1, CasX, and CasY are complex and diverse, while C2c1 recognizes the stringent 5’-TTN, so its target sites are easier to predict than other systems, thereby reducing potential off-target effects.
[0006] Chinese Invention Patent CN111757889B discloses a Cas protein, Cas12f.4, and also discloses that this protein can perform gene editing in eukaryotic cells, but there is a potential off-target effect. The present application optimizes this protein to reduce its off-target effect during gene editing. SUMMARY
[0007] The inventors of the present application have made a large number of experiments and repeated groping, and by site-directed mutagenesis of Cas12f.4 (which is referred to as Cas12i3 or Cas12i.3 in the present application) protein, the editing activity thereof is improved, the off-target effect thereof in the process of gene editing is reduced, and the application range thereof is expanded.
[0008] Cas effector protein
[0009] In one aspect, the present application provides a Cas mutant protein with reduced off-target effect, which has a mutation at the amino acid site corresponding to the following amino acid site in the amino acid sequence shown in SEQ ID No. 1, compared with the amino acid sequence of the parent Cas protein: position 876 and / or position 890.
[0010] In one embodiment, the Cas mutant protein has a mutation at the above-mentioned position 876 amino acid site; further, on the basis of the mutation of the amino acid at position 876, the amino acid site at position 890 is also mutated.
[0011] In one embodiment, the Cas mutant protein has a mutation at the above-mentioned position 890 amino acid site; further, on the basis of the mutation of the amino acid at position 890, the amino acid site at position 876 is also mutated.
[0012] In one embodiment, the amino acid at position 876 is mutated to an amino acid other than D, for example, A, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, R.
[0013] In one embodiment, the amino acid at position 890 is mutated to an amino acid other than A, for example, D, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, R.
[0014] Further, the Cas mutant protein has a mutation at any one or any several of the following amino acid sites corresponding to the amino acid sequence shown in SEQ ID No. 1, compared with the amino acid sequence of the parent Cas protein: position 7, position 233, position 267, position 369, position 433.
[0015] In one embodiment, the amino acid at position 7 is mutated to an amino acid other than S, for example, A, V, G, L, Q, F, W, Y, D, K, E, N, M, T, C, P, H, R, I; preferably, R, H, K, M, F, P, A, W, I, V, L, Q, C or Y, preferably, R.
[0016] In one embodiment, the amino acid at position 233 or 267 is mutated to an amino acid other than D, e.g., A, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, the amino acid at position 233 or 267 is mutated to R.
[0017] In one embodiment, the amino acid at position 369 is mutated to an amino acid other than N, e.g., A, V, G, L, Q, F, W, Y, D, S, E, K, M, T, C, P, H, R, I; preferably, R.
[0018] In one embodiment, the amino acid at position 433 is mutated to an amino acid other than S, e.g., A, V, G, L, Q, F, W, Y, D, N, E, K, M, T, C, P, H, R, I; preferably, R.
[0019] In one embodiment, the amino acid at position 876 is mutated to an amino acid other than D, e.g., A, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, R.
[0020] In one embodiment, the amino acid at position 890 is mutated to an amino acid other than A, e.g., D, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, R.
[0021] In some embodiments, the parent Cas protein is a naturally-occurring wild-type Cas protein; in other embodiments, the parent Cas protein is an engineered Cas protein.
[0022] Cas proteins or Casl2i proteins from a variety of organisms can be used as parent Cas proteins, in some embodiments, the parent Cas protein or Casl2i protein has nuclease activity. In some embodiments, the parent Cas protein is a nuclease, i.e., cleaves both strands of a target double-helical nucleic acid (e.g., double-helical DNA). In some embodiments, the parent Cas protein is a nickase, i.e., cleaves a single strand of a target double-helical nucleic acid (e.g., double-helical DNA).
[0023] In one embodiment, the parent Cas protein is a Cas protein of the Casl2 family, preferably a Cas protein of the Casl2i family, e.g., Casl2il, Casl2i2, Casl2i3, etc.
[0024] In one embodiment, the Cas protein of the Cas12 family has an amino acid sequence that has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%, or 100% sequence identity to SEQ ID No. 1.
[0025] In one embodiment, the Cas protein of the Cas12 family has an amino acid sequence that has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%, or 100% sequence identity to SEQ ID No. 1.
[0026] In one embodiment, the Cas mutant protein is selected from any one of the following groups I-III:
[0027] I. a Cas mutant protein resulting from mutating the amino acid sequence set forth in SEQ ID No. 1 at any one or any number of the following amino acid positions: position 876, position 879; optionally, the Cas mutant protein further has a mutation at any one or any number of the following amino acid positions of the amino acid sequence set forth in SEQ ID No. 1: position 7, position 233, position 267, position 369, position 433;
[0028] II. a Cas mutant protein having the mutation sites described in I; and, having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the Cas mutant protein described in I;
[0029] III. has the mutation site described in I; and has one or more substitutions, deletions, or additions of amino acids compared to the Cas mutant protein described in I; the one or more amino acids include 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 substitutions, deletions, or additions of amino acids.
[0030] In another aspect, the present application also provides a Cas mutant protein with reduced off-target efficiency, which has a mutation at the amino acid site corresponding to position 876 and / or position 890 of the amino acid sequence shown in SEQ ID No. 1 compared to the amino acid sequence of the parent Cas protein; further, the Cas mutant protein also includes any 1, any 2, any 3, any 4, or any 5 amino acid site mutations selected from the group consisting of position 7, position 233, position 267, position 369, or position 433 of the amino acid sequence shown in SEQ ID No. 1.
[0031] In one embodiment, the Cas mutant protein has a mutation at the above-mentioned position 876 and position 890 amino acid sites; further, based on the above-mentioned position 876 and position 890 amino acid mutations, it also includes any one or any several (for example, any 2, any 3, any 4, or any 5) amino acid mutations selected from the group consisting of position 7, position 233, position 267, position 369, or position 433. For example, the position 876, position 890, and position 7 are simultaneously mutated, the position 876, position 890, position 7, and position 233 are simultaneously mutated, the position 876, position 890, position 7, position 233, and position 267 are simultaneously mutated, the position 876, position 890, position 7, position 233, position 267, and position 369 are simultaneously mutated, the position 876, position 890, position 7, position 233, position 267, position 369, and position 433 are simultaneously mutated.
[0032] In one embodiment, the position 7 amino acid is mutated to an amino acid other than S, for example, A, V, G, L, Q, F, W, Y, D, K, E, N, M, T, C, P, H, R, I; preferably, R, H, K, M, F, P, A, W, I, V, L, Q, C, or Y, preferably, R.
[0033] In one embodiment, the position 233 amino acid or position 267 amino acid is mutated to an amino acid other than D, for example, A, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, the position 233 amino acid or position 267 amino acid is mutated to R.
[0034] In one embodiment, the amino acid at position 369 is mutated to an amino acid other than N, for example, A, V, G, L, Q, F, W, Y, D, S, E, K, M, T, C, P, H, R, I; preferably, R.
[0035] In one embodiment, the amino acid at position 433 is mutated to an amino acid other than S, for example, A, V, G, L, Q, F, W, Y, D, N, E, K, M, T, C, P, H, R, I; preferably, R.
[0036] In one embodiment, the amino acid at position 876 is mutated to an amino acid other than D, for example, A, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, R.
[0037] In one embodiment, the amino acid at position 890 is mutated to an amino acid other than A, for example, D, V, G, L, Q, F, W, Y, N, S, E, K, M, T, C, P, H, R, I; preferably, R.
[0038] In one embodiment, the amino acid sequence of the parent Cas protein is set forth in SEQ ID No. 1.
[0039] In some embodiments, the parent Cas protein is a naturally occurring wild-type Cas protein; in other embodiments, the parent Cas protein is an engineered Cas protein.
[0040] Cas proteins or Cas12i proteins from a variety of organisms can be used as parent Cas proteins, in some embodiments, the parent Cas protein or Cas12i protein has nuclease activity. In some embodiments, the parent Cas protein is a nuclease, i.e., cleaves both strands of a target double-helical nucleic acid (e.g., double-helical DNA). In some embodiments, the parent Cas protein is a nickase, i.e., cleaves a single strand of a target double-helical nucleic acid (e.g., double-helical DNA).
[0041] In one embodiment, the parent Cas protein is a Cas protein of the Cas12 family, preferably, a Cas protein of the Cas12i family, for example, Cas12i1, Cas12i2, Cas12i3, etc.
[0042] In one embodiment, the amino acid sequence of the Cas protein of the Cas12 family has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%, or 100% sequence identity to SEQ ID No. 1.
[0043] In one embodiment, the amino acid sequence of the parent Cas protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to SEQ ID No. 1.
[0044] The present application finds that when the above amino acid sites are mutated to positively charged amino acids such as R, H or K, or, to polar uncharged amino acids such as M, F, P, A, W, I, V, L, the editing activity of the Cas protein can be significantly improved; when mutated to partially non-polar uncharged amino acids such as Q, C or Y, the editing activity of the Cas protein can also be significantly improved.
[0045] It is clear to those skilled in the art that the structure of a protein can be changed without adversely affecting its activity and functionality, e.g., one or more conservative amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples of conservative amino acid substitutions and embodiments are clear to those skilled in the art. Specifically, an amino acid residue can be replaced with another amino acid residue belonging to the same group, i.e., a nonpolar amino acid residue is replaced with another nonpolar amino acid residue, a polar uncharged amino acid residue is replaced with another polar uncharged amino acid residue, a basic amino acid residue is replaced with another basic amino acid residue, and an acidic amino acid residue is replaced with another acidic amino acid residue. Such substituted amino acid residues can or can not be encoded by the genetic code. Conservative substitutions in which one amino acid is replaced with another amino acid from the same group are within the scope of the present application, provided that the substitution does not result in the inactivation of the biological activity of the protein. Thus, the proteins of the present application can contain one or more conservative substitutions in the amino acid sequence, which are preferably made according to Table 1. In addition, the present application also encompasses proteins which further contain one or more other non-conservative substitutions, provided that the non-conservative substitutions do not significantly affect the desired function and biological activity of the proteins of the present application.
[0046] Conservative amino acid substitutions can be made at one or more predicted nonessential amino acid residues. A "nonessential" amino acid residue is a residue that can be altered without a significant loss of activity, i.e., a loss of activity of less than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A "conservative amino acid substitution" is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain. The amino acid substitution can be made in the non-conserved regions of the Cas mutant proteins described above. In general, such substitutions will not be made in conservative amino acid residues, or in amino acid residues that are within a conserved motif, where such residues are required for protein activity. However, it will be appreciated by one skilled in the art that functional variants can have fewer conservative or non-conservative changes in the conserved regions.
[0047] Table 1
[0048]
[0049]
[0050] It is well known in the art that one or more amino acid residues can be altered (substituted, deleted, truncated, or inserted) from the N and / or C terminus of a protein while still retaining its functional activity. Thus, proteins having one or more amino acid residues altered from the N and / or C terminus of a Cas protein while retaining its desired functional activity are also within the scope of the present application. These alterations can include alterations introduced by modern molecular methods, such as PCR, including PCR amplification of protein-encoding sequences by means of inclusion of amino acid-encoding sequences in the oligonucleotides used in the PCR amplification to alter or extend the protein-encoding sequence.
[0051] It is recognized that proteins can be altered in a variety of ways including amino acid substitutions, deletions, truncations and insertions, and that methods for such manipulations are generally known in the art. For example, amino acid sequence variants of the above-described proteins can be prepared by mutations in the DNA. They can also be prepared by other forms of mutagenesis and / or by directed evolution, for example, using known mutagenic, recombination and / or shuffling methods in combination with relevant screening methods to make single or multiple amino acid substitutions, deletions and / or insertions.
[0052] As will be appreciated by those of ordinary skill in the art, these minor amino acid changes in the Cas proteins of the application can occur (e.g., naturally-occurring mutations) or be produced (e.g., using r-DNA technology) without loss of protein function or activity. If the mutations occur in the catalytic domain, active site, or other functional domain of the protein, the properties of the polypeptide can change, but the polypeptide can retain its activity. If the mutations present are not near the catalytic domain, active site, or other functional domain, less impact can be expected.
[0053] As will be appreciated by those of ordinary skill in the art, essential amino acids of the Cas mutant proteins of the application can be identified according to methods known in the art, such as by site-directed mutagenesis or protein evolution or bioinformatic analysis. The catalytic domain, active site, or other functional domain of the protein can also be determined by physical analysis of the structure, such as by nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling in combination with mutations of putative key site amino acids.
[0054] In the present application, the amino acid residues can be represented by single letter or three letter, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamine (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine (Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), arginine (Arg, R).
[0055] The term "AxxB" means that the amino acid A at the xxth position is changed to amino acid B. If not otherwise specified, it means that the amino acid A at the xxth position from the N-terminus is changed to amino acid B. For example, D876R means that the D at the 876th position is mutated to R. When multiple amino acid positions are mutated simultaneously, the mutations can be represented in the form of D876R-A890R or D876R / A890R, for example, D876R-A890R means that the D at the 876th position is mutated to R and the A at the 890th position is mutated to R.
[0056] The specific amino acid positions (numbering) in the protein according to the present application are determined by aligning the amino acid sequence of the protein of interest with SEQ ID No. 1 using standard sequence alignment tools, for example, by aligning the two sequences using the Smith-Waterman algorithm or the CLUSTALW2 algorithm, wherein the sequences are considered to be aligned when the alignment score is highest. The alignment score can be calculated according to the method described in Wilbur, W. J. and Lipman, D. J. (1983) Rapid similarity searches of nucleic acid and protein data banks. Proc. Natl. Acad. Sci. USA, 80:726-730. In the ClustalW2 (1.82) algorithm, the default parameters are preferably used: Protein gap open penalty = 10.0; Protein gap extension penalty = 0.2; Protein matrix = Gonnet; Protein / DNA end gap = -1; Protein / DNA GAP DIST = 4. The AlignX program (part of the vector NTI suite) is preferably used to determine the positions of the specific amino acids in the protein according to the present application by aligning the amino acid sequence of the protein with SEQ ID No. 1 using the default parameters suitable for multiple alignments (gap open penalty: 10, gap extension penalty 0.05).
[0057] The skilled person can compare and align the amino acid sequence of any parent Cas protein with SEQ ID No. 1 or 3 for sequence identity using software commonly used in the art, such as Clustal Omega, and further obtain the amino acid positions in the parent Cas protein that correspond to the amino acid positions defined based on SEQ ID No. 1 or 3 as described herein.
[0058] The biological functions of the Cas protein include, but are not limited to, the activity of binding to a guide RNA, endonuclease activity, the activity of binding to a specific site of a target sequence and cleaving under the guidance of a guide RNA, including but not limited to Cis cleavage activity and Trans cleavage activity.
[0059] In the present application, the "Cas mutant protein" can also be referred to as a mutated Cas protein, or a Cas protein variant.
[0060] The present application also provides a fusion protein comprising a Cas mutant protein as described above and other modified moieties.
[0061] In one embodiment, the modified moiety is selected from the group consisting of another protein or polypeptide, a detectable label, or any combination thereof.
[0062] In one embodiment, the modified moiety is selected from the group consisting of an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), and a domain having an activity selected from the group consisting of a nucleotide deaminase, a cytidine deaminase, an adenosine deaminase, a methylase activity, a demethylase, a transcriptional activation activity, a transcriptional repression activity, a transcriptional release factor activity, a histone modification activity, a nuclease activity, a single-stranded RNA cleavage activity, a double-stranded RNA cleavage activity, a single-stranded DNA cleavage activity, a double-stranded DNA cleavage activity, and a nucleic acid binding activity; and any combination thereof. The NLS sequence is well known to those skilled in the art, examples of which include but are not limited to the SV40 large T antigen, EGL-13, c-Myc, and TUS proteins.
[0063] In one embodiment, the NLS sequence is located at, near or close to the end (e.g., N-terminus, C-terminus or both) of the Cas protein of the present application.
[0064] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and other suitable epitope tags (e.g., purification, detection or tracking) can be selected by those skilled in the art.
[0065] The reporter gene sequence is well known to those skilled in the art, examples of which include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0066] In one embodiment, the fusion protein of the present application comprises a domain capable of binding to a DNA molecule or intracellular molecule, such as maltose binding protein (MBP), DNA binding domain (DBD) of Lex A, DBD of GAL4, etc.
[0067] In one embodiment, the fusion protein of the present application comprises a detectable label, such as a fluorescent dye, e.g., FITC or DAPI.
[0068] In one embodiment, the Cas protein of the present application is coupled, conjugated or fused to the modification moiety, optionally through a linker.
[0069] In one embodiment, the modification moiety is directly linked to the N- or C-terminus of the Cas protein of the present application.
[0070] In one embodiment, the modification moiety is linked to the N- or C-terminus of the Cas protein of the present application through a linker. Such linkers are well known in the art, examples of which include, but are not limited to, a linker comprising one or more (e.g., 1, 2, 3, 4 or 5) amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA or Ava), or PEG, etc.
[0071] The Cas protein, protein derivative or fusion protein of the present application is not limited by the way it is produced, e.g., it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.
[0072] Nucleic acid of a Cas protein
[0073] In another aspect, the present application provides an isolated polynucleotide comprising:
[0074] (a) a polynucleotide sequence encoding the Cas mutein or fusion protein of the present application;
[0075] or a polynucleotide complementary to the polynucleotide of (a).
[0076] In one embodiment, the nucleotide sequence is codon-optimized for expression in a prokaryotic cell. In one embodiment, the nucleotide sequence is codon-optimized for expression in a eukaryotic cell.
[0077] In one embodiment, the cell is an animal cell, e.g., a mammalian cell.
[0078] In one embodiment, the cell is a human cell.
[0079] In one embodiment, the cell is a plant cell, e.g., a cell of a cultivated plant (such as cassava, corn, sorghum, wheat, or rice), an alga, a tree, or a vegetable.
[0080] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.
[0081] Guide RNA (gRNA)
[0082] In another aspect, the present application provides a gRNA comprising a first segment and a second segment; the first segment is also referred to as a "backbone region", a "protein-binding segment", a "protein-binding sequence", or a "direct repeat sequence"; the second segment is also referred to as a "targeting sequence of a targeting nucleic acid" or a "targeting segment of a targeting nucleic acid", or a "guide sequence targeting a target sequence".
[0083] The first segment of the gRNA is capable of interacting with a Cas protein of the present application, thereby allowing the Cas protein and the gRNA to form a complex.
[0084] In a preferred embodiment, the first segment is a direct repeat sequence as described above.
[0085] The targeting sequence of a targeting nucleic acid or the targeting segment of a targeting nucleic acid of the present application comprises a nucleotide sequence that is complementary to a sequence in a target nucleic acid. In other words, the targeting sequence of a targeting nucleic acid or the targeting segment of a targeting nucleic acid of the present application interacts with a target nucleic acid in a sequence-specific manner via hybridization (i.e., base pairing). Thus, the targeting sequence of a targeting nucleic acid or the targeting segment of a targeting nucleic acid can be altered, or can be modified to hybridize to any desired sequence within a target nucleic acid. The nucleic acid is selected from DNA or RNA.
[0086] The percentage of complementarity between the targeting sequence of a targeting nucleic acid or the targeting segment of a targeting nucleic acid and the target sequence of a target nucleic acid can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).
[0087] The "scaffold region", "protein-binding segment", "protein-binding sequence", or "direct repeat sequence" of the gRNA of the present application can interact with a CRISPR protein (or, Cas protein). The gRNA of the present application, through the targeting sequence targeting a nucleic acid, directs the Cas protein it interacts with to a specific nucleotide sequence within the target nucleic acid.
[0088] Preferably, the guide RNA comprises a first segment and a second segment from 5' to 3' direction.
[0089] In the present application, the second segment can also be understood as a guide sequence hybridizing with the target sequence.
[0090] The gRNA of the present application can form a complex with the Cas protein.
[0091] Vector
[0092] The present application also provides a vector comprising the Cas mutant protein, the isolated nucleic acid molecule or polynucleotide as described above; preferably, it further comprises a regulatory element operably linked thereto.
[0093] In one embodiment, the regulatory element is selected from one or more of the group consisting of enhancer, transposon, promoter, terminator, leader sequence, polyadenylation sequence, marker gene.
[0094] In one embodiment, the vector comprises a cloning vector, an expression vector, a shuttle vector, an integration vector.
[0095] In some embodiments, the vector included in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a bacteriophage, and the like, which are well known to those skilled in the art.
[0096] CRISPR system
[0097] The present application provides an engineered non-naturally occurring vector system, or a CRISPR-Cas system, comprising a Cas mutant protein or a nucleic acid sequence encoding the Cas mutant protein and a nucleic acid encoding one or more guide RNAs.
[0098] In one embodiment, the nucleic acid sequence encoding the Cas mutant protein and the nucleic acid encoding one or more guide RNAs are artificially synthesized.
[0099] In one embodiment, the nucleic acid sequence encoding the Cas mutant protein and the nucleic acid encoding one or more guide RNAs are not naturally co-existing.
[0100] The one or more guide RNAs target one or more target sequences in the cell. The one or more target sequences hybridize to a genomic locus of a DNA molecule encoding one or more gene products, and direct the Cas protein to the genomic locus of the DNA molecule encoding the one or more gene products, where the Cas protein modifies, edits, or cleaves the target sequence upon reaching the target sequence location, whereby the expression of the one or more gene products is altered or modified.
[0101] The cells of the present application include one or more of an animal, a plant, or a microorganism.
[0102] In some embodiments, the Cas protein is codon-optimized for expression in the cell.
[0103] In some embodiments, the Cas protein directs cleavage of one or both strands at the target sequence location.
[0104] The present application also provides an engineered, non-naturally occurring vector system, which can include one or more vectors, the one or more vectors comprising:
[0105] a) a first regulatory element operably linked to a gRNA,
[0106] b) a second regulatory element operably linked to the Cas protein;
[0107] wherein components (a) and (b) are located on the same or different vectors of the system.
[0108] The first and second regulatory elements include promoters (e.g., constitutive promoters or inducible promoters), enhancers (e.g., 35S promoter or 35S enhanced promoter), internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences).
[0109] In some embodiments, the vectors in the system are viral vectors (e.g., retroviral vectors, lentiviral vectors, adenoviral vectors, adeno-associated vectors, and herpes simplex vectors), and can also be plasmids, viruses, cosmids, phages, and the like, which are well known to those skilled in the art.
[0110] In some embodiments, the systems provided herein are in a delivery system. In some embodiments, the delivery system is a nanoparticle, a liposome, an exosome, a microvesicle, and a gene gun.
[0111] In one embodiment, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell. In one embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0112] In one embodiment, the target sequence is present within a cell. In one embodiment, the target sequence is present within a nucleus or within a cytoplasm (e.g., organelle). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.
[0113] In one embodiment, the Cas protein is linked to one or more NLS sequences. In one embodiment, the fusion protein comprises one or more NLS sequences. In one embodiment, the NLS sequence is linked to the N- or C-terminus of the protein. In one embodiment, the NLS sequence is fused to the N- or C-terminus of the protein.
[0114] In another aspect, the present application relates to an engineered CRISPR system comprising the above-mentioned Cas protein and one or more guide RNAs, wherein the guide RNA comprises a spacer sequence capable of hybridizing to a target nucleic acid and a direct repeat sequence, and the Cas protein is capable of binding the guide RNA and targeting a target nucleic acid sequence complementary to the spacer sequence.
[0115] Protein-nucleic acid complex / composition
[0116] In another aspect, the present application provides a complex or composition comprising:
[0117] (i) a protein component selected from the group consisting of: the above-mentioned Cas protein, derivatized protein or fusion protein, and any combination thereof; and
[0118] (ii) a nucleic acid component comprising (a) a guide sequence capable of hybridizing to a target sequence; and (b) a direct repeat sequence capable of binding to the Cas protein of the present application.
[0119] The protein component and the nucleic acid component are associated with each other to form a complex.
[0120] In one embodiment, the nucleic acid component is a guide RNA in a CRISPR-Cas system.
[0121] In one embodiment, the complex or composition is non-naturally occurring or modified. In one embodiment, at least one component in the complex or composition is non-naturally occurring or modified. In one embodiment, the first component is non-naturally occurring or modified; and / or, the second component is non-naturally occurring or modified.
[0122] Activated CRISPR complex
[0123] In another aspect, the present application also provides an activated CRISPR complex comprising: (1) a protein component selected from the group consisting of: a Cas protein, a derivatized protein or a fusion protein of the present application, and any combination thereof; (2) a gRNA comprising (a) a guide sequence capable of hybridizing to a target sequence; and (b) a direct repeat sequence capable of binding to a Cas protein of the present application; and (3) a target sequence bound to the gRNA. Preferably, the binding is by the target sequence of the targeting nucleic acid on the gRNA to the target nucleic acid.
[0124] The term "activated CRISPR complex", "activated complex" or "ternary complex" as used herein refers to the complex of a Cas protein, a gRNA and a target nucleic acid in the CRISPR system after binding or modification.
[0125] The Cas protein and the gRNA of the present application can form a binary complex which is activated upon binding to a nucleic acid substrate that is complementary to the spacer sequence in the gRNA (or alternatively, the guide sequence hybridized to the target nucleic acid) to form an activated CRISPR complex. In some embodiments, the spacer sequence of the gRNA is perfectly matched to the target substrate. In other embodiments, the spacer sequence of the gRNA is partially (contiguously or non-contiguously) matched to the target substrate.
[0126] In preferred embodiments, the activated CRISPR complex can exhibit a collateral nuclease cleavage activity, which refers to the non-specific cleavage activity or the random cleavage activity exhibited by the activated CRISPR complex to a single-stranded nucleic acid, also known in the art as trans cleavage activity.
[0127] Delivery and delivery composition
[0128] The Cas protein, the gRNA, the fusion protein, the nucleic acid molecule, the vector, the system, the complex and the composition of the present application can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipofection, nucleofection, microinjection, sonoporation, biolistics, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection, magnetofection, lipofection, perforation transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial virions, and the like.
[0129] Accordingly, in another aspect, the present application provides a delivery composition comprising a delivery vehicle, and one or any combination of: a Cas protein, a fusion protein, a nucleic acid molecule, a vector, a system, a complex and a composition of the present application.
[0130] In one embodiment, the delivery vehicle is a particle.
[0131] In one embodiment, the delivery vehicle is selected from a lipid particle, a sugar particle, a metal particle, a protein particle, a liposome, an exosome, a microvesicle, a gene gun, or a viral vector (e.g., a replication-defective retrovirus, a lentivirus, an adenovirus, or an adeno-associated virus).
[0132] Host cell
[0133] The present application also relates to a cell or cell line, or progeny thereof, in vitro, ex vivo, or in vivo, comprising a Cas protein, a fusion protein, a nucleic acid molecule, a protein-nucleic acid complex, an activated CRISPR complex, a vector, a delivery composition of the present application.
[0134] In certain embodiments, the cell is a prokaryotic cell.
[0135] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, e.g., a cell of a non-human primate, a bovine, an ovine, a porcine, a canine, a monkey, a rabbit, a rodent (e.g., a rat or a mouse). In certain embodiments, the cell is a non-mammalian eukaryotic cell, e.g., a cell of an avian bird (e.g., a chicken), a fish, or a crustacean (e.g., a clam, a shrimp). In certain embodiments, the cell is a plant cell, e.g., a cell of a monocotyledonous or dicotyledonous plant or of a cultivated plant or a food crop such as cassava, corn, sorghum, soybean, wheat, oat, or rice, e.g., a cell of an alga, a tree, or a plant producing a fruit or a vegetable (e.g., a tree such as a citrus tree, a nut tree; a solanaceous plant, cotton, tobacco, tomato, grape, coffee, cacao, etc.).
[0136] In certain embodiments, the cell is a stem cell or a stem cell line.
[0137] In certain cases, the host cell of the present application comprises a modification of a gene or genome that is not present in its wild type.
[0138] Methods and uses of gene editing
[0139] The Cas mutant protein, nucleic acid, the above-described composition, the above-described CRISPR / Cas system, the above-described vector system, the above-described delivery composition, or the above-described activated CRISPR complex or the above-described host cell of the present invention can be used for any or more of the following purposes: targeting and / or editing target nucleic acids; cleaving double-stranded DNA, single-stranded DNA, or single-stranded RNA; non-specifically cleaving and / or degrading side-branched nucleic acids; non-specifically cleaving single-stranded nucleic acids; nucleic acid detection; detecting nucleic acids in target samples; specifically editing double-stranded nucleic acids; base editing double-stranded nucleic acids; base editing single-stranded nucleic acids; reducing off-target efficiency during gene editing. In other embodiments, they can also be used to prepare reagents or kits for any or more of the above purposes.
[0140] The present invention also provides the use of the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition, or the above-mentioned activated CRISPR complex in gene editing, gene targeting, or gene cutting; or, in the preparation of reagents or kits for gene editing, gene targeting, or gene cutting.
[0141] In one embodiment, the gene editing, gene targeting, or gene cutting is performed intracellularly and / or extracellularly.
[0142] The present invention also provides a method for editing, targeting, or cleaving a target nucleic acid, the method comprising contacting the target nucleic acid with the aforementioned Cas protein, nucleic acid, the aforementioned composition, the aforementioned CIRSPR / Cas system, the aforementioned vector system, the aforementioned delivery composition, or the aforementioned activated CRISPR complex. In one embodiment, the method comprises editing, targeting, or cleaving the target nucleic acid intracellularly or extracellularly.
[0143] The gene editing or editing of target nucleic acids includes modifying genes, knocking out genes, altering the expression of gene products, repairing mutations, and / or inserting polynucleotides, and gene mutations.
[0144] The editing can be performed in prokaryotic and / or eukaryotic cells.
[0145] On the other hand, the present invention also provides a method for reducing off-target effects / off-target efficiency of gene editing, the method comprising contacting the target nucleic acid with the aforementioned Cas protein, nucleic acid, the aforementioned composition, the aforementioned CIRSPR / Cas system, the aforementioned vector system, the aforementioned delivery composition, or the aforementioned activated CRISPR complex. In one embodiment, the method comprises editing, targeting, or cleaving the target nucleic acid intracellularly or extracellularly.
[0146] In another aspect, the present application also provides use of the above-mentioned Cas protein, nucleic acid, above-mentioned composition, above-mentioned CIRSPR / Cas system, above-mentioned vector system, above-mentioned delivery composition, or above-mentioned activated CRISPR complex in nucleic acid detection, or in preparation of a reagent or kit for nucleic acid detection.
[0147] In another aspect, the present application also provides a method for cleaving single-stranded nucleic acid, comprising contacting a population of nucleic acid with the above-mentioned Cas protein and gRNA, wherein the population of nucleic acid comprises a target nucleic acid and a plurality of non-target single-stranded nucleic acid, and the Cas protein cleaves the plurality of non-target single-stranded nucleic acid.
[0148] The gRNA is capable of binding to the Cas protein.
[0149] The gRNA is capable of targeting the target nucleic acid.
[0150] The contacting can be in vitro, ex vivo, or inside a cell in vivo.
[0151] Preferably, the cleaving of single-stranded nucleic acid is non-specific cleaving.
[0152] In another aspect, the present application also provides use of the above-mentioned Cas protein, nucleic acid, above-mentioned composition, above-mentioned CIRSPR / Cas system, above-mentioned vector system, above-mentioned delivery composition, or above-mentioned activated CRISPR complex in non-specific cleaving of single-stranded nucleic acid, or in preparation of a reagent or kit for non-specific cleaving of single-stranded nucleic acid.
[0153] In another aspect, the present application also provides a kit for gene editing, gene targeting, or gene cleaving, comprising the above-mentioned Cas protein, gRNA, nucleic acid, above-mentioned composition, above-mentioned CIRSPR / Cas system, above-mentioned vector system, above-mentioned delivery composition, above-mentioned activated CRISPR complex, or above-mentioned host cell.
[0154] In another aspect, the present application also provides a kit for reducing off-target effect / off-target efficiency in gene editing, comprising the above-mentioned Cas protein, gRNA, nucleic acid, above-mentioned composition, above-mentioned CIRSPR / Cas system, above-mentioned vector system, above-mentioned delivery composition, above-mentioned activated CRISPR complex, or above-mentioned host cell.
[0155] In another aspect, the present application also provides a kit for detecting a target nucleic acid in a sample, the kit comprising: (a) a Cas protein, or a nucleic acid encoding the Cas protein; (b) a guide RNA, or a nucleic acid encoding the guide RNA, or a precursor RNA comprising the guide RNA, or a nucleic acid encoding the precursor RNA; and (c) a single-stranded nucleic acid detector that is single-stranded and does not hybridize to the guide RNA.
[0156] It is known in the art that the precursor RNA can be cleaved or processed into the above-mentioned mature guide RNA.
[0157] In another aspect, the present application provides use of the above-mentioned Cas protein, nucleic acid, above-mentioned composition, above-mentioned CIRSPR / Cas system, above-mentioned vector system, above-mentioned delivery composition, above-mentioned activated CRISPR complex, or above-mentioned host cell in the manufacture of a medicament or a kit for:
[0158] (i) gene or genome editing;
[0159] (ii) target nucleic acid detection and / or diagnosis;
[0160] (iii) modifying an organism or a non-human organism by editing a target sequence in a target locus;
[0161] (iv) treatment of a disease;
[0162] (iv) targeting a target gene.
[0163] Preferably, the above-mentioned gene or genome editing is performed in or outside a cell.
[0164] Preferably, the above-mentioned target nucleic acid detection and / or diagnosis is performed in vitro.
[0165] Preferably, the above-mentioned treatment of a disease is treatment of a disorder caused by a defect in a target sequence in a target locus.
[0166] In another aspect, the present application provides a method of detecting a target nucleic acid in a sample, the method comprising contacting the sample with a Cas protein, a gRNA (guide RNA) and a single-stranded nucleic acid detector, the gRNA comprising a region that binds to the Cas protein and a guide sequence that hybridizes to the target nucleic acid; detecting a detectable signal generated by cleavage of the single-stranded nucleic acid detector by the Cas protein, thereby detecting the target nucleic acid; the single-stranded nucleic acid detector does not hybridize to the gRNA.
[0167] Methods of specifically modifying target nucleic acids
[0168] In another aspect, the present application also provides a method of specifically modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with the above-mentioned Cas protein, nucleic acid, above-mentioned composition, above-mentioned CIRSPR / Cas system, above-mentioned vector system, above-mentioned delivery composition, or above-mentioned activated CRISPR complex.
[0169] The specific modification can occur in vivo or in vitro.
[0170] The specific modification can occur in vivo or in vitro.
[0171] In some cases, the cell is selected from a prokaryotic cell or a eukaryotic cell, for example, an animal cell, a plant cell, or a microbial cell.
[0172] In one embodiment, the modification refers to a break of the target sequence, such as a single / double strand break of DNA, or a single strand break of RNA.
[0173] In some cases, the method further comprises contacting the target nucleic acid with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of the copy of the donor polynucleotide is integrated into the target nucleic acid.
[0174] In one embodiment, the modification further comprises inserting an editing template (e.g., an exogenous nucleic acid) into the break.
[0175] In one embodiment, the method further comprises: contacting an editing template with the target nucleic acid, or delivering into a cell comprising the target nucleic acid. In this embodiment, the method repairs the broken target gene by homologous recombination with the exogenous template polynucleotide; in some embodiments, the repair results in a mutation, including an insertion, a deletion, or a substitution of one or more nucleotides of the target gene, and in other embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.
[0176] Detection (non-specific cleavage)
[0177] In another aspect, the present application provides a method of detecting a target nucleic acid in a sample, the method comprising: contacting the sample with the above-mentioned Cas protein, nucleic acid, above-mentioned composition, above-mentioned CIRSPR / Cas system, above-mentioned vector system, above-mentioned delivery composition, or above-mentioned activated CRISPR complex and a single-stranded nucleic acid detector; detecting a detectable signal resulting from cleavage of the single-stranded nucleic acid detector by the Cas protein, thereby detecting the target nucleic acid.
[0178] In the present application, the target nucleic acid comprises ribonucleotides or deoxyribonucleotides; and includes single-stranded nucleic acids, double-stranded nucleic acids, such as single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA.
[0179] In one embodiment, the target nucleic acid is derived from a sample of virus, bacteria, microorganism, soil, water source, human, animal, plant, etc. Preferably, the target nucleic acid is the product of PCR, NASBA, RPA, SDA, LAMP, HAD, NEAR, MDA, RCA, LCR, RAM, etc.
[0180] In one embodiment, the target nucleic acid is viral nucleic acid, bacterial nucleic acid, specific nucleic acid associated with disease, such as specific mutation site or SNP site or nucleic acid different from control; preferably, the virus is plant virus or animal virus, for example, papillomavirus, hepadnavirus, herpesvirus, adenovirus, poxvirus, parvovirus, coronavirus; preferably, the virus is coronavirus, preferably, SARS, SARS-CoV2 (COVID-19), HCoV-229E, HCoV-OC43, HCoV-NL63, HCoV-HKU1, Mers-Cov.
[0181] In the present application, the gRNA has at least 50% matching degree with the target sequence on the target nucleic acid, preferably at least 60%, preferably at least 70%, preferably at least 80%, preferably at least 90%.
[0182] In one embodiment, when the target sequence contains one or more characteristic sites (such as specific mutation site or SNP), the characteristic site is completely matched with the gRNA.
[0183] In one embodiment, the detection method can contain one or more gRNAs with different guide sequences targeting different target sequences.
[0184] In the present application, the single-stranded nucleic acid detector includes but is not limited to single-stranded DNA, single-stranded RNA, DNA-RNA hybrid, nucleic acid analog, base modifier, and single-stranded nucleic acid detector containing abasic spacer, etc.; "nucleic acid analog" includes but is not limited to: locked nucleic acid, bridged nucleic acid, morpholino nucleic acid, glycol nucleic acid, hexitol nucleic acid, threose nucleic acid, arabino nucleic acid, 2'oxymethyl RNA, 2'methoxyacetyl RNA, 2'fluoro RNA, 2'amino RNA, 4'sulfur RNA and combinations thereof, including optional ribonucleotide or deoxyribonucleotide residues.
[0185] In the present application, the detectable signal is realized by the following ways: visual detection, sensor-based detection, color detection, fluorescence signal-based detection, gold nanoparticle-based detection, fluorescence polarization, colloidal phase change / dispersion, electrochemical detection and semiconductor-based detection.
[0186] In the present application, preferably, the single-stranded nucleic acid detector is provided with a fluorescent group and a quencher group at its two ends, respectively, and can exhibit a detectable fluorescent signal after the single-stranded nucleic acid detector is cleaved. The fluorescent group is selected from one or any of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red or LC RED460; and the quencher group is selected from one or any of BHQ1, BHQ2, BHQ3, Dabcy1 or Tamra.
[0187] In other embodiments, the 5' end and the 3' end of the single-stranded nucleic acid detector are respectively provided with different labeled molecules, and the colloidal gold detection method is used to detect the colloidal gold test results of the single-stranded nucleic acid detector before and after being cleaved by the Cas protein; the single-stranded nucleic acid detector before and after being cleaved by the Cas protein will exhibit different color development results on the colloidal gold detection line and the quality control line.
[0188] In some embodiments, the method of detecting a target nucleic acid can further comprise comparing the level of the detectable signal to a reference signal level, and determining the amount of the target nucleic acid in the sample based on the level of the detectable signal.
[0189] In some embodiments, the method of detecting a target nucleic acid can further comprise using RNA reporter nucleic acids and DNA reporter nucleic acids (e.g., fluorescent colors) on different channels, and determining the level of the detectable signal by measuring the signal level of the RNA and DNA reporter molecules, and by measuring the amount of the target nucleic acid in the RNA and DNA reporter molecules, based on the combination (e.g., using the minimum or product) of the level of the detectable signal to sample.
[0190] In one embodiment, the target gene is present in a cell.
[0191] In one embodiment, the cell is a prokaryotic cell.
[0192] In one embodiment, the cell is a eukaryotic cell.
[0193] In one embodiment, the cell is an animal cell.
[0194] In one embodiment, the cell is a human cell.
[0195] In one embodiment, the cell is a plant cell, such as a cell of a cultivated plant (e.g., cassava, corn, sorghum, wheat or rice), an alga, a tree or a vegetable.
[0196] In one embodiment, the target gene is present in a nucleic acid molecule (e.g., a plasmid) outside the cell.
[0197] In one embodiment, the target gene is present in a plasmid.
[0198] Definitions of terms
[0199] In the present application, the scientific and technical terms used herein have the meanings commonly understood by one of ordinary skill in the art, unless otherwise indicated. Also, the molecular genetic, nucleic acid chemical, chemical, molecular biological, biochemical, cell culture, microbiological, cell biological, genomic, and recombinant DNA, among other, operational steps described herein are well-known and routine procedures in the corresponding fields. Also, for better understanding of the present application, the definitions and explanations of the relevant terms are provided below.
[0200] Nucleic acid cleavage or cleaving a nucleic acid herein includes: a DNA or RNA break in a target nucleic acid produced by a Cas enzyme described herein (Cis cleavage), a DNA or RNA break in a side branch nucleic acid substrate (single stranded nucleic acid substrate) (i.e. non-specific or non-targeted, Trans cleavage). In some embodiments, the cleavage is a double stranded DNA break. In some embodiments, the cleavage is a single stranded DNA break or a single stranded RNA break.
[0201] CRISPR system
[0202] As used herein, the term "Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meaning commonly understood by one of ordinary skill in the art, which generally includes a transcription product or other element related to the expression of a CRISPR-associated ("Cas") gene, or a transcription product or other element capable of directing the activity of the Cas gene.
[0203] CRISPR / Cas complex
[0204] As used herein, the term "CRISPR / Cas complex" refers to a complex formed by the binding of a guide RNA or mature crRNA to a Cas protein, which includes a direct repeat sequence that is hybridized to a guide sequence of a target sequence and bound to a Cas protein, the complex being capable of recognizing and cleaving a polynucleotide that can hybridize to the guide RNA or mature crRNA.
[0205] Guide RNA (guide RNA, gRNA)
[0206] As used herein, the terms "guide RNA" (gRNA), "mature crRNA," "guide sequence" are used interchangeably and have the meaning generally understood by one of skill in the art. Generally, a guide RNA can comprise or essentially consist of or consist of a direct repeat sequence and a guide sequence.
[0207] In certain instances, a guide sequence is any polynucleotide sequence that has sufficient complementarity to a target sequence to hybridize to the target sequence and direct specific binding of a CRISPR / Cas complex to the target sequence. In one embodiment, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned, is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining optimal alignment is within the capabilities of a person of ordinary skill in the art. For example, there are published and commercially available alignment algorithms and programs such as, but not limited to, ClustalW, Smith-Waterman in matlab, Bowtie, Geneious, Biopython, and SeqMan.
[0208] Target sequence
[0209] A "target sequence" refers to a polynucleotide targeted by a guide sequence in a gRNA, e.g., a sequence that has complementarity to the guide sequence, wherein hybridization between the target sequence and the guide sequence will facilitate formation of a CRISPR / Cas complex (including a Cas protein and a gRNA). Perfect complementarity is not required, so long as there is sufficient complementarity to cause hybridization and facilitate formation of a CRISPR / Cas complex.
[0210] A target sequence can comprise any polynucleotide, such as DNA or RNA. In certain instances, the target sequence is located within a cell or outside of a cell. In certain instances, the target sequence is located in the nucleus or cytoplasm of a cell. In certain instances, the target sequence can be located within an organelle of a eukaryotic cell, e.g., a mitochondrion or a chloroplast. A sequence or template that can be used for recombination into a target locus comprising the target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence." In one embodiment, the editing template is an exogenous nucleic acid. In one embodiment, the recombination is homologous recombination.
[0211] In the present invention, the "target sequence" or "target polynucleotide" or "target nucleic acid" can be any endogenous or exogenous polynucleotide to a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).
[0212] Off-target effects
[0213] As used herein, the term "off-target effects" or "off-target efficiency" refers to the phenomenon that, when using CRISPR / Cas technology system to edit genes, the designed specific guide RNA (gRNA) recognizes and locates the target site, and the mismatch between DNA and gRNA also allows the Cas-gRNA complex to recognize and cut the genomic sequence similar to the target site, resulting in off-target phenomenon, i.e., if the gRNA is completely matched with the target site and successfully edited, it is called precise positioning, and if the gRNA forms a mismatch with the non-target DNA sequence, it will cause unexpected mutations in the gene site, which is called off-target effect.
[0214] There are two forms of local mismatch of gRNA at non-target sites: (1) the off-target DNA sequence only has a base mismatch with the gRNA, and the length is consistent; (2) the off-target DNA sequence has a base mismatch with the gRNA by inserting or deleting DNA sequences to make the length different. Studies have shown that off-target not only reduces the efficiency of genome editing, but also causes chromosomal rearrangement and even destroys the imperfectly matched genes. In addition, it may also cause functional genes to lose activity.
[0215] Single-stranded nucleic acid detector
[0216] The single-stranded nucleic acid detector described in the present invention refers to a sequence containing 2-200 nucleotides, preferably 2-150 nucleotides, preferably 3-100 nucleotides, preferably 3-30 nucleotides, preferably 4-20 nucleotides, and more preferably 5-15 nucleotides. Preferably, it is a single-stranded DNA molecule, a single-stranded RNA molecule or a single-stranded DNA-RNA hybrid.
[0217] The single-stranded nucleic acid detector described at both ends includes different reporter groups or labeled molecules, which do not present a reporter signal when it is in the initial state (i.e., the uncut state), and presents a detectable signal after being cut, i.e., the detectable difference between before and after cutting.
[0218] In one embodiment, the reporter group or label molecule comprises a fluorescent group selected from one or any of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red or LC RED460; and a quencher group selected from one or any of BHQ1, BHQ2, BHQ3, Dabcy1 or Tamra.
[0219] In one embodiment, the single stranded nucleic acid detector has a first molecule (e.g. FAM or FITC) attached to the 5' end and a second molecule (e.g. biotin) attached to the 3' end. The reaction system containing the single stranded nucleic acid detector is used in conjunction with a flow strip to detect the target nucleic acid (preferably, a colloidal gold detection method). The flow strip is designed to have two capture lines, with an antibody to the first molecule (i.e. first molecule antibody) at the sample contact end (colloidal gold), an antibody to the first molecule antibody at the first line (control line), and an antibody to the second molecule (i.e. second molecule antibody, e.g. avidin) at the second line (test line). As the reaction flows along the strip, the first molecule antibody binds to the first molecule carrying the cleaved or uncleaved oligonucleotide to the capture lines, the cleaved reporter will bind to the antibody to the first molecule antibody at the first capture line, and the uncleaved reporter will bind to the second molecule antibody at the second capture line. The binding of the reporter group at each line will result in a strong readout / signal (e.g. color). As more reporter is cleaved, more signal will accumulate at the first capture line, and less signal will appear at the second line. In certain aspects, the present application relates to the use of a flow strip as described herein for detecting a nucleic acid. In certain aspects, the present application relates to a method of detecting a nucleic acid using a flow strip as defined herein, e.g. a (lateral) flow test or a (lateral) flow immuno-chromatographic assay. In certain aspects, the molecules in the single stranded nucleic acid detector can be replaced by each other, or the position of the molecules can be changed, as long as the reporting principle is the same or similar to the present application, and the improved ways are also included in the present application.
[0220] The detection method described in the present application can be used for quantitative detection of the target nucleic acid to be detected. The quantitative detection index can be quantified according to the signal strength of the reporter group, such as the luminescence intensity of the fluorescent group, or the width of the color development strip, etc.
[0221] Wild type
[0222] As used herein, the term "wild type" has the meaning generally understood by those skilled in the art, which means the typical form of an organism, strain, gene or characteristic that distinguishes it from a mutant or variant form when it occurs in nature, which can be isolated from a source in nature and has not been intentionally modified by man.
[0223] Derivatization
[0224] As used herein, the term "derivatized" refers to a chemical modification of an amino acid, polypeptide, or protein, in which one or more substituents have been covalently attached to the amino acid, polypeptide, or protein. The substituents can also be referred to as side chains.
[0225] A derivatized protein is a derivative of the protein, typically, derivatization of a protein does not adversely affect the desired activity of the protein (e.g., the activity of binding to a guide RNA, the activity of an endonuclease, the activity of binding to and cleaving a target sequence at a specific site under the direction of a guide RNA), that is, the derivative of the protein has the same activity as the protein.
[0226] Derivatized protein
[0227] Also referred to as "protein derivatives," refers to modified forms of a protein, e.g., in which one or more amino acids of the protein can be deleted, inserted, modified, and / or substituted.
[0228] Non-naturally occurring
[0229] As used herein, the terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of humans. When these terms are used to describe nucleic acid molecules or polypeptides, they indicate that the nucleic acid molecule or polypeptide is at least substantially isolated from at least another component with which it is associated in nature or as found in nature.
[0230] Orthologue
[0231] As used herein, the term "orthologue" has the meaning generally understood by those skilled in the art. As a further guide, an "orthologue" of a protein as described herein refers to a protein belonging to a different species that performs the same or a similar function as the protein of which it is an orthologue.
[0232] Identity
[0233] As used herein, the term "identity" is used in reference to the matching of sequences between two polypeptides or between two nucleic acids. When a position in each of two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percentage of identity" between two sequences is a function of the number of matching positions shared by the sequences divided by the number of positions compared x 100. For example, if 6 of 10 positions in two sequences are matched then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (3 of 6 positions are matched). Typically, the comparison is made over the length of the sequences being compared, after aligning the two sequences to produce maximum identity. Such alignment can be achieved by methods such as that of Needleman et al. (1970) J. Mol. Biol. 48:443-453, which can be conveniently performed by computer program such as the Align program (DNAstar, Inc.). Percentage identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl. Biosci., 4:11-17 (1988)) as integrated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. In addition, percentage identity between two amino acid sequences can be determined using the algorithm of Needleman and Wunsch (J MoI Biol. 48:444-453 (1970)) as implemented in the GAP program in the GCG software package (available at www.gcg.com), using either a Blossum 62 matrix or a PAM250 matrix, and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a gap length weight of 1, 2, 3, 4, 5, or 6.
[0234] Vector
[0235] The term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules that comprise one or more free ends, no free ends (e.g., circular), nucleic acid molecules that comprise DNA, RNA, or both; and other varieties of nucleic acids known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection, and the resulting genetically modified host cell can be used to express the genetic material elements carried by the vector. A vector can be introduced into a host cell to thereby produce a transcript, protein, or peptide, including a protein, fusion protein, isolated nucleic acid molecule, etc. (e.g., a CRISPR transcript, such as a nucleic acid transcript, protein, or enzyme) as described herein. A vector can contain a variety of control elements, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, a vector can contain a replication origin.
[0236] One type of vector is a "plasmid," which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques.
[0237] Another type of vector is a viral vector, wherein virally-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by a virus for transfection into a host cell. Certain vectors (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors) are capable of autonomous replication in a host cell into which they are introduced.
[0238] Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors."
[0239] Host cell
[0240] As used herein, the term "host cell" refers to a cell that can be used to introduce a vector, including but not limited to, prokaryotic cells such as E. coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.
[0241] Those of skill in the art will appreciate that the design of the expression vector can depend on such factors as the choice of the host cell to be transformed, the level of expression of the desired gene, etc.
[0242] Regulatory element
[0243] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences), for which detailed description can be found in Goeddel, *Gene Expression Technology: Methods in Enzymology*, 185, Academic Press, San Diego, California (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In some cases, regulatory elements can also be directed to express in a time-dependent manner (such as in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue or cell type specific. In some cases, the term "regulatory element" covers enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).
[0244] Promoter
[0245] As used herein, the term "promoter" has its art-understood meaning and refers to a non-coding nucleotide sequence located upstream of a gene that initiates expression of the downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell essentially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell essentially only when the cell is of the tissue type to which the promoter corresponds.
[0246] NLS
[0247] A "nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" a protein for import into the nucleus by nuclear transport, i.e., a protein with an NLS is transported to the nucleus. Typically, an NLS comprises positively charged Lys or Arg residues that are exposed on the surface of the protein. Exemplary nuclear localization sequences include, but are not limited to, NLS from SV40 large T antigen, EGL-13, c-Myc, and TUS protein. In some embodiments, the NLS comprises a PKKKRKV sequence. In some embodiments, the NLS comprises an AVKRPAATKKAGQAKKKKLD sequence. In some embodiments, the NLS comprises a PAAKRVKLD sequence. In some embodiments, the NLS comprises a MSRRRKANPTKLSENAKKLAKEVEN sequence. In some embodiments, the NLS comprises a KLKIKRPVK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP Al, the sequence KIPIK in the yeast transcriptional repressor Matα2, and PY-NLS.
[0248] Operably linked
[0249] As used herein, the term "operably linked" is intended to mean that a nucleotide sequence of interest is linked to the one or more regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or, when the vector is introduced into a host cell, in the host cell).
[0250] Complementarity
[0251] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence by virtue of traditional Watson-Crick or other non-traditional types of bonding. Percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 is 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Perfect complementarity" indicates that all consecutive residues of a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. "Substantially complementary" as used herein refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to hybridization under stringent conditions of two nucleic acids.
[0252] Stringent conditions
[0253] As used herein, "stringent conditions" for hybridization refer to conditions under which a nucleic acid having complementarity to a target sequence will hybridize primarily to that target sequence and not to non-target sequences. Stringent conditions are often sequence dependent, and are varied depending on many factors. In general, the longer the sequence, the higher the temperature at which the sequence will specifically hybridize to its target sequence.
[0254] Hybridization
[0255] The terms "hybridize" or "complementary" or "substantially complementary" refer to the non-covalent binding of nucleotides of a nucleic acid (e.g., RNA, DNA) to another nucleic acid by means of complementary base pairing and / or G / U base pairing, "annealing" or "hybridizing" in a sequence-specific, anti-parallel fashion (i.e., nucleic acid specifically binds to complementary nucleic acid).
[0256] Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases can occur. Suitable conditions for hybridization between two nucleic acids depend on the length and complementarity of the nucleic acids, which are variables known in the art. Typically, the length of the hybridizable nucleic acid is 8 nucleotides or more (e.g., 10 nucleotides or more, 12 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 22 nucleotides or more, 25 nucleotides or more, or 30 nucleotides or more).
[0257] It is understood that the sequence of a polynucleotide need not be 100% complementary to the sequence of its target nucleic acid to hybridize specifically thereto. A polynucleotide can comprise 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% complementarity to the sequence of the target region in the sequence of the target nucleic acid to which it hybridizes.
[0258] Hybridization of a target sequence to a gRNA means that at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize, forming a complex; or that at least 12, 15, 16, 17, 18, 19, 20, 21, 22, or more bases of the nucleic acid sequences of the target sequence and the gRNA can base pair, hybridizing to form a complex.
[0259] Expression
[0260] As used herein, the term "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., transcribed into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. Transcripts and encoded polypeptides can be collectively referred to as "gene product." If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell.
[0261] Linker
[0262] As used herein, the term "linker" refers to a linear polypeptide formed by the linkage of a plurality of amino acid residues via peptide bonds. The linkers of the present application can be artificial synthetic amino acid sequences, or naturally occurring polypeptide sequences, such as polypeptides having a hinge region function. Such linker polypeptides are well known in the art (see, e.g., Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, R.J. et al. (1994) Structure 2:1121-1123).
[0263] Treatment
[0264] As used herein, the term "treatment" refers to the treatment or cure of a disorder, the delay of onset of symptoms of a disorder, and / or the delay of progression of a disorder.
[0265] Subject
[0266] As used herein, the term "subject" includes, but is not limited to, various animals, plants, and microorganisms.
[0267] Animal
[0268] For example, a mammal, for example, a bovine, equine, ovine, porcine, canine, feline, leporid, rodent (e.g., a mouse or rat), non-human primate (e.g., a macaque or cynomolgus monkey), or a human. In certain embodiments, the subject (e.g., a human) has a disorder (e.g., a disorder resulting from a disease-associated gene defect).
[0269] Plant
[0270] The term "plant" is to be understood as any differentiated multicellular organism capable of performing photosynthesis, including crop plants, in particular monocotyledonous or dicotyledonous plants, at any stage of maturation or development, vegetable crops, including artichokes, Brussels sprouts, cress, leeks, asparagus, lettuce (e.g. head lettuce, leaf lettuce, long-leaf lettuce), bok choy, yellow fleshed yam, melons (e.g. muskmelons, watermelons, crenshaw melons, white cantaloupes, Roman melons), oilseed crops (e.g. Brussels sprouts, cabbage, cauliflower, broccoli, collard greens, curly kale, Chinese cabbage, bok choy), artichokes, carrots, napa, okra, onions, celery, parsley, chickpeas, parsnips, chicory, peppers, potatoes, gourds (e.g. zucchini, cucumbers, small zucchini, squash, pumpkin), radishes, dry bulb onions, rutabaga, eggplant (also known as aubergine), burdock, endive, green onions, endive, garlic, spinach, green onions, squash, greens, sugar beets (sugar beet and fodder beet), sweet potatoes, Swiss chard, wasabi, tomatoes, turnips, and spices; fruits and / or vine crops, such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, aronia, almonds, chestnuts, hazelnuts, pecans, pistachios, walnuts, citrus, blueberries, boysenberries, cranberries, currants, goji berries, raspberries, strawberries, blackberries, grapes, avocados, bananas, kiwis, persimmons, pomegranates, pineapples, tropical fruits, pomes, melons, mangoes, papayas, and lychees; field crops, such as clover, alfalfa, milkweed, meadow grass, corn / maize (fodder corn, sweet corn, popcorn), hops, jojoba, peanuts, rice, safflower, small grain crops (barley, oats, rye, wheat, etc.), sorghum, tobacco, kapok, legumes (beans, lentils, peas, soybeans), oil plants (rape, mustard, poppy, olive, sunflower, coconut, castor oil plant, cocoa, groundnuts), Arabidopsis, fiber plants (cotton, flax, hemp, jute), lauraceae (cinnamon, camphor), or a plant such as coffee, sugar cane, tea, and natural rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants, and trees such as forests (broad-leaved trees and evergreens, such as conifers), fruit trees, ornamental trees, and nut-bearing trees, and shrubs and other young plants.
[0271] Beneficial effects of the invention
[0272] The present application reduces the off-target effect of Cas12i3 protein by mutation, which has a wide application prospect.
[0273] Embodiments of the present application will be described in detail below with reference to the attached drawings and examples, but those skilled in the art will understand that the following drawings and examples are only used to illustrate the present application, and are not a limitation on the scope of the present application. According to the following detailed description of the drawings and preferred embodiments, various objects and advantages of the present application will become apparent to those skilled in the art. BRIEF DESCRIPTION OF DRAWINGS
[0274] Figure 1 Detection of off-target effects of Cas proteins at different target sites.
[0275] Figure 2 Structural diagram of Cas12i3Max protein.
[0276] Figure 3 Verification of editing efficiency of Cas proteins with mutations at different amino acid sites.
[0277] Figure 4 Testing of off-target effects of Cas proteins with mutations at different amino acid sites.
[0278] Figure 5 Testing of mismatch tolerance of different mutant proteins at T8 site.
[0279] Figure 6 Detection of targeting editing efficiency and off-target efficiency of Cas12i3Max protein and AP44 protein.
[0280] Figure 7 Testing of off-target effects of Cas proteins with amino acid site combination mutations. DETAILED DESCRIPTION
[0281] The following examples merely illustrate the application and are not intended to limit the application in any way. The experiments and methods described in the examples were performed essentially as described in the art and as described in various references unless otherwise indicated. For example, the general techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, among others, used in the present application can be found in Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd ed. (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F. M. Ausubel et al. eds., (1987)); the series METHODS IN ENZYMOLOGY (Academic Press, Inc.): PCR 2: A PRACTICAL APPROACH (M. J. MacPherson, B. D. Hames, and G. R. Taylor eds. (1995), Harlow and Lane, ANTIBODIES, A LABORATORY MANUAL, (1988), and ANIMAL CELL CULTURE (R. I. Freshney, ed. (1987)).
[0282] In addition, where particular conditions are not specified in the examples, those conditions were performed under routine conditions or as suggested by the manufacturer. Where the manufacturer of reagents or instruments is not indicated, it is intended that the reagents or instruments used were of a routine variety available from a commercial vendor. Those skilled in the art will recognize that the examples describe the present application in terms of preferred embodiments and are not intended to limit the scope of the application as claimed. All publications and other references mentioned herein are incorporated by reference in their entirety.
[0283] Example 1. Obtaining Cas mutant proteins
[0284] For the known Cas protein (Cas12f.4 in CN111757889B, which is referred to as Cas12i3 in the present embodiment), the applicant predicts key amino acid sites that can affect its biological function through bioinformatics, and mutates the amino acid sites to obtain a Cas mutant protein with improved editing activity. Specifically, the Cas12i3 coding sequence is codon-optimized and synthesized, the amino acid sequence of the wild-type Cas12i3 is shown as SEQ ID No. 1, and the nucleic acid sequence is shown as SEQ ID No. 2. The potential amino acids of Cas12i3 combined with the target sequence are subjected to site-directed mutagenesis by bioinformatics methods.
[0285] SEQ ID No. 1:
[0286]
[0287] SEQ ID No. 2:
[0288]
[0289] The variants of Cas protein were generated by PCR-based site-directed mutagenesis. The specific method is to design the DNA sequence of Cas12i3 protein into two parts with the mutation site as the center, design two pairs of primers to amplify the two parts of DNA sequence, and introduce the sequence to be mutated on the primer, and finally load the two fragments into pcDNA3.3-eGFP vector by Gibson cloning. The combination of mutants is achieved by splitting the DNA of Cas12i3 protein into multiple segments, using PCR, Gibson clone to construct. Fragment amplification kit: TransStart FastPfu DNA Polymerase (containing 2.5 mM dNTPs), see the instruction for specific experimental procedures. Gel recovery kit: Gel DNA Extraction MiniKit, see the instruction for specific experimental procedures. Kit for vector construction: pEASY-Basic Seamless Cloning and Assembly Kit (CU201-03), see the instruction for specific experimental procedures.
[0290] After screening, Cas12i3Max (also known as Cas-SF01) with significantly improved editing activity was obtained, which has mutations of R at positions 7, 233, 267, 369, and 433 from the N terminus, respectively, relative to the sequence shown in SEQ ID No. 1 (the Cas12i3Max is also recorded in the Chinese invention patent application, patent application number: 202310088437.4), and the amino acid sequence of Cas12i3Max is shown in SEQ ID No. 3.
[0291] SEQ ID No. 3:
[0292]
[0293] Example 2. Verification of off-target effects of Cas mutant proteins
[0294] A fluorescent reporter system suitable for verifying Cas12i3 cleavage was constructed with reference to the literature (Yang, Yi, et al. "Highly efficient and rapid detection of the cleavage activity of Cas9 / gRNA via a fluorescent reporter." Applied biochemistry and biotechnology 180.4 (2016): 655-667.), and target points T5, T6, and T8 were designed in GFFP. A mismatched crRNA vector was constructed with a 2bp base window, and the off-target effect of Cas12i3 Max was verified. The mutation principle was A→C, T→A, C→G, and G→A. The crRNA containing 2bp mutations was annealed and ligated to the BbsI site to obtain the pCas12i3Max_GFFP_crRNA_Mxx plasmid containing GFFP and mCherry fluorescent tags. Plating: HEK293T cells were plated to 80% confluence and plated into 24-well plates. Mixing: 1.5ug of plasmid was added to 100ul of opti-MEM and mixed. The diluted Hieff Trans TM Lipofectamine 3000 was mixed with the diluted plasmid and incubated at room temperature for 15 min. The incubated mixture was added to the cell culture medium in the plated cells for transfection. After 24 hours of transfection, each transfection well was replaced with 1000ul of DMEM medium. Then, the cells were incubated for another 24 hours before collecting the cells for flow cytometry detection. After removing the culture medium, the transfection cells in each well were digested with 300ul of 0.25% trypsin. The digestion was terminated by adding 300ul of medium, and after centrifugation at 100rpm for 3 min, the cells were resuspended in 500ul of medium. The cells were then pipetted into 5ml round-bottom polystyrene test tubes with cell filter caps and refrigerated. Cells with GFP signals were sorted using a flow cytometer (FACS), and GFFP did not emit light. After editing, the recombinant GFP emitted fluorescence. Y610 (561 / 610nm) was used to detect mCherry fluorescence, and B525 (488 / 525nm) was used to detect GFP fluorescence. The editing efficiency was calculated by dividing the number of mCherry and GFP cells by the number of GFP cells.
[0295] The results of the designed target points and off-target effect detection are shown in Table 1. Figure 1As shown, T5, T6, and T8 target sites are designed at the corresponding positions of GFFP, the underlined region is the PAM site, Mis12 represents the 1st and 2nd mismatch sites at the PAM end, Mis34 represents the 3rd and 4th mismatch sites at the PAM end, Mis56 represents the 5th and 6th mismatch sites at the PAM end, Mis78 represents the 7th and 8th mismatch sites at the PAM end, Mis910 represents the 9th and 10th mismatch sites at the PAM end, Mis1112 represents the 11th and 12th mismatch sites at the PAM end, Mis1314 represents the 13th and 14th mismatch sites at the PAM end, Mis1516 represents the 15th and 16th mismatch sites at the PAM end, Mis1718 represents the 17th and 18th mismatch sites at the PAM end, Mis1920 represents the 19th and 20th mismatch sites at the PAM end, the right side from 0-1 represents the off-target effect, the darker the color, the higher the off-target effect of the target site, and the left side from 0-1 represents the editing efficiency, the darker the color, the higher the editing efficiency of the target site. Figure 1 As shown, it can be seen that Cas12i3Max has obvious off-target effect at different mismatch target sites.
[0296] The Cas12i3Max protein is further optimized to improve the editing activity while reducing the off-target effect.
[0297] Example 3. Screening of mutation sites to reduce off-target effects of Cas proteins
[0298] Using the Cas12i3Max protein structure diagram in Figure 2 , the predicted amino acid sites are mutated using a rapid mutation kit, a total of 80 sites are obtained, and the mutant fragments are cloned and inserted into the Cas12i3Max-T5 / T8 vector by Afe I and Pst I restriction endonuclease. The editing efficiency of the 80 mutants is screened, and the screening method is referred to Example 2. The GFFP fluorescent reporter system is used at the T5 target site, and sites with editing efficiency reduced by no more than 3.5% are selected.
[0299] The screening results are shown in Figure 3 , the single amino acid mutation sites are numbered as AP01-AP80 according to the amino acid sequence order, a total of 16 sites with editing efficiency not lower than 3.5% compared with the control (Cas12i3Max in the figure) are screened, which are AP5, AP6, AP20, AP27, AP31, AP35, AP37, AP39, AP44, AP48, AP49, AP54, AP58, AP70, AP74, AP76, and the rest of the mutants have obvious decrease in editing activity. The above 16 mutants are verified by off-target experiment.
[0300] The target points T8, T8mis78 and T8mis910 were selected for off-target effect experiment detection, and the detection results are shown in Figure 4 As shown in the figure, T8 represents the editing efficiency at the target point T8, T8mis78 represents the editing efficiency at the target point T8 with mismatch at the 7th and 8th positions at the PAM end, and T8mis910 represents the editing efficiency at the target point T8 with mismatch at the 9th and 10th positions at the PAM end. The results show that in AP44 (based on Cas12i3Max, corresponding to the 876th amino acid of SEQ ID No. 1 is mutated to R; that is, the 876th amino acid of the sequence shown in SEQ ID No. 3 is mutated to R) and AP58 (based on Cas12i3Max, corresponding to the 890th amino acid of SEQ ID No. 1 is mutated to R; that is, the 890th amino acid of the sequence shown in SEQ ID No. 3 is mutated to R), the editing efficiency is significantly reduced in the case of mismatch, that is, the mutation of the 876th and 890th amino acid sites can significantly reduce the mismatch tolerance, indicating that they can reduce the off-target effect of the Cas protein, while not affecting the editing efficiency.
[0301] Example 4. Off-target effects and editing efficiency test experiments of mutant proteins
[0302] The off-target effect and editing efficiency of the AP44 protein (also known as Cas-SF01HiFi) screened in the above embodiment were tested, and the mismatch sites were designed at the target point T8 of GFFP in Reference Example 2. The results are shown in Figure 5 As shown in the figure, M12 represents that the gRNA is designed to have mismatch sites at the 1st and 2nd positions at the PAM end, M34 represents that the gRNA is designed to have mismatch sites at the 3rd and 4th positions at the PAM end, M56 represents that the gRNA is designed to have mismatch sites at the 5th and 6th positions at the PAM end, M78 represents that the gRNA is designed to have mismatch sites at the 7th and 8th positions at the PAM end, M910 represents that the gRNA is designed to have mismatch sites at the 9th and 10th positions at the PAM end, M1112 represents that the gRNA is designed to have mismatch sites at the 11th and 12th positions at the PAM end, M1314 represents that the gRNA is designed to have mismatch sites at the 13th and 14th positions at the PAM end, M1516 represents that the gRNA is designed to have mismatch sites at the 15th and 16th positions at the PAM end, M1718 represents that the gRNA is designed to have mismatch sites at the 17th and 18th positions at the PAM end, M1920 represents that the gRNA is designed to have mismatch sites at the 19th and 20th positions at the PAM end, Cas-SF01 represents the mutant protein Cas12i3Max, and Cas-SF01HiFi represents the mutant protein AP44, according to Figure 5As shown, the editing efficiency of the mutant protein AP44 with 3-16 mismatches at the T8 site is lower than that of the mutant protein Cas12i3Max, i.e., the mismatch tolerance of AP44 with 3-16 mismatches at the T8 site is lower than that of Cas12i3Max, indicating that the mutation of the amino acid at position 876 can significantly reduce the mismatch tolerance and reduce the off-target effect of the Cas protein.
[0303] The target editing efficiency and off-target efficiency of Cas12i3Max and AP44 were compared by designing a target in the TTR gene, TTR-target11 Figure 6 as shown in ON in the middle) and the off-target sites detected by GUIDE-seq Figure 6 as shown in OFF in the middle), the sequence information of the TTR-target11 target site is: 5'-ATTTGTAGAAGGGATATACAAAGT-3', and the sequence information of the off-target site is: 5'-ATTTGTAGAAGGAATATACAAGTA-3', and the italic part is the PAM sequence; the detection results are as shown in Figure 6 Each dot represents 2 repeats, and both the mutant protein Cas-SF01HiFi (i.e., the mutant protein AP44) and Cas-SF01 (i.e., the mutant protein Cas12i3Max) exhibit high target editing efficiency at the target site, but the editing efficiency of Cas-SF01HiFi (i.e., the mutant protein AP44) at the off-target site is significantly reduced.
[0304] Further study of the target for reducing the off-target effect of Cas12i3Max, the mutant protein AP44 and AP588 were combined (both D876R and A890R corresponding to the protein sequence position of SEQ ID No. 1 were mutated based on Cas12i3Max) to obtain AP44+AP58 (i.e., the 876th and 890th amino acids of the sequence shown in SEQ ID No. 3 were both mutated to R), which was verified for off-target effect at the target T8 mis78, and the results are as shown in Figure 7 As shown, the editing efficiency of AP44, AP58, and AP44+AP58 is significantly reduced compared with the control Cas12i3Max (I3max in the figure), and the editing efficiency of the combined off-target site of AP44+AP58 is more significantly reduced; this further indicates that the combined mutation of the 876th and 890th amino acids of the sequence shown in SEQ ID No. 1 from the N terminus can further reduce the mismatch tolerance and further reduce the off-target effect of the Cas protein.
[0305] Meanwhile, the 876th and / or 890th sites of the wild-type Cas12i3 (SEQ ID No. 1) were mutated to verify the off-target effect. The 876th and 890th sites of the wild-type Cas12i3 (amino acid sequence shown as SEQ ID No. 1) were respectively mutated and combinedly mutated, and the off-target effect experiment was detected at the target points T8, T8mis78 and T8mis910 according to the above experimental method, and the results showed that when the wild-type Cas2i3 had mutations at the 876th or 890th sites, i.e. D876R (corresponding to the mutation of the 876th site of the protein sequence of SEQ ID No. 1 to R) or A890R (corresponding to the mutation of the 890th site of the protein sequence of SEQ ID No. 1 to R), the editing efficiency was obviously reduced when the gRNA was mismatched with the target point, and when the Cas12i3 had D876R-A890R (corresponding to the simultaneous mutation of the 876th and 890th sites of the protein sequence of SEQ ID No. 1 to R), the editing efficiency was further reduced when the gRNA was mismatched with the target point.
[0306] Therefore, it can be seen that the mutation of the 876th and / or 890th sites of the Cas12i3 protein can significantly reduce the off-target effect of the Cas protein caused by the mismatch of the gRNA.
[0307] Although the specific embodiments of the present application have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details according to all the teachings disclosed herein, and these changes are within the scope of protection of the present application. The entire scope of the present application is given by the appended claims and any equivalents thereof.
Claims
1. A Cas mutant protein, wherein the mutant protein has mutations at amino acid positions corresponding to positions 876 and / or 890 of the amino acid sequence set forth in SEQ ID No. 1, wherein the amino acid at each of the positions 876 and 890 is mutated to R, as compared to the amino acid sequence of a parent Cas protein. The Cas mutant protein further has mutations at amino acid positions corresponding to positions 7, 233, 267, 369, and 433 of the amino acid sequence set forth in SEQ ID No. 1, wherein the amino acid at each of the positions 7, 233, 267, 369, and 433 is mutated to R, as compared to the amino acid sequence of the parent Cas protein. The parent Cas protein is Cas12i3.
2. A fusion protein comprising the Cas mutant protein of claim 1 and a further modification moiety, wherein the modification moiety is a detectable label.
3. An isolated polynucleotide, comprising, The polynucleotide is a polynucleotide sequence encoding the Cas mutant protein of claim 1, or a polynucleotide sequence encoding the fusion protein of claim 2.
4. A vector, characterized by, The vector comprises the polynucleotide of claim 3 and regulatory elements operably linked thereto.
5. A CRISPR-Cas system, characterized in that, The system comprises the Cas mutant protein of claim 1 and at least one gRNA, wherein the gRNA is capable of binding to the Cas mutant protein of claim 1.
6. A composition characterized in that, The composition comprises: (i) a protein component selected from the group consisting of: the Cas mutant protein of claim 1 or the fusion protein of claim 2; (ii) a nucleic acid component, which is a gRNA, wherein the gRNA comprises a direct repeat sequence capable of binding to the Cas mutant protein of claim 1 and a guide sequence capable of targeting a target sequence; The protein component and the nucleic acid component are capable of binding to each other to form a complex.
7. An engineered host cell, which is a non-motile, plant cell, characterized in that, The host cell comprises the Cas mutant protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6.
8. Use of the Cas mutant protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7, in any one or more of the following applications selected from the group consisting of: gene editing, gene targeting, gene cleavage, cleaving double-stranded DNA, single-stranded DNA, single-stranded RNA, non-specific cleavage and / or degradation of branched nucleic acids, non-specific cleavage of single-stranded nucleic acids, nucleic acid detection, specific editing of double-stranded nucleic acids, base editing of double-stranded nucleic acids, base editing of single-stranded nucleic acids, reducing off-target efficiency of gene editing; or, in the manufacture of a kit for the applications, wherein the applications are selected from the group consisting of: Gene or genome editing, target nucleic acid detection and / or diagnosis, editing target sequence in target locus to modify organism, treatment of disease, targeting target gene, cleaving target gene.
9. A method for editing target nucleic acid, targeting target nucleic acid or cleaving target nucleic acid or reducing off-target efficiency in gene editing, the method is a method for non-disease diagnosis and treatment purposes, the method comprises contacting target nucleic acid with the Cas mutant protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7.
10. A kit for gene editing, gene targeting or gene cleavage or reducing off-target efficiency in gene editing, the kit comprises the Cas mutant protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7.
Citation Information
Patent Citations
Novel CRISPR / Cas12f enzymes and systems
CN111757889B
Cas protein with improved editing activity and application thereof
CN116004573A
TARGETED VIRAL-MEDIATED PLANT GENOME EDITING USING CRISPR / Cas9
US20170114351A1