CRISPR / Cas effector protein HT001, system and application

By optimizing the Cas protein and sgRNA molecules in the CRISPR/Cas complex, the problems of low editing efficiency and off-target effects in plants are solved, and more efficient and accurate gene editing effects are achieved.

CN118773168BActive Publication Date: 2025-07-25HUAZHI RICE BIO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410811882.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-06-21
Publication Date
2025-07-25
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

The existing CRISPR/Cas editing system has off-target effects, low editing efficiency and PAM motif limitations in plant gene editing, which affects its accuracy, flexibility and safety, hinders its widespread application in plants.

Method used

A novel CRISPR/Cas complex is provided, containing optimized Cas effector proteins and sgRNA molecules, which improves editing efficiency and exhibits efficient nuclease activity in bacteria, plants and animals by modifying the amino acid sequence and guide RNA structure of Cas enzymes.

Benefits of technology

It improves the gene editing efficiency of the CRISPR/Cas system in plants, enhances the accuracy and flexibility of editing, and expands its application scope.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GDA0005239428800000321
    Figure GDA0005239428800000321
  • Figure GDA0005239428800000331
    Figure GDA0005239428800000331
  • Figure GDA0005239428800000332
    Figure GDA0005239428800000332
Patent Text Reader

Abstract

The present invention provides a CRISPR / Cas effector protein and system. The present invention provides a protein, its ortholog, homolog, variant or functional fragment, wherein the protein is selected from: (a1) a protein containing the amino acid sequence shown in SEQ ID NO: 1; (a2) a protein containing an amino acid sequence with one or more amino acid substitutions, deletions or additions compared with the amino acid sequence shown in SEQ ID NO: 1, and this protein retains the biological function of the protein shown in SEQ ID NO: 1; or (a3) a protein containing an amino acid sequence having at least 80% sequence identity with the amino acid sequence shown in SEQ ID NO: 1, and this protein retains the biological function of the protein shown in SEQ ID NO: 1; the ortholog, homolog, variant or functional fragment retains the biological function of the protein shown in SEQ ID NO: 1.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of nucleic acid editing, and more specifically, the present invention relates to CRISPR / Cas effector proteins and systems. Background Art

[0002] Clustered regularly interspaced short palindromic repeats (CRISPR-associated, CRISPR-Cas) are important immune defense systems for archaea and bacteria against viral and plasmid infections. The CRISPR / Cas system can recognize exogenous DNA or RNA, cut them, and silence the expression of exogenous genes. It is precisely because of this precise targeting function that the CRISPR / Cas system has been developed into a highly efficient gene editing tool. CRISPR / Cas gene editing uses RNA to specifically bind to target sequences on the genome and cut DNA to produce double-strand breaks, and uses biological non-homologous end joining or homologous recombination for site-directed gene editing.

[0003] The commonly used CRISPR editing system CRISPR / Cas9 system mainly consists of endonuclease Cas9, crRNA and tracrRNA, which recognize the PAM motif of NGG. The mature crRNA unit is composed of a guide (Direct Repeat, DR) sequence and a spacer sequence (spacer). In practical applications, crRNA and tracrRNA form a chimeric single-molecule guide RNA (single guide RNA, sgRNA), and Cas9 and sgRNA form a binary complex in the cell to perform gene editing. Another commonly used CRISPR / Cas12a system belongs to Type V, which recognizes the PAM motif of TTTV, which is located at the 5' end of the spacer and produces a 4-5bp 5' overhanging end after cutting the target sequence. This system only requires Cas12a protein and crRNA to produce cutting at a specific site. At the same time, in addition to the endonuclease function, the Cas12a protein also has RNase activity, which can process the precursor crRNA (pre-crRNA) into a single mature crRNA for gene editing, so it is more convenient in the design of multi-gene editing. The two types of CRISPR / Cas systems have different advantages and are currently widely used in different gene editing studies.

[0004] Current CRISPR / Cas editing systems are limited by inherent off-target effects, and the PAM (Positive Aspect-Oriented Mechanism) limitation on editable genes restricts their application in plant gene editing. Furthermore, the editing efficiency of CRISPR / Cas systems in many plants needs further improvement. Many factors influence editing efficiency, including Cas gene expression, guide RNA expression, and the number and location of nuclear localization signals (NLS). These issues affect the accuracy, flexibility, controllability, and safety of CRISPR editing systems during application, hindering further expansion of their functionality and application scope. Summary of the Invention

[0005] The present invention aims to provide a Cas effector protein, a fusion protein comprising such a protein, and a nucleic acid molecule encoding them. The invention also relates to CRISPR / Cas complexes and compositions for nucleic acid editing (e.g., gene or genome editing) comprising the Cas effector protein or fusion protein of the present invention, or a nucleic acid molecule encoding them. The invention further relates to methods for nucleic acid editing (e.g., gene or genome editing) using a Cas effector protein or fusion protein comprising the present invention.

[0006] In a first aspect of the invention, a protein, its ortholog, homolog, variant, or functional fragment is provided, wherein the protein is selected from:

[0007] (a1) a protein comprising the amino acid sequence shown in SEQ ID NO: 1;

[0008] (a2) A protein containing an amino acid sequence having one or more substitutions, deletions, or additions compared to the amino acid sequence shown in SEQ ID NO:1, wherein the protein retains the biological function of the protein shown in SEQ ID NO:1; or

[0009] (a3) A protein containing an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the amino acid sequence shown in SEQ ID NO:1, wherein the protein retains the biological function of the protein shown in SEQ ID NO:1;

[0010] Wherein, the orthologs, homologs, variants or functional fragments retain the biological function of the protein shown in SEQ ID NO: 1.

[0011] In one or more embodiments, the biological function of the protein represented by SEQ ID NO: 1 includes or is the activity of a Cas enzyme.

[0012] In one or more embodiments, in (a2), an amino acid substitution, deletion or addition occurs at position 362 and / or position 464, preferably G at position 362 and / or D at position 464 of the amino acid sequence shown in SEQ ID NO: 1; more preferably, position 362 of the amino acid sequence shown in SEQ ID NO: 1 is mutated from G to a basic amino acid, most preferably to R, H or K; and position 464 is mutated from D to a basic amino acid, most preferably to R, H or K.

[0013] In one or more embodiments, the protein, its orthologs, homologs, variants, or functional fragments further contain functional units, wherein the functional units include or are selected from: epitope tags, reporter gene sequences, nuclear localization signal sequences, targeting portions, transcriptional activation domains, transcriptional repression domains, nuclease domains, and domains having activities including or selected from: methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity, and any combination thereof.

[0014] In one or more embodiments, the functional unit is directly connected to the N-terminus or C-terminus of the sequence (a1) to (a3), or is connected to the N-terminus or C-terminus through a linker; more preferably, the linker comprises or is selected from MA, GIHGVPAA.

[0015] In one or more embodiments, the functional unit is a nuclear localization signal sequence; more preferably, the nuclear localization signal sequence is selected from nuclear localization signal sequences from: SV40 virus large T antigen, EGL 13, cMyc and TUS protein; most preferably, the nuclear localization signal sequence is selected from: PKKKRKV, AVKRPAATKKAGQAKKKKLD, PAAKRVKLD, MSRRRKANPTKLSENAKKLAKEVEN, KLKIKRPVK, the acidic M9 domain of hnRNP A1, the sequence KIPIK in the yeast transcriptional repressor Matα2, and PY NLS.

[0016] In one or more embodiments, the protein is selected from:

[0017] (b1) A protein containing the amino acid sequence shown in SEQ ID NO:22;

[0018] (b2) Containing one or more amino acid substitutions, deletions, or additions compared to the amino acid sequence shown in SEQ ID NO:22, the protein retains the biological function of the protein shown in SEQ ID NO:22; or

[0019] (b3) Containing sequence identity with the sequence shown in SEQ ID NO:22 (at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%), the protein retains the biological function of the protein shown in SEQ ID NO:22;

[0020] Wherein, the ortholog, homolog, variant or functional fragment retains the biological function of the protein shown in SEQ ID NO: 22.

[0021] In one or more embodiments, the biological function of the protein shown in SEQ ID NO: 22 includes or is Cas enzyme activity.

[0022] In a second aspect of the present invention, an sgRNA molecule is provided, which contains a direct repeat sequence (crRNA) and optionally a trans-acting crRNA (tracrRNA).

[0023] In one or more embodiments, the direct repeat sequence (crRNA) is selected from the following sequences, or consists of the following sequences:

[0024] (c1) the sequence shown in SEQ ID NO: 3 or 4;

[0025] (c2) the sequence represented by bases 5 to 27 of SEQ ID NO: 3 or 4, or the sequence represented by bases 173 to 195 of SEQ ID NO: 6;

[0026] (c3) is a sequence that has one or more substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 base substitutions, deletions, or additions) compared to the sequence shown in (c1) or (c2).

[0027] (c4) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in (c1) or (c2);

[0028] (c5) a sequence that hybridizes under stringent conditions to the sequence described in any one of (c1) to (c4); or

[0029] The complementary sequence of any of the sequences described in (c6)(c1)-(c4);

[0030] Furthermore, the sequence of any one of (c2)-(c6) substantially retains the activity of the sequence from which it is derived as a direct repeat sequence in the CRISPR-Cas system.

[0031] In one or more embodiments, the trans-acting crRNA (tracrRNA) is selected from the following sequences, or consists of a sequence selected from the following sequences:

[0032] (d1) The sequence shown in SEQ ID NO:5;

[0033] (d2) The sequence shown by bases 1-172 in SEQ ID NO:5, or the sequence shown by bases 1-172 in SEQ ID NO:6;

[0034] (d3) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in SEQ ID NO:5;

[0035] (d4) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with the sequence shown in SEQ ID NO:5;

[0036] (d5) a sequence that hybridizes under stringent conditions to the sequence described in any one of (d1) to (d4); or

[0037] The complementary sequence of any of the sequences described in (d6)(d1)-(d4);

[0038] Furthermore, the sequence of any one of (d2)-(d6) substantially retains the biological function of the sequence from which it is derived, the activity of the sequence as a trans-acting crRNA (tracrRNA) in the CRISPR-Cas system.

[0039] In one or more embodiments, the sgRNA molecule comprises or consists of a sequence selected from the following:

[0040] (e1) The sequence shown in any one of SEQ ID NO: 14-19 or 28-31;

[0041] (e2) A sequence having one or more base substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in any of SEQ ID NO:14-19, 28-31;

[0042] (e3) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with any of the sequences shown in SEQ ID NO: 14-19, 28-31;

[0043] (e4) A sequence that hybridizes under stringent conditions with any of the sequences described in (e1)-(e3); or

[0044] The complementary sequence of any of the sequences described in (e5)(e1)-(e3);

[0045] Furthermore, any one of (e2)-(e5) retains substantially the biological function of the sequence from which it is derived, namely, the activity of a single-molecule guide RNA in the CRISPR-Cas system.

[0046] In one or more embodiments, the sgRNA molecule comprises one or more stem-loops or optimized secondary structures.

[0047] In one or more embodiments, the sequence in any one of (e2)-(e5) retains one or more (e.g., 2, 3 or all 4) stem-loops or secondary structures of the sequence from which it originates.

[0048] In one or more embodiments, the sequence in any one of (e2)-(e5) retains one or more (e.g., 2, 3 or all 4) stem-loop or secondary structures of the sequence from which it is derived, and is missing 10 to 50 bases.

[0049] In one or more embodiments, the stem-loop structure is a stem-loop structure formed by bases at positions 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) in the sequence shown in SEQ ID NO:6.

[0050] In a third aspect of the invention, a guide RNA molecule is provided, comprising the sgRNA molecule and guide sequence described herein from the 5' to 3' direction, wherein the guide sequence comprises a complementary sequence to the target sequence.

[0051] In one or more embodiments, the guiding sequence comprises, or consists of, sequences selected from, the following:

[0052] (f1) The sequence shown in any one of SEQ ID NO: 7, 20, 23, 24, 25, 26, 27;

[0053] (f2) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in any of SEQ ID NO: 7, 20, 23, 24, 25, 26, 27.

[0054] (f3) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with any of the sequences shown in SEQ ID NO: 7, 20, 23, 24, 25, 26, or 27;

[0055] (f4) a sequence that hybridizes under stringent conditions to the sequence described in any one of (f1) to (f3); or

[0056] (f5) A complementary sequence of the sequence described in any one of (f1) to (f3);

[0057] Furthermore, the sequence in any one of (f2)-(f5) substantially retains the biological function of the sequence from which it is derived; more preferably, the biological function of the sequence refers to its activity as a guide sequence in the CRISPR-Cas system.

[0058] In one or more embodiments, the guide RNA molecule comprises, or is composed of, sequences selected from, the following:

[0059] (g1) the sequence shown in any one of SEQ ID NOs: 9-13;

[0060] (g2) The sequence shown by bases 36-250 in SEQ ID NO:9, the sequence shown by bases 36-234 in SEQ ID NO:10, the sequence shown by bases 36-220 in SEQ ID NO:11, the sequence shown by bases 36-195 in SEQ ID NO:12, or the sequence shown by bases 36-213 in SEQ ID NO:13;

[0061] (g3) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 9-13;

[0062] (g4) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 9-13;

[0063] (g5) a sequence that hybridizes under stringent conditions to the sequence described in any one of (g1) to (g4); or

[0064] (g6) A complementary sequence of the sequence described in any one of (g1) to (g4);

[0065] Furthermore, the sequence described in any one of (g2) to (g6) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a guide RNA sequence in the CRISPR-Cas system.

[0066] In one or more embodiments, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell, or alternatively, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0067] In one or more embodiments, the target sequence is present in a cell, or alternatively, the target sequence is present in a nucleic acid molecule in vitro.

[0068] In one or more embodiments, the cell is a bacterial, plant, or animal cell.

[0069] In one or more embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer sequence adjacent to the continued PAM, and the PAM has a sequence represented by 5'-AAN, wherein N is selected from A, G, T, and C.

[0070] In a fourth aspect of the present invention, a composition is provided, comprising:

[0071] (h1) a protein component selected from the group consisting of: a protein according to the present invention, its orthologs, homologs, variants or functional fragments, and any combination thereof; and

[0072] (h2) a nucleic acid component comprising the sgRNA molecule described in the present invention; or a guide RNA molecule described in the present invention;

[0073] In this composition, the protein component and the nucleic acid component exist independently or are combined to form a complex.

[0074] Therefore, it can also be said that the present invention provides a complex, which comprises the protein component and the nucleic acid component, which are combined with each other to form a complex.

[0075] In one or more embodiments, the nucleic acid component comprises the sgRNA molecule of the present invention and the guide sequence of the present invention from the 5' to the 3' direction.

[0076] In one or more embodiments, the nucleic acid component further comprises a promoter, which is connected to the 5' end of the nucleic acid molecule; preferably, the promoter is the J23119 promoter; more preferably, the nucleotide sequence of the J23119 promoter comprises or is positions 1-35 of the nucleotide sequence shown in any one of SEQ ID NOs: 9-13.

[0077] In one or more embodiments, the nucleic acid component further comprises a poly-T sequence, more preferably the poly-T sequence is continuous with more than 3, more than 4, more than 5, more than 6, more than 7 or more Ts.

[0078] In a fifth aspect of the present invention, an isolated nucleic acid molecule is provided, comprising:

[0079] (i1) a nucleotide sequence encoding the protein of the present invention;

[0080] (i2) a nucleotide sequence encoding the sgRNA molecule of the present invention;

[0081] (i3) a nucleotide sequence encoding the guide RNA molecule of the present invention; and / or,

[0082] (i4) contains the nucleotide sequence of (i1)-(i3).

[0083] In one or more embodiments, the nucleotide sequence described in any one of (i1)-(i4) is codon-optimized for expression in a prokaryotic or eukaryotic cell.

[0084] In a sixth aspect of the present invention, a vector, a composition containing the vector, a host cell or a kit is provided, wherein:

[0085] The vector comprises the isolated nucleic acid molecule of the present invention;

[0086] The carrier-containing composition comprises one or more carriers, and the one or more carriers comprise:

[0087] (j1) a first nucleic acid, which is a nucleotide sequence encoding the protein of the present invention, its ortholog, homolog, variant or functional fragment; optionally, the first nucleic acid is operably linked to a first regulatory element; and

[0088] (j2) a second nucleic acid encoding a nucleotide sequence comprising an sgRNA molecule of the present invention, or encoding a nucleotide sequence of a guide RNA molecule of the present invention; optionally, the second nucleic acid is operably linked to a second regulatory element;

[0089] wherein the first nucleic acid and the second nucleic acid are present on the same or different vectors;

[0090] The host cell comprises the isolated nucleic acid molecule of the present invention, and / or expresses the protein of the present invention, its ortholog, homolog, variant or functional fragment, the sgRNA molecule of the present invention, the guide sequence of the present invention, the guide RNA molecule of the present invention, and / or contains the composition of the present invention;

[0091] The kit contains the protein of the present invention, its ortholog, homolog, variant or functional fragment, the sgRNA molecule of the present invention, the guide sequence of the present invention, the guide RNA molecule of the present invention, and / or the composition of the present invention.

[0092] In a seventh aspect of the invention, applications are provided for the protein described herein, its orthologs, homologs, variants, or functional fragments, the sgRNA molecule described herein, the guide RNA molecule described herein, the complex or composition described herein, the isolated nucleic acid molecule described herein, the vector described herein, the vector composition, the delivery composition, the host cell, or the kit, wherein the application is selected from:

[0093] (k1) is used for nucleic acid editing or modification;

[0094] (k2) is used to prepare formulations for nucleic acid editing or modification;

[0095] (k3) In vitro or ex vivo DNA detection;

[0096] (k4) Preparations for in vitro or ex vivo DNA detection;

[0097] (k5) Used to modify cells, cell lines or organisms by changing one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

[0098] Other aspects of the invention will be apparent to those skilled in the art from the disclosure herein. Attached Figure Description

[0099] Figure 1 The sgRNA-WT secondary structure of HT001.

[0100] Figure 2, in vitro cutting results of HT001 guided by crRNA, tracrRNA, crRNA-tracrRNA mixture and sgRNA-WT.

[0101] Figure 3 , the PAM domain of HT001.

[0102] Figure 4 A schematic diagram of the stem-loop structure of HT001's sgRNA.

[0103] Figure 5 Schematic diagram of the colony growth of sgRNA-WT and its four individual stem-loop structure deletions.

[0104] Figure 6 , Schematic diagram of the engineering modification of HT001 truncated sgRNA (SL4).

[0105] Figure 7 , in vitro cutting electrophoresis diagram of HT001sgRNA after engineering modification.

[0106] Figure 8 , schematic diagram of the HT001 cleavage site.

[0107] Figure 9 , sgRNAWT-HT001, and sgRNAR1-HT001 editing efficiency in tobacco. Deletion 2, Deletion 4, Deletion 5, Deletion 7, and Deletion 9 represent 2, 4, 5, 6, and 9 bp deletions, respectively.

[0108] Figure 10 , sgRNAWT-HT001 and sgRNAR1-HT001 editing efficiency in human cells.

[0109] Figure 11 , editing efficiency of HT001 protein wild type (WT) single mutation and double mutation in tobacco. DETAILED DESCRIPTION

[0110] Through extensive experimentation and repeated exploration, the inventors discovered a novel endonuclease (Cas enzyme), the sequence of which is shown in SEQ ID NO:1. By predicting the higher-order structure of the repetitive sequence and conducting experimental verification, the crRNA and tracrRNA sequences of this Cas enzyme were further obtained. Based on this discovery, the inventors developed a novel CRISPR / Cas complex that exhibits nuclease activity in bacteria, plants, and animals, and further optimized it to obtain a more efficient CRISPR / sgRNAR1-HT001 gene editing system.

[0111] Definition of terms

[0112] Unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, procedures in molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA used herein are conventional procedures widely used in the relevant fields. To facilitate a better understanding of the present invention, definitions and explanations of relevant terms are provided below.

[0113] The terms "Cas protein", "Cas enzyme", "Cas effector protein", "CRISPR / Cas effector protein" and "CRISPR enzyme" are used interchangeably. Nucleic acid cleavage or cleavage of nucleic acids herein include: DNA or RNA breakage in the target nucleic acid (Cis cleavage) produced by the Cas enzymes described herein, breakage of DNA or RNA in a side branch nucleic acid substrate (single-stranded nucleic acid substrate) (i.e., non-specific or non-targeted, Trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.

[0114] The term "homolog" has the meaning commonly understood by those skilled in the art, and generally refers to substances that have similarities in sequence or structure and are derived from a common ancestral molecule through chemotaxis or evolution. A "homolog" of a protein herein refers to a protein that has similarities in sequence or structure to the protein.

[0115] The terms "ortholog," "orthologue," and "ortholog" are used interchangeably and have the meanings commonly understood by those skilled in the art. As a further guide, an "ortholog" of a protein refers to a protein from a different species that performs the same or a similar function as the protein to which it is an ortholog.

[0116] The term "variant" or "mutant" has the meaning generally understood by those skilled in the art, which is different from the "wild type" and refers to an atypical form of an organism, strain, or gene or a form that is different from its natural form, which is usually intentionally modified by humans.

[0117] The term "functional fragment" has the meaning commonly understood by those skilled in the art and refers to a fragment having a biological function. For example, a "functional fragment" of a protein refers to a fragment of the protein that can exert its biological function (e.g., the enzymatic activity of the protein).

[0118] The terms “clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system” and “CRISPR system” are used interchangeably and have the meanings commonly understood by those skilled in the art, which generally include transcripts or other elements associated with the expression of CRISPR-associated (“Cas”) genes, or transcripts or other elements capable of directing the activity of the Cas genes.

[0119] The term "CRISPR / Cas complex" refers to a complex formed by the binding of guide RNA (guide RNA) and Cas protein, which includes a guide sequence that hybridizes to the target sequence and an sgRNA bound to the Cas protein. The complex is capable of recognizing and cleaving polynucleotides that can hybridize to the guide RNA.

[0120] The terms “guide RNA,” “guide RNA,” “gRNA,” “guide sequence,” or “guide sequence” are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, a guide RNA may comprise sgRNA and a guide sequence, or consist essentially of or comprise of sgRNA and a guide sequence. In some cases, the guide sequence is any polynucleotide sequence that is sufficiently complementary to a target sequence to hybridize with said target sequence and guide the specific binding of the CRISPR / Cas complex to said target sequence.

[0121] The terms "single-molecule guide RNA," "sgRNA," or "single-guide RNA" are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, sgRNAs can interact with Cas proteins. In some cases, sgRNAs contain only or consist of direct-repetitive sequences (crRNAs). In others, sgRNAs contain both direct-repetitive sequences (crRNAs) and trans-acting crRNAs (tracrRNAs). In still others (e.g., due to modification or optimization of the sgRNA), sgRNAs contain a segment of a direct-repetitive sequence (crRNA) and a segment of a trans-acting crRNA (tracrRNA).

[0122] The term "direct repeat" can refer to the DNA coding sequence in the CRISPR locus, or to the RNA encoded by crRNA.

[0123] Therefore, in the context of RNA, when referring to a guide RNA, sgRNA, direct repeat sequence, or guide sequence, each T should be understood to represent a U.

[0124] The term "target sequence" refers to a polynucleotide targeted by a guide sequence in a gRNA, such as a sequence having complementarity with the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex (including Cas proteins and gRNA). Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex. The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the nucleus or cytoplasm of the cell. In some cases, the target sequence can be located in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast. A sequence or template that can be used to recombine into a target locus comprising the target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence". In the present invention, a "target sequence" or "target polynucleotide" or "target nucleic acid" can be any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell).

[0125] The term "vector" refers to a nucleic acid molecule capable of delivering another nucleic acid molecule linked to it. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. Vectors can be introduced into host cells through transformation, transduction, or transfection, thereby enabling the expression of the genetic material elements they carry in the host cells. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, the vector may contain a replication initiation site. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop in which another DNA fragment can be inserted, for example, using standard molecular cloning techniques. Another type of vector is the viral vector, in which a virus-derived DNA or RNA sequence is present in a vector used to package the virus (e.g., retrovirus, replication-defective retrovirus, adenovirus, replication-defective adenovirus, and adeno-associated virus). Viral vectors also contain polynucleotides carried by a virus for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episodic mammalian vectors) are capable of autonomous replication in the host cell into which they are introduced. Other vectors (e.g., non-episodic mammalian vectors) integrate into the host cell's genome after introduction and thereby replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of genes they are operatively linked to. Such vectors are referred to herein as "expression vectors."

[0126] The term "expression" refers to the process by which a DNA template is transcribed into polynucleotides (such as mRNA or other RNA transcripts) and / or the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides are derived from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.

[0127] The term "5'UTR" refers to the 5' untranslated region, which has potential regulatory functions.

[0128] The term "3'UTR" refers to the 3' untranslated region, which may have a regulatory role

[0129] The term "promoter" has a meaning well-known to those skilled in the art, referring to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell essentially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell essentially only when the cell is a cell of the tissue type corresponding to that promoter.

[0130] The term "codon" refers to the pattern in which three adjacent nucleotides in a messenger RNA molecule are grouped together to represent a certain amino acid during protein synthesis.

[0131] The term "codon optimization" refers to a new technique that improves protein expression levels in organisms by increasing the translation efficiency of target genes.

[0132] The terms "nuclear localization signal," "nuclear localization sequence," or "NLS" are amino acid sequences that "tag" a protein to be transported into the cell nucleus via nuclear transport; that is, proteins with NLS are transported to the cell nucleus. Typically, NLS contain positively charged Lys or Arg residues exposed on the protein surface.

[0133] Cas effector protein

[0134] The present invention provides a protein having the amino acid sequence shown in SEQ ID NO:1 or its orthologs, homologs, variants or functional fragments thereof; wherein the orthologs, homologs, variants or functional fragments substantially retain the biological function of the sequence from which they are derived.

[0135] In this invention, the biological functions of the above sequences include, but are not limited to, binding activity with guide RNA, endonuclease activity, and binding and cleaving activity with specific sites of the target sequence under the guidance of guide RNA.

[0136] In some embodiments, the orthologs, homologs, and variants have at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in SEQ ID NO:1, and substantially retain the biological functions of the protein shown in SEQ ID NO:1 (e.g., activity to bind to guide RNA, endonuclease activity, and activity to bind to and cleave a target sequence at a specific site guided by guide RNA).

[0137] In some implementations, the protein is an effector protein in a CRISPR / Cas system.

[0138] In some embodiments, the protein of the present invention comprises, or consists of, sequences selected from, the following:

[0139] (i) The sequence shown in SEQ ID NO:1;

[0140] (ii) a sequence having one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 1;

[0141] (iii) a sequence that has at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to the sequence set forth in SEQ ID NO: 1; or

[0142] (iv) Orthologs, homologs, variants or functional fragments of the sequence shown in SEQ ID NO:1, while retaining the biological function of the protein shown in SEQ ID NO:1.

[0143] In certain embodiments, the substitution, deletion or addition of amino acid in (ii) may occur at position 362 and / or position 464 of the amino acid sequence shown in SEQ ID NO: 1, preferably the substitution, deletion or addition of amino acid in (ii) occurs at position 362 G and / or position 464 D.

[0144] In some embodiments, in (ii), position 362 of the amino acid sequence shown in SEQ ID NO:1 is mutated from G to a basic amino acid (including R, H, K), for example, mutated to R.

[0145] In certain embodiments, in (ii), position 464 of the amino acid sequence shown in SEQ ID NO: 1 is mutated from D to a basic amino acid (including R, H, K), for example, to K.

[0146] In certain embodiments, in (ii), position 362 of the amino acid sequence shown in SEQ ID NO: 1 is mutated from G to a basic amino acid (including R, H, K, preferably R), and position 464 is also mutated from D to a basic amino acid (including R, H, K, preferably K). In certain specific embodiments, in (ii), position 362 of the amino acid sequence shown in SEQ ID NO: 1 is mutated from G to R, and / or position 464 is mutated from D to K.

[0147] In certain embodiments, the primers shown in SEQ ID NOs: 32 and 33 are used to mutate position 362 of the amino acid sequence shown in SEQ ID NO: 1 from G to R.

[0148] In certain embodiments, the primers shown in SEQ ID NOs: 34 and 35 are used to mutate position 464 of the amino acid sequence shown in SEQ ID NO: 1 from D to K.

[0149] In certain embodiments, the protein of the present invention has the amino acid sequence shown in SEQ ID NO:1.

[0150] In certain embodiments, the protein shown in SEQ ID NO: 1 can be obtained by transcription and translation of the nucleotide sequence shown in SEQ ID NO: 2.

[0151] Derived Cas effector proteins

[0152] The proteins of the present invention can be derivatized, for example, by being linked to another molecule (e.g., another polypeptide or protein). Generally, protein derivatization (e.g., labeling) does not adversely affect the desired activity of the protein (e.g., activity binding to guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by guide RNA). Therefore, the proteins of the present invention are also intended to include such derivatized forms. For example, the proteins of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage, or other means) to one or more other molecular groups, such as another protein or polypeptide, a detection reagent, a pharmaceutical reagent, etc.

[0153] Specifically, the Cas protein of the present invention can be linked to a functional unit. In this document, the functional unit may be selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting portions, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB or SID domains), nuclease domains (e.g., Fok1), and domains having activities selected from: methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof.

[0154] In some embodiments, the functional unit is directly linked to the N-terminus or C-terminus of the protein of the present invention, or is linked to the N-terminus or C-terminus of the protein of the present invention via a linker. Such linkers are well known in the art, and examples include, but are not limited to, linkers comprising one or more (e.g., 1, 2, 3, 4, or 5) amino acids (such as Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc. In some embodiments, the linker comprises the sequence shown by MA. In some embodiments, the linker comprises the sequence shown by GHIHGVPAA.

[0155] In some embodiments, the functional unit is an NLS. The NLS enhances the ability of the Cas protein of the present invention to enter the cell nucleus. In these embodiments, the Cas protein of the present invention may comprise one or more NLS sequences. The NLS sequences may be located at, near, or adjacent to the ends (e.g., the N-terminus or C-terminus) of the Cas protein of the present invention. In some embodiments, the Cas protein comprises two NLS sequences located at, near, or adjacent to the N-terminus and C-terminus of the Cas protein of the present invention, respectively.

[0156] Exemplary NLS sequences include, but are not limited to, NLSs derived from: SV40 viral large T antigen, EGL 13, cMyc, and TUS protein. In some embodiments, the NLS contains the PKKKRKV sequence. In some embodiments, the NLS contains the AVKRPAATKKAGQAKKKKLD sequence. In some embodiments, the NLS contains the PAAKRVKLD sequence. In some embodiments, the NLS contains the MSRRRKANPTKLSENAKKLAKEVEN sequence. In some embodiments, the NLS contains the KLKIKRPVK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the KIPIK sequence in the yeast transcriptional repressor Matα2, and the PY NLS. In some exemplary embodiments, the NLS sequence is as shown in SEQ ID NO:21 or SEQ ID NO:36. In some embodiments, the NLS sequence is located at, near, or close to the end (e.g., N-terminus or C-terminus) of the protein of the present invention.

[0157] In certain embodiments, the functional unit is an epitope tag, including but not limited to an antibody tag (e.g., Myc, HA), a fluorescent tag (e.g., FLAG), an enzyme tag (e.g., GST), a biotin tag, a metal tag (e.g., His), or a sugar tag. In certain embodiments, the epitope tag is a FLAG tag. In certain examples, the FLAG tag comprises the sequence MDYKDHDGDYKDHDIDYKDDDDK.

[0158] In some embodiments, the derived Cas effector protein of the present invention comprises, or is composed of, sequences selected from, the following:

[0159] (i) the sequence shown in SEQ ID NO: 22;

[0160] (ii) a sequence having one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 22; or

[0161] (iii) a sequence having at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to the sequence set forth in SEQ ID NO: 22;

[0162] (iv) a direct homologue, homologue, variant or functional fragment of the sequence shown in SEQ ID NO: 22, while retaining the biological function of the protein shown in SEQ ID NO: 22.

[0163] In certain embodiments, the amino acid substitution, deletion or addition in (ii) may occur at position 401 and / or position 503 of the amino acid sequence shown in SEQ ID NO: 22, preferably the amino acid substitution, deletion or addition in (ii) occurs at position 401 G and / or position 503 D.

[0164] In some embodiments, in (ii), position 401 of the amino acid sequence shown in SEQ ID NO:22 is mutated from G to a basic amino acid (including R, H, K), for example, mutated to R.

[0165] In certain embodiments, in (ii), position 503 of the amino acid sequence shown in SEQ ID NO: 22 is mutated from D to a basic amino acid (including R, H, K), for example, to K.

[0166] In certain embodiments, in (ii), position 401 of the amino acid sequence shown in SEQ ID NO: 22 is mutated from G to a basic amino acid (including R, H, K, preferably R), and position 503 is also mutated from D to a basic amino acid (including R, H, K, preferably K).

[0167] In some embodiments, the protein of the present invention has the amino acid sequence shown in SEQ ID NO:22.

[0168] In some embodiments, the Cas protein of the present invention can be linked to a targeting moiety to enable the Cas protein to be targeted. For example, it can be linked to a detectable label to facilitate detection of the Cas protein of the present invention. For example, it can be linked to an epitope tag to facilitate expression, detection, tracing and / or purification of the Cas protein of the present invention.

[0169] The derivative Cas effector protein of the present invention can be produced by genetic engineering methods (recombinant technology) or chemical synthesis methods.

[0170] Single-molecule guide RNA (sgRNA) molecules

[0171] The present invention also provides an isolated nucleic acid molecule comprising or consisting of a sequence selected from the following:

[0172] (i) a sequence shown in any one of SEQ ID NOs: 4-6, 14-19, and 28-31;

[0173] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 4-6, 14-19 and 28-31;

[0174] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 4-6, 14-19, and 28-31;

[0175] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or

[0176] (v) a complementary sequence of the sequence described in any one of (i) to (iii);

[0177] Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence is the activity as a single-molecule guide RNA in the CRISPR-Cas system.

[0178] In certain embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:

[0179] (a) the nucleotide sequence shown in any one of SEQ ID NOs: 4-6, 14-19, and 28-31;

[0180] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or

[0181] (c) The complementary sequence of the sequence described in (a).

[0182] In certain embodiments, the isolated nucleic acid molecule is RNA.

[0183] In certain embodiments, the isolated nucleic acid molecule is a single-molecule guide RNA in a CRISPR-Cas system.

[0184] In certain embodiments, the isolated nucleic acid molecule comprises a direct repeat sequence (crRNA).

[0185] In certain embodiments, the isolated nucleic acid molecule comprises a trans-acting crRNA (tracrRNA).

[0186] In some embodiments, the isolated nucleic acid molecule comprises a direct repeat sequence (crRNA) and a trans-acting crRNA (tracrRNA).

[0187] In certain embodiments, the direct repeat sequence (crRNA) comprises or consists of a sequence selected from the group consisting of:

[0188] (i) The sequence shown in SEQ ID NO:3 or 4;

[0189] (ii) the sequence represented by bases 5 to 27 of SEQ ID NO: 3 or 4, or the sequence represented by bases 173 to 195 of SEQ ID NO: 6;

[0190] (iii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 3 or 4;

[0191] (iv) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in SEQ ID NO: 3 or 4;

[0192] (v) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iv); or

[0193] Complementary sequences to any of the sequences described in (vi)(i)-(iv);

[0194] Furthermore, the sequence described in any one of (ii) to (vi) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a direct repeat sequence in the CRISPR-Cas system.

[0195] In some embodiments, the direct repeat sequence (crRNA) comprises or consists of the sequence shown in SEQ ID NO:3 or 4.

[0196] In certain embodiments, the direct repeat sequence (crRNA) comprises or is the sequence shown in bases 5-27 of SEQ ID NO: 3 or 4.

[0197] In certain embodiments, the trans-acting crRNA (tracrRNA) comprises or consists of a sequence selected from the group consisting of:

[0198] (i) The sequence shown in SEQ ID NO:5;

[0199] (ii) the sequence represented by bases 1 to 172 of SEQ ID NO: 5, or the sequence represented by bases 1 to 172 of SEQ ID NO: 6;

[0200] (iii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 5;

[0201] (iv) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in SEQ ID NO: 5;

[0202] (v) A sequence that hybridizes under stringent conditions with any of the sequences described in (i)-(iv); or

[0203] (vi) a complementary sequence of the sequence described in any one of (i) to (iv);

[0204] Furthermore, the sequence of any one of (ii) to (vi) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a trans-acting crRNA (tracrRNA) in the CRISPR-Cas system.

[0205] In certain embodiments, the isolated nucleic acid molecule consists of a direct repeat sequence (crRNA) and a trans-acting crRNA (tracrRNA), for example, it comprises the sequence shown in SEQ ID NO: 6 or consists of the sequence shown in SEQ ID NO: 6.

[0206] In certain embodiments, the trans-acting crRNA (tracrRNA) comprises or is the sequence shown in bases 1-172 of SEQ ID NO:5.

[0207] In certain embodiments, the single-molecule guide RNA sequence is in a truncated form, for example, it lacks 10 to 50 (e.g., 10 to 40) bases compared to the sequence of SEQ ID NO: 6, and retains the stem-loop structure of the single-molecule guide RNA.

[0208] In certain embodiments, the single-molecule guide RNA lacks the stem-loop structure formed by bases 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) in the sequence shown in SEQ ID NO:6, respectively.

[0209] In some embodiments, the single-molecule guide RNA sequence is in truncated form, with bases 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) missing, respectively, compared to the sequence shown in SEQ ID NO:6.

[0210] In certain embodiments, the stem-loop structure of the single-molecule guide RNA is a stem-loop structure formed by bases 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) in the sequence shown in SEQ ID NO:6.

[0211] In certain embodiments, the structure of the single-molecule guide RNA is as follows Figure 3 shown.

[0212] In some embodiments, the truncated single-molecule guide RNA comprises, or consists of, sequences selected from, a group of sequences selected from:

[0213] (i) the sequence shown in any one of SEQ ID NOs: 28-31;

[0214] (ii) A sequence having one or more base substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 base substitutions, deletions, or additions) compared to any of the sequences shown in SEQ ID NO:28-31;

[0215] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with any of the sequences shown in SEQ ID NO:28-31;

[0216] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or

[0217] (v) a complementary sequence of the sequence described in any one of (i) to (iii);

[0218] Furthermore, the sequences described in any one of (ii)-(v) substantially retain the biological function of the sequences from which they are derived, namely, their activity as single-molecule guide RNA in the CRISPR-Cas system.

[0219] In some embodiments, the truncated single-molecule guide RNA comprises, or consists of, sequences selected from, a group of sequences selected from:

[0220] (a) the nucleotide sequence shown in any one of SEQ ID NOs: 28-31;

[0221] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or

[0222] The complementary sequence of the sequence described in (c)(a).

[0223] In some embodiments, the truncated single-molecule guide RNA retains only a portion of the stem-loop structure formed by bases at positions 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) of the sequence shown in SEQ ID NO:6.

[0224] In certain embodiments, the structure of the truncated form of the single-molecule guide RNA is as follows Figure 4 shown.

[0225] In some embodiments, the truncated single-molecule guide RNA (sgRNA) comprises, or consists of, sequences selected from, the following:

[0226] (i) The sequence shown in any one of SEQ ID NO: 14-19;

[0227] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in any of SEQ ID NO:14-19;

[0228] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with any of the sequences shown in SEQ ID NO:14-19;

[0229] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or

[0230] (v) a complementary sequence of the sequence described in any one of (i) to (iii);

[0231] Furthermore, the sequences described in any one of (ii)-(v) substantially retain the biological function of the sequences from which they are derived, namely, their activity as single-molecule guide RNA in the CRISPR-Cas system.

[0232] Guide sequence

[0233] The present invention also provides a guide sequence that can hybridize with a target sequence.

[0234] In some implementations, the guide sequence is attached to the 3' end of the nucleic acid molecule (e.g., sgRNA).

[0235] In certain embodiments, the guide sequence comprises the complement of the target sequence.

[0236] In some embodiments, the guiding sequence comprises, or consists of, sequences selected from, the following:

[0237] (i) The sequence shown in any one of SEQ ID NO: 7, 20, 23, 24, 25, 26, 27;

[0238] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in any of SEQ ID NO: 7, 20, 23, 24, 25, 26, 27;

[0239] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with any of the sequences shown in SEQ ID NO: 7, 20, 23, 24, 25, 26, or 27;

[0240] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or

[0241] (v) a complementary sequence of the sequence described in any one of (i) to (iii);

[0242] Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a guide sequence in the CRISPR-Cas system.

[0243] Guide RNA sequence

[0244] The present invention also provides a guide RNA sequence comprising, from the 5' to the 3' direction, the sgRNA sequence described herein and a guide sequence capable of hybridizing with a target sequence.

[0245] In some embodiments, the guide RNA sequence comprises a sequence selected from the following:

[0246] (a) an sgRNA sequence comprising or consisting of a sequence selected from the following:

[0247] (i) The sequence shown in any one of SEQ ID NO: 4-6, 14-19 and 28-31;

[0248] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in any one of SEQ ID NO: SEQ ID NO: 4-6, 14-19, and 28-31;

[0249] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with any of the sequences shown in SEQ ID NO: SEQ ID NO: 4-6, 14-19, and 28-31;

[0250] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or

[0251] (v) a complementary sequence of the sequence described in any one of (i) to (iii);

[0252] Furthermore, the sequence described in any one of (ii)-(v) substantially retains the biological function of the sequence from which it is derived, the biological function of which refers to the activity of the sgRNA sequence in the CRISPR-Cas system.

[0253] (b) A guide sequence that hybridizes with the target sequence, comprising, or consisting of, sequences selected from, the following:

[0254] (i) The sequence shown in any one of SEQ ID NO: 7, 20, 23, 24, 25, 26, 27;

[0255] (ii) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in any of SEQ ID NO: 7, 20, 23, 24, 25, 26, 27;

[0256] (iii) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with any of the sequences shown in SEQ ID NO: 7, 20, 23, 24, 25, 26, or 27;

[0257] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or

[0258] (v) a complementary sequence of the sequence described in any one of (i) to (iii);

[0259] Furthermore, the sequence described in any one of (ii)-(v) substantially retains the biological function of the sequence from which it is derived, the biological function of which refers to its activity as a guide sequence in the CRISPR-Cas system.

[0260] In some embodiments, the guide RNA sequence comprises, or consists of, sequences selected from, the following:

[0261] (i) The sequence shown in any one of SEQ ID NO: 9-13;

[0262] (ii) The sequence shown by bases 36-250 in SEQ ID NO:9, the sequence shown by bases 36-234 in SEQ ID NO:10, the sequence shown by bases 36-220 in SEQ ID NO:11, the sequence shown by bases 36-195 in SEQ ID NO:12, or the sequence shown by bases 36-213 in SEQ ID NO:13;

[0263] (iii) A sequence having one or more base substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to any of the sequences shown in SEQ ID NO: 9-13;

[0264] (iv) A sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with any of the sequences shown in SEQ ID NO: 9-13;

[0265] (v) A sequence that hybridizes under stringent conditions with any of the sequences described in (i)-(iv); or

[0266] (vi) a complementary sequence of the sequence described in any one of (i) to (iv);

[0267] Furthermore, the sequence described in any one of (ii) to (vi) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a guide RNA sequence in the CRISPR-Cas system.

[0268] In certain embodiments, the guide RNA sequence comprises or consists of a sequence selected from the group consisting of:

[0269] (a) The nucleotide sequence shown in any one of SEQ ID NO: 9-13;

[0270] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or

[0271] The complementary sequence of the sequence described in (c)(a).

[0272] CRISPR / Cas complex

[0273] The present invention also provides a composite comprising:

[0274] (i) a protein component selected from: a Cas effector protein or a derived Cas effector protein according to the present invention; and

[0275] (ii) A nucleic acid component comprising, from 5' to 3', the sgRNA sequence described in this invention and a guide sequence capable of hybridizing with a target sequence;

[0276] Wherein, the protein component and the nucleic acid component combine with each other to form a complex.

[0277] In certain embodiments, the nucleic acid component is RNA.

[0278] In certain embodiments, the nucleic acid component is a guide RNA in a CRISPR-Cas system, such as a guide RNA sequence as defined in any embodiment of the invention.

[0279] In certain embodiments, the nucleic acid component further comprises a promoter (such as, but not limited to, the J23119 promoter) linked to the 5' end of the nucleic acid molecule.

[0280] In some specific embodiments, the nucleotide sequence of the J23119 promoter includes or is the nucleotide sequence of any one of SEQ ID NO: 9-13, positions 1-35.

[0281] In certain embodiments, the nucleic acid component further comprises a poly-T sequence, for example, more than 3, more than 4, more than 5, more than 6, more than 7 or more consecutive Ts.

[0282] Encoding nucleic acid, vector and host cell

[0283] The present invention also provides an isolated nucleic acid molecule comprising:

[0284] (i) The nucleotide sequences encoding the Cas effector protein and derived Cas effector protein of the present invention;

[0285] (ii) a nucleotide sequence encoding an sgRNA sequence or guide RNA as described in the present invention; or

[0286] (iii) Nucleotide sequences containing (i) and (ii).

[0287] In some embodiments, the nucleic acid molecule encodes a protein described herein having an amino acid sequence as shown in SEQ ID NO: 1 or 22 or an ortholog, homolog, variant or functional fragment thereof.

[0288] In some embodiments, the nucleic acid molecule comprises, or is composed of, sequences selected from, the following:

[0289] (I) the sequence shown in SEQ ID NO: 2;

[0290] (II) A sequence having one or more substitutions, deletions, or additions (e.g., substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases) compared to the sequence shown in SEQ ID NO:2;

[0291] (II) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in SEQ ID NO: 2;

[0292] (IV) a sequence that hybridizes under stringent conditions to the sequence described in any one of (I) to (III); or

[0293] The complementary sequence of the sequence described in any of (V)(I)-(III).

[0294] In certain embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:

[0295] (a) The nucleotide sequence shown in SEQ ID NO:2;

[0296] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or

[0297] (c) The complementary sequence of the sequence described in (a).

[0298] In certain embodiments, the nucleotide sequence described in any one of (i)-(iii) is codon-optimized for expression in prokaryotes. In certain embodiments, the nucleotide sequence described in any one of (i)-(iii) is codon-optimized for expression in eukaryotic cells.

[0299] The present invention also provides a vector comprising isolated nucleic acid molecules as described herein. The vector of the present invention can be a cloning vector or an expression vector. In some embodiments, the vector of the present invention is, for example, a plasmid, a granulocyte, a bacteriophage, a Cosmid, etc. In some preferred embodiments, the vector is capable of expressing the Cas effector protein and its derivatives, isolated nucleic acid molecules, or CRISPR / Cas complexes of the present invention in bacteria, plants, and animals.

[0300] The present invention also provides host cells comprising the isolated nucleic acid molecules or vectors as described above. Such host cells include, but are not limited to, prokaryotic cells such as *Escherichia coli* cells, and eukaryotic cells such as yeast cells, insect cells, plant cells (such as *Arabidopsis thaliana* cells, tobacco cells), and animal cells (such as mammalian cells, such as mouse cells, human cells, etc.). The cells of the present invention can also be cell lines, such as HEK293 cells.

[0301] Composition and carrier composition

[0302] The present invention also provides a composition comprising:

[0303] (i) a first component selected from the group consisting of: a Cas effector protein of the present invention (including a derivatized form thereof), a nucleotide sequence encoding the Cas effector protein or a derivative protein, and any combination thereof; and

[0304] (ii) a second component, which is a nucleotide sequence comprising a guide RNA, or a coding sequence thereof;

[0305] The guide RNA comprises an sgRNA sequence and a guide sequence from the 5' to the 3' direction, and the guide sequence can hybridize with the target sequence.

[0306] In certain embodiments, the first component and the second component of the composition exist independently.

[0307] In some embodiments, the guide RNA is capable of forming a complex with the protein or derived protein described in (i).

[0308] In some embodiments, the sgRNA sequence is an isolated nucleic acid molecule as defined in any embodiment of the present invention.

[0309] In some embodiments, the guide RNA sequence is ligated to the 3' end of the sgRNA sequence. In some embodiments, the guide RNA sequence contains a complementary sequence to the target sequence.

[0310] The present invention also provides a carrier composition comprising one or more carriers, wherein the one or more carriers comprise:

[0311] (i) a first nucleic acid, which is a nucleotide sequence encoding the Cas effector protein or fusion protein of the present invention; optionally, the first nucleic acid is operatively linked to a first regulatory element; and

[0312] (ii) a second nucleic acid encoding a nucleotide sequence comprising a guide RNA sequence; optionally, the second nucleic acid is operably linked to a second regulatory element;

[0313] in:

[0314] The first nucleic acid and the second nucleic acid may exist on the same or different vectors;

[0315] The guide RNA can form a complex with the effector protein or fusion protein described in (i).

[0316] In certain embodiments, the sgRNA sequence is an isolated nucleic acid molecule as defined in any embodiment of the present invention, including truncated forms thereof.

[0317] In certain embodiments, the guide RNA sequence is a guide RNA sequence as defined in any embodiment of the present invention.

[0318] In certain embodiments, the composition is non-naturally occurring or modified. In certain embodiments, at least one component in the composition is non-naturally occurring or modified.

[0319] In certain embodiments, the first regulatory element is a promoter, such as a Pol II type promoter (including but not limited to: ZmUbi, Actin, CmYLCV, UBQ, 35S, SPL), a tissue-specific promoter (including but not limited to: YAO, CDC45, rbcS), an inducible promoter (including but not limited to: XEV).

[0320] In certain embodiments, the second regulatory element is a promoter, such as a Pol II type or a Pol III type promoter (including but not limited to: ZmU6, OsU3, OsU6a, OsU6b, OsU6c, AtU6, Actin, 35S, Ubi, UBQ, SPL, CmYLCV), a tissue-specific promoter (including but not limited to: YAO, CDC45, rbcS), an inducible promoter (including but not limited to: XEV).

[0321] In some implementations, one type of vector is a plasmid, which refers to a circular double-stranded DNA loop in which additional DNA fragments can be inserted, for example, using standard molecular cloning techniques. Another type of vector is a viral vector, in which a virus-derived DNA or RNA sequence is present in a vector used to package viruses (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also contain polynucleotides carried by a virus for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episodic plant vectors) are capable of autonomous replication in the host cells in which they are introduced. Other vectors (e.g., non-episodic plant vectors) integrate into the genome of the host cell after introduction and thereby replicate along with the host genome. Moreover, some vectors are capable of directing the expression of genes they are operatively linked to. Such vectors are referred to herein as “expression vectors.” Common expression vectors used in recombinant DNA technology are typically in plasmid form.

[0322] Recombinant expression vectors may contain the nucleic acid molecules of the present invention in a form suitable for nucleic acid expression in a host cell, which means that these recombinant expression vectors contain one or more regulatory elements selected based on the host cell to be used for expression, which are operably linked to the nucleic acid sequence to be expressed.

[0323] Delivery and delivery compositions

[0324] The Cas effector proteins, derived proteins, sgRNA sequences, guide RNA sequences, CRISPR / Cas complexes, encoding nucleic acids, vectors, compositions, and / or vector compositions of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic pore effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.

[0325] Therefore, the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: a Cas effector protein, a derivative protein, an sgRNA sequence, a guide RNA sequence, a CRISPR / Cas complex, an encoding nucleic acid, a vector, a composition and / or a vector composition of the present invention.

[0326] In some implementations, the delivery carrier is a particle.

[0327] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, a gene gun, or a viral vector (e.g., a replication-defective retrovirus, a lentivirus, an adenovirus, or an adeno-associated virus).

[0328] Reagent test kit

[0329] The present invention provides a kit comprising one or more of the components described above. In certain embodiments, the kit comprises one or more components selected from the following: a Cas effector protein, a derivative protein, an sgRNA sequence, a guide RNA sequence, a CRISPR / Cas complex, an encoding nucleic acid, a vector, a composition, and / or a vector composition of the present invention.

[0330] In some embodiments, the kit also includes instructions for using the composition and / or carrier composition.

[0331] In some embodiments, the components contained in the kit of the present invention can be provided in any suitable container.

[0332] In certain embodiments, the kit further comprises one or more buffers. The buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer and combinations thereof. In certain embodiments, the buffer is alkaline. In certain embodiments, the buffer has a pH from about 7 to about 10.

[0333] In certain embodiments, the kit further comprises one or more oligonucleotides corresponding to a targeting sequence for insertion into a vector so as to operably link the targeting sequence and regulatory elements. In certain embodiments, the kit comprises a homologous recombination template polynucleotide.

[0334] Methods and uses

[0335] The present invention provides a method for modifying a target gene, comprising: contacting the target gene with a CRISPR / Cas complex as described in the present invention, a composition as described in the present invention, or a vector composition, or delivering it to a cell containing the target gene; wherein the target sequence is present in the target gene.

[0336] In some embodiments, the method is used to modify a target gene in vitro or ex vivo. In some embodiments, the method is not a method for treating humans or animals as a therapy. In some embodiments, the method does not include the step of modifying human germline genetic characteristics.

[0337] In certain embodiments, the target gene is present in a cell. In certain embodiments, the cell is a prokaryotic cell. In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a human cell (including primary cells and cell lines, such as human HEK293 cell lines). In certain embodiments, the cell is a plant cell, such as a cell of a cultivated plant (such as tobacco, Arabidopsis thaliana, cassava, corn, sorghum, wheat or rice), algae, tree or vegetable. In certain embodiments, the cell is a bacterial cell. In certain specific embodiments, the target gene can be a VEGF gene, a DNMT1 gene or a PMT1 gene of an animal (such as a human) or a plant (such as tobacco).

[0338] In certain embodiments, the target gene is present in a nucleic acid molecule (e.g., a plasmid) in vitro. In certain embodiments, the target gene is present in a plasmid.

[0339] In some implementations, the modification refers to a break in the target sequence, such as a double-strand break in DNA or a single-strand break in RNA.

[0340] In certain embodiments, the disruption results in decreased transcription of the target gene.

[0341] In certain embodiments, the method further comprises: contacting the editing template with the target gene, or delivering it to a cell comprising the target gene. In such embodiments, the method repairs the broken target gene by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides of the target gene. In certain embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.

[0342] Thus, in certain embodiments, the modification further comprises inserting an editing template (eg, an exogenous nucleic acid) into the break.

[0343] In certain embodiments, the protein, derived protein, isolated nucleic acid molecule, complex, vector or composition is contained in a delivery vehicle.

[0344] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, and viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0345] In certain embodiments, the methods are used to modify a cell, cell line, or organism by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

[0346] The present invention provides a method for altering the expression of a gene product, comprising: contacting a CRISPR / Cas complex, a composition, or a vector composition as described in the present invention with a nucleic acid molecule encoding the gene product, or delivering it to a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

[0347] In some embodiments, the method is used to alter the expression of a gene product in vitro or in vitro. In some embodiments, the method is not a method for treating humans or animals as a therapy. In some embodiments, the method does not include the step of modifying human germline genetic characteristics.

[0348] In certain embodiments, the nucleic acid molecule is present in a nucleic acid molecule (e.g., a plasmid) in vitro. In certain embodiments, the nucleic acid molecule is present in a plasmid.

[0349] In some embodiments, the expression of the gene product is altered (e.g., enhanced or reduced). In some embodiments, the expression of the gene product is enhanced. In some embodiments, the expression of the gene product is reduced.

[0350] In certain embodiments, the gene product is a protein.

[0351] In another aspect, the present invention relates to the use of the Cas effector protein, sgRNA sequence, guide RNA sequence, CRISPR / Cas complex, encoding nucleic acid, vector, composition and / or vector composition, kit or delivery composition described herein for nucleic acid editing (e.g., in vitro or ex vivo nucleic acid editing), or for use in the preparation of formulations for nucleic acid editing.

[0352] In some embodiments, the nucleic acid to be edited is present inside a cell. In some embodiments, the cell is a prokaryotic or eukaryotic cell. In some embodiments, the nucleic acid to be edited is present in an in vitro nucleic acid molecule (e.g., a plasmid).

[0353] In certain embodiments, the nucleic acid editing includes gene or genome editing, such as modifying a gene, knocking out a gene, changing the expression of a gene product, repairing a mutation, and / or inserting a polynucleotide. In certain embodiments, the gene or genome editing does not include a step of modifying human germline genetic characteristics. In certain embodiments, the use is not a method for treating a human or animal by therapy.

[0354] In some embodiments, the use also includes repairing an edited target sequence by homologous recombination with a foreign template polynucleotide, wherein the repair may produce a mutation in the target sequence, including the insertion, deletion, or substitution of one or more nucleotides.

[0355] In another aspect, the present invention relates to the use of the Cas effector protein, derivative protein, sgRNA sequence, guide RNA sequence, CRISPR / Cas complex, encoding nucleic acid, vector, composition and / or vector composition, kit or delivery composition of the present invention in preparing a preparation for: (i) in vitro or ex vivo DNA detection; (ii) editing the target sequence in the target locus to modify an organism or non-human organism (e.g., a prokaryotic organism).

[0356] In certain embodiments, the preparation is used for the detection of single-stranded DNA or double-stranded DNA (eg, the detection of single-stranded or double-stranded DNA in prokaryotic cells).

[0357] In another aspect, the present invention also relates to a method for detecting a target DNA in a sample, comprising the following steps:

[0358] (1) contacting the sample with the following components: the CRISPR / Cas complex, encoding nucleic acid, vector, composition and / or vector composition as described in the present invention, and single-stranded DNA with a label; wherein,

[0359] The guide RNA sequence contained in the CRISPR / Cas complex or composition is capable of hybridizing with the target DNA, and

[0360] The single-stranded DNA does not hybridize with the guide RNA sequence;

[0361] (2) Detect target DNA by measuring the detectable signal generated by the Cas effector protein contained in the CRISPR / Cas complex or composition cleaving labeled single-stranded DNA.

[0362] In certain embodiments, the target DNA is viral DNA or bacterial DNA.

[0363] In certain embodiments, the target DNA is tumor cell DNA.

[0364] In certain embodiments, the target DNA is single-stranded or double-stranded.

[0365] In some embodiments, the detectable signal is determined by one or more methods selected from: imaging-based detection, sensor-based detection, color detection, gold nanoparticle-based detection, fluorescence polarization, colloidal phase transition / dispersion, electrochemical detection, and semiconductor-based sensing.

[0366] In certain embodiments, the method further comprises the step of amplifying the target DNA in the sample.

[0367] The advantages of the present invention are:

[0368] 1. The Cas protein of the present invention contains 521 amino acids, is relatively small, has high flexibility and controllability, and is easy to deliver in vivo.

[0369] 2. The present invention adds a nuclear localization signal (NLS) to the Cas protein, thereby enhancing the gene editing efficiency of the CRISPR / Cas system.

[0370] 3. The PAM site recognized by the Cas protein in this invention is AAN (N is A, T, C, G), which enriches the existing PAM recognition sites of the CRISPR / Cas system and improves the range of target selection.

[0371] 4. Mutating the Cas protein of this invention further improves gene editing efficiency.

[0372] 5. The present invention optimizes the guide RNA and the sgRNA sequence therein, further improving the gene editing efficiency of the CRISPR / Cas system by at least 1.60 times, enhancing the accuracy and flexibility of its gene editing.

[0373] 6. This invention successfully applied the CRISPR / Cas system to Escherichia coli, tobacco cells, and human HEK293 cells, demonstrating the stability of the system and its feasibility for application in bacteria, plants, and animals.

[0374] The present invention will be further described below with reference to specific embodiments, which are intended to illustrate the present invention but are not intended to limit the present invention.

[0375] Unless otherwise specified, the experiments and methods described in the examples are generally performed in accordance with conventional methods well known in the art and described in various references. For example, conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA used in this invention can be found in Sambrook, Fritsch, and Maniatis, *Molecular Cloning: A Laboratory Manual*, 2nd edition (1989); *Current Protocols in Molecular Biology* (edited by FM. Ausubel et al., (1987)); and the *Methods in Enzymology* series (academic publishing company): *PCR 2: A PRACTICAL*. APPROACH (edited by MJ MacPherson, BD Hames and GR Taylor (1995)), Harlow and Lane (1988) Antibodies, A Laboratory Manual, and Animal Cell Culture (edited by R.R. Freshney (1987)).

[0376] In addition, for conditions not specifically specified in the examples, standard conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all commercially available products.

[0377] Those skilled in the art will appreciate that the embodiments described herein are by way of example only and are not intended to limit the scope of protection claimed herein. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.

[0378] The sources of some of the reagents involved in the following examples are as follows:

[0379] LB liquid medium: 10g tryptone, 5g yeast extract, 10g NaCl, bring to a final volume of 1L, and sterilize. If antibiotics are required, add them after the medium has cooled, at a final concentration of 50μg / ml.

[0380] Chloroform / isoamyl alcohol: Add 10 ml of isoamyl alcohol to 240 ml of chloroform and mix well.

[0381] RNP buffer: 100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg / ml BSA, pH 7.9.

[0382] The prokaryotic expression vector pET-28a was purchased from Beijing Quanshijin Biotechnology Co., Ltd.

[0383] E. coli competent cells BL21 were purchased from Epicentre.

[0384] Example 1. Acquisition of HT001 gene and HT001 guide RNA

[0385] 1. Phage genome assembly and annotation: phage sequences from public databases (marine microbiome, global marine virome) were collected and assembled using metaSPAdes, followed by sequence annotation.

[0386] 2. Protein filtering: De-redundancy of annotated proteins is performed through sequence consistency, and proteins with completely identical sequences are removed.

[0387] 3. Obtaining CRISPR-related proteins: Collect CRISPR protein family sequences from NCBI, Uniprot, and literature, perform protein sequence alignment using DIAMOND BLASTP, and output alignment results with Evalue < 1E-5 to obtain suspected CRISPR protein sequences in bacteriophages.

[0388] 4. Clustering of CRISPR proteins: Multiple sequence alignment of CRISPR protein family sequences and phage-derived CRISPR protein sequences was performed using Maft. Phylogenetic trees were constructed using IQTREE to infer the types of CRISPR proteins derived from phages.

[0389] 5. Identification of CRISPR sites (Repeats and Spacers): Use MinCED and CRISPRDetect to identify the CRISPR sites of phage-derived CRISPR proteins.

[0390] 6. Protein Sequence Similarity Analysis and Domain Annotation: DIAMOND BLASTP was used to align Cas proteins from the Genbank database, NR database, and published patents, as well as phage-derived Cas proteins. Cas proteins with an Evalue < 1E-5 and a similarity of less than 80% were selected as new CRISPR / Cas protein families. Multiple sequence alignment of each CRISPR / Cas family protein was performed using Maft, followed by conserved domain analysis using JPred and HHpred to identify protein families containing the RuvC domain. Based on this, the inventors obtained a novel Cas effector protein, named HT001 (SEQ ID NO:1), whose encoding DNA sequence is shown in SEQ ID NO:2.

[0391] 7. To obtain the guide RNA for HT001, CRISPR-RNA (crRNA) repeats from the phage-encoded CRISPR locus were identified using MinCED and CRISPRDetect. Repeats were compared using the Needleman-Wunsch algorithm followed by pairwise similarity scores generated using BROSSNeedle. The prototype direct repeat sequence of HT001 (SEQ ID NO: 3) was obtained.

[0392] 8. Synthesize a DNA sequence encoding the HT001 protein (SEQ ID NO: 22) with a nuclear localization signal, connect the double-stranded DNA molecule to the prokaryotic expression vector pET-28a, and obtain the recombinant plasmid pET-28a-HT001. Sequence the recombinant plasmid pET-28a-CRISPR / HT001. Introduce the recombinant plasmid pET-28a-HT001 into Escherichia coli BL21 (DE3) to obtain a recombinant bacterium, which is named BL21 / pET-28a-HT001. Pick a single clone of BL21 / pET-28a-HT001 and inoculate it into 100 mL LB liquid medium (containing 100 μg / mL kanamycin). Culture at 37°C and 200 rpm for 12 hours to obtain a culture solution. The culture was inoculated into 50 mL of LB liquid medium (containing 100 μg / mL kanamycin) at a volume ratio of 1:100. The cells were shaken at 37°C, 200 rpm, and cultured to an OD600nm value of 0.6. IPTG was then added to a concentration of 1 mM and cultured at 28°C, 220 rpm, and shaken for 4 h. The cells were centrifuged at 10,000 rpm at 4°C for 10 min, and the pellet was collected. The pellet was then resuspended in 100 mL of 100 mM Tris-HCl buffer, pH 8.0, and ultrasonically disrupted (ultrasonic power 600 W, cycle: 4 s break, 6 s pause, for a total of 20 min). The pellet was then centrifuged at 10,000 rpm at 4°C for 10 min, and supernatant A was collected. Supernatant A was then centrifuged at 12,000 rpm at 4°C for 10 min, and supernatant B was collected. Supernatant B was purified using a nickel column produced by GE (for the specific purification steps, refer to the instructions of the nickel column), and then the Cas protein was quantified using a protein quantification kit produced by Thermo Fisher Scientific.

[0393] 9. Design a template for guide RNA transcription. The structure of the transcription template is: T7 promoter + HT001 prototype direct repeat sequence (SEQ ID NO: 3) + guide sequence (SEQ ID NO: 20). Primers were designed using Primer 5.0 software, ensuring that the forward primer and reward primer had at least 18 bp of overlapping sequence. Prepare the reaction system, gently pipette to mix, and briefly centrifuge. Place the reaction in a PCR instrument for PCR amplification. Use the MinElute PCR Purification Kit to purify the template by extracting with phenol: chloroform: isoamyl alcohol (25:24:1) to remove DNase I in the system. Determine the concentration of the purified crRNA using Nanodrop, and uniformly dilute to 250 ng / μl. Aliquot the mixture into 200 μl PCR tubes and store at -80°C until needed. Establish a double-stranded DNA digestion system: Prepare the reaction system, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 15 min; add 300 ng substrate DNA (100 ng / μl), 3 μl, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 h; add RNase, incubate at 37°C for 15 min to fully digest RNA impurities; add proteinase K, incubate at 58°C for 15 min to digest HT001 protein; gel electrophoresis results showed that no DNA cleavage activity was detected in the purified HT001 protein and the corresponding crRNA. Figure 1 ), indicating that HT001 requires additional components for successful dsDNA cleavage.

[0394] 10. A p15a-based plasmid containing the HT001 site and a non-targeting spacer was transformed into a chemically competent E. coli Top10 strain. Total RNA was extracted from 5 mL of cell culture using the EasyPure RNA kit (TransGen Biotech). 15 mg of RNA was treated with 2 units of DNase I (NEB) at 37°C for 30 minutes to remove DNA. 1 mM ATP was added, and the sample was incubated at 37°C for 1 hour for 5'-phosphorylation, followed by phenol / chloroform extraction. T4 PNK-treated RNA was treated with 5 units of RppH (NEB) at 37°C for 1 hour to hydrolyze 5'-pyrophosphate, followed by purification by phenol / chloroform extraction. RNA purity was determined by agarose gel electrophoresis. Small RNA libraries were constructed using the VAHTS Small RNA Library Prep Kit for Illumina (Vazyme Biotech) according to the manufacturer's instructions and subjected to Illumina HiSeq sequencing at the HaploX Genomics Center. The raw sequencing data were processed to remove adapters and sequencing artifacts and maintain high-quality reads. Sequencing reads were aligned to the reference sequence using BWA, and mapped unique reads were analyzed using the SAM tool. Analysis showed that a 200nt RNA transcript mapped to the 3' end of the HT001 gene was enriched in the HT001 system. This RNA transcript contained an anti-repeat sequence that could potentially base-pair with the short direct repeats of the corresponding crRNA, suggesting that this could serve as tracrRNA (SEQ ID NO: 4) targeted by the HT001 nuclease DNA.

[0395] 11. Design a template for guide RNA transcription. The template structure is: T7 promoter + HT001 tracrRNA sequence (SEQ ID NO:5) + guide sequence (SEQ ID NO:20). Primers were designed using Primer 5.0 software, ensuring at least 18 bp overlap between the forward and reward primers. Prepare the reaction mixture, gently pipette to mix, briefly centrifuge, and place in a PCR instrument for PCR amplification. Purify the template using the MinElute PCR Purification Kit, extracting with phenol:chloroform:isoamyl alcohol (25:24:1) to remove DNAse I from the system. Measure the concentration of the purified tracrRNA using Nanodrop and dilute to 250 ng / μl. Aliquot into 200 μl PCR centrifuge tubes and store at -80℃ for later use. Prepare the reaction mixture, gently pipette to mix, and briefly centrifuge. Incubate at 37℃ for 15 min. For the DNA cleavage reaction, add 300 ng of substrate DNA (100 ng / μl), 3 μl, gently pipette to mix, and briefly centrifuge. The cells were placed at 37°C for 8 hours; RNAse was added and the cells were placed at 37°C for 15 minutes to fully digest the RNA impurities in the system; Proteinase K was added and the cells were placed at 58°C for 15 minutes to digest the HT001 protein; gel electrophoresis results showed that the purified HT001 protein had cleavage activity when both crRNA and putative tracrRNA were supplemented, whereas the same DNA substrate could not be cleaved when either crRNA or tracrRNA was provided alone ( Figure 1 ).

[0396] 12. The secondary structure of the guide sequence is predicted to obtain the secondary structure of the guide RNA. By artificially removing the 3' end 4-nt of tracrRNA and the 5' end 4-nt of mature crRNA, the tracrRNA and mature crRNA are further connected and fused into sgRNA ( Figure 1A template for guide RNA transcription was designed, with the following structure: T7 promoter + HT001 sgRNA sequence (SEQ ID NO:6) + guide sequence (SEQ ID NO:20). Primers were designed using Primer 5.0 software, ensuring at least 18 bp overlap between the forward primer and the reward primer. The reaction mixture was prepared, gently mixed, and briefly centrifuged. PCR amplification was performed using a PCR amplification instrument. Template purification was performed using the MinElute PCR Purification Kit, with phenol:chloroform:isoamyl alcohol (25:24:1) extraction to remove DNAseI from the system. The concentration of purified sgRNA was measured using Nanodrop, and the mixture was uniformly diluted to 250 ng / μl, aliquoted into 200 μl PCR centrifuge tubes, and stored at -80℃ for later use. The double-stranded DNA digestion system was established: the reaction mixture was prepared, gently mixed, and briefly centrifuged. Incubate at 37°C for 15 min; for DNA cleavage, add 300 ng substrate DNA (100 ng / μl), 3 μl, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 h; add RNase, incubate at 37°C for 15 min to fully digest RNA impurities; add proteinase K, incubate at 58°C for 15 min to digest HT001 protein; gel electrophoresis results showed that the chimeric sgRNA-WT exhibited dsDNA cleavage activity comparable to that of isolated crRNA and tracrRNA. Figure 2 Therefore, the expression of the CRISPR-HT001 system can be simplified for practical applications.

[0397] Example 2: Identification of the PAM domain of the HT001 gene

[0398] 1. The recombinant plasmid pET-28a+CRISPR / HT001+sgRNA was constructed and sequenced. Based on the sequencing results, the structure of the recombinant plasmid pET-28a+CRISPR / HT001+sgRNA was described as follows: The small fragment between the restriction endonuclease HindIII and EcoRI recognition sequences of the vector pET-28a was replaced with the sequence shown in SEQ ID NO:2, and the nucleotide sequence encoding the prototype sgRNA of HT001 shown in SEQ ID NO:6 and the guide sequence identified by the PAM domain of HT001 shown in SEQ ID NO:7 were ligated to the SphⅠ restriction site of the vector.

[0399] 2. Obtaining recombinant Escherichia coli: The recombinant plasmid pET-28a+CRISPR / HT001+sgRNA was introduced into Escherichia coli BL21 to obtain recombinant Escherichia coli, named BL21 / pET-28a+CRISPR / HT001+sgRNA. The recombinant plasmid pET-28a was then introduced into Escherichia coli BL21 to obtain recombinant Escherichia coli, named BL21 / pET-28a.

[0400] 3. Construction of the PAM library: The sequence shown in SEQ ID NO:8 was synthesized artificially and ligated into the pUC19 vector, wherein the sequence shown in SEQ ID NO:5 includes eight random bases (NNNNNNNN) and the target sequence. The plasmid was transformed into *E. coli* containing the CRISPR / HT001 locus (BL21 / pET-28a+CRISPR / HT001+sgRNA) and *E. coli* without the CRISPR / HT001 locus (BL21 / pET-28a), respectively. After incubation at 37°C for 1 hour, the plasmid was extracted, and the PAM region sequence was amplified by PCR and sequenced.

[0401] 4. Obtaining the PAM library domain: The occurrence frequency of PAM sequences in 65,536 combinations was counted in both the experimental and control groups, and normalized using the total number of PAM sequences in each group. For any PAM sequence, a PAM sequence was considered significantly consumed if log2 (normalized value of control group / normalized value of experimental group) was greater than 3.5. Significantly consumed PAM sequences were obtained from all PAM sequences. Furthermore, Weblogo was used to predict the significantly consumed PAM sequences, revealing that the PAM domain of the HT001 protein is 5'-AAN-3' (…). Figure 3 ), N is any amino acid.

[0402] Example 3: Validation of key structural elements of guide RNA

[0403] Because the predicted secondary structure of the guide RNA is complex and contains multiple stem-loop structures, the inventors conducted a systematic truncation method to verify the key structural elements. Based on sgRNA-WT, the loss of four individual stem-loop structures was used to assess the importance of the key structural elements of sgRNA.

[0404] according to Figure 4SL1, SL2, SL3 and SL4 were partially deleted as shown, and the sgRNA WT expression cassettes of the sequences shown in SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12 and SEQ ID NO: 13, SL1-deleted sgRNASL1 expression cassette, SL2-deleted sgRNASL2 expression cassette, SL3-deleted sgRNASL3 expression cassette, and SL4-deleted sgRNASL4 expression cassette were synthesized according to the pattern of J23119 promoter + sgRNA + guide RNA + polyT, respectively, wherein the housekeeping gene gapA of Escherichia coli BL21 was selected as the target site. Each gRNA expression cassette was ligated into the pESL28(a) vector through the SphⅠ restriction site, and the small fragment between the restriction endonuclease HindⅢ and EcoR I recognition sequences of the vector pET-28a was replaced with the sequence shown in SEQ ID NO: 2 to construct recombinant plasmids pESL28(a)+HT001+J23119-sgRNAWT, pESL28(a)+HT001+J23119-sgRNASL1, pESL28(a)+HT001+J23119-sgRNASL2, pESL28(a)+HT001+J23119-sgRNASL3 and pESL28(a)+HT001+J23119-sgRNASL4.

[0405] The correctly sequenced recombinant plasmids were transformed into Escherichia coli BL21 and named BL21 / pESL28(a)+HT001+J23119-sgRNAWT, BL21 / pESL28(a)+HT001+J23119-sgRNASL1, BL21 / pESL28(a)+HT001+J23119-sgRNASL2, BL21 / pESL28(a)+HT001+J23119-sgRNASL3 and BL21 / pESL28(a)+HT001+J23119-sgRNASL4. The plasmids were plated on LB solid medium containing kanamycin resistance through dilution gradient concentration. The growth rate and survival rate of Escherichia coli BL21 / pESL28(a)+HT001+J23119-sgRNASL4 lacking SL4 were higher than those of the other Escherichia coli ( Figure 5 This demonstrates that the stem-loop structure sgRNA-SL4 is a key element in sgRNA-WT.

[0406] Example 4: Engineering of Guide RNA

[0407] 1. Optimization of sgRNA

[0408] 1. Since the cleavage activity of HT001 / sgRNA-WT is low, in order to improve the activity of the HT001 system, the inventors engineered the key component sgRNA-SL4 in sgRNA-WT, and synthesized R1, R2, R3, R4, R5 and R6 according to the sequences shown in SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18 and SEQ ID NO: 19, respectively. Figure 6 shown.

[0409] 2. Transcription and purification of HT001 protein guide RNA:

[0410] 1. Design a guide RNA transcription template. The structure of the transcription template is: T7 promoter + mature sgRNA of HT001 (SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19) + guide sequence (SEQ ID NO: 20). Primers were designed using Primer 5.0 software to ensure that the forward primer and reward primer had at least 18 bp of overlapping sequence.

[0411] 2. Prepare the reaction system, pipette gently to mix, centrifuge briefly, and place in a PCR instrument for annealing.

[0412] 3. The template was purified using the MinElute PCR Purification Kit (purchased from Takara, catalog number 9761). The steps are as follows:

[0413] 1) Add 5 volumes of PB to the PCR product, place a MinElute column on a 2 ml collection tube, incubate at room temperature for 2 min, and centrifuge at 12,000 g for 2 min.

[0414] 2) Discard the waste solution and add 750 μl of Buffer PE (add ethanol before use) and incubate at 12,000 g / 2 min.

[0415] 3) Discard the waste liquid, add 350 μl of Buffer PE, incubate at 12000 g for 2 min, discard the waste liquid, incubate at 12000 g for 2 min;

[0416] 4) Transfer the MinElute column to a new 1.5 ml centrifuge tube, open the tube lid, and incubate at 65°C for 2 minutes.

[0417] 5) Add 20 μl of preheated EB solution, let it stand for 2 minutes, and then centrifuge at 12000g / 2min. To improve the recovery rate, pass the contents of the centrifuge tube 2-3 times through the MinElute centrifuge column;

[0418] 6) Determine the concentration using Nanodrop and store at -20℃ for later use.

[0419] 4. Purification of guide RNA: Extraction with phenol:chloroform:isoamyl alcohol (25:24:1) to remove DNAse I from the system;

[0420] 1) Add 80 μl RNA-free H2O to the post-transcription reaction system and adjust the volume to 100;

[0421] 2) Take 2 ml of Phase Lock Gel (PLG) Heavy, centrifuge at 15000g for 2 min, add 100 μl of phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μl of RNA digested with DNAse I, gently tap the Phase-Lock tube 5-10 times to mix it evenly, and then centrifuge at 15℃ / 16000g for 12 min;

[0422] 3) Take a new RNA-free 1.5ml centrifuge tube, aspirate the supernatant from the previous centrifugation into the centrifuge tube, being careful not to aspirate the gel, add an equal volume of isopropanol and one-tenth volume of sodium acetate solution, mix well with a pipette tip, and place in a -20℃ refrigerator for 1 hour or overnight.

[0423] 4) Centrifuge at 4℃ / 16000g for 30 min, discard the supernatant, add 75% pre-cooled ethanol, mix the precipitate by suction and beating, centrifuge at 4℃ / 16000g for 12 min, discard the supernatant, let stand in a fume hood for 2-3 min to dry the ethanol on the surface of the RNA, add 100μl of RNAfree H2O, and mix by suction and beating.

[0424] 5) Determine the concentration of the purified crRNA using Nanodrop, dilute to 250 ng / μl, dispense into 200 μl PCR centrifuge tubes, and freeze at -80°C until use.

[0425] III. In vitro expression and purification of HT001 protein

[0426] The steps for in vitro expression and purification of HT001 protein are as follows:

[0427] 1. A synthetic DNA sequence encoding the HT001 protein (SEQ ID NO:22) with a nuclear localization signal.

[0428] 2. The double-stranded DNA molecule synthesized in step 1 was ligated to the prokaryotic expression vector pET-28a to obtain the recombinant plasmid pET-28a-HT001. The recombinant plasmid pET-28a-CRISPR / HT001 was sequenced. Sequencing results showed that the recombinant plasmid pET-28a-CRISPR / HT001 expressed the HT001 protein with a nuclear localization signal as shown in SEQ ID NO:22.

[0429] 3. The recombinant plasmid pET-28a-HT001 was introduced into Escherichia coli BL21(DE3) to obtain the recombinant bacteria, which was named BL21 / pET-28a-HT001. Single colonies of BL21 / pET-28a-HT001 were picked and inoculated into 100 mL of LB liquid medium (containing 100 μg / mL kanamycin) and cultured at 37℃ with shaking at 200 rpm for 12 h to obtain the bacterial culture.

[0430] 4. Take the cultured bacterial solution and inoculate it into 50 mL of LB liquid medium (containing 100 μg / mL kanamycin) at a volume ratio of 1:100. Incubate at 37℃ and 200 rpm with shaking until the OD600nm value is 0.6. Then add IPTG to make the concentration 1 mM, incubate at 28℃ and 220 rpm with shaking for 4 h, centrifuge at 4℃ and 10000 rpm for 10 min, and collect the bacterial pellet.

[0431] 5. Take the bacterial cell pellet, add 100 mL of pH 8.0, 100 mM Tris-HCl buffer, resuspend, and then sonicate to disrupt (ultrasonic power 600 W, cycle program: disrupt for 4 s, pause for 6 s, for a total of 20 min). Then centrifuge at 4 °C and 10,000 rpm for 10 min and collect the supernatant A.

[0432] 6. Take the supernatant A, centrifuge at 4℃ and 12000rpm for 10min, and collect the supernatant B.

[0433] 7. The supernatant B was purified using a nickel column manufactured by GE (refer to the nickel column instruction manual for specific purification steps), and then the Cas protein was quantified using a protein quantification kit manufactured by Thermo Fisher Scientific.

[0434] 8. Establishment of double-stranded DNA enzyme digestion system:

[0435] (1) Prepare the reaction mixture, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 15 min; DNA cleavage reaction.

[0436] (2) Add 300 ng of substrate DNA (100 ng / μl), 3 μl, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 h;

[0437] (3) Add RNase, place at 37°C for 15 min to fully digest RNA impurities in the system;

[0438] (4) Proteinase K was added and incubated at 58°C for 15 min to digest the HT001 protein;

[0439] (5) Agarose gel testing.

[0440] Gel electrophoresis results showed ( Figure 7 HT001 / sgRNAR1-R4 all have certain cleavage activity, while R5 and R6 have no cleavage activity. The system with the highest cleavage activity is R1, and this system is called sgRNAR1-HT001.

[0441] Example 5: Identification of the cleavage pattern of the CRISPR / HT001 system

[0442] 1. In vitro expression and purification of HT001 protein

[0443] 1. Synthesize a DNA sequence encoding the HT001 protein (SEQ ID NO: 22) with a nuclear localization signal.

[0444] 2. The double-stranded DNA molecule synthesized in step 1 was ligated to the prokaryotic expression vector pET-28a to obtain the recombinant plasmid pET-28a-HT001. The recombinant plasmid pET-28a-CRISPR / HT001 was sequenced. Sequencing results showed that the recombinant plasmid pET-28a-CRISPR / HT001 expressed the HT001 protein with a nuclear localization signal as shown in SEQ ID NO:22.

[0445] 3. The recombinant plasmid pET-28a-HT001 was introduced into Escherichia coli BL21(DE3) to obtain the recombinant bacteria, which was named BL21 / pET-28a-HT001. Single colonies of BL21 / pET-28a-HT001 were picked and inoculated into 100 mL of LB liquid medium (containing 100 μg / mL kanamycin) and cultured at 37℃ with shaking at 200 rpm for 12 h to obtain the bacterial culture.

[0446] 4. Take the cultured bacterial solution and inoculate it into 50 mL of LB liquid medium (containing 100 μg / mL kanamycin) at a volume ratio of 1:100. Incubate at 37℃ and 200 rpm with shaking until the OD600nm value is 0.6. Then add IPTG to make the concentration 1 mM, incubate at 28℃ and 220 rpm with shaking for 4 h, centrifuge at 4℃ and 10000 rpm for 10 min, and collect the bacterial pellet.

[0447] 5. Take the bacterial cell pellet, add 100 mL of pH 8.0, 100 mM Tris-HCl buffer, resuspend, and then sonicate to disrupt (ultrasonic power 600 W, cycle program: disrupt for 4 s, pause for 6 s, for a total of 20 min). Then centrifuge at 4 °C and 10,000 rpm for 10 min and collect the supernatant A.

[0448] 6. Take the supernatant A, centrifuge at 4℃ and 12000rpm for 10min, and collect the supernatant B.

[0449] 7. The supernatant B was purified using a nickel column manufactured by GE (refer to the nickel column instruction manual for specific purification steps), and then the Cas protein was quantified using a protein quantification kit manufactured by Thermo Fisher Scientific.

[0450] 2. Transcription and purification of HT001 protein guide RNA:

[0451] 1. Design a template for guide RNA transcription. The structure of the transcription template is: T7 promoter + sgRNA-R1 (optimal sgRNA after HT001 engineering modification) + guide sequence (SEQ ID NO:20). Primers were designed using Primer 5.0 software to ensure that the forward primer and the reward primer have at least 18 bp of overlapping sequence.

[0452] 2. Prepare the reaction system, gently pipette to mix, centrifuge briefly, and place in a PCR instrument for PCR amplification.

[0453] 3. Purify the template using the MinElute PCR Purification Kit, following these steps:

[0454] 1) Add 5 volumes of PB to the PCR product, place a MinElute column on a 2ml collection tube, let stand at room temperature for 2 minutes, 12000g / 2min;

[0455] 2) Discard the waste solution and add 750 μl of Buffer PE (add ethanol before use) and incubate at 12,000 g / 2 min.

[0456] 3) Discard the waste liquid, add 350 μl Buffer PE, centrifuge at 12000 g / 2 min, discard the waste liquid, centrifuge at 12000 g for 2 min;

[0457] 4) Transfer the MinElute column to a new 1.5ml centrifuge tube, open the cap, and incubate at 65℃ for 2 minutes;

[0458] 5) Add 20 μl of preheated EB solution, let it stand for 2 minutes, and then centrifuge at 12000g / 2min. To improve the recovery rate, pass the contents of the centrifuge tube 2-3 times through the MinElute centrifuge column;

[0459] 6) Determine the concentration using Nanodrop and store at -20℃ for later use.

[0460] 4. Purification of guide RNA: Extraction with phenol:chloroform:isoamyl alcohol (25:24:1) to remove DNAse I from the system;

[0461] 1) Add 80 μl RNA-free H2O to the post-transcription reaction system and adjust the volume to 100;

[0462] 2) Take 2 ml of Phase Lock Gel (PLG) Heavy, centrifuge at 15000g for 2 min, add 100 μl of phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μl of RNA digested with DNAse I, gently tap the Phase-Lock tube 5-10 times to mix it evenly, and then centrifuge at 15℃ / 16000g for 12 min;

[0463] 3) Take a new RNA-free 1.5ml centrifuge tube, aspirate the supernatant from the previous centrifugation into the centrifuge tube, being careful not to aspirate the gel, add an equal volume of isopropanol and one-tenth volume of sodium acetate solution, mix well with a pipette tip, and place in a -20℃ refrigerator for 1 hour or overnight.

[0464] 4) Centrifuge at 4℃ / 16000g for 30 min, discard the supernatant, add 75% pre-cooled ethanol, mix the precipitate by suction and beating, centrifuge at 4℃ / 16000g for 12 min, discard the supernatant, let stand in a fume hood for 2-3 min to dry the ethanol on the surface of the RNA, add 100μl of RNAfree H2O, and mix by suction and beating.

[0465] 5) Determine the concentration of the purified crRNA using Nanodrop, dilute to 250 ng / μl, dispense into 200 μl PCR centrifuge tubes, and freeze at -80°C until use.

[0466] 5. Establishment of double-stranded DNA enzyme digestion system:

[0467] (1) Prepare the reaction mixture, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 15 min; DNA cleavage reaction.

[0468] (2) Add 300 ng of substrate DNA (100 ng / μl), 3 μl, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 h;

[0469] (3) Add RNase, place at 37°C for 15 min to fully digest RNA impurities in the system;

[0470] (4) Proteinase K was added and incubated at 58°C for 15 min to digest the HT001 protein;

[0471] (5) First-generation sequencing results showed that HT001 was cleaved at positions 21-22 bp downstream of the target strand PAM and 27-28 bp downstream of the non-target strand PAM. Figure 8 ).

[0472] Example 6: HT001 has editing activity in plant cells

[0473] Using the modified sgRNA-R1 of Example 4 (sgRNA-R1 is the optimal sgRNA after modification), three target sequences (i.e., PMT1-Target1, PMT1-Target2, and PMT1-Target3) shown in SEQ ID NO: 23, SEQ ID NO: 24, and SEQ ID NO: 25 were designed for the tobacco PMT1 gene. The target sequences were synthesized according to the J23119 promoter sequence + the engineered sgRNA-R1 encoding nucleotide sequence + the tobacco target sequence (SEQ ID NO: 23, SEQ ID NO: 24, and SEQ ID NO: 25) + poly T sequence, and loaded into the CR2 vector through the restriction site Pac I. The DNA sequence encoding the HT001 protein (SEQ ID NO: 22) with a nuclear localization signal was artificially synthesized, and the synthesized double-stranded DNA molecule was ligated into the CR2 vector by double enzyme digestion and located in the same expression cassette as the ZmUbi promoter to form target 1 knockout vector, target 2 knockout vector, and target 3 knockout vector. Plasmids were extracted from the correctly sequenced knockout vectors and transformed into Agrobacterium gv3101.

[0474] Conducting genetic transformation experiments in tobacco:

[0475] (1) Incubate Agrobacterium tumefaciens in liquid LB containing kanamycin (50 mg / mL) and rifampin (50 mg / mL) at 200 rpm and 28 °C for 2 days;

[0476] (2) Take 1 mL of activated Agrobacterium tumefaciens culture and incubate it in 50 mL of liquid LB medium containing kanamycin and rifampin at 200 rpm and 28 °C.

[0477] (3) The bacterial culture is ready when the OD600 is between 0.6 and 0.8. The bacterial culture is collected by centrifuging at 4000 rpm at room temperature for 10 minutes and the supernatant is discarded.

[0478] (4) Resuspend the bacterial culture in 100 mL of liquid MS medium without added sucrose;

[0479] (5) Cut tobacco leaves and sterilize them with 70% alcohol for 2 min in a clean bench, then sterilize them with 5% sodium hypochlorite solution for 10 min, rinse them with sterile water 3-4 times, and absorb the moisture with filter paper;

[0480] (6) Cut off the edges of the tobacco leaves with scissors and cut the leaves into 1cm x 1cm square leaves (try to make sure that the cut leaves do not contain veins).

[0481] (7) Soak the square leaves in an Erlenmeyer flask containing Agrobacterium tumefaciens solution and shake gently for 10-15 minutes;

[0482] (8) Rinse with sterile water 3-4 times to clean the bacterial solution, and place it on filter paper to absorb the dryness;

[0483] (9) Place the leaves on a co-culture medium (6-BA: 0.5 mg / L; NAA: 2 mg / L) with filter paper and co-culture at 25°C in the dark for 2 days;

[0484] (10) The leaves were transferred to callus induction medium (6-BA: 0.5 mg / L; NAA: 2 mg / L; Timentin: 200 mg / L; Basta: 3 mg / L) at 25°C, sealed, and cultured with alternating light periods of 16 h and dark periods of 8 h. The medium was replaced every 2 weeks until callus tissue grew.

[0485] (11) Transfer the callus tissue to the resistance selection differentiation medium (6-BA: 0.5 mg / L; NAA: 2 mg / L; termethin: 200 mg / L; basta: 3 mg / L; kan: 50 mg / L) until regenerated seedlings grow.

[0486] (12) The gene editing of regenerated seedlings was detected. The results are shown in Table 1. The modified sgRNAR1-HT001 has gene editing activity in tobacco. The results of sgRNA-R1 were compared with those of sgRNA-WT. The editing efficiency of sgRNAR1 was 1.60-1.79 times that of the original one ( Figure 9 ).

[0487] Table 1

[0488]

[0489] Example 7: HT001 exhibits editing activity in human cells.

[0490] The sequence containing the U6 promoter and guide RNA (the guide sequence for eukaryotic editing in human cells after engineering sgRNA-R1 and SEQ ID NO: 27) and the HT001 gene were constructed on the mCherry red fluorescent protein expression vector pEE6.4-mCherry, which can express the HT001 protein in mammalian cells. That is, the crRNA was transcribed and the Cas protein was expressed by a single vector, and the red fluorescence indicated that the vector was successfully transfected into the host cell. The expression vector was transfected into the human HEK293 cell line using the PEI transfection method. After 48 hours of culture, cells positive for mCherry red fluorescence were obtained by flow cytometry sorting. DNA from all cells was extracted, and a 700bp sequence containing the target site was amplified. The PCR product was connected to a B-simple vector for first-generation sequencing. The sequencing was completed by Huazhi Biotechnology Co., Ltd. The sequencing results were aligned with the VEGFA gene in the human genome, and the cleavage method of HT001 at the target site was identified. In addition, its editing efficiency for VEGFA was identified. The PCR product was used to construct a second-generation sequencing library using Tn5. The sequencing was completed by Huazhi Biotechnology Co., Ltd., and the editing efficiency of HT001 for genes such as VEGFA and DNMT1 was identified ( Figure 10 As shown in Table 2, the overall editing efficiency of sgRNAR1-HT001 is 2.29-2.43 times that before optimization.

[0491] Table 2

[0492]

[0493] Example 8: Enhancing the cleavage activity of HT001 protein

[0494] To enhance the cleavage activity of the HT001 protein, HT001 mutants bearing the G362R (primer sequences: SEQ ID NOs: 32-33), D464K (primer sequences: SEQ ID NOs: 34-35), and G362R + D464K (primer sequences: SEQ ID NOs: 32-35) were constructed. Their cleavage activity compared to the wild-type protein was tested in tobacco cells. These proteins were synthesized using the J23119 promoter sequence, the engineered sgRNA-R1 nucleotide sequence, the tobacco target sequence (SEQ ID NOs: 23, 24, and 25), and the polyT sequence. These proteins were then inserted into the CR2 vector via the PacI restriction site. HT001 proteins bearing the G362R, D464K, or G362R + D464K mutations were then ligated into the CR2 vector via double restriction enzyme digestion and located in the same expression cassette as the ZmUbi promoter, creating six knockout vectors targeting the tobacco PMT1 gene. The knockout vectors with correct sequencing were extracted and the plasmids were transferred into Agrobacterium gv3101 for tobacco genetic transformation, and finally regenerated seedlings were obtained.

[0495] Detecting gene editing in regenerated seedlings ( Figure 11 PCR amplification was performed on the PMT1 gene locus. As shown in Table 3, sequencing analysis revealed that the G362R and D464K point mutations increased the cleavage activity by approximately 1.10-1.98 times compared to the wild-type protein, with G362R+D464K resulting in the highest increase in cleavage activity, approximately 2.38 times.

[0496] Table 3

[0497]

[0498] The sequences involved in the embodiments are shown below:

[0499] SEQ ID NO: 1 Amino acid sequence of HT001 (the bold underlined parts are G362 and D464 respectively)

[0500]

[0501] SEQ ID NO: 2 HT001 coding nucleotide sequence (wherein the underlined parts are the codons encoding G362 and D464 respectively)

[0502] GGT TCTAATCGCCGTAAGGTAACCGTCCGTAAGCTGGCAAAGACCCACCATATCGTAGCGCTACGCCGCGAAACCGGACTGCATCAGCTCACTAAAAATCTAACCACCCGGTACCCGCTCATTGGCATTGAAGACCTTAACGTCGCCGGAATGACTGCATCAGCATCAGGAACTATCGAAAACCCCGGTAAGAACGTAGCCCAGAAAGCCGGACTCAACCGGGCAGTACTTGACGTAGCTTTCGGTACATTCCGAAACCAGCTCGAATACAAAGCCGCTTGGTATGGTTCAGCCGTTCAGGTCATC GAC CGTTACTACCCGTCATCACAGACCTGTAGCAACTGTGGCAAACGACCTGACGCTAAGCTTGCCCTAAACGACCGTGTATATAAGTGTGGACACTGCCACGCTGTGATTGACCGCGACCTCAACGCCGCAATCAATATCCGCCGTGAAGCAGAACGCCTACATGCGGAAGCATAA

[0503] Prototype direct repeat sequence (crRNA) of SEQ ID NO:3 HT001

[0504] AUAUUCCCCGUACACGCGGGGAUGGCUCG

[0505] Coding nucleotide sequence of the prototype direct repeat sequence of SEQ ID NO: .4 HT001

[0506] ATATTCCCCGTACACGCGGGGATGGCTCG

[0507] Trans-activating CRISPR RNA (tracrRNA) of SEQ ID NO:5 HT001

[0508] AGCCGTTCAGGTCATCGACCGTTACTACCCGTCATCACAGACCTGTAGCAACTGTGGCAAACGACCTGACGCTAAGCTTGCCCTAAACGACCGTGTATATAAGTGTGGACACTGCCACGCTGTGATTGACCGCGACCTCAACGCCGCAATCAATATCCGCCGTGAAGCAGAACGCC

[0509] SEQ ID NO:6HT001 prototype sgRNA encoding nucleotide sequence (underlined is tracrRNA, italics are crRNA)

[0510]

[0511] Guide sequence for PAM domain identification of SEQ ID NO:7HT001

[0512] ACATAACGCACGGCTCGAACGTC

[0513] SEQ ID NO:8 PAM library sequence (the underlined part is the random base, the italic part is the guide sequence, and the bold part is the target sequence)

[0514]

[0515] The gRNA expression cassette synthesis sequence of the prototype sgRNA of SEQ ID NO:9HT001 (where positions 1-35 are the J23119 promoter sequence, positions 36-230 are the nucleotide sequence encoding the prototype sgRNA, the bolded part is the guide sequence, and ttttttt indicates the poly-T sequence).

[0516]

[0517] The gRNA expression cassette synthesis sequence of SL1-deleted sgRNA SEQ ID NO:10HT001 (where positions 1-35 are the J23119 promoter sequence, positions 36-214 are the SL1-deleted sgRNA sequence, the bolded part is the guide RNA sequence, and ttttttt indicates the poly-T sequence)

[0518]

[0519] Synthetic gRNA expression cassette sequence of SL2-deficient sgRNA SEQ ID NO:11HT001 (where positions 1-35 are the J23119 promoter sequence, positions 36-200 are the SL2-deficient sgRNA sequence, the bolded part is the guide RNA sequence, and ttttttt indicates the poly-T sequence)

[0520]

[0521] The synthetic sequence of the gRNA expression cassette of the SL3-deficient sgRNA of SEQ ID NO:12HT001 (where positions 1-35 are the J23119 promoter sequence, positions 36-175 are the SL3-deficient sgRNA sequence, the bolded part is the guide RNA sequence, and ttttttt indicates the poly-T sequence)

[0522]

[0523] The gRNA expression cassette synthesis sequence of SL4-deficient sgRNA of SEQ ID NO:13HT001 (where positions 1-35 are the J23119 promoter sequence, positions 36-193 are the SL4-deficient sgRNA sequence, the bolded part is the guide RNA sequence, and ttttttt indicates the poly-T sequence)

[0524]

[0525] SEQ ID NO:14 The engineered sgRNA-R1 encoding nucleotide sequence

[0526] ATCCGCCGTGAAGCAGAATCCCCGTACACGCGGGGAT

[0527] SEQ ID NO: 15 Engineered sgRNA-R2 encoding nucleotide sequence

[0528] ATCCGCCGTGAATCCACGCGGGGAT

[0529] SEQ ID NO: 16 Engineered sgRNA-R3 encoding nucleotide sequence

[0530] ATCCGCGAATCCGGGGAT

[0531] SEQ ID NO:17 The engineered sgRNA-R4 encoding nucleotide sequence

[0532] ATCCGCGAATCCGCGGAT

[0533] SEQ ID NO: 18 Engineered sgRNA-R5 encoding nucleotide sequence

[0534] ATCCGAATCCGGAT

[0535] SEQ ID NO:19 The engineered sgRNA-R6 encoding nucleotide sequence

[0536] ATCAATCGAT

[0537] SEQ ID NO:20 In vitro cutting guide sequence

[0538] acctatctcagcgatctgtc

[0539] SEQ ID NO:21 NLS sequence-1

[0540] PKKKRKV

[0541] The amino acid sequence of the SEQ ID NO:22HT001-NLS fusion protein (where positions 1-23 are 3xFLAG tags, the underlined part is the NLS sequence, the italicized part is the linker amino acid, and the bolded part is the amino acid sequence of the HT001 protein).

[0542]

[0543] SEQ ID NO: 23 Target sequence of tobacco target 1 (the underlined portion is the PAM sequence)

[0544] AAT TTTAACGATCCTCGTGTAAC

[0545] SEQ ID NO: 24 Target sequence of tobacco target 2 (the underlined portion is the PAM sequence)

[0546] AAC TGTCGTCAAGTCTTTAAGGG

[0547] SEQ ID NO:25 Target sequence of target site 3 in tobacco (where the underlined portion is the PAM sequence)

[0548] AAC TTTCAAGGTATTGTGTTTAA

[0549] Guide sequence in Escherichia coli SEQ ID NO:26

[0550] TATGACTCCACTCACGGCCG

[0551] SEQ ID NO:27 Guide sequence for eukaryotic editing in human cells

[0552] CUAGGAAUAUUGAAGGGGGGC

[0553] sgRNA with SL1 deletion SEQ ID NO:28HT001

[0554] AGCCGTTCAGGTCATCGACCGTTACTACCCGTCATCACAGACCTGTAGCAACTGTGGCAAACGACCTGACGCTAAGCTTGCCCTAAACGACCGTGTATATAAGCTGTGATTGACCGCGACCTCAACGCCGCAATCAATATCCGCCGTGAAGCAGAATCCCCGTACACGCGGGGATGGCT

[0555] SEQ ID NO: 29HT001 SL2-deleted sgRNA

[0556] AGCCGTTCAGGTCATCGACCGTTACTACCCGTCATCACAGACCTGTAGCAACTGTGGCAAACGACCTGACGCTAAGCTTGCCCTAAACGACCGTGTATATAAGTGTGGACACTGCCACGCTGATATCCGCCGTGAAGCAGAATCCCCGTACACGCGGGGATGGCT

[0557] sgRNA with SL3 deletion in SEQ ID NO:30HT001

[0558] AGCCGTTCAGCTGACGCTAAGCTTGCCCTAAACGACCGTGTATATAAGTGTGGACACTGCCACGCTGTGATTGACCGCGACCTCAACGCCGCAATCAATATCCGCCGTGAAGCAGAATCCCCGTACACGCGGGGATGGCT

[0559] sgRNA with SL4 deletion SEQ ID NO:31HT001

[0560] AGCCGTTCAGGTCATCGACCGTTACTACCCGTCATCACAGACCTGTAGCAACTGTGGCAAACGACCTGACGCTAAGCTTGCCCTAAACGACCGTGTATATAAGTGTGGACACTGCCACGCTGTGATTGACCGCGACCTCAACGCCGCAATCAATGGCT

[0561] SEQ ID NO:32HT001 F-direction primer with G mutated to R at amino acid position 362 (the bold underlined portion indicates the nucleotide to be mutated)

[0562]

[0563] SEQ ID NO: 33HT001 R-direction primer with G at amino acid position 362 mutated to R (the bold underlined portion indicates the nucleotide to be mutated)

[0564]

[0565] SEQ ID NO:34HT001 amino acid position 464 D mutated to K F direction primer (the bold underlined part is the nucleotide to be mutated)

[0566]

[0567] SEQ ID NO: 35HT001 R-direction primer with D at amino acid position 464 mutated to K (the bold underlined portion indicates the nucleotide to be mutated)

[0568]

[0569] SEQ ID NO:36NLS sequence-2

[0570] KRPAATKKAGQAKKKK

[0571] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims. Furthermore, all documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference.

Claims

1. A protein, characterized in that, The protein is selected from: (a1) a protein having the amino acid sequence shown in SEQ ID NO:1; (a2) a protein having an amino acid sequence with one or two amino acid substitutions compared to the amino acid sequence shown in SEQ ID NO:1, and the protein retains the biological function of the protein shown in SEQ ID NO:1; wherein, at position 362 of the amino acid sequence shown in SEQ ID NO:1, G is mutated to R, or at position 464, D is mutated to K, or at position 362, G is mutated to R and at position 464, D is mutated to K.

2. The protein according to claim 1, wherein The protein further contains a nuclear localization signal sequence, and the nuclear localization signal sequence is PKKKRKV or KRPAATKKAGQAKKKK.

3. The protein according to claim 2, characterized in that, The nuclear localization signal sequence is connected to the C-terminus of the sequence of (a1) or (a2) through a linker, or is connected to the N-terminus of the sequence of (a1) or (a2) through a linker.

4. The protein according to claim 3, wherein The linker is GIHGVPAA.

5. The protein according to any one of claims 1-4, characterized in that, The protein is selected from: (b1) a protein having the amino acid sequence shown in SEQ ID NO:22; (b2) a protein having an amino acid sequence with one or two amino acid substitutions compared to the amino acid sequence shown in SEQ ID NO:22, and the protein retains the biological function of the protein shown in SEQ ID NO:22; wherein, at position 401 of the amino acid sequence shown in SEQ ID NO:22, G is mutated to R, or at position 503, D is mutated to K, or at position 401, G is mutated to R and at position 503, D is mutated to K.

6. A composition, comprising: (h1) a protein component, which is selected from the proteins described in any one of claims 1-5; and (h2) a nucleic acid component, which contains an sgRNA molecule; or is a guide RNA molecule, and the guide RNA molecule contains the sgRNA molecule and a guide sequence in the 5'-to-3' direction, wherein the guide sequence contains a complementary sequence of a target sequence, wherein: (1) the sgRNA molecule is composed of a direct repeat sequence (crRNA) and a trans-acting crRNA (tracrRNA), and the direct repeat sequence (crRNA) is composed of a sequence selected from the following: (c1) the sequence shown in SEQ ID NO:3 or 4; or (c2) the sequence shown by the bases at positions 5-27 in SEQ ID NO:3 or 4; And, the sequence of (c2) retains the activity of the sequence from which it is derived as a direct repeat sequence in the CRISPR-Cas system; The trans-acting crRNA (tracrRNA) is composed of a sequence selected from the following: (d1) the sequence shown in SEQ ID NO:5; or (d2) the sequence shown by the bases at positions 1-172 in SEQ ID NO:5, or the sequence shown by the bases at positions 1-172 in SEQ ID NO:6; And, the sequence of (d2) retains the biological function of the sequence from which it is derived, and the activity of the sequence as a trans-acting crRNA (tracrRNA) in the CRISPR-Cas system; or, (2) The sgRNA molecule consists of a sequence selected from the following: (e1) The sequence shown in any one of SEQ ID NO: 14-17, 28-30; Wherein, the protein component and the nucleic acid component in the composition exist independently or combine with each other to form a complex.

7. The composition according to claim 6, wherein The sgRNA molecule consists of a sequence selected from the following: (e1) The sequence shown in SEQ ID NO:

6.

8. The composition according to claim 6, characterized in that, The sgRNA molecule contains one or more stem-loops or optimized secondary structures; the stem-loop structure is the stem-loop structure formed by the bases at positions 11-65, 104-119, 123-152, and 155-191 in the sequence shown in SEQ ID NO:

6.

9. The composition according to claim 6, wherein The guide sequence consists of a sequence selected from the following: (f1) The sequence shown in any one of SEQ ID NO: 7, 20, 23, 24, 25, 26, 27; 10. The composition according to claim 6, wherein The target sequence is a DNA or RNA sequence from a prokaryotic cell or a eukaryotic cell, or, the target sequence is a DNA or RNA sequence that does not naturally exist.

11. The composition according to claim 10, wherein The target sequence exists in a cell, or, the target sequence exists in a nucleic acid molecule in vitro.

12. The composition according to claim 11, wherein, The cell is a bacterial, plant or animal cell.

13. The composition according to claim 10, characterized in that, When the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif PAM, and the PAM has a sequence shown as 5'-AAN, wherein, N is selected from A, G, T, C.

14. The composition according to claim 6, characterized in that, The guide RNA molecule consists of a sequence selected from the following: (g1) The sequence shown in any one of SEQ ID NO: 9-12; or (g2) The sequence shown by the bases at positions 36-250 in SEQ ID NO: 9, the sequence shown by the bases at positions 36-234 in SEQ ID NO: 10, the sequence shown by the bases at positions 36-220 in SEQ ID NO: 11, or the sequence shown by the bases at positions 36-195 in SEQ ID NO: 12; And, the sequence of (g2) retains the biological function of the sequence from which it is derived, and the biological function of the sequence refers to the activity as a guide RNA sequence in the CRISPR-Cas system.

15. The composition according to any one of claims 6-14, characterized in that, The nucleic acid component further contains a promoter, and the promoter is connected to the 5' end of the nucleic acid molecule.

16. The composition according to claim 15, wherein, The promoter is the J23119 promoter.

17. The composition according to claim 16, wherein The nucleotide sequence of the J23119 promoter contains or is the first 35 positions of the nucleotide sequence shown in any one of SEQ ID NO: 9-12.

18. The composition according to any one of claims 6-14, characterized in that, The nucleic acid component further contains a poly-T sequence.

19. The composition according to claim 18, wherein, The poly-T sequence is 3 or more, 4 or more, 5 or more, 6 or more, or 7 or more consecutive Ts.

20. An isolated nucleic acid molecule, which is as follows (i1), or as follows (i1) and (i2), or as follows (i1) and (i3), or as follows (i1), (i2) and (i3): (i1) The nucleotide sequence encoding the protein according to any one of claims 1-5; (i2) The nucleotide sequence encoding the sgRNA molecule defined in any one of claims 6-8; (i3)Encoding the nucleotide sequence of the guide RNA molecule defined in any one of claims 6-14.

21. The nucleic acid molecule according to claim 20, wherein (i1)-(i3) The nucleotide sequence described in any one of the above is codon-optimized for expression in prokaryotic or eukaryotic cells.

22. A carrier, characterized in that, The vector contains the isolated nucleic acid molecule described in claim 20 or 21.

23. A carrier composition containing the carrier described in claim 22, characterized in that, The vector composition contains one or more vectors, and the one or more vectors contain: (j1) A first nucleic acid, which is the nucleotide sequence encoding the protein described in any one of claims 1-5; and (j2) A second nucleic acid, which encodes the nucleotide sequence of the sgRNA molecule defined in any one of claims 6-8, or encodes the nucleotide sequence of the guide RNA molecule defined in any one of claims 6-14; Wherein the first nucleic acid and the second nucleic acid are present on the same or different vectors.

24. The carrier composition according to claim 23, wherein, The first nucleic acid is operably linked to a first regulatory element.

25. The carrier composition according to claim 23 or 24, characterized in that, The second nucleic acid is operably linked to a second regulatory element.

26. A host cell, characterized in that: (1) The host cell contains the isolated nucleic acid molecule described in claim 20 or 21, or (2) The host cell expresses the protein described in any one of claims 1-5, the sgRNA molecule defined in any one of claims 6-8, the guide sequence defined in any one of claims 9-13, the guide RNA molecule defined in any one of claims 6-14, or (3) The host cell contains the composition described in any one of claims 6-19.

27. A kit, characterized in that, The kit contains: The protein described in any one of claims 1-5; or The protein described in any one of claims 1-5 and the sgRNA molecule defined in any one of claims 6-8; or The protein described in any one of claims 1-5, the sgRNA molecule defined in any one of claims 6-8, and the guide sequence defined in any one of claims 9-13; or The protein described in any one of claims 1-5 and the guide RNA molecule defined in any one of claims 6-14; or The composition described in any one of claims 6-19.

28. Use of the protein described in any one of claims 1-5, the composition described in any one of claims 6-19, the isolated nucleic acid molecule described in claim 20 or 21, the vector described in claim 22, the vector composition described in any one of claims 23-25, the host cell described in claim 26, or the kit described in claim 27, wherein the use is selected from: (k1) For nucleic acid editing or modification for non-therapeutic purposes; (k2) For preparing a preparation for nucleic acid editing or modification; (k3) For in vitro or ex vivo DNA detection for non-therapeutic purposes; (k4) For preparing a preparation for in vitro or ex vivo DNA detection; (k5) For non-therapeutic purposes to modify a cell, cell line, or organism by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

Citation Information

Patent Citations

  • Novel CRISPR-Cas12N enzyme and system

    CN113930412A

  • Class II V-type CRISPR system

    CN116096876A