Crispr / cas effector protein and system
By optimizing Cas protein and guide RNA, the gene editing efficiency of the CRISPR/Cas system is improved, and the problem of low efficiency of existing systems in plant gene editing is solved, and the accuracy and flexibility of the application are enhanced.
Patent Information
- Application Number
- PCT/CN2024/100739
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-06-21
- Publication Date
- 2025-05-30
AI Technical Summary
The existing CRISPR/Cas editing system has off-target effects and PAM limitations in plant gene editing, resulting in low editing efficiency and affecting the accuracy, flexibility and safety of the application.
A CRISPR/Cas complex and composition comprising a specific Cas effector protein, including optimized Cas proteins and guide RNA molecules, is provided to improve the accuracy and efficiency of gene editing through these components.
By optimizing Cas protein and guide RNA, the gene editing efficiency of the CRISPR/Cas system is improved, and the accuracy and flexibility of its application in plants is enhanced.
Smart Images

Figure PCTCN2024100739-FTAPPB-I100001 
Figure PCTCN2024100739-FTAPPB-I100002 
Figure PCTCN2024100739-FTAPPB-I100003
Abstract
Description
CRISPR / Cas effector proteins and systems
[0001] This application claims priority to the Chinese patent application filed on November 24, 2023, entitled “CRISPR / Cas effector proteins and systems” and application number 202311586038.7, the entire contents of which are incorporated herein. Technical Field
[0002] The present invention belongs to the field of nucleic acid editing, and more specifically, the present invention relates to CRISPR / Cas effector proteins and systems. Background Art
[0003] Clustered regularly interspaced short palindromic repeats (CRISPR-associated, CRISPR-Cas) are important immune defense systems used by archaea and bacteria to resist viral and plasmid infections. The CRISPR / Cas system can recognize exogenous DNA or RNA, cut them, and silence the expression of exogenous genes. Precisely because of this precise targeting function, the CRISPR / Cas system has been developed into a highly efficient gene editing tool. CRISPR / Cas gene editing uses RNA to specifically bind to target sequences on the genome and cut DNA to produce double-strand breaks, and then uses biological non-homologous end joining or homologous recombination for site-directed gene editing.
[0004] The commonly used CRISPR editing system CRISPR / Cas9 system mainly consists of endonuclease Cas9, crRNA and tracrRNA, which recognize the PAM motif of NGG. The mature crRNA unit is composed of a guide (Direct Repeat, DR) sequence and a spacer sequence (spacer). In practical applications, crRNA and tracrRNA form a chimeric single-molecule guide RNA (single guide RNA, sgRNA), and Cas9 and sgRNA form a binary complex in the cell to perform gene editing. Another commonly used CRISPR / Cas12a system belongs to Type V, which recognizes the PAM motif of TTTV, which is located at the 5' end of the spacer and produces a 4-5bp 5' overhanging end after cutting the target sequence. This system only requires Cas12a protein and crRNA to produce cutting at a specific site. At the same time, in addition to the endonuclease function, the Cas12a protein also has RNase activity, which can process the precursor crRNA (pre-crRNA) into a single mature crRNA for gene editing, so it is more convenient in the design of multi-gene editing. The two types of CRISPR / Cas systems have different advantages and are currently widely used in different gene editing studies.
[0005] Existing CRISPR / Cas editing systems are limited by inherent off-target effects. Furthermore, the PAM restriction of editable genes hinders their application in plant gene editing. Furthermore, the editing efficiency of CRISPR / Cas systems in many plants needs to be further improved. Numerous factors influence editing efficiency, including Cas gene expression, guide RNA expression, and the number and location of nuclear localization signals (NLSs). These issues impact the precision, flexibility, controllability, and safety of CRISPR editing systems during application, hindering the further expansion of their functionality and application.
[0006] Summary of the Invention
[0007] The present invention aims to provide a Cas effector protein, a fusion protein comprising such a protein, and nucleic acid molecules encoding them. The present invention also relates to CRISPR / Cas complexes and compositions for nucleic acid editing (e.g., gene or genome editing), comprising the Cas effector protein or fusion protein of the present invention, or nucleic acid molecules encoding them. The present invention also relates to a method for nucleic acid editing (e.g., gene or genome editing), which uses a Cas effector protein or fusion protein comprising the present invention.
[0008] In a first aspect of the present invention, a protein, a direct homologue, a homologue, a variant or a functional fragment thereof is provided, wherein the protein is selected from:
[0009] (a1) a protein comprising the amino acid sequence shown in SEQ ID NO: 1;
[0010] (a2) a protein comprising an amino acid sequence having one or more amino acid substitutions, deletions or additions compared to the amino acid sequence of SEQ ID NO: 1, which protein retains the biological function of the protein of SEQ ID NO: 1; or
[0011] (a3) a protein comprising an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO: 1, wherein the protein retains the biological function of the protein of SEQ ID NO: 1;
[0012] Wherein, the orthologs, homologs, variants or functional fragments retain the biological function of the protein shown in SEQ ID NO: 1.
[0013] In one or more embodiments, the biological function of the protein represented by SEQ ID NO: 1 includes or is the activity of a Cas enzyme.
[0014] In one or more embodiments, in (a2), an amino acid substitution, deletion or addition occurs at position 362 and / or position 464, preferably G at position 362 and / or D at position 464, of the amino acid sequence shown in SEQ ID NO: 1; more preferably, position 362 of the amino acid sequence shown in SEQ ID NO: 1 is mutated from G to a basic amino acid, most preferably to R, H or K; and position 464 is mutated from D to a basic amino acid, most preferably to R, H or K.
[0015] In one or more embodiments, the protein, its orthologs, homologs, variants or functional fragments further contain a functional unit, wherein the functional unit includes or is selected from: an epitope tag, a reporter gene sequence, a nuclear localization signal sequence, a targeting portion, a transcription activation domain, a transcription repression domain, a nuclease domain, a domain having an activity including or selected from the following: methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity and nucleic acid binding activity, and any combination thereof.
[0016] In one or more embodiments, the functional unit is directly connected to the N-terminus or C-terminus of the sequence (a1) to (a3), or is connected to the N-terminus or C-terminus through a linker; more preferably, the linker comprises or is selected from MA, GIHGVPAA.
[0017] In one or more embodiments, the functional unit is a nuclear localization signal sequence; more preferably, the nuclear localization signal sequence is selected from the following nuclear localization signal sequences: SV40 virus large T antigen, EGL 13, c Myc and TUS protein; most preferably, the nuclear localization signal sequence is selected from: PKKKRKV, AVKRPAATKKAGQAKKKKLD, PAAKRVKLD, MSRRRKANPTKLSENAKKLAKEVEN, KLKIKRPVK, the acidic M9 domain of hnRNP A1, the sequence KIPIK in the yeast transcription repressor Matα2, and PY NLS.
[0018] In one or more embodiments, the protein is selected from the group consisting of:
[0019] (b1) a protein comprising the amino acid sequence shown in SEQ ID NO: 22;
[0020] (b2) contains one or more amino acid substitutions, deletions, or additions compared to the amino acid sequence of SEQ ID NO: 22, and the protein retains the biological function of the protein of SEQ ID NO: 22; or
[0021] (b3) having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence of SEQ ID NO: 22, and the protein retains the biological function of the protein of SEQ ID NO: 22;
[0022] Wherein, the ortholog, homolog, variant or functional fragment retains the biological function of the protein shown in SEQ ID NO: 22.
[0023] In one or more embodiments, the biological function of the protein shown in SEQ ID NO: 22 includes or is the activity of a Cas enzyme.
[0024] In a second aspect of the present invention, an sgRNA molecule is provided, which contains a direct repeat sequence (crRNA) and optionally a trans-acting crRNA (tracrRNA).
[0025] In one or more embodiments, the direct repeat sequence (crRNA) is selected from the following sequences, or consists of the following sequences:
[0026] (c1) the sequence shown in SEQ ID NO: 3 or 4;
[0027] (c2) the sequence represented by bases 5 to 27 of SEQ ID NO: 3 or 4, or the sequence represented by bases 173 to 195 of SEQ ID NO: 6;
[0028] (c3) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in (c1) or (c2);
[0029] (c4) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in (c1) or (c2);
[0030] (c5) a sequence that hybridizes under stringent conditions to the sequence described in any one of (c1) to (c4); or
[0031] (c6) A complementary sequence of the sequence described in any one of (c1) to (c4);
[0032] Furthermore, the sequence of any one of (c2)-(c6) substantially retains the activity of the sequence from which it is derived as a direct repeat sequence in the CRISPR-Cas system.
[0033] In one or more embodiments, the trans-acting crRNA (tracrRNA) is selected from the following sequences, or consists of a sequence selected from the following sequences:
[0034] (d1) the sequence shown in SEQ ID NO: 5;
[0035] (d2) the sequence represented by bases 1 to 172 of SEQ ID NO: 5, or the sequence represented by bases 1 to 172 of SEQ ID NO: 6;
[0036] (d3) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 5;
[0037] (d4) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in SEQ ID NO: 5;
[0038] (d5) a sequence that hybridizes under stringent conditions to the sequence described in any one of (d1) to (d4); or
[0039] (d6) A complementary sequence of the sequence described in any one of (d1) to (d4);
[0040] Furthermore, the sequence of any one of (d2)-(d6) substantially retains the biological function of the sequence from which it is derived, the activity of the sequence as a trans-acting crRNA (tracrRNA) in the CRISPR-Cas system.
[0041] In one or more embodiments, the sgRNA molecule comprises or consists of a sequence selected from the following:
[0042] (e1) the sequence shown in any one of SEQ ID NOs: 14-19, 28-31;
[0043] (e2) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 14-19, 28-31;
[0044] (e3) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 14-19, 28-31;
[0045] (e4) a sequence that hybridizes under stringent conditions to the sequence described in any one of (e1) to (e3); or
[0046] (e5) A complementary sequence of the sequence described in any one of (e1) to (e3);
[0047] Furthermore, the sequence described in any one of (e2) to (e5) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a single-molecule guide RNA in the CRISPR-Cas system.
[0048] In one or more embodiments, the sgRNA molecule comprises one or more stem-loops or optimized secondary structures.
[0049] In one or more embodiments, the sequence of any of (e2)-(e5) retains one or more (eg, 2, 3, or all 4) stem-loops or secondary structures of the sequence from which it is derived.
[0050] In one or more embodiments, the sequence of any one of (e2)-(e5) retains one or more (e.g., 2, 3, or all 4) stem-loops or secondary structures of the sequence of SEQ ID NO: 6 from which it is derived, and lacks 10 to 50 bases.
[0051] In one or more embodiments, the stem-loop structure is a stem-loop structure formed by bases 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) in the sequence shown in SEQ ID NO: 6.
[0052] In a third aspect of the present invention, a guide RNA molecule is provided, which comprises, from 5' to 3' direction, the sgRNA molecule of the present invention and a guide sequence, wherein the guide sequence comprises a complementary sequence to the target sequence.
[0053] In one or more embodiments, the targeting sequence comprises or consists of a sequence selected from the group consisting of:
[0054] (f1) the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0055] (f2) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0056] (f3) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0057] (f4) a sequence that hybridizes under stringent conditions to the sequence described in any one of (f1) to (f3); or
[0058] (f5) A complementary sequence of the sequence described in any one of (f1) to (f3);
[0059] Furthermore, the sequence described in any one of (f2) to (f5) substantially retains the biological function of the sequence from which it is derived; more preferably, the biological function of the sequence refers to the activity as a guide sequence in the CRISPR-Cas system.
[0060] In one or more embodiments, the guide RNA molecule comprises or consists of a sequence selected from the group consisting of:
[0061] (g1) the sequence shown in any one of SEQ ID NOs: 9-13;
[0062] (g2) the sequence set forth at bases 36 to 250 of SEQ ID NO:9, the sequence set forth at bases 36 to 234 of SEQ ID NO:10, the sequence set forth at bases 36 to 220 of SEQ ID NO:11, the sequence set forth at bases 36 to 195 of SEQ ID NO:12, or the sequence set forth at bases 36 to 213 of SEQ ID NO:13;
[0063] (g3) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 9-13;
[0064] (g4) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 9-13;
[0065] (g5) a sequence that hybridizes under stringent conditions to the sequence described in any one of (g1) to (g4); or
[0066] (g6) A complementary sequence of the sequence described in any one of (g1) to (g4);
[0067] Furthermore, the sequence described in any one of (g2) to (g6) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a guide RNA sequence in the CRISPR-Cas system.
[0068] In one or more embodiments, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell, or alternatively, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0069] In one or more embodiments, the target sequence is present in a cell, or alternatively, the target sequence is present in a nucleic acid molecule in vitro.
[0070] In one or more embodiments, the cell is a bacterial, plant, or animal cell.
[0071] In one or more embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer sequence adjacent to the continued PAM, and the PAM has a sequence represented by 5'-AAN, wherein N is selected from A, G, T, and C.
[0072] In a fourth aspect of the present invention, a composition is provided, comprising:
[0073] (h1) a protein component selected from the group consisting of: a protein according to the present invention, its orthologs, homologs, variants or functional fragments, and any combination thereof; and
[0074] (h2) a nucleic acid component comprising the sgRNA molecule described in the present invention; or a guide RNA molecule described in the present invention;
[0075] Wherein, the protein component and the nucleic acid component in the composition exist independently, or are combined with each other to form a complex.
[0076] Therefore, it can also be said that the present invention provides a complex, which comprises the protein component and the nucleic acid component, which are combined with each other to form a complex.
[0077] In one or more embodiments, the nucleic acid component comprises the sgRNA molecule of the present invention and the guide sequence of the present invention from the 5' to the 3' direction.
[0078] In one or more embodiments, the nucleic acid component further comprises a promoter, which is connected to the 5' end of the nucleic acid molecule; preferably, the promoter is the J23119 promoter; more preferably, the nucleotide sequence of the J23119 promoter comprises or is positions 1-35 of the nucleotide sequence shown in any one of SEQ ID NOs: 9-13.
[0079] In one or more embodiments, the nucleic acid component further comprises a poly-T sequence, more preferably the poly-T sequence is continuous with more than 3, more than 4, more than 5, more than 6, more than 7 or more Ts.
[0080] In a fifth aspect of the present invention, an isolated nucleic acid molecule is provided, comprising:
[0081] (i1) a nucleotide sequence encoding the protein of the present invention;
[0082] (i2) a nucleotide sequence encoding the sgRNA molecule of the present invention;
[0083] (i3) a nucleotide sequence encoding the guide RNA molecule of the present invention; and / or,
[0084] (i4) contains the nucleotide sequence of (i1)-(i3).
[0085] In one or more embodiments, the nucleotide sequence described in any one of (i1)-(i4) is codon-optimized for expression in a prokaryotic or eukaryotic cell.
[0086] In a sixth aspect of the present invention, a vector, a composition containing the vector, a host cell or a kit is provided, wherein:
[0087] The vector comprises the isolated nucleic acid molecule of the present invention;
[0088] The carrier-containing composition comprises one or more carriers, and the one or more carriers comprise:
[0089] (j1) a first nucleic acid, which is a nucleotide sequence encoding the protein of the present invention, its ortholog, homolog, variant or functional fragment; optionally, the first nucleic acid is operably linked to a first regulatory element; and
[0090] (j2) a second nucleic acid encoding a nucleotide sequence comprising an sgRNA molecule of the present invention, or encoding a nucleotide sequence of a guide RNA molecule of the present invention; optionally, the second nucleic acid is operably linked to a second regulatory element;
[0091] wherein the first nucleic acid and the second nucleic acid are present on the same or different vectors;
[0092] The host cell comprises the isolated nucleic acid molecule of the present invention, and / or expresses the protein of the present invention, its ortholog, homolog, variant or functional fragment, the sgRNA molecule of the present invention, the guide sequence of the present invention, the guide RNA molecule of the present invention, and / or contains the composition of the present invention;
[0093] The kit contains the protein of the present invention, its ortholog, homolog, variant or functional fragment, the sgRNA molecule of the present invention, the guide sequence of the present invention, the guide RNA molecule of the present invention, and / or the composition of the present invention.
[0094] In a seventh aspect, the present invention provides a use of the protein, its ortholog, homolog, variant or functional fragment thereof, the sgRNA molecule, the guide RNA molecule, the complex or composition, the isolated nucleic acid molecule, the vector, the vector composition, the delivery composition, the host cell or the kit according to the present invention, wherein the use is selected from:
[0095] (k1) for nucleic acid editing or modification;
[0096] (k2) Preparations for nucleic acid editing or modification;
[0097] (k3) in vitro or ex vivo DNA testing;
[0098] (k4) Preparations for in vitro or ex vivo DNA detection;
[0099] (k5) Used to modify cells, cell lines or organisms by changing one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.
[0100] Other aspects of the invention will be apparent to those skilled in the art in view of the disclosure herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] Figure 1. Secondary structure of sgRNA-WT of HT001.
[0102] Figure 2. In vitro cleavage results of HT001 guided by crRNA, tracrRNA, crRNA-tracrRNA mixture, and sgRNA-WT.
[0103] Figure 3. PAM domain structure of HT001.
[0104] Figure 4. Schematic diagram of the sgRNA stem-loop structure of HT001.
[0105] Figure 5. Schematic diagram of colony growth of sgRNA-WT and its four individual stem-loop structure deletions.
[0106] Figure 6. Schematic diagram of the engineering modification of HT001 truncated sgRNA (SL4).
[0107] Figure 7. In vitro cleavage electrophoresis of HT001sgRNA after engineering modification.
[0108] Figure 8. Schematic diagram of the HT001 cleavage site.
[0109] Figure 9. Editing efficiency of sgRNAWT-HT001 and sgRNAR1-HT001 in tobacco. Deletion 2, Deletion 4, Deletion 5, Deletion 7, and Deletion 9 represent 2, 4, 5, 6, and 9 bp deletions, respectively.
[0110] Figure 10. Editing efficiency of sgRNAWT-HT001 and sgRNAR1-HT001 in human cells.
[0111] Figure 11. Editing efficiency of HT001 protein wild type (WT) single mutation and double mutation in tobacco. DETAILED DESCRIPTION
[0112] After a lot of experiments and repeated exploration, the inventors discovered a new type of nuclease (Cas enzyme), the sequence of which is shown in SEQ ID NO: 1. By predicting the high-level structure of the repetitive sequence and experimental verification, the crRNA and tracrRNA sequences of the Cas enzyme were further obtained. Based on this discovery, the inventors developed a new CRISPR / Cas complex that exhibits nuclease activity in bacteria, plants, and animals, and obtained a more efficient CRISPR / sgRNAR1-HT001 gene editing system through optimization.
[0113] Definition of terms
[0114] Unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, procedures in molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA used herein are conventional procedures widely used in the relevant fields. To facilitate a better understanding of the present invention, definitions and explanations of relevant terms are provided below.
[0115] The terms "Cas protein", "Cas enzyme", "Cas effector protein", "CRISPR / Cas effector protein" and "CRISPR enzyme" are used interchangeably. Nucleic acid cleavage or cleavage of nucleic acids herein include: DNA or RNA breakage in the target nucleic acid (Cis cleavage) produced by the Cas enzymes described herein, breakage of DNA or RNA in a side branch nucleic acid substrate (single-stranded nucleic acid substrate) (i.e., non-specific or non-targeted, Trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.
[0116] The term "homolog" has the meaning commonly understood by those skilled in the art, and generally refers to substances that have similarities in sequence or structure and are derived from a common ancestral molecule through chemotaxis or evolution. A "homolog" of a protein herein refers to a protein that has similarities in sequence or structure to the protein.
[0117] The terms "ortholog," "orthologue," and "ortholog" are used interchangeably and have the meanings commonly understood by those skilled in the art. As a further guide, an "ortholog" of a protein refers to a protein from a different species that performs the same or a similar function as the protein to which it is an ortholog.
[0118] The term "variant" or "mutant" has the meaning generally understood by those skilled in the art, which is different from the "wild type" and refers to an atypical form of an organism, strain, or gene or a form that is different from its natural form, which is usually intentionally modified by humans.
[0119] The term "functional fragment" has the meaning commonly understood by those skilled in the art and refers to a fragment having a biological function. For example, a "functional fragment" of a protein refers to a fragment of the protein that can exert its biological function (e.g., the enzymatic activity of the protein).
[0120] The terms “clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system” and “CRISPR system” are used interchangeably and have the meanings commonly understood by those skilled in the art, which generally include transcripts or other elements associated with the expression of CRISPR-associated (“Cas”) genes, or transcripts or other elements capable of directing the activity of the Cas genes.
[0121] The term "CRISPR / Cas complex" refers to a complex formed by the binding of guide RNA (guide RNA) and Cas protein, which includes a guide sequence that hybridizes to the target sequence and an sgRNA bound to the Cas protein. The complex is capable of recognizing and cleaving polynucleotides that can hybridize to the guide RNA.
[0122] The terms "guide RNA," "guide RNA," "guide RNA," "gRNA," "guide sequence," or "guide sequence" are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, a guide RNA can comprise, consist essentially of, or consist of an sgRNA and a guide sequence. In certain instances, a guide sequence is any polynucleotide sequence that has sufficient complementarity to a target sequence to hybridize to the target sequence and direct specific binding of a CRISPR / Cas complex to the target sequence.
[0123] The terms "single molecule guide RNA", "sgRNA" or "single guide RNA" are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally speaking, sgRNA is capable of binding to the Cas protein. In some cases, the sgRNA contains only or consists of direct repeat sequences (crRNA). In some cases, the sgRNA contains direct repeat sequences (crRNA) and also contains trans-acting crRNA (tracrRNA). In some cases (e.g., when the sgRNA is modified or optimized), the sgRNA contains a sequence in the direct repeat sequence (crRNA) and also contains a sequence in the trans-acting crRNA (tracrRNA).
[0124] The term "direct repeat" can refer to the DNA coding sequence in the CRISPR locus, or to the RNA encoded by crRNA.
[0125] Therefore, in the context of RNA, when referring to a guide RNA, sgRNA, direct repeat sequence, or guide sequence, each T should be understood to represent a U.
[0126] The term "target sequence" refers to a polynucleotide targeted by a guide sequence in a gRNA, such as a sequence having complementarity with the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex (including Cas proteins and gRNA). Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex. The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the nucleus or cytoplasm of the cell. In some cases, the target sequence can be located in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast. A sequence or template that can be used to recombine into a target locus comprising the target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence". In the present invention, a "target sequence" or "target polynucleotide" or "target nucleic acid" can be any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell).
[0127] The term "vector" refers to a nucleic acid molecule that is capable of transporting another nucleic acid molecule to which it is attached. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends, or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other various polynucleotides known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection so that the genetic material elements it carries are expressed in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. In addition, a vector may also contain a replication initiation site. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences are present in a vector for packaging viruses (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by the virus for transfection into a host cell. Some vectors (e.g., bacterial vectors and additional mammalian vectors with bacterial replication origins) can replicate autonomously in the host cell into which they are introduced. Other vectors (e.g., non-additional mammalian vectors) are integrated into the genome of the host cell after introducing the host cell, and thus replicate together with the host genome. Moreover, some vectors can instruct the expression of the gene that they are operably connected. Such vectors are referred to as "expression vectors" at this.
[0128] The term "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0129] The term "5'UTR" refers to the 5' untranslated region, which may have a regulatory role.
[0130] The term "3'UTR" refers to the 3' untranslated region, which may have a regulatory role
[0131] The term "promoter" has a meaning well known to those skilled in the art and refers to a non-coding nucleotide sequence located upstream of a gene that can initiate expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell essentially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell essentially only when the cell is a cell of the tissue type corresponding to the promoter.
[0132] The term "codon" refers to the pattern in which three adjacent nucleotides in a messenger RNA molecule are grouped together to represent a certain amino acid during protein synthesis.
[0133] The term "codon optimization" refers to a new technology that improves protein expression levels in organisms by increasing the translation efficiency of target genes.
[0134] The term "nuclear localization signal," "nuclear localization sequence," or "NLS" is an amino acid sequence that "tags" a protein for import into the cell nucleus via nuclear transport, i.e., proteins with an NLS are transported to the cell nucleus. Typically, an NLS comprises a positively charged Lys or Arg residue exposed on the surface of the protein.
[0135] Cas effector proteins
[0136] The present invention provides a protein having an amino acid sequence as shown in SEQ ID NO: 1 or a direct homologue, homologue, variant or functional fragment thereof; wherein the direct homologue, homologue, variant or functional fragment substantially retains the biological function of the sequence from which it is derived.
[0137] In the present invention, the biological functions of the above sequences include, but are not limited to, the activity of binding to guide RNA, endonuclease activity, and the activity of binding to and cutting specific sites of the target sequence under the guidance of the guide RNA.
[0138] In certain embodiments, the orthologs, homologs, and variants have at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in SEQ ID NO: 1, and substantially retain the biological function of the protein shown in SEQ ID NO: 1 (e.g., activity of binding to a guide RNA, endonuclease activity, activity of binding to and cutting a specific site of a target sequence under the guidance of a guide RNA).
[0139] In certain embodiments, the protein is an effector protein in a CRISPR / Cas system.
[0140] In certain embodiments, the protein of the present invention comprises or consists of a sequence selected from the group consisting of:
[0141] (i) the sequence shown in SEQ ID NO: 1;
[0142] (ii) a sequence having one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 1;
[0143] (iii) a sequence having at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to the sequence set forth in SEQ ID NO: 1; or
[0144] (iv) a direct homologue, homologue, variant or functional fragment of the sequence shown in SEQ ID NO: 1, while retaining the biological function of the protein shown in SEQ ID NO: 1.
[0145] In certain embodiments, the substitution, deletion or addition of amino acid in (ii) may occur at position 362 and / or position 464 of the amino acid sequence shown in SEQ ID NO: 1, preferably the substitution, deletion or addition of amino acid in (ii) occurs at position 362 G and / or position 464 D.
[0146] In certain embodiments, in (ii), position 362 of the amino acid sequence shown in SEQ ID NO: 1 is mutated from G to a basic amino acid (including R, H, K), for example, to R.
[0147] In certain embodiments, in (ii), position 464 of the amino acid sequence shown in SEQ ID NO: 1 is mutated from D to a basic amino acid (including R, H, K), for example, to K.
[0148] In certain embodiments, in (ii), position 362 of the amino acid sequence set forth in SEQ ID NO: 1 is mutated from G to a basic amino acid (including R, H, K, preferably R), and position 464 is also mutated from D to a basic amino acid (including R, H, K, preferably K). In certain specific embodiments, in (ii), position 362 of the amino acid sequence set forth in SEQ ID NO: 1 is mutated from G to R, and / or position 464 is mutated from D to K.
[0149] In certain embodiments, the primers shown in SEQ ID NOs: 32 and 33 are used to mutate position 362 of the amino acid sequence shown in SEQ ID NO: 1 from G to R.
[0150] In certain embodiments, the primers shown in SEQ ID NOs: 34 and 35 are used to mutate position 464 of the amino acid sequence shown in SEQ ID NO: 1 from D to K.
[0151] In certain embodiments, the protein of the present invention has the amino acid sequence shown in SEQ ID NO:1.
[0152] In certain embodiments, the protein shown in SEQ ID NO: 1 can be obtained by transcription and translation of the nucleotide sequence shown in SEQ ID NO: 2.
[0153] Derived Cas effector proteins
[0154] The protein of the present invention can be derivatized, for example, by being linked to another molecule (e.g., another polypeptide or protein). Typically, the derivatization (e.g., labeling) of the protein will not adversely affect the desired activity of the protein (e.g., activity binding to a guide RNA, endonuclease activity, activity binding to and cutting a specific site of a target sequence under the guidance of a guide RNA). Therefore, the protein of the present invention is also intended to include such derivatized forms. For example, the protein of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage or other means) to one or more other molecular groups, such as another protein or polypeptide, a detection reagent, a pharmaceutical agent, etc.
[0155] In particular, the Cas protein of the present invention can be connected to a functional unit.Herein, the functional unit can be selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting moieties, transcription activation domains (e.g., VP64), transcription repression domains (e.g., KRAB domains or SID domains), nuclease domains (e.g., Fok1), with a domain selected from the following activities: methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity and nucleic acid binding activity; and any combination thereof.
[0156] In certain embodiments, the functional unit is directly connected to the N-terminus or C-terminus of the protein of the present invention, or is connected to the N-terminus or C-terminus of the protein of the present invention through a linker. Such linkers are well known in the art, and examples include but are not limited to linkers comprising one or more (e.g., 1, 2, 3, 4, or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc. In certain embodiments, the linker comprises the sequence shown in MA. In certain embodiments, the linker comprises the sequence shown in GIHGVPAA.
[0157] In some embodiments, the functional unit is an NLS. NLS can improve the ability of the Cas protein of the present invention to enter the cell nucleus. In these embodiments, the Cas protein of the present invention may comprise one or more NLS sequences. The NLS sequence may be located at, near or close to the end (e.g., N-terminus or C-terminus) of the Cas protein of the present invention. In certain embodiments, the Cas protein comprises two NLS sequences, which are respectively located at, near or close to the N-terminus and C-terminus of the Cas protein of the present invention.
[0158] Exemplary NLS sequences include, but are not limited to, NLSs from the following: SV40 virus large T antigen, EGL 13, cMyc, and TUS protein. In certain embodiments, the NLS comprises the PKKKRKV sequence. In certain embodiments, the NLS comprises the AVKRPAATKKAGQAKKKKLD sequence. In certain embodiments, the NLS comprises the PAAKRVKLD sequence. In certain embodiments, the NLS comprises the MSRRRKANPTKLSENAKKLAKEVEN sequence. In certain embodiments, the NLS comprises the KLKIKRPVK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the sequence KIPIK and PY NLS in the yeast transcription repressor Mata2. In certain exemplary embodiments, the NLS sequence is as shown in SEQ ID NO: 21 or SEQ ID NO: 36. In certain embodiments, the NLS sequence is located at, near, or close to a terminus (e.g., N-terminus or C-terminus) of the protein of the present invention.
[0159] In certain embodiments, the functional unit is an epitope tag, including but not limited to an antibody tag (e.g., Myc, HA), a fluorescent tag (e.g., FLAG), an enzyme tag (e.g., GST), a biotin tag, a metal tag (e.g., His), or a sugar tag. In certain embodiments, the epitope tag is a FLAG tag. In certain examples, the FLAG tag comprises the sequence MDYKDHDGDYKDHDIDYKDDDDK.
[0160] In certain embodiments, the derivative Cas effector protein of the invention comprises or consists of a sequence selected from the group consisting of:
[0161] (i) the sequence shown in SEQ ID NO: 22;
[0162] (ii) a sequence having one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 22; or
[0163] (iii) a sequence having at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to the sequence set forth in SEQ ID NO: 22;
[0164] (iv) an ortholog, homolog, variant or functional fragment of the sequence shown in SEQ ID NO: 22, while retaining the biological function of the protein shown in SEQ ID NO: 22.
[0165] In certain embodiments, the amino acid substitution, deletion or addition in (ii) may occur at position 401 and / or position 503 of the amino acid sequence shown in SEQ ID NO: 22, preferably the amino acid substitution, deletion or addition in (ii) occurs at position 401 G and / or position 503 D.
[0166] In certain embodiments, in (ii), position 401 of the amino acid sequence shown in SEQ ID NO: 22 is mutated from G to a basic amino acid (including R, H, K), for example, to R.
[0167] In certain embodiments, in (ii), position 503 of the amino acid sequence shown in SEQ ID NO: 22 is mutated from D to a basic amino acid (including R, H, K), for example, to K.
[0168] In certain embodiments, in (ii), position 401 of the amino acid sequence shown in SEQ ID NO: 22 is mutated from G to a basic amino acid (including R, H, K, preferably R), and position 503 is also mutated from D to a basic amino acid (including R, H, K, preferably K).
[0169] In certain embodiments, the protein of the present invention has the amino acid sequence shown in SEQ ID NO:22.
[0170] In some embodiments, the Cas protein of the present invention can be linked to a targeting moiety to enable the Cas protein to be targeted. For example, it can be linked to a detectable label to facilitate detection of the Cas protein of the present invention. For example, it can be linked to an epitope tag to facilitate expression, detection, tracing and / or purification of the Cas protein of the present invention.
[0171] The derivative Cas effector protein of the present invention can be produced by genetic engineering methods (recombinant technology) or chemical synthesis methods.
[0172] Single-molecule guide RNA (sgRNA) molecules
[0173] The present invention also provides an isolated nucleic acid molecule comprising or consisting of a sequence selected from the following:
[0174] (i) a sequence shown in any one of SEQ ID NOs: 4-6, 14-19, and 28-31;
[0175] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 4-6, 14-19 and 28-31;
[0176] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 4-6, 14-19, and 28-31;
[0177] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or
[0178] (v) a complementary sequence of the sequence described in any one of (i) to (iii);
[0179] Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence is the activity as a single-molecule guide RNA in the CRISPR-Cas system.
[0180] In certain embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:
[0181] (a) the nucleotide sequence shown in any one of SEQ ID NOs: 4-6, 14-19, and 28-31;
[0182] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or
[0183] (c) The complementary sequence of the sequence described in (a).
[0184] In certain embodiments, the isolated nucleic acid molecule is RNA.
[0185] In certain embodiments, the isolated nucleic acid molecule is a single-molecule guide RNA in a CRISPR-Cas system.
[0186] In certain embodiments, the isolated nucleic acid molecule comprises a direct repeat sequence (crRNA).
[0187] In certain embodiments, the isolated nucleic acid molecule comprises a trans-acting crRNA (tracrRNA).
[0188] In certain embodiments, the isolated nucleic acid molecule comprises a direct repeat sequence (crRNA) and a trans-acting crRNA (tracrRNA).
[0189] In certain embodiments, the direct repeat sequence (crRNA) comprises or consists of a sequence selected from the group consisting of:
[0190] (i) the sequence shown in SEQ ID NO: 3 or 4;
[0191] (ii) the sequence represented by bases 5 to 27 of SEQ ID NO: 3 or 4, or the sequence represented by bases 173 to 195 of SEQ ID NO: 6;
[0192] (iii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 3 or 4;
[0193] (iv) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence set forth in SEQ ID NO: 3 or 4;
[0194] (v) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iv); or
[0195] (vi) a complementary sequence of the sequence described in any one of (i) to (iv);
[0196] Furthermore, the sequence described in any one of (ii) to (vi) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a direct repeat sequence in the CRISPR-Cas system.
[0197] In certain embodiments, the direct repeat sequence (crRNA) comprises or consists of the sequence shown in SEQ ID NO: 3 or 4.
[0198] In certain embodiments, the direct repeat sequence (crRNA) comprises or is the sequence shown in bases 5 to 27 of SEQ ID NO: 3 or 4.
[0199] In certain embodiments, the trans-acting crRNA (tracrRNA) comprises or consists of a sequence selected from the group consisting of:
[0200] (i) the sequence shown in SEQ ID NO: 5;
[0201] (ii) the sequence represented by bases 1 to 172 of SEQ ID NO: 5, or the sequence represented by bases 1 to 172 of SEQ ID NO: 6;
[0202] (iii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 5;
[0203] (iv) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence set forth in SEQ ID NO: 5;
[0204] (v) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iv); or
[0205] (vi) a complementary sequence of the sequence described in any one of (i) to (iv);
[0206] Furthermore, the sequence of any one of (ii) to (vi) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a trans-acting crRNA (tracrRNA) in the CRISPR-Cas system.
[0207] In certain embodiments, the isolated nucleic acid molecule consists of a direct repeat sequence (crRNA) and a trans-acting crRNA (tracrRNA), for example, it comprises the sequence shown in SEQ ID NO: 6 or consists of the sequence shown in SEQ ID NO: 6.
[0208] In certain embodiments, the trans-acting crRNA (tracrRNA) comprises or is the sequence shown in bases 1-172 of SEQ ID NO:5.
[0209] In certain embodiments, the single-molecule guide RNA sequence is in a truncated form, for example, it lacks 10 to 50 (eg, 10 to 40) bases compared to the sequence of SEQ ID NO: 6, and retains the stem-loop structure of the single-molecule guide RNA.
[0210] In certain embodiments, the single-molecule guide RNA lacks the stem-loop structure formed by bases 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) in the sequence shown in SEQ ID NO: 6, respectively.
[0211] In certain embodiments, the single-molecule guide RNA sequence is in a truncated form, and compared with the sequence shown in SEQ ID NO: 6, bases 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) are missing, respectively.
[0212] In certain embodiments, the stem-loop structure of the single-molecule guide RNA is a stem-loop structure formed by bases 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) in the sequence shown in SEQ ID NO: 6.
[0213] In certain embodiments, the structure of the single-molecule guide RNA is shown in FIG3 .
[0214] In certain embodiments, the truncated form of the single-molecule guide RNA comprises or consists of a sequence selected from the group consisting of:
[0215] (i) the sequence shown in any one of SEQ ID NOs: 28-31;
[0216] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 28-31;
[0217] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 28-31;
[0218] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or
[0219] (v) a complementary sequence of the sequence described in any one of (i) to (iii);
[0220] Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence is the activity as a single-molecule guide RNA in the CRISPR-Cas system.
[0221] In certain embodiments, the truncated form of the single-molecule guide RNA comprises or consists of a sequence selected from the group consisting of:
[0222] (a) the nucleotide sequence shown in any one of SEQ ID NOs: 28-31;
[0223] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or
[0224] (c) The complementary sequence of the sequence described in (a).
[0225] In certain embodiments, the truncated form of the single-molecule guide RNA retains only a portion of the stem-loop structure formed by bases 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) in the sequence shown in SEQ ID NO: 6.
[0226] In certain embodiments, the structure of the truncated form of the single-molecule guide RNA is shown in FIG4 .
[0227] In certain embodiments, the truncated form of the single-molecule guide RNA (sgRNA) comprises or consists of a sequence selected from the group consisting of:
[0228] (i) the sequence shown in any one of SEQ ID NOs: 14-19;
[0229] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 14-19;
[0230] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 14-19;
[0231] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or
[0232] (v) a complementary sequence of the sequence described in any one of (i) to (iii);
[0233] Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence is the activity as a single-molecule guide RNA in the CRISPR-Cas system.
[0234] Guide sequence
[0235] The present invention also provides a guide sequence capable of hybridizing to a target sequence.
[0236] In certain embodiments, the guide sequence is linked to the 3' end of the nucleic acid molecule (eg, sgRNA).
[0237] In certain embodiments, the guide sequence comprises the complement of the target sequence.
[0238] In some embodiments, the targeting sequence comprises or consists of a sequence selected from the group consisting of:
[0239] (i) the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0240] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0241] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0242] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or
[0243] (v) a complementary sequence of the sequence described in any one of (i) to (iii);
[0244] Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a guide sequence in the CRISPR-Cas system.
[0245] Guide RNA sequence
[0246] The present invention also provides a guide RNA sequence, which comprises, from the 5' to the 3' direction, the sgRNA sequence of the present invention and a guide sequence capable of hybridizing with a target sequence.
[0247] In certain embodiments, the guide RNA sequence comprises a sequence selected from the group consisting of:
[0248] (a) an sgRNA sequence comprising or consisting of a sequence selected from the following:
[0249] (i) SEQ ID NO: a sequence shown in any one of SEQ ID NOs: 4-6, 14-19, and 28-31;
[0250] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 4-6, 14-19 and 28-31;
[0251] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence set forth in any one of SEQ ID NOs: 4-6, 14-19, and 28-31;
[0252] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or
[0253] (v) a complementary sequence of the sequence described in any one of (i) to (iii);
[0254] Furthermore, the sequence of any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence is the activity as an sgRNA sequence in the CRISPR-Cas system;
[0255] (b) a guide sequence that hybridizes to the target sequence, comprising or consisting of a sequence selected from the group consisting of:
[0256] (i) the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0257] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0258] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27;
[0259] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or
[0260] (v) a complementary sequence of the sequence described in any one of (i) to (iii);
[0261] Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a guide sequence in the CRISPR-Cas system.
[0262] In certain embodiments, the guide RNA sequence comprises or consists of a sequence selected from the group consisting of:
[0263] (i) the sequence shown in any one of SEQ ID NOs: 9-13;
[0264] (ii) the sequence set forth at bases 36 to 250 of SEQ ID NO:9, the sequence set forth at bases 36 to 234 of SEQ ID NO:10, the sequence set forth at bases 36 to 220 of SEQ ID NO:11, the sequence set forth at bases 36 to 195 of SEQ ID NO:12, or the sequence set forth at bases 36 to 213 of SEQ ID NO:13;
[0265] (iii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 9-13;
[0266] (iv) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 9-13;
[0267] (v) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iv); or
[0268] (vi) a complementary sequence of the sequence described in any one of (i) to (iv);
[0269] Furthermore, the sequence described in any one of (ii) to (vi) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a guide RNA sequence in the CRISPR-Cas system.
[0270] In certain embodiments, the guide RNA sequence comprises or consists of a sequence selected from the group consisting of:
[0271] (a) the nucleotide sequence shown in any one of SEQ ID NOs: 9-13;
[0272] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or
[0273] (c) The complementary sequence of the sequence described in (a).
[0274] CRISPR / Cas complex
[0275] The present invention also provides a composite comprising:
[0276] (i) a protein component selected from: a Cas effector protein or a derived Cas effector protein according to the present invention; and
[0277] (ii) a nucleic acid component comprising, from 5' to 3' direction, the sgRNA sequence of the present invention and a guide sequence capable of hybridizing with the target sequence;
[0278] Wherein, the protein component and the nucleic acid component combine with each other to form a complex.
[0279] In certain embodiments, the nucleic acid component is RNA.
[0280] In certain embodiments, the nucleic acid component is a guide RNA in a CRISPR-Cas system, such as a guide RNA sequence as defined in any embodiment of the invention.
[0281] In certain embodiments, the nucleic acid component further comprises a promoter (such as, but not limited to, the J23119 promoter) linked to the 5' end of the nucleic acid molecule.
[0282] In certain specific embodiments, the nucleotide sequence of the J23119 promoter comprises or is: positions 1-35 of the nucleotide sequence shown in any one of SEQ ID NOs: 9-13.
[0283] In certain embodiments, the nucleic acid component further comprises a poly-T sequence, for example, more than 3, more than 4, more than 5, more than 6, more than 7 or more consecutive Ts.
[0284] Encoding nucleic acid, vector and host cell
[0285] The present invention also provides an isolated nucleic acid molecule comprising:
[0286] (i) a nucleotide sequence encoding the Cas effector protein or a derived Cas effector protein of the present invention;
[0287] (ii) a nucleotide sequence encoding an sgRNA sequence or guide RNA as described in the present invention; or
[0288] (iii) a nucleotide sequence comprising (i) and (ii).
[0289] In some embodiments, the nucleic acid molecule encodes a protein described herein having an amino acid sequence as shown in SEQ ID NO: 1 or 22, or an ortholog, homolog, variant, or functional fragment thereof.
[0290] In some embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:
[0291] (I) the sequence shown in SEQ ID NO: 2;
[0292] (II) a sequence having one or more base substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 base substitutions, deletions, or additions) compared to the sequence shown in SEQ ID NO: 2;
[0293] (II) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence set forth in SEQ ID NO: 2;
[0294] (IV) a sequence that hybridizes under stringent conditions to the sequence described in any one of (I) to (III); or
[0295] (V) A complementary sequence of the sequence described in any one of (I) to (III).
[0296] In certain embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:
[0297] (a) the nucleotide sequence shown in SEQ ID NO: 2;
[0298] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or
[0299] (c) The complementary sequence of the sequence described in (a).
[0300] In certain embodiments, the nucleotide sequence described in any one of (i)-(iii) is codon-optimized for expression in prokaryotes. In certain embodiments, the nucleotide sequence described in any one of (i)-(iii) is codon-optimized for expression in eukaryotic cells.
[0301] The present invention also provides a vector comprising an isolated nucleic acid molecule as described herein. The vector of the present invention can be a cloning vector or an expression vector. In certain embodiments, the vector of the present invention is, for example, a plasmid, a cosmid, a phage, a cosmid, or the like. In certain preferred embodiments, the vector is capable of expressing the Cas effector protein and its derivatives, isolated nucleic acid molecule, or CRISPR / Cas complex of the present invention in bacteria, plants, and animals.
[0302] The present invention also provides a host cell comprising the isolated nucleic acid molecule or vector as described above. Such host cells include, but are not limited to, prokaryotic cells such as Escherichia coli cells, and eukaryotic cells such as yeast cells, insect cells, plant cells (such as Arabidopsis cells, tobacco cells) and animal cells (such as mammalian cells, such as mouse cells, human cells, etc.). The cells of the present invention can also be cell lines, such as HEK293 cells.
[0303] Compositions and carrier compositions
[0304] The present invention also provides a composition comprising:
[0305] (i) a first component selected from the group consisting of: a Cas effector protein of the present invention (including a derivatized form thereof), a nucleotide sequence encoding the Cas effector protein or a derivative protein, and any combination thereof; and
[0306] (ii) a second component, which is a nucleotide sequence comprising a guide RNA, or a coding sequence thereof;
[0307] The guide RNA comprises an sgRNA sequence and a guide sequence from the 5' to the 3' direction, and the guide sequence can hybridize with the target sequence.
[0308] In certain embodiments, the first component and the second component of the composition exist independently.
[0309] In certain embodiments, the guide RNA is capable of forming a complex with the protein or derived protein described in (i).
[0310] In certain embodiments, the sgRNA sequence is an isolated nucleic acid molecule as defined in any embodiment of the present invention.
[0311] In certain embodiments, the guide RNA sequence is linked to the 3' end of the sgRNA sequence. In certain embodiments, the guide RNA sequence comprises a complementary sequence to the target sequence.
[0312] The present invention also provides a carrier composition comprising one or more carriers, wherein the one or more carriers comprise:
[0313] (i) a first nucleic acid, which is a nucleotide sequence encoding a Cas effector protein or fusion protein of the present invention; optionally, the first nucleic acid is operably linked to a first regulatory element; and
[0314] (ii) a second nucleic acid encoding a nucleotide sequence comprising a guide RNA sequence; optionally, the second nucleic acid is operably linked to a second regulatory element;
[0315] in:
[0316] The first nucleic acid and the second nucleic acid are present on the same or different vectors;
[0317] The guide RNA is capable of forming a complex with the effector protein or fusion protein described in (i).
[0318] In certain embodiments, the sgRNA sequence is an isolated nucleic acid molecule as defined in any embodiment of the present invention, including truncated forms thereof.
[0319] In certain embodiments, the guide RNA sequence is a guide RNA sequence as defined in any embodiment of the present invention.
[0320] In certain embodiments, the composition is non-naturally occurring or modified. In certain embodiments, at least one component in the composition is non-naturally occurring or modified.
[0321] In certain embodiments, the first regulatory element is a promoter, such as a Pol II type promoter (including but not limited to: ZmUbi, Actin, CmYLCV, UBQ, 35S, SPL), a tissue-specific promoter (including but not limited to: YAO, CDC45, rbcS), an inducible promoter (including but not limited to: XEV).
[0322] In certain embodiments, the second regulatory element is a promoter, such as a Pol II type or a Pol III type promoter (including but not limited to: ZmU6, OsU3, OsU6a, OsU6b, OsU6c, AtU6, Actin, 35S, Ubi, UBQ, SPL, CmYLCV), a tissue-specific promoter (including but not limited to: YAO, CDC45, rbcS), an inducible promoter (including but not limited to: XEV).
[0323] In certain embodiments, a type of vector is a plasmid, which refers to a circular double-stranded DNA loop in which other DNA fragments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which virally derived DNA or RNA sequences are present in a vector for packaging viruses (for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by viruses for transfection into a host cell. Some vectors (for example, bacterial vectors and additional plant vectors with bacterial replication origins) can replicate autonomously in the host cell into which they are introduced. Other vectors (for example, non-additional plant vectors) are integrated into the genome of the host cell after introducing the host cell, and thus replicate together with the host genome. Moreover, some vectors can instruct the expression of the genes that they are operably connected. Such a vector is referred to as an "expression vector" at this. The common expression vector used in recombinant DNA technology is typically a plasmid form.
[0324] Recombinant expression vectors may contain the nucleic acid molecules of the present invention in a form suitable for nucleic acid expression in a host cell, which means that these recombinant expression vectors contain one or more regulatory elements selected based on the host cell to be used for expression, which are operably linked to the nucleic acid sequence to be expressed.
[0325] Delivery and delivery compositions
[0326] The Cas effector proteins, derivative proteins, sgRNA sequences, guide RNA sequences, CRISPR / Cas complexes, encoding nucleic acids, vectors, compositions, and / or vector compositions of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipofection, nucleofection, microinjection, sonoporation, gene guns, calcium phosphate-mediated transfection, cationic transfection, lipofection, dendritic transfection, heat shock transfection, nucleofection, magnetofection, lipofection, puncture transfection, optical transfection, agent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial virions, and the like.
[0327] Therefore, the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: a Cas effector protein, a derivative protein, an sgRNA sequence, a guide RNA sequence, a CRISPR / Cas complex, an encoding nucleic acid, a vector, a composition and / or a vector composition of the present invention.
[0328] In certain embodiments, the delivery vehicle is a particle.
[0329] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, a gene gun, or a viral vector (e.g., a replication-defective retrovirus, a lentivirus, an adenovirus, or an adeno-associated virus).
[0330] Reagent test kit
[0331] The present invention provides a kit comprising one or more of the components described above. In certain embodiments, the kit comprises one or more components selected from the following: a Cas effector protein, a derivative protein, an sgRNA sequence, a guide RNA sequence, a CRISPR / Cas complex, an encoding nucleic acid, a vector, a composition, and / or a vector composition of the present invention.
[0332] In certain embodiments, the kit further comprises instructions for using the composition and / or vector composition.
[0333] In certain embodiments, the components included in the kits of the present invention may be provided in any suitable container.
[0334] In certain embodiments, the kit further comprises one or more buffers. The buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer and combinations thereof. In certain embodiments, the buffer is alkaline. In certain embodiments, the buffer has a pH from about 7 to about 10.
[0335] In certain embodiments, the kit further comprises one or more oligonucleotides corresponding to a targeting sequence for insertion into a vector so as to operably link the targeting sequence and regulatory elements. In certain embodiments, the kit comprises a homologous recombination template polynucleotide.
[0336] Methods and uses
[0337] The present invention provides a method for modifying a target gene, comprising: contacting the CRISPR / Cas complex, the composition or the vector composition as described in the present invention with the target gene, or delivering it into a cell containing the target gene; the target sequence is present in the target gene.
[0338] In certain embodiments, the method is used to modify a target gene in vitro or ex vivo. In certain embodiments, the method is not a method for treating a human or animal by therapy. In certain embodiments, the method does not include a step of modifying human germline genetic characteristics.
[0339] In certain embodiments, the target gene is present in a cell. In certain embodiments, the cell is a prokaryotic cell. In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a human cell (including primary cells and cell lines, such as human HEK293 cell lines). In certain embodiments, the cell is a plant cell, such as a cell of a cultivated plant (such as tobacco, Arabidopsis thaliana, cassava, corn, sorghum, wheat or rice), algae, tree or vegetable. In certain embodiments, the cell is a bacterial cell. In certain specific embodiments, the target gene can be a VEGF gene, a DNMT1 gene or a PMT1 gene of an animal (such as a human) or a plant (such as tobacco).
[0340] In certain embodiments, the target gene is present in a nucleic acid molecule (e.g., a plasmid) in vitro. In certain embodiments, the target gene is present in a plasmid.
[0341] In certain embodiments, the modification refers to a break in the target sequence, such as a double-strand break in DNA or a single-strand break in RNA.
[0342] In certain embodiments, the disruption results in decreased transcription of the target gene.
[0343] In certain embodiments, the method further comprises: contacting the editing template with the target gene, or delivering it to a cell comprising the target gene. In such embodiments, the method repairs the broken target gene by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides of the target gene. In certain embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.
[0344] Thus, in certain embodiments, the modification further comprises inserting an editing template (eg, an exogenous nucleic acid) into the break.
[0345] In certain embodiments, the protein, derived protein, isolated nucleic acid molecule, complex, vector or composition is contained in a delivery vehicle.
[0346] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).
[0347] In certain embodiments, the methods are used to modify a cell, cell line, or organism by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.
[0348] The present invention provides a method for altering the expression of a gene product, comprising: contacting the CRISPR / Cas complex, the composition or the vector composition of the present invention with a nucleic acid molecule encoding the gene product, or delivering the nucleic acid molecule to a cell comprising the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.
[0349] In certain embodiments, the method is used to modify the expression of a gene product in vitro or in vitro. In certain embodiments, the method is not a method for treating a human or animal by therapy. In certain embodiments, the method does not include a step of modifying human germline genetic characteristics.
[0350] In certain embodiments, the nucleic acid molecule is present in a nucleic acid molecule (e.g., a plasmid) in vitro. In certain embodiments, the nucleic acid molecule is present in a plasmid.
[0351] In certain embodiments, the expression of the gene product is altered (e.g., enhanced or reduced). In certain embodiments, the expression of the gene product is enhanced. In certain embodiments, the expression of the gene product is reduced.
[0352] In certain embodiments, the gene product is a protein.
[0353] In another aspect, the present invention relates to the Cas effector protein, sgRNA sequence, guide RNA sequence, CRISPR / Cas complex, encoding nucleic acid, vector, composition and / or vector composition, kit or delivery composition of the present invention, for use in nucleic acid editing (e.g., in vitro or ex vivo nucleic acid editing), or for use in preparing a preparation for nucleic acid editing.
[0354] In certain embodiments, the nucleic acid to be edited is present in a cell. In certain embodiments, the cell is a prokaryotic cell or a eukaryotic cell. In certain embodiments, the nucleic acid to be edited is present in an in vitro nucleic acid molecule (e.g., a plasmid).
[0355] In certain embodiments, the nucleic acid editing includes gene or genome editing, such as modifying a gene, knocking out a gene, changing the expression of a gene product, repairing a mutation, and / or inserting a polynucleotide. In certain embodiments, the gene or genome editing does not include a step of modifying human germline genetic characteristics. In certain embodiments, the use is not a method for treating a human or animal by therapy.
[0356] In certain embodiments, the use further comprises repairing the edited target sequence by homologous recombination with an exogenous template polynucleotide, wherein the repair can produce a mutation of the target sequence, including insertion, deletion or substitution of one or more nucleotides.
[0357] In another aspect, the present invention relates to the use of the Cas effector protein, derivative protein, sgRNA sequence, guide RNA sequence, CRISPR / Cas complex, encoding nucleic acid, vector, composition and / or vector composition, kit or delivery composition of the present invention in preparing a preparation for: (i) in vitro or ex vivo DNA detection; (ii) editing the target sequence in the target locus to modify an organism or non-human organism (e.g., a prokaryotic organism).
[0358] In certain embodiments, the preparation is used for the detection of single-stranded DNA or double-stranded DNA (eg, the detection of single-stranded or double-stranded DNA in prokaryotic cells).
[0359] In another aspect, the present invention also relates to a method for detecting a target DNA in a sample, comprising the following steps:
[0360] (1) contacting the sample with the following components: the CRISPR / Cas complex, encoding nucleic acid, vector, composition and / or vector composition as described in the present invention, and single-stranded DNA with a label; wherein,
[0361] The guide RNA sequence contained in the CRISPR / Cas complex or composition is capable of hybridizing with the target DNA, and
[0362] The single-stranded DNA does not hybridize to the guide RNA sequence;
[0363] (2) Measuring a detectable signal generated by the Cas effector protein contained in the CRISPR / Cas complex or composition cleaving the labeled single-stranded DNA, thereby detecting the target DNA.
[0364] In certain embodiments, the target DNA is viral DNA or bacterial DNA.
[0365] In certain embodiments, the target DNA is tumor cell DNA.
[0366] In certain embodiments, the target DNA is single-stranded or double-stranded.
[0367] In certain embodiments, the detectable signal is determined by one or more methods selected from the group consisting of imaging-based detection, sensor-based detection, color detection, gold nanoparticle-based detection, fluorescence polarization, colloidal phase transition / dispersion, electrochemical detection, and semiconductor-based sensing.
[0368] In certain embodiments, the method further comprises the step of amplifying the target DNA in the sample.
[0369] The advantages of the present invention are:
[0370] 1. The Cas protein of the present invention contains 521 amino acids, is relatively small, has high flexibility and controllability, and is easy to deliver in vivo.
[0371] 2. The present invention adds a nuclear localization signal (NLS) to the Cas protein, thereby enhancing the gene editing efficiency of the CRISPR / Cas system.
[0372] 3. The PAM site recognized by the Cas protein of the present invention is AAN (N is A, T, C, G), which enriches the PAM recognition sites of the existing CRISPR / Cas system and increases the range of target selection.
[0373] 4. Mutating the Cas protein of the present invention further improves the efficiency of gene editing.
[0374] 5. The present invention optimizes the guide RNA and the sgRNA sequence therein, further improving the gene editing efficiency of the CRISPR / Cas system by at least 1.60 times, enhancing the accuracy and flexibility of its gene editing.
[0375] 6. The present invention successfully applied the CRISPR / Cas system in Escherichia coli, tobacco cells, and human HEK293 cells, demonstrating the stability of the system and the feasibility of its application in bacteria, plants, and animals.
[0376] The present invention will be further described below with reference to specific embodiments, which are intended to illustrate the present invention but are not intended to limit the present invention.
[0377] Unless otherwise indicated, the experiments and methods described in the examples are carried out substantially according to conventional methods well known in the art and described in various references. For example, conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA used in the present invention can be found in Sambrook, Fritsch and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd edition (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (FM Ausubel et al., eds., (1987)); METHODS IN ENZYMOLOGY series (Academic Publishing Company): PCR 2: A PRACTICAL METHOD. APPROACH) (MJ MacPherson, BD Hames and GR Taylor, eds. (1995)), Harlow and Lane, eds. (1988) ANTIBODIES, A LABORATORY MANUAL, and ANIMAL CELL CULTURE (RI Freshney, ed. (1987)).
[0378] In addition, the specific conditions not specified in the examples were carried out according to conventional conditions or the conditions recommended by the manufacturers. The reagents or instruments used without indicating the manufacturer were all conventional products that can be purchased from the market.
[0379] Those skilled in the art will appreciate that the embodiments describe the present invention by way of example and are not intended to limit the scope of the invention as claimed. All publications and other references mentioned herein are incorporated herein by reference in their entirety.
[0380] The sources of some of the reagents involved in the following examples are as follows:
[0381] LB liquid medium: 10g tryptone, 5g yeast extract, 10g NaCl, dilute to 1L, and sterilize. If antibiotics are needed, add them after the medium has cooled, to a final concentration of 50μg / ml.
[0382] Chloroform / isoamyl alcohol: Add 10 ml of isoamyl alcohol to 240 ml of chloroform and mix well.
[0383] RNP buffer: 100 mM sodium chloride, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg / ml BSA, pH 7.9.
[0384] The prokaryotic expression vector pET-28a was purchased from Beijing Quanshijin Biotechnology Co., Ltd.
[0385] Escherichia coli competent BL21 was purchased from Epicentre
[0386] Example 1. Acquisition of HT001 gene and HT001 guide RNA
[0387] 1. Phage genome assembly and annotation: phage sequences from public databases (marine microbiome, global marine virome) were collected and assembled using metaSPAdes, followed by sequence annotation.
[0388] 2. Protein filtering: De-redundancy of annotated proteins is performed through sequence consistency, and proteins with completely identical sequences are removed.
[0389] 3. Acquisition of CRISPR-related proteins: CRISPR protein family sequences were collected from NCBI, Uniprot, and literature. Protein sequences were aligned using DIAMOND BLASTP. The alignment results with an Evalue < 1E-5 were output to obtain protein sequences suspected of being CRISPR in phages.
[0390] 4. Clustering of CRISPR proteins: Use Mafft to perform multiple sequence alignment of CRISPR protein family sequences and suspected phage CRISPR protein sequences, and use IQTREE to construct an evolutionary tree to infer the type of phage-derived CRISPR protein.
[0391] 5. Identification of CRISPR sites (Repeats and Spacers): Use MinCED and CRISPRDetect to identify the CRISPR sites of phage-derived CRISPR proteins.
[0392] 6. Protein sequence similarity analysis and domain annotation: DIAMOND BLASTP was used to align the Genbank database, NR database, and Cas proteins from published patents and phage-derived Cas proteins, and Cas proteins with Evalue <1E-5 alignment results and a similarity of less than 80% were screened as new CRISPR / Cas protein families. Mafft was used to perform multiple sequence alignments on each CRISPR / Cas family protein, and then JPred and HHpred were used for conserved domain analysis to identify protein families containing the RuvC domain. On this basis, the inventors obtained a new Cas effector protein, named HT001 (SEQ ID NO: 1), whose encoding DNA sequence is shown in SEQ ID NO: 2.
[0393] 7. To obtain the guide RNA for HT001, CRISPR-RNA (crRNA) repeats from the phage-encoded CRISPR locus were identified using MinCED and CRISPRDetect. Repeats were compared by using the Needleman-Wunsch algorithm followed by BROSS Needle to generate pairwise similarity scores. The prototype direct repeat sequence of HT001 was obtained (SEQ ID NO: 3).
[0394] 8. Synthesize a DNA sequence encoding the HT001 protein (SEQ ID NO: 22) with a nuclear localization signal, and ligate the double-stranded DNA molecule to the prokaryotic expression vector pET-28a to obtain the recombinant plasmid pET-28a-HT001. Sequence the recombinant plasmid pET-28a-CRISPR / HT001. Introduce the recombinant plasmid pET-28a-HT001 into Escherichia coli BL21 (DE3) to obtain a recombinant bacterium, which is named BL21 / pET-28a-HT001. Pick a single clone of BL21 / pET-28a-HT001 and inoculate it into 100 mL of LB liquid medium (containing 100 μg / mL kanamycin). Incubate at 37°C and 200 rpm with shaking for 12 h to obtain a culture solution. The culture was inoculated into 50 mL of LB liquid medium (containing 100 μg / mL kanamycin) at a volume ratio of 1:100 and cultured at 37°C, 200 rpm, with shaking until the OD600nm value reached 0.6. IPTG was then added to a concentration of 1 mM and cultured at 28°C, 220 rpm, with shaking for 4 h. The pellet was centrifuged at 10,000 rpm at 4°C for 10 min, and the cell pellet was collected. The pellet was resuspended in 100 mL of 100 mM Tris-HCl buffer, pH 8.0, and ultrasonically disrupted (ultrasonic power 600 W, cycle: 4 s disruption, 6 s pause, for a total of 20 min). The pellet was then centrifuged at 10,000 rpm at 4°C for 10 min, and supernatant A was collected. Supernatant A was then centrifuged at 12,000 rpm at 4°C for 10 min, and supernatant B was collected. Supernatant B was purified using a nickel column produced by GE (for the specific purification steps, refer to the instructions of the nickel column), and then the Cas protein was quantified using a protein quantification kit produced by Thermo Fisher Scientific.
[0395] 9. Design a template for guide RNA transcription. The structure of the transcription template is: T7 promoter + HT001 prototype direct repeat sequence (SEQ ID NO: 3) + guide sequence (SEQ ID NO: 20). Primers were designed using Primer 5.0 software, ensuring that the forward primer and reward primer had at least 18 bp of overlapping sequence. Prepare the reaction system, gently pipette to mix, and briefly centrifuge. Place the reaction in a PCR instrument for PCR amplification. Use the MinElute PCR Purification Kit to purify the template. Extract with phenol:chloroform:isoamyl alcohol (25:24:1) to remove DNase I in the system. Determine the concentration of the purified crRNA using Nanodrop, and uniformly dilute to 250 ng / μl. Aliquot the mixture into 200 μl PCR tubes and store at -80°C until needed. Establish a double-stranded DNA digestion system: Prepare the reaction system, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 15 minutes; add 300 ng of substrate DNA (100 ng / μl) and 3 μl, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 hours; add RNAse and incubate at 37°C for 15 minutes to fully digest RNA impurities in the system; add proteinase K and incubate at 58°C for 15 minutes to digest the HT001 protein. Gel electrophoresis results showed that no DNA cleavage activity was detected after the purified HT001 protein and the corresponding crRNA (Figure 1), indicating that HT001 requires additional components for successful dsDNA cleavage.
[0396] 10. A p15a-based plasmid containing the HT001 site and a non-targeting spacer was transformed into a chemically competent E. coli Top10 strain. Total RNA was extracted from 5 mL of cell culture using the EasyPure RNA kit (TransGen Biotech). 15 mg of RNA was treated with 2 units of DNase I (NEB) at 37°C for 30 minutes to remove DNA. 1 mM ATP was added, and the sample was incubated at 37°C for 1 hour for 5'-phosphorylation, followed by phenol / chloroform extraction. T4PNK-treated RNA was treated with 5 units of RppH (NEB) at 37°C for 1 hour to hydrolyze 5'-pyrophosphate and subsequently purified by phenol / chloroform extraction. RNA purity was determined by agarose gel electrophoresis. Small RNA libraries were constructed using the VAHTS Small RNA Library Prep Kit for Illumina (Vazyme Biotech) according to the manufacturer's instructions and subjected to Illumina HiSeq sequencing at the HaploX Genomics Center. The raw sequencing data were processed to remove adapters and sequencing artifacts and maintain high-quality reads. Sequencing reads were aligned to the reference sequence using BWA, and mapped unique reads were analyzed using the SAM tool. Analysis showed that a 200nt RNA transcript mapped to the 3' end of the HT001 gene was enriched in the HT001 system. This RNA transcript contained an anti-repeat sequence that could potentially base-pair with a short direct repeat of the corresponding crRNA, suggesting that this could serve as tracrRNA (SEQ ID NO: 4) targeted by the HT001 nuclease DNA.
[0397] 11. Design a template for guide RNA transcription. The template structure is: T7 promoter + HT001 tracrRNA sequence (SEQ ID NO: 5) + guide sequence (SEQ ID NO: 20). Primers were designed using Primer 5.0 software, ensuring that the forward and reward primers overlap by at least 18 bp. The reaction mixture was prepared, gently pipetted to mix, and then briefly centrifuged. PCR amplification was performed in a PCR instrument. The template was purified using a MinElute PCR Purification Kit using phenol:chloroform:isoamyl alcohol (25:24:1) extraction to remove DNAse I. The purified tracrRNA concentration was determined using Nanodrop and diluted to 250 ng / μl. Aliquots were distributed into 200 μl PCR tubes and stored frozen at -80°C until further use. The reaction mixture was prepared, gently pipetted to mix, and then briefly centrifuged. Incubate at 37°C for 15 min. For the DNA cleavage reaction, 300 ng of substrate DNA (100 ng / μl) was added, 3 μl of the tube was added, gently pipetted to mix, and then briefly centrifuged. Place at 37°C for 8 hours; add RNAse and place at 37°C for 15 minutes to fully digest RNA impurities in the system; add proteinase K and place at 58°C for 15 minutes to digest HT001 protein; gel electrophoresis results showed that the purified HT001 protein had cleavage activity when both crRNA and putative tracrRNA were supplemented, while the same DNA substrate could not be cleaved when crRNA or tracrRNA was provided alone (Figure 1).
[0398] 12. Predict the secondary structure of the guide sequence and determine the secondary structure of the guide RNA. By artificially removing the 3'-terminal 4-nt of the tracrRNA and the 5'-terminal 4-nt of the mature crRNA, the tracrRNA and mature crRNA were further ligated and fused to form sgRNA (Figure 1). A template for guide RNA transcription was designed. The structure of the transcription template was: T7 promoter + HT001 sgRNA sequence (SEQ ID NO: 6) + guide sequence (SEQ ID NO: 20). Primers were designed using Primer 5.0 software, ensuring that the forward primer and reward primer overlapped by at least 18 bp. The reaction system was prepared, gently pipetted to mix, and then briefly centrifuged. PCR amplification was performed in a PCR instrument. The template was purified using the MinElute PCR Purification Kit. DNAseI in the system was removed by extraction with phenol:chloroform:isoamyl alcohol (25:24:1). The purified sgRNA concentration was determined using Nanodrop and diluted to 250 ng / μl. Aliquots were dispensed into 200 μl PCR tubes and stored at -80°C until further use. To establish a double-stranded DNA digestion system: Prepare the reaction system, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 15 minutes. Add 300 ng of substrate DNA (100 ng / μl) and 3 μl to the DNA cleavage reaction, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 hours. Add RNAse and incubate at 37°C for 15 minutes to fully digest RNA impurities in the system. Add proteinase K and incubate at 58°C for 15 minutes to digest the HT001 protein. Gel electrophoresis results showed that the chimeric sgRNA-WT exhibited comparable dsDNA cleavage activity to isolated crRNA and tracrRNA (Figure 2), thus simplifying the CRISPR-HT001 expression system for practical applications.
[0399] Example 2: Identification of the PAM domain of the HT001 gene
[0400] 1. Construction and sequencing of the recombinant plasmid pET-28a+CRISPR / HT001+sgRNA. Based on the sequencing results, the structure of the recombinant plasmid pET-28a+CRISPR / HT001+sgRNA was described as follows: the small fragment between the restriction endonuclease HindIII and EcoRI recognition sequences of the pET-28a vector was replaced with the sequence shown in SEQ ID NO: 2, and the prototype sgRNA encoding nucleotide sequence of HT001 shown in SEQ ID NO: 6 and the guide sequence identified by the PAM domain of HT001 shown in SEQ ID NO: 7 were ligated to the SphI restriction site of the vector.
[0401] 2. Obtaining recombinant E. coli: The recombinant plasmid pET-28a+CRISPR / HT001+sgRNA was introduced into E. coli BL21 to obtain recombinant E. coli, named BL21 / pET-28a+CRISPR / HT001+sgRNA. The recombinant plasmid pET-28a was introduced into E. coli BL21 to obtain recombinant E. coli, named BL21 / pET-28a.
[0402] 3. Construction of the PAM library: The sequence shown in SEQ ID NO: 8 was artificially synthesized and ligated into the pUC19 vector, wherein the sequence shown in SEQ ID NO: 5 includes eight random bases (NNNNNNNN) and the target sequence. The plasmids were respectively transformed into E. coli containing the CRISPR / HT001 locus (BL21 / pET-28a+CRISPR / HT001+sgRNA) and E. coli without the CRISPR / HT001 locus (BL21 / pET-28a). After treatment at 37°C for 1 hour, the plasmids were extracted, and the PAM region sequence was PCR amplified and sequenced.
[0403] 4. Obtaining PAM library domains: The number of occurrences of 65,536 PAM sequence combinations in the experimental and control groups was counted and normalized using the total number of PAM sequences in each group. For any PAM sequence, if the log2 (normalized value of the control group / normalized value of the experimental group) was greater than 3.5, the PAM was considered significantly depleted. Significantly depleted PAM sequences were identified from all PAM sequences. Furthermore, Weblogo was used to predict these significantly depleted PAM sequences, revealing the PAM domain of the HT001 protein to be 5'-AAN-3' (Figure 3), where N is any amino acid.
[0404] Example 3: Verification of key structural elements of guide RNA
[0405] Since the predicted secondary structure of the guide RNA is complex and contains multiple stem-loop structures, in order to verify the key structural elements, the inventors conducted verification by systematic truncation, and deleted four separate stem-loop structures based on sgRNA-WT to evaluate the importance of the key structural elements of sgRNA.
[0406] 4 , SL1, SL2, SL3, and SL4 were partially deleted, and the sgRNAWT expression cassettes of sequences shown in SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, and SEQ ID NO: 13, the sgRNASL1 expression cassette with SL1 deletion, the sgRNASL2 expression cassette with SL2 deletion, the sgRNASL3 expression cassette with SL3 deletion, and the sgRNASL4 expression cassette with SL4 deletion were synthesized according to the pattern of J23119 promoter + sgRNA + guide RNA + polyT, respectively. The target site was the housekeeping gene gapA of Escherichia coli BL21. Each gRNA expression cassette was ligated into the pESL28(a) vector through the SphI restriction site, and the small fragment between the restriction endonuclease HindIII and EcoRI recognition sequences of the vector pET-28a was replaced with the sequence shown in SEQ ID NO: 2 to construct recombinant plasmids pESL28(a)+HT001+J23119-sgRNAWT, pESL28(a)+HT001+J23119-sgRNASL1, pESL28(a)+HT001+J23119-sgRNASL2, pESL28(a)+HT001+J23119-sgRNASL3 and pESL28(a)+HT001+J23119-sgRNASL4.
[0407] The correctly sequenced recombinant plasmids were transformed into Escherichia coli BL21 and named BL21 / pESL28(a)+HT001+J23119-sgRNAWT, BL21 / pESL28(a)+HT001+J23119-sgRNASL1, BL21 / pESL28(a)+HT001+J23119-sgRNASL2, BL21 / pESL28(a)+HT001+J23119-sgRNASL3 and BL21 / pESL28(a)+HT001+J23119-sgRNASL4. The plasmids were plated on LB solid medium containing kanamycin resistance through dilution gradient concentration. Among them, the growth rate and survival rate of Escherichia coli BL21 / pESL28(a)+HT001+J23119-sgRNASL4 lacking SL4 were higher than those of the other E. coli (Figure 5). It was demonstrated that the stem-loop structure sgRNA-SL4 is a key element in sgRNA-WT.
[0408] Example 4: Engineering of guide RNA
[0409] 1. Optimization of sgRNA
[0410] 1. Since the cleavage activity of HT001 / sgRNA-WT is low, in order to improve the activity of the HT001 system, the inventors engineered the key component sgRNA-SL4 in sgRNA-WT. After truncation according to the stem-loop structure of sgRNA-SL4, R1, R2, R3, R4, R5 and R6 were synthesized according to the sequences shown in SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18 and SEQ ID NO: 19, as shown in Figure 6.
[0411] 2. Transcription and purification of HT001 protein guide RNA:
[0412] 1. Design a guide RNA transcription template. The structure of the transcription template is: T7 promoter + mature sgRNA of HT001 (SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19) + guide sequence (SEQ ID NO: 20). Primers were designed using Primer 5.0 software to ensure that the forward primer and reward primer had at least 18 bp of overlapping sequence.
[0413] 2. Prepare the reaction system, pipette gently to mix, centrifuge briefly, and place in a PCR instrument for annealing.
[0414] 3. Use MinElute PCR Purification Kit (purchased from Takara, catalog number 9761) to purify the template. The steps are as follows:
[0415] 1) Add 5 volumes of PB to the PCR product, place a MinElute column on a 2 ml collection tube, incubate at room temperature for 2 min, and centrifuge at 12,000 g for 2 min.
[0416] 2) Discard the waste solution and add 750 μl of Buffer PE (add ethanol before use) and incubate at 12,000 g / 2 min.
[0417] 3) Discard the waste liquid, add 350 μl Buffer PE, centrifuge at 12000 g / 2 min, discard the waste liquid, centrifuge at 12000 g for 2 min;
[0418] 4) Transfer the MinElute column to a new 1.5 ml centrifuge tube, open the tube lid, and incubate at 65°C for 2 minutes.
[0419] 5) Add 20 μl of preheated EB solution, let it stand for 2 minutes, and then centrifuge at 12000g / 2min. To improve the recovery rate, pass the contents of the centrifuge tube 2-3 times through the MinElute centrifuge column;
[0420] 6) Measure the concentration using Nanodrop and store at -20°C until use.
[0421] 4. Purification of guide RNA: Phenol: chloroform: isoamyl alcohol (25:24:1) extraction to remove DNAse I in the system;
[0422] 1) Add 80 μl RNA-free H2O to the post-transcription reaction system and adjust the volume to 100;
[0423] 2) Remove 2 ml of Phase Lock Gel (PLG) Heavy and centrifuge at 15,000 g for 2 min. Add 100 μl of phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μl of RNA digested with DNAse I. Gently flick the Phase-Lock tube 5-10 times to mix thoroughly. Centrifuge at 15°C / 16,000 g for 12 min.
[0424] 3) Take a new RNA-free 1.5ml centrifuge tube and aspirate the supernatant from the previous step into the tube, being careful not to aspirate the gel. Add an equal volume of isopropanol and one-tenth the volume of sodium acetate solution to the supernatant. Mix thoroughly by pipetting with a pipette and place in a -20°C refrigerator for 1 hour or overnight.
[0425] 4) Centrifuge at 4°C / 16,000 g for 30 min, discard the supernatant, add 75% pre-chilled ethanol, pipette and mix the precipitate, centrifuge at 4°C / 16,000 g for 12 min, discard the supernatant, let stand in a fume hood for 2-3 min to dry the ethanol on the RNA surface, add 100 μl of RNA-free H2O, and pipette and mix.
[0426] 5) Determine the concentration of the purified crRNA using Nanodrop, dilute to 250 ng / μl, dispense into 200 μl PCR centrifuge tubes, and freeze at -80°C until use.
[0427] 3. In vitro expression and purification of HT001 protein
[0428] The steps for in vitro expression and purification of HT001 protein are as follows:
[0429] 1. Artificially synthesize a DNA sequence encoding the HT001 protein (SEQ ID NO: 22) with a nuclear localization signal.
[0430] 2. Ligate the double-stranded DNA molecule synthesized in step 1 with the prokaryotic expression vector pET-28a to generate the recombinant plasmid pET-28a-HT001. Sequencing of the recombinant plasmid pET-28a-CRISPR / HT001 was performed. Sequencing results showed that the recombinant plasmid pET-28a-CRISPR / HT001 expressed the HT001 protein with a nuclear localization signal as shown in SEQ ID NO: 22.
[0431] 3. Introduce the recombinant plasmid pET-28a-HT001 into Escherichia coli BL21(DE3) to obtain a recombinant bacterium, named BL21 / pET-28a-HT001. Select a single colony of BL21 / pET-28a-HT001 and inoculate it into 100 mL of LB liquid medium (containing 100 μg / mL kanamycin). Incubate the culture at 37°C, 200 rpm, and shake for 12 h to obtain a culture solution.
[0432] 4. Take the culture solution and inoculate it into 50 mL LB liquid medium (containing 100 μg / mL kanamycin) at a volume ratio of 1:100. Incubate the culture at 37°C and 200 rpm with shaking until the OD600nm value reaches 0.6. Then add IPTG and adjust its concentration to 1 mM. Incubate the culture at 28°C and 220 rpm with shaking for 4 h. Centrifuge at 4°C and 10,000 rpm for 10 min to collect the bacterial precipitate.
[0433] 5. Take the bacterial pellet, add 100 mL of pH 8.0, 100 mM Tris-HCl buffer, resuspend and ultrasonically disrupt (ultrasonic power 600 W, cycle program: disruption 4 s, pause 6 s, total 20 min), then centrifuge at 4°C, 10000 rpm for 10 min, and collect supernatant A.
[0434] 6. Take supernatant A, centrifuge at 4°C, 12000 rpm for 10 min, and collect supernatant B.
[0435] 7. Supernatant B was purified using a nickel column produced by GE (refer to the instructions of the nickel column for the specific purification steps), and then the Cas protein was quantified using a protein quantification kit produced by Thermo Fisher Scientific.
[0436] 8. Establishment of double-stranded DNA enzyme digestion system:
[0437] (1) Prepare the reaction system, gently pipette to mix, and then centrifuge briefly. Place at 37°C for 15 minutes; DNA cleavage reaction
[0438] (2) Add 300 ng of substrate DNA (100 ng / μl) to 3 μl of the tube, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 h.
[0439] (3) Add RNAse and incubate at 37°C for 15 min to fully digest RNA impurities in the system;
[0440] (4) Proteinase K was added and incubated at 58°C for 15 min to digest the HT001 protein;
[0441] (5) Agarose gel testing.
[0442] Gel electrophoresis results showed (Figure 7) that HT001 / sgRNAR1-R4 all had certain cleavage activity, while R5 and R6 had no cleavage activity, and the system cleavage activity of R1 was the highest. This system was called sgRNAR1-HT001.
[0443] Example 5: Identification of the cleavage pattern of the CRISPR / HT001 system
[0444] 1. In vitro expression and purification of HT001 protein
[0445] 1. Synthesize a DNA sequence encoding the HT001 protein (SEQ ID NO: 22) with a nuclear localization signal.
[0446] 2. Ligate the double-stranded DNA molecule synthesized in step 1 with the prokaryotic expression vector pET-28a to generate the recombinant plasmid pET-28a-HT001. Sequencing of the recombinant plasmid pET-28a-CRISPR / HT001 was performed. Sequencing results showed that the recombinant plasmid pET-28a-CRISPR / HT001 expressed the HT001 protein with a nuclear localization signal as shown in SEQ ID NO: 22.
[0447] 3. Introduce the recombinant plasmid pET-28a-HT001 into Escherichia coli BL21(DE3) to obtain a recombinant strain, named BL21 / pET-28a-HT001. Select a single colony of BL21 / pET-28a-HT001 and inoculate it into 100 mL of LB liquid medium (containing 100 μg / mL kanamycin). Incubate with shaking at 37°C and 200 rpm for 12 h to obtain a culture solution.
[0448] 4. Take the culture solution and inoculate it into 50 mL LB liquid medium (containing 100 μg / mL kanamycin) at a volume ratio of 1:100. Incubate the culture at 37°C and 200 rpm with shaking until the OD600nm value reaches 0.6. Then add IPTG and adjust its concentration to 1 mM. Incubate the culture at 28°C and 220 rpm with shaking for 4 h. Centrifuge at 4°C and 10,000 rpm for 10 min to collect the bacterial precipitate.
[0449] 5. Take the bacterial pellet, add 100 mL of pH 8.0, 100 mM Tris-HCl buffer, resuspend and ultrasonically disrupt (ultrasonic power 600 W, cycle program: disruption 4 s, pause 6 s, total 20 min), then centrifuge at 4°C, 10000 rpm for 10 min, and collect supernatant A.
[0450] 6. Take supernatant A, centrifuge at 4°C, 12000 rpm for 10 min, and collect supernatant B.
[0451] 7. Supernatant B was purified using a nickel column produced by GE (refer to the instructions of the nickel column for the specific purification steps), and then the Cas protein was quantified using a protein quantification kit produced by Thermo Fisher Scientific.
[0452] 2. Transcription and purification of HT001 protein guide RNA:
[0453] 1. Design a guide RNA transcription template. The structure of the transcription template is: T7 promoter + sgRNA-R1 (optimal sgRNA after HT001 engineering modification) + guide sequence (SEQ ID NO: 20). Primers were designed using Primer 5.0 software to ensure that the forward primer and reward primer have at least 18 bp of overlapping sequence.
[0454] 2. Prepare the reaction system, gently pipette to mix, centrifuge briefly, and place in a PCR instrument for PCR amplification.
[0455] 3. Use MinElute PCR Purification Kit to purify the template. The steps are as follows:
[0456] 1) Add 5 volumes of PB to the PCR product, place a MinElute column on a 2 ml collection tube, incubate at room temperature for 2 min, and centrifuge at 12,000 g for 2 min.
[0457] 2) Discard the waste liquid and add 750 μl of Buffer PE (add ethanol before use) at 12,000 g / 2 min.
[0458] 3) Discard the waste liquid, add 350 μl Buffer PE, centrifuge at 12000 g / 2 min, discard the waste liquid, centrifuge at 12000 g for 2 min;
[0459] 4) Transfer the MinElute column to a new 1.5 ml centrifuge tube, open the tube lid, and incubate at 65°C for 2 minutes.
[0460] 5) Add 20 μl of preheated EB solution, let it stand for 2 minutes, and then centrifuge at 12000g / 2min. To improve the recovery rate, pass the contents of the centrifuge tube 2-3 times through the MinElute centrifuge column;
[0461] 6) Measure the concentration using Nanodrop and store at -20°C until use.
[0462] 4. Purification of guide RNA: Phenol:chloroform:isoamyl alcohol (25:24:1) extraction to remove DNAse I in the system;
[0463] 1) Add 80 μl RNA-free H2O to the post-transcription reaction system and adjust the volume to 100;
[0464] 2) Remove 2 ml of Phase Lock Gel (PLG) Heavy and centrifuge at 15,000 g for 2 minutes. Add 100 μl of phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μl of RNA digested with DNAse I. Gently flick the Phase-Lock tube 5-10 times to mix thoroughly. Centrifuge at 15°C / 16,000 g for 12 minutes.
[0465] 3) Take a new RNA-free 1.5ml centrifuge tube and aspirate the supernatant from the previous step into the tube, being careful not to aspirate the gel. Add an equal volume of isopropanol and one-tenth the volume of sodium acetate solution to the supernatant. Mix thoroughly by pipetting with a pipette and place in a -20°C refrigerator for 1 hour or overnight.
[0466] 4) Centrifuge at 4°C / 16,000 g for 30 min, discard the supernatant, add 75% pre-chilled ethanol, pipette and mix the precipitate, centrifuge at 4°C / 16,000 g for 12 min, discard the supernatant, let stand in a fume hood for 2-3 min to dry the ethanol on the RNA surface, add 100 μl of RNA-free H2O, and pipette and mix.
[0467] 5) Determine the concentration of the purified crRNA using Nanodrop, dilute to 250 ng / μl, dispense into 200 μl PCR centrifuge tubes, and freeze at -80°C until use.
[0468] 5. Establishment of double-stranded DNA enzyme digestion system:
[0469] (1) Prepare the reaction system, gently pipette to mix, and then centrifuge briefly. Place at 37°C for 15 minutes; DNA cleavage reaction
[0470] (2) Add 300 ng of substrate DNA (100 ng / μl) to 3 μl of the tube, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 h.
[0471] (3) Add RNAse and incubate at 37°C for 15 min to fully digest RNA impurities in the system;
[0472] (4) Proteinase K was added and incubated at 58°C for 15 min to digest the HT001 protein;
[0473] (5) The first generation sequencing results showed that HT001 was cut at the position 21-22 bp downstream of the PAM of the target chain and at the position 27-28 bp downstream of the PAM of the non-target chain (Figure 8).
[0474] Example 6: HT001 has editing activity in plant cells
[0475] Using the modified sgRNA-R1 of Example 4 (sgRNA-R1 is the optimal sgRNA after modification), three target sequences (i.e., PMT1-Target1, PMT1-Target2, and PMT1-Target3) shown in SEQ ID NO: 23, SEQ ID NO: 24, and SEQ ID NO: 25 were designed for the tobacco PMT1 gene. The target sequences were synthesized according to the J23119 promoter sequence + the engineered sgRNA-R1 encoding nucleotide sequence + the tobacco target sequence (SEQ ID NO: 23, SEQ ID NO: 24, and SEQ ID NO: 25) + poly T sequence, and loaded into the CR2 vector through the enzyme cutting site PacI. The DNA sequence encoding the HT001 protein (SEQ ID NO: 22) with a nuclear localization signal was artificially synthesized, and the synthesized double-stranded DNA molecule was ligated into the CR2 vector by double enzyme digestion and located in the same expression cassette as the ZmUbi promoter to form target 1 knockout vector, target 2 knockout vector, and target 3 knockout vector. The knockout vectors with correct sequencing were extracted and transformed into Agrobacterium gv3101.
[0476] Conducting genetic transformation experiments in tobacco:
[0477] (1) Agrobacterium culture was cultured in liquid LB containing kanamycin (50 mg / mL) and rifampicin (50 mg / mL) at 200 rpm and 28°C for 2 days;
[0478] (2) 1 mL of activated Agrobacterium tumefaciens culture was added to 50 mL of liquid LB medium containing kanamycin and rifampicin and cultured at 200 rpm and 28°C.
[0479] (3) Cultivate the bacterial solution until the OD600 is between 0.6 and 0.8. Centrifuge the solution at 4000 rpm for 10 minutes at room temperature to collect the bacteria and discard the supernatant.
[0480] (4) Resuspend the bacterial solution in 100 mL of liquid MS medium without sucrose;
[0481] (5) Cut tobacco leaves and sterilize them with 70% alcohol for 2 min in a clean bench, then sterilize them with 5% sodium hypochlorite solution for 10 min, rinse them with sterile water 3-4 times, and absorb the moisture with filter paper;
[0482] (6) Cut the edges of the tobacco leaves with scissors and cut the leaves into 1 cm x 1 cm square leaves (try to make sure the cut leaves do not contain veins);
[0483] (7) Soak the square leaf in a conical flask containing Agrobacterium solution and shake gently for 10-15 minutes;
[0484] (8) Rinse with sterile water 3-4 times to clean the bacterial solution and place it on filter paper to dry;
[0485] (9) Place the leaves on a co-culture medium (6-BA: 0.5 mg / L; NAA: 2 mg / L) with filter paper, seal and culture in the dark at 25°C for 2 days;
[0486] (10) The leaves were transferred to callus induction medium (6-BA: 0.5 mg / L; NAA: 2 mg / L; Timentin: 200 mg / L; Basta: 3 mg / L) at 25°C, sealed, and cultured with alternating light periods of 16 h and dark periods of 8 h. The medium was changed every 2 weeks until callus tissue grew.
[0487] (11) The callus tissue was transferred to the differentiation medium for resistance screening (6-BA: 0.5 mg / L; NAA: 2 mg / L; timentin: 200 mg / L; basta: 3 mg / L; kan: 50 mg / L) until regenerated seedlings grew.
[0488] (12) The gene editing status of the regenerated seedlings was detected. The results are shown in Table 1. The modified sgRNAR1-HT001 has gene editing activity in tobacco. Comparing the results of sgRNA-R1 with those of sgRNA-WT, the editing efficiency of sgRNAR1 was 1.60-1.79 times that before optimization (Figure 9).
[0489] Table 1
[0490] Example 7: HT001 has editing activity in human cells
[0491] The HT001 gene and the HT001 gene were constructed on the mCherry red fluorescent protein expression vector pEE6.4-mCherry, which is capable of expressing the HT001 protein in mammalian cells. This vector transcribes crRNA and expresses the Cas protein, with red fluorescence indicating successful transfection of the vector into host cells. The expression vector was transfected into the human HEK293 cell line using the PEI transfection method. After 48 hours of culture, cells positive for mCherry red fluorescence were isolated by flow cytometry. DNA from all cells was extracted, and a 700-bp sequence containing the target site was amplified. The PCR products were connected to a B-simple vector for first-generation sequencing, which was performed by Huazhi Biotechnology Co., Ltd. The sequencing results were aligned to the VEGFA gene in the human genome, and the cleavage mode of HT001 at the target site was identified. In addition, its editing efficiency for VEGFA was identified. The PCR products were used to construct a second-generation sequencing library using Tn5, and sequencing was performed by Huazhi Biotechnology Co., Ltd. The editing efficiency of HT001 for genes such as VEGFA and DNMT1 was identified (Figure 10). As shown in Table 2, the overall editing efficiency of sgRNAR1-HT001 was 2.29-2.43 times that before optimization.
[0492] Table 2
[0493] Example 8: Improving the cleavage activity of HT001 protein
[0494] To enhance the cleavage activity of the HT001 protein, HT001 mutants bearing the G362R (primer sequences: SEQ ID NOs: 32-33), D464K (primer sequences: SEQ ID NOs: 34-35), and G362R+D464K (primer sequences: SEQ ID NOs: 32-35) were constructed. Their cleavage activity compared to the wild-type protein was tested in tobacco cells. These proteins were synthesized using the J23119 promoter sequence, the engineered sgRNA-R1 nucleotide sequence, the tobacco target sequence (SEQ ID NOs: 23, 24, and 25), and the polyT sequence. These proteins were then inserted into the CR2 vector via the Pac I restriction enzyme site. HT001 proteins bearing the G362R, D464K, or G362R+D464K mutations were then ligated into the CR2 vector via double enzyme digestion and located in the same expression cassette as the ZmUbi promoter, creating six knockout vectors targeting the tobacco PMT1 gene. The knockout vectors with correct sequencing were extracted and the plasmids were transferred into Agrobacterium gv3101 for tobacco genetic transformation, and finally regenerated seedlings were obtained.
[0495] To examine gene editing in regenerated seedlings (Figure 11), PCR amplification of the PMT1 gene locus was performed. As shown in Table 3, sequencing analysis revealed that the G362R and D464K point mutations resulted in approximately 1.10-1.98-fold increases in cleavage activity compared to the wild-type protein, with G362R+D464K leading to the highest increase in cleavage activity, approximately 2.38-fold, compared to the wild-type protein.
[0496] Table 3
[0497] The sequences involved in the examples are as follows:
[0498] SEQ ID NO: 1 Amino acid sequence of HT001 (the bold underlined parts are G362 and D464 respectively)
[0499] SEQ ID NO: 2 HT001 coding nucleotide sequence (wherein the underlined parts are the codons encoding G362 and D464 respectively)
[0500] SEQ ID NO: 3 Prototype direct repeat sequence (crRNA) of HT001
[0501] SEQ ID NO: 4 Nucleotide sequence encoding the prototype direct repeat sequence of HT001
[0502] SEQ ID NO: 5 Transactivating CRISPR RNA (tracrRNA) of HT001
[0503] SEQ ID NO: 6 Prototype sgRNA encoding nucleotide sequence of HT001 (tracrRNA is underlined, crRNA is in italics)
[0504] SEQ ID NO: 7 Targeting sequence identified by the PAM domain of HT001
[0505] SEQ ID NO: 8 PAM library sequence (the underlined part is a random base, the italic part is the guide sequence, and the bold part is the target sequence)
[0506] SEQ ID NO: 9 Synthetic sequence of the gRNA expression cassette of the prototype sgRNA of HT001 (wherein positions 1-35 are the J23119 promoter sequence, positions 36-230 are the prototype sgRNA encoding nucleotide sequence, the bold portion is the guide sequence, and ttttttt represents the poly T sequence)
[0507] SEQ ID NO: 10 Synthetic sequence of the gRNA expression cassette for the SL1-deleted sgRNA of HT001 (wherein positions 1-35 are the J23119 promoter sequence, positions 36-214 are the SL1-deleted sgRNA sequence, the bold portion is the guide RNA sequence, and ttttttt represents a poly-T sequence)
[0508] SEQ ID NO: 11 Synthetic sequence of the gRNA expression cassette for the SL2-deleted sgRNA of HT001 (wherein positions 1-35 are the J23119 promoter sequence, positions 36-200 are the SL2-deleted sgRNA sequence, the bold portion is the guide RNA sequence, and ttttttt represents the poly T sequence)
[0509] SEQ ID NO: 12 Synthetic sequence of the gRNA expression cassette for the SL3-deleted sgRNA of HT001 (wherein positions 1-35 are the J23119 promoter sequence, positions 36-175 are the SL3-deleted sgRNA sequence, the bold portion is the guide RNA sequence, and ttttttt represents the poly T sequence)
[0510] SEQ ID NO: 13 Synthetic sequence of the gRNA expression cassette for the SL4-deleted sgRNA of HT001 (wherein positions 1-35 are the J23119 promoter sequence, positions 36-193 are the SL4-deleted sgRNA sequence, the bold portion is the guide RNA sequence, and ttttttt represents the poly T sequence)
[0511] SEQ ID NO: 14 Engineered sgRNA-R1 encoding nucleotide sequence
[0512] SEQ ID NO: 15 Engineered sgRNA-R2 encoding nucleotide sequence
[0513] SEQ ID NO: 16 Engineered sgRNA-R3 encoding nucleotide sequence
[0514] SEQ ID NO: 17 Engineered sgRNA-R4 encoding nucleotide sequence
[0515] SEQ ID NO: 18 Engineered sgRNA-R5 encoding nucleotide sequence
[0516] SEQ ID NO: 19 Engineered sgRNA-R6 encoding nucleotide sequence
[0517] SEQ ID NO: 20 In vitro cleavage guide sequence
[0518] SEQ ID NO: 21 NLS sequence-1
[0519] SEQ ID NO: 22 Amino acid sequence of HT001-NLS fusion protein (wherein positions 1-23 are 3xFLAG tags, the underlined portion is the NLS sequence, the italicized portion is the linker amino acids, and the bold portion is the amino acid sequence of the HT001 protein)
[0520] SEQ ID NO: 23 Target sequence of tobacco target 1 (the underlined portion is the PAM sequence)
[0521] SEQ ID NO: 24 Target sequence of tobacco target 2 (the underlined portion is the PAM sequence)
[0522] SEQ ID NO: 25 Target sequence of target 3 of tobacco (the underlined portion is the PAM sequence)
[0523] SEQ ID NO: 26 Targeting sequence in Escherichia coli
[0524] SEQ ID NO: 27 Target sequence for eukaryotic editing in human cells
[0525] SEQ ID NO: 28 sgRNA for SL1 deletion of HT001
[0526] SEQ ID NO: 29 sgRNA for SL2 deletion of HT001
[0527] SEQ ID NO: 30 sgRNA for SL3 deletion of HT001
[0528] SEQ ID NO: 31 sgRNA for SL4 deletion of HT001
[0529] SEQ ID NO: 32 F-direction primer for mutation of G to R at amino acid position 362 of HT001 (the bold underlined portion indicates the nucleotide to be mutated)
[0530] SEQ ID NO: 33 R-direction primer for mutation of G to R at amino acid position 362 of HT001 (the bold underlined portion indicates the nucleotide to be mutated)
[0531] SEQ ID NO: 34 F-direction primer for mutation of D to K at amino acid position 464 of HT001 (the bold underlined portion indicates the nucleotide to be mutated)
[0532] SEQ ID NO: 35 R-direction primer for mutation of D to K at amino acid position 464 of HT001 (the bold underlined portion indicates the nucleotide to be mutated)
[0533] SEQ ID NO: 36 NLS sequence-2
[0534] The above-described embodiments merely represent several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make several modifications and improvements without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be based on the appended claims. At the same time, all documents mentioned in this application are cited as references in this application, just as if each document was cited as a reference individually.
Claims
1. A protein, its ortholog, homolog, variant or functional fragment, characterized in that: The protein is selected from: (a1) a protein comprising the amino acid sequence shown in SEQ ID NO: 1; (a2) a protein comprising an amino acid sequence having one or more amino acid substitutions, deletions or additions compared to the amino acid sequence of SEQ ID NO: 1, which protein retains the biological function of the protein of SEQ ID NO: 1; or (a3) a protein comprising an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO: 1, wherein the protein retains the biological function of the protein of SEQ ID NO: 1; Wherein, the ortholog, homolog, variant or functional fragment retains the biological function of the protein shown in SEQ ID NO: 1; Preferably, in (a2), an amino acid substitution, deletion or addition occurs at position 362 and / or position 464, more preferably at position 362G and / or position 464D of the amino acid sequence shown in SEQ ID NO: 1; more preferably, position 362 of the amino acid sequence shown in SEQ ID NO: 1 mutates from G to a basic amino acid, most preferably to R, H or K; position 464 mutates from D to a basic amino acid, most preferably to R, H or K.
2. The protein according to claim 1, its ortholog, homolog, variant or functional fragment, characterized in that: The protein, its orthologs, homologs, variants or functional fragments further contain functional units, wherein the functional units include: epitope tags, reporter gene sequences, nuclear localization signal sequences, targeting moieties, transcription activation domains, transcription inhibition domains, nuclease domains, domains having the following activities: methylase activity, demethylase activity, transcription activation activity, transcription inhibition activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity and nucleic acid binding activity, and any combination thereof; Preferably, the functional unit is directly connected to the N-terminus or C-terminus of the sequence (a1) to (a3), or connected to the N-terminus or C-terminus via a linker; more preferably, the linker comprises or is selected from MA, GIHGVPAA; Preferably, the functional unit is a nuclear localization signal sequence; more preferably, the nuclear localization signal sequence is selected from the following nuclear localization signal sequences: SV40 virus large T antigen, EGL 13, c Myc and TUS protein; most preferably, the nuclear localization signal sequence is selected from: PKKKRKV, AVKRPAATKKAGQAKKKKLD, PAAKRVKLD, MSRRRKANPTKLSENAKKLAKEVEN, KLKIKRPVK, the acidic M9 domain of hnRNP A1, the sequence KIPIK in the yeast transcription repressor Matα2, and PY NLS.
3. The protein according to claim 1, its ortholog, homolog, variant or functional fragment, characterized in that: The protein is selected from: (b1) a protein comprising the amino acid sequence shown in SEQ ID NO: 22; (b2) contains one or more amino acid substitutions, deletions or additions compared to the amino acid sequence of SEQ ID NO: 22, and the protein retains the biological function of the protein of SEQ ID NO: 22; or (b3) having a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% to the sequence shown in SEQ ID NO: 22, and the protein retains the biological function of the protein shown in SEQ ID NO: 22; Wherein, the ortholog, homolog, variant or functional fragment retains the biological function of the protein shown in SEQ ID NO:
22.
4. A sgRNA molecule comprising a direct repeat sequence (crRNA) and an optional trans-acting crRNA (tracrRNA), characterized in that: The direct repeat sequence (crRNA) is selected from the following sequences, or consists of the following sequences: (c1) the sequence shown in SEQ ID NO: 3 or 4; (c2) the sequence represented by bases 5 to 27 of SEQ ID NO: 3 or 4, or the sequence represented by bases 173 to 195 of SEQ ID NO: 6; (c3) a sequence having one or more base substitutions, deletions or additions compared to the sequence shown in (c1) or (c2); (c4) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in (c1) or (c2); (c5) a sequence that hybridizes to the sequence described in any one of (c1) to (c4) under stringent conditions; or (c6) A complementary sequence of the sequence described in any one of (c1) to (c4); Furthermore, the sequence described in any one of (c2) to (c6) substantially retains the activity of the sequence from which it is derived as a direct repeat sequence in the CRISPR-Cas system; The trans-acting crRNA (tracrRNA) is selected from the following sequences, or consists of the following sequences: (d1) the sequence shown in SEQ ID NO: 5; (d2) the sequence shown at bases 1 to 172 of SEQ ID NO: 5, or the sequence shown at bases 1 to 172 of SEQ ID NO: 6; (d3) a sequence having one or more base substitutions, deletions or additions compared to the sequence shown in SEQ ID NO: 5; (d4) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity to the sequence shown in SEQ ID NO:5; (d5) a sequence that hybridizes to the sequence described in any one of (d1) to (d4) under stringent conditions; or (d6) A complementary sequence of the sequence described in any one of (d1) to (d4); Furthermore, the sequence described in any one of (d2) to (d6) substantially retains the biological function of the sequence from which it is derived, the activity of the sequence as a trans-acting crRNA (tracrRNA) in the CRISPR-Cas system; Preferably, the sgRNA molecule comprises a sequence selected from the following, or consists of a sequence selected from the following: (e1) a sequence shown in any one of SEQ ID NOs: 14-19, 28-31; (e2) a sequence having one or more base substitutions, deletions or additions compared to the sequence shown in any one of SEQ ID NOs: 14-19, 28-31; (e3) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 14-19, 28-31; (e4) a sequence that hybridizes to the sequence described in any one of (e1) to (e3) under stringent conditions; or (e5) A complementary sequence of the sequence described in any one of (e1) to (e3); Furthermore, the sequence described in any one of (e2)-(e5) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a single-molecule guide RNA in the CRISPR-Cas system.
5. The sgRNA molecule according to claim 4, characterized in that The sgRNA molecule comprises one or more stem-loops or optimized secondary structures; Preferably, the sequence of any one of (e2) to (e5) retains one or more stem-loops or secondary structures of the sequence from which it is derived; More preferably, the sequence described in any one of (e2) to (e5) retains one or more stem-loops or secondary structures of the sequence of SEQ ID NO: 6 from which it is derived, and lacks 10 to 50 bases; Most preferably, the stem-loop structure is a stem-loop structure formed by bases at positions 11-65 (SL3), 104-119 (SL1), 123-152 (SL2), and 155-191 (SL4) in the sequence shown in SEQ ID NO:
6.
6. A guide RNA molecule, comprising, from 5' to 3' direction, the sgRNA molecule of claim 4 or 5 and a guide sequence, wherein the guide sequence comprises a complementary sequence to the target sequence; Preferably, the guide sequence comprises a sequence selected from the following, or consists of a sequence selected from the following: (f1) a sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27; (f2) a sequence having one or more base substitutions, deletions or additions compared to the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, and 27; (f3) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity to the sequence shown in any one of SEQ ID NOs: 7, 20, 23, 24, 25, 26, 27; (f4) a sequence that hybridizes to the sequence described in any one of (f1) to (f3) under stringent conditions; or (f5) A complementary sequence of the sequence described in any one of (f1) to (f3); Furthermore, the sequence described in any one of (f2) to (f5) substantially retains the biological function of the sequence from which it is derived; more preferably, the biological function of the sequence refers to the activity as a guide sequence in the CRISPR-Cas system; Preferably, the guide RNA molecule comprises a sequence selected from the following, or consists of a sequence selected from the following: (g1) a sequence shown in any one of SEQ ID NOs: 9-13; (g2) the sequence shown at bases 36 to 250 of SEQ ID NO:9, the sequence shown at bases 36 to 234 of SEQ ID NO:10, the sequence shown at bases 36 to 220 of SEQ ID NO:11, the sequence shown at bases 36 to 195 of SEQ ID NO:12, or the sequence shown at bases 36 to 213 of SEQ ID NO:13; (g3) a sequence having one or more base substitutions, deletions or additions compared to the sequence shown in any one of SEQ ID NOs: 9-13; (g4) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity to any one of SEQ ID NOs: 9-13; (g5) a sequence that hybridizes to the sequence described in any one of (g1) to (g4) under stringent conditions; or (g6) A complementary sequence of the sequence described in any one of (g1) to (g4); Furthermore, the sequence described in any one of (g2) to (g6) substantially retains the biological function of the sequence from which it is derived, and the biological function of the sequence refers to the activity as a guide RNA sequence in the CRISPR-Cas system; Preferably, the target sequence is a DNA or RNA sequence from a prokaryotic cell or a eukaryotic cell, or the target sequence is a non-naturally occurring DNA or RNA sequence; more preferably, the target sequence exists in a cell, or the target sequence exists in a nucleic acid molecule in vitro; more preferably, the cell is a bacterial, plant or animal cell; most preferably, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer sequence adjacent to the continued PAM, and the PAM has the sequence shown in 5'-AAN, wherein, N is selected from A, G, T, C.
7. A composition comprising: (h1) a protein component selected from the group consisting of: a protein according to any one of claims 1 to 3, a direct homologue, homologue, variant or functional fragment thereof, and any combination thereof; and (h2) a nucleic acid component comprising the sgRNA molecule of claim 4 or 5; or the guide RNA molecule of claim 6; in, The protein component and the nucleic acid component in the composition exist independently, or are combined with each other to form a complex; Preferably, the nucleic acid component comprises the sgRNA molecule of claim 4 or 5 and the guide sequence defined in claim 6 from 5' to 3' direction; Preferably, the nucleic acid component further comprises a promoter, and the promoter is connected to the 5' end of the nucleic acid molecule; more preferably, the promoter is the J23119 promoter; most preferably, the nucleotide sequence of the J23119 promoter comprises or is positions 1-35 of the nucleotide sequence shown in any one of SEQ ID NOs: 9-13; Preferably, the nucleic acid component further comprises a poly-T sequence, and more preferably the poly-T sequence is more than 3, more than 4, more than 5, more than 6, more than 7 or more consecutive Ts.
8. An isolated nucleic acid molecule comprising: (i1) encoding the protein according to any one of claims 1 to 3, or its ortholog, homolog, variant or functional fragment Nucleotide sequence; (i2) a nucleotide sequence encoding the sgRNA molecule according to claim 4 or 5; (i3) a nucleotide sequence encoding the guide RNA molecule of claim 6; and / or, (i4) a nucleotide sequence comprising (i1) to (i3); Preferably, the nucleotide sequence described in any one of (i1) to (i4) is codon-optimized for expression in prokaryotic cells or eukaryotic cells.
9. A vector, a composition, a host cell or a kit containing the vector, characterized in that: The vector comprises the isolated nucleic acid molecule of claim 8; The carrier-containing composition comprises one or more carriers, and the one or more carriers comprise: (j1) a first nucleic acid, which is a nucleotide sequence encoding the protein according to any one of claims 1 to 3, its ortholog, homolog, variant or functional fragment; optionally, the first nucleic acid is operably linked to a first regulatory element; and (j2) a second nucleic acid encoding a nucleotide sequence comprising the sgRNA molecule of claim 4 or 5, or encoding a nucleotide sequence of the guide RNA molecule of claim 6; optionally, the second nucleic acid is operably linked to a second regulatory element; wherein the first nucleic acid and the second nucleic acid are present on the same or different vectors; The host cell comprises the isolated nucleic acid molecule of claim 8, and / or expresses the protein of any one of claims 1 to 3, its ortholog, homolog, variant or functional fragment, the sgRNA molecule of claim 4 or 5, the guide sequence defined in claim 6, the guide RNA molecule of claim 6, and / or contains the composition of claim 7; The kit contains the protein according to any one of claims 1 to 3, its ortholog, homolog, variant or functional fragment, the sgRNA molecule according to claim 4 or 5, the guide sequence defined in claim 6, the guide RNA molecule according to claim 6, and / or the composition according to claim 7.
10. Use of the protein according to any one of claims 1 to 3, its ortholog, homolog, variant or functional fragment, the sgRNA molecule according to claim 4 or 5, the guide RNA molecule according to claim 6, the composition according to claim 7, the isolated nucleic acid molecule according to claim 8, the vector, vector composition, delivery composition, host cell or kit according to claim 9, wherein the use is selected from: (k1) for nucleic acid editing or modification; (k2) Preparations for nucleic acid editing or modification; (k3) in vitro or ex vivo DNA testing; (k4) Preparation for in vitro or ex vivo DNA detection; (k5) Used to modify cells, cell lines or organisms by changing one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.
Citation Information
Patent Citations
Crispr / cas effector protein and system
CN112004932A
Type 2 CRISPR / Cas9 gene editing system and application thereof
CN114075559A
Novel Cas effector proteins, gene editing systems and their applications
CN114934031A
CasD protein, CRISPR / CasD gene editing system and application of CRISPR / CasD gene editing system in plant gene editing
CN116286742A
NOVEL Cas ENZYME AND SYSTEM, AND USE THEREOF
US20220186206A1