CRISPR / Cas effect protein and system
Patent Information
- Application Number
- CN202480013823.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-21
- Filing Date
- 2024-02-21
- Publication Date
- 2025-10-24
AI Technical Summary
The existing CRISPR/Cas system is difficult to achieve large fragment deletions and chromosome elimination in genome editing, and the length of fragment deletions produced by the Type I system is uncertain, limiting its applicability.
A Type I-A CRISPR-Cas system was developed, which contains a specific Cas protein and guide RNA, which can recognize specific PAM sequences and achieve precise large-fragment deletion and editing of target genes or genomes.
It achieves precise deletion and editing of large fragments of the genome, improves the accuracy and controllability of editing, and overcomes the difficulties of existing systems in deleting large fragments and eliminating chromosomes.
Smart Images

Figure 00000114_0000 
Figure 00000114_0001 
Figure 00000115_0000
Abstract
Description
CRISPR / Cas effector proteins and systems
[0001] This application is based on and claims priority to an application with CN application number 202310187307.6 and filing date February 21, 2023. The contents of the CN application are hereby incorporated into this application as a whole. Technical Field
[0002] The present invention relates to the technical field of regularly clustered interspaced short palindromic repeats (CRISPR). Specifically, the present invention relates to Type IA CRISPR-Cas effector proteins and systems, fusion proteins comprising such proteins, and nucleic acid molecules encoding them. The present invention also relates to complexes and compositions for nucleic acid editing (e.g., gene or genome editing, large fragment deletions, single base editing, genomic structural variation), which comprise proteins or fusion proteins of the present invention, or nucleic acid molecules encoding them. The present invention also relates to methods for nucleic acid editing (e.g., gene or genome editing, large fragment deletions, single base editing, genomic structural variation), which use proteins or fusion proteins comprising the present invention. Background Art
[0003] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to specifically bind to target sequences on the genome and cut DNA to produce double-strand breaks, and then performs site-specific editing of the genome through biological non-homologous end joining or homologous recombination repair methods. Currently, based on the classification of existing CRISPR systems, they can be divided into two major categories: class 1 and class 2 (Liu and Doudna 2020). Among them, the class 2 system is mainly composed of a single effector protein, and the widely used CRISPR / Cas9 system belongs to the type II family in the class 2 system. Although the technical application of the CRISPR / Cas9 system in the field of gene editing is very mature, since the types of edits produced by CRISPR / Cas9 after genome editing are mainly small deletion fragments, it is still very difficult to use the CRISPR / Cas9 system for large genome fragment deletion or chromosome elimination.
[0004] Class 1 systems are primarily composed of multiple effector proteins and are currently divided into three families: type I, type II, and type III. The most maturely studied are type E systems within the type I family. Similar to class 2 systems, class 1 systems, guided by a guide RNA, invade target sequences by recognizing a PAM motif, thereby binding and cleaving the substrate DNA. Type IE systems primarily consist of two components: the nuclease-active Cas3 protein and the Cascade complex, which includes Cas5, Cas6, Cas7, Cas8e, and Cas11. The guide RNA binds to the Cascade complex to recognize the substrate DNA and then recruits the Cas3 protein to cleave it. Reports on editing human 293T cells using the type IE system have found that it primarily induces long, long-range deletions in the genome. However, the random length of these deletions limits their application in production. Furthermore, there are few reports on genome editing in eukaryotic organisms using other class 1 families.
[0005] Therefore, given the current CRISPR / Cas system's limitations in terms of the length of deletions produced by genome editing and the limitations of random fragment deletions produced by type I system editing, it is of great significance to develop a more robust CRISPR / Cas system that can achieve precise large-fragment deletions of the genome.
[0006] Summary of the Invention
[0007] After extensive experiments and repeated exploration, the inventors of the present application unexpectedly developed a new Type IA CRISPR-Cas system or vector system and a method for applying the system, which can be used to achieve precise large-scale deletion of target genes or genomes and / or other target nucleic acid editing (such as modifying genes, knocking out genes, changing the expression of gene products, repairing mutations, inserting polynucleotides, and / or single-base mutations, etc.).
[0008] I. Type IA system - protein part
[0009] In one aspect, the present application provides a Type IA CRISPR-Cas system comprising:
[0010] (1) a Cas5a protein or a nucleotide sequence encoding a Cas5a protein, wherein the Cas5a protein has an amino acid sequence as shown in any one of SEQ ID NOs: 2, 8, 14, 20, or an ortholog, homolog, variant, or functional fragment thereof;
[0011] (2) a Cas8a protein or a nucleotide sequence encoding a Cas8a protein, wherein the Cas8a protein has an amino acid sequence as shown in any one of SEQ ID NOs: 3, 9, 15, and 21, or an ortholog, homolog, variant, or functional fragment thereof;
[0012] (3) a Cas7 protein or a nucleotide sequence encoding a Cas7 protein, wherein the Cas7 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 4, 10, 16, and 22, or an ortholog, homolog, variant, or functional fragment thereof;
[0013] (4) a Cas6 protein or a nucleotide sequence encoding a Cas6 protein, wherein the Cas6 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 5, 11, 17, 23, or an ortholog, homolog, variant, or functional fragment thereof; and
[0014] (5) a csa5 protein or a nucleotide sequence encoding a csa5 protein, wherein the csa5 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 6, 12, 18, and 24, or an ortholog, homolog, variant, or functional fragment thereof;
[0015] Wherein, in any one of (1) to (5), the ortholog, homolog, variant or functional fragment substantially retains the biological function of the sequence from which it is derived.
[0016] In certain embodiments, the orthologs, homologs, variants have one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence from which they are derived, or have at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity and substantially retain the biological function of the sequence from which they are derived.
[0017] In certain embodiments, the system further comprises: (6) a Cas3 protein or a nucleotide sequence encoding a Cas3 protein, wherein the Cas3 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 1, 7, 13, 19, or an ortholog, homolog, variant, or functional fragment thereof;
[0018] Wherein, the orthologs, homologs, variants or functional fragments substantially retain the biological function of the sequence from which they are derived.
[0019] In certain embodiments, the orthologs, homologs, variants have one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence from which they are derived, or have at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity and substantially retain the biological function of the sequence from which they are derived.
[0020] In the present invention, the biological function of the above sequence refers to the activity of the Cas effector protein, including but not limited to the activity of binding to the guide RNA, the endonuclease activity, and the activity of binding to and cutting a specific site of the target sequence or its complementary sequence under the guidance of the guide RNA.
[0021] The protein of the present invention can be derivatized, for example, by being linked to another molecule (e.g., another polypeptide or protein). Typically, the derivatization (e.g., labeling) of the protein will not adversely affect the desired activity of the protein (e.g., activity of binding to a guide RNA, endonuclease activity, activity of binding to and cutting a specific site of a target sequence or its complementary sequence under the guidance of a guide RNA). Therefore, the protein of the present invention is also intended to include such derivatized forms. For example, the protein of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage or other means) to one or more other molecular groups, such as another protein or polypeptide, a detection reagent, a pharmaceutical agent, etc.
[0022] In particular, the protein of the present invention can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to improve the ability of the protein of the present invention to enter the cell nucleus. For example, it can be linked to a targeting moiety to make the protein of the present invention targeted. For example, it can be linked to a detectable label to facilitate detection of the protein of the present invention. For example, it can be linked to an epitope tag to facilitate expression, detection, tracing and / or purification of the protein of the present invention.
[0023] In certain embodiments, any one of the cas proteins in the system optionally comprises an additional protein or polypeptide selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), an adenosine deaminase (e.g., TadA8e), a cytosine deaminase (e.g., APOBEC3), a domain having an activity selected from the following: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity (e.g., single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity) and nucleic acid binding activity; and any combination thereof.
[0024] In certain embodiments, at least one cas protein in the system comprises the additional protein or polypeptide; for example, the protein described in each of (1)-(6) comprises the additional protein or polypeptide.
[0025] In certain embodiments, the additional protein or polypeptide is an NLS sequence. In certain embodiments, the protein described in each of (1) to (6) comprises an NLS sequence.
[0026] In certain embodiments, the NLS sequence is as shown in SEQ ID NO:65.
[0027] In certain embodiments, the NLS sequence is located at or near a terminus (eg, the N-terminus or the C-terminus) of the protein.
[0028] In certain embodiments, the additional protein or polypeptide is an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3). In certain embodiments, one of the proteins described in any one of (1) to (5) comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3).
[0029] In certain embodiments, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (eg, N-terminus or C-terminus) of the protein (eg, Cas8a protein).
[0030] In certain embodiments, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
[0031] In certain embodiments, the additional protein or polypeptide is linked to the protein with or without a linker.
[0032] In certain embodiments, the linker is a peptide linker or a non-peptide linker.
[0033] In certain embodiments, the peptide linker sequence is shown in SEQ ID NO: 66, 67 or 95.
[0034] In certain embodiments, the proteins of the present invention comprise an epitope tag. Such epitope tags are well known to those skilled in the art, and examples thereof include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag according to the desired purpose (e.g., purification, detection, or tracing).
[0035] In certain embodiments, the protein of the present invention comprises a reporter gene sequence. Such reporter genes are well known to those skilled in the art, and examples thereof include but are not limited to GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0036] In certain embodiments, the system does not comprise a Cas3 protein or a nucleotide sequence encoding a Cas3 protein.
[0037] In certain embodiments, a Cas protein in the system (e.g., Cas5a protein, Cas8a protein, Cas7 protein, Cas6 protein, or Csa5 protein) comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the Cas protein.
[0038] In certain embodiments, the cas8a protein in the system comprises an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
[0039] In certain embodiments, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
[0040] In certain embodiments, the adenosine deaminase or cytosine deaminase is linked to the protein via a linker or without a linker.
[0041] In certain embodiments, the linker is a peptide linker or a non-peptide linker;
[0042] In certain embodiments, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; in certain embodiments, the peptide linker sequence is as shown in SEQ ID NO: 95.
[0043] In certain embodiments, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in any one of SEQ ID NOs: 96-99.
[0044] II. Type IA-1 system
[0045] In certain embodiments, in the system, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein respectively comprise the amino acid sequences shown in SEQ ID NOs: 2-6.
[0046] In certain embodiments, one or more (e.g., all 5) of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are connected to an NLS sequence (e.g., a sequence shown in SEQ ID NO: 65) via a linker or not.
[0047] In certain embodiments, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein connected to the NLS sequence respectively comprise the amino acid sequences shown in SEQ ID NOs: 69-73.
[0048] In certain embodiments, the system further comprises: (6) a cas3 protein or a nucleotide sequence encoding a cas3 protein;
[0049] Wherein, the cas3 protein comprises the amino acid sequence shown in SEQ ID NO:1.
[0050] In certain embodiments, the Cas3 protein is connected to an NLS sequence (eg, the sequence shown in SEQ ID NO: 65) with or without a linker.
[0051] In certain embodiments, the cas3 protein linked to the NLS sequence comprises the amino acid sequence shown in SEQ ID NO:68.
[0052] In certain embodiments, the system does not comprise a Cas3 protein or a nucleotide sequence encoding a Cas3 protein.
[0053] In certain embodiments, a Cas protein in the system (e.g., Cas5a protein, Cas8a protein, Cas7 protein, Cas6 protein, or Csa5 protein) comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the Cas protein.
[0054] In certain embodiments, the cas8a protein in the system comprises an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
[0055] In certain embodiments, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
[0056] In certain embodiments, the adenosine deaminase or cytosine deaminase is linked to the protein via a linker or without a linker.
[0057] For example, the linker is a peptide linker or a non-peptide linker.
[0058] For example, the peptide linker sequence is shown in SEQ ID NO: 66, 67 or 95; in certain embodiments, the peptide linker sequence is shown in SEQ ID NO: 95.
[0059] For example, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in SEQ ID NO:96.
[0060] I-II.Type IA-2 system
[0061] In certain embodiments, in the system, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein respectively comprise the amino acid sequences shown in SEQ ID NOs: 8-12.
[0062] In certain embodiments, one or more (e.g., all 5) of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are connected to an NLS sequence (e.g., a sequence shown in SEQ ID NO: 65) via a linker or not.
[0063] In certain embodiments, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein connected to the NLS sequence respectively comprise the amino acid sequences shown in SEQ ID NOs: 75-79.
[0064] In certain embodiments, the system further comprises: (6) a cas3 protein or a nucleotide sequence encoding a cas3 protein;
[0065] Wherein, the cas3 protein comprises the amino acid sequence shown in SEQ ID NO:7.
[0066] In certain embodiments, the Cas3 protein is connected to an NLS sequence (eg, the sequence shown in SEQ ID NO: 65) with or without a linker.
[0067] In certain embodiments, the cas3 protein linked to the NLS sequence comprises the amino acid sequence shown in SEQ ID NO:74.
[0068] In certain embodiments, the system does not comprise a Cas3 protein or a nucleotide sequence encoding a Cas3 protein.
[0069] In certain embodiments, a Cas protein in the system (e.g., Cas5a protein, Cas8a protein, Cas7 protein, Cas6 protein, or Csa5 protein) comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the Cas protein.
[0070] In certain embodiments, the cas8a protein in the system comprises an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
[0071] In certain embodiments, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
[0072] In certain embodiments, the adenosine deaminase or cytosine deaminase is linked to the protein via a linker or without a linker.
[0073] In certain embodiments, the linker is a peptide linker or a non-peptide linker.
[0074] In certain embodiments, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; in certain embodiments, the peptide linker sequence is as shown in SEQ ID NO: 95.
[0075] In certain embodiments, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in SEQ ID NO:97.
[0076] I-III.Type IA-3 system
[0077] In certain embodiments, in the system, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein respectively comprise the amino acid sequences shown in SEQ ID NOs: 14-18.
[0078] In certain embodiments, one or more (e.g., all 5) of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are connected to an NLS sequence (e.g., a sequence shown in SEQ ID NO: 65) via a linker or not.
[0079] In certain embodiments, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein connected to the NLS sequence respectively comprise the amino acid sequences shown in SEQ ID NOs: 81-85.
[0080] In certain embodiments, the system further comprises: (6) a cas3 protein or a nucleotide sequence encoding a cas3 protein;
[0081] Wherein, the cas3 protein comprises the amino acid sequence shown in SEQ ID NO:13.
[0082] In certain embodiments, the Cas3 protein is connected to an NLS sequence (eg, the sequence shown in SEQ ID NO: 65) with or without a linker.
[0083] In certain embodiments, the cas3 protein linked to the NLS sequence comprises the amino acid sequence shown in SEQ ID NO:80.
[0084] In certain embodiments, the system does not comprise a Cas3 protein or a nucleotide sequence encoding a Cas3 protein.
[0085] In certain embodiments, a Cas protein in the system (e.g., Cas5a protein, Cas8a protein, Cas7 protein, Cas6 protein, or Csa5 protein) comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the Cas protein.
[0086] In certain embodiments, the cas8a protein in the system comprises an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
[0087] In certain embodiments, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
[0088] In certain embodiments, the adenosine deaminase or cytosine deaminase is linked to the protein via a linker or without a linker.
[0089] In certain embodiments, the linker is a peptide linker or a non-peptide linker.
[0090] In certain embodiments, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; in certain embodiments, the peptide linker sequence is as shown in SEQ ID NO: 95.
[0091] In certain embodiments, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in SEQ ID NO:98.
[0092] I-IV. Type IA-4 systems
[0093] In certain embodiments, in the system, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein respectively comprise the amino acid sequences shown in SEQ ID NOs: 20-24.
[0094] In certain embodiments, one or more (e.g., all 5) of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are connected to an NLS sequence (e.g., a sequence shown in SEQ ID NO: 65) via a linker or not.
[0095] In certain embodiments, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein connected to the NLS sequence respectively comprise the amino acid sequences shown in SEQ ID NOs: 87-91.
[0096] In certain embodiments, the system further comprises: (6) a cas3 protein or a nucleotide sequence encoding a cas3 protein;
[0097] Wherein, the cas3 protein comprises the amino acid sequence shown in SEQ ID NO:19.
[0098] In certain embodiments, the Cas3 protein is connected to an NLS sequence (eg, the sequence shown in SEQ ID NO: 65) with or without a linker.
[0099] In certain embodiments, the cas3 protein connected to the NLS sequence comprises the amino acid sequence shown in SEQ ID NO:86.
[0100] In certain embodiments, the system does not comprise a Cas3 protein or a nucleotide sequence encoding a Cas3 protein.
[0101] In certain embodiments, a Cas protein in the system (e.g., Cas5a protein, Cas8a protein, Cas7 protein, Cas6 protein, or Csa5 protein) comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the Cas protein.
[0102] In certain embodiments, the cas8a protein in the system comprises an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
[0103] In certain embodiments, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
[0104] In certain embodiments, the adenosine deaminase or cytosine deaminase is linked to the protein via a linker or without a linker.
[0105] In certain embodiments, the linker is a peptide linker or a non-peptide linker.
[0106] In certain embodiments, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; in certain embodiments, the peptide linker sequence is as shown in SEQ ID NO: 95.
[0107] In certain embodiments, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in SEQ ID NO:99.
[0108] II. Type IA system - protein and guide RNA
[0109] In certain embodiments, the system further comprises a guide RNA of a Type IA CRISPR-Cas system or a nucleotide sequence encoding the guide RNA; wherein the guide RNA comprises a direct repeat sequence and a guide sequence capable of hybridizing with a target sequence.
[0110] In certain embodiments, the direct repeat sequence comprises a stem-loop structure.
[0111] In certain embodiments, the direct repeat sequence is capable of binding to one or more cas proteins in the system; for example, the direct repeat sequence is capable of binding to one or more proteins selected from cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein; for example, the guide RNA is capable of binding to the Cascade complex formed by cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein.
[0112] In certain embodiments, when the target sequence is DNA, the protospacer adjacent motif (PAM) recognized by the system has a sequence represented by 5'CCN-. In certain embodiments, the PAM has a sequence represented by 5'CCT- or 5'CCC-.
[0113] In certain embodiments, the direct repeat sequence comprises a first region and a second region, and the first region comprises a stem-loop structure.
[0114] In certain embodiments, the first region is located 5' to the second region.
[0115] In certain embodiments, there may or may not be extra nucleotides between the first region and the second region.
[0116] In certain embodiments, the guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a guide sequence located between the first copy of the direct repeat sequence and the second copy of the direct repeat sequence.
[0117] In certain embodiments, the guide RNA comprises the second region of the first copy of the direct repeat sequence, a guide sequence, and the first region of the second copy of the direct repeat sequence.
[0118] In certain embodiments, the targeting sequence is located between the second region of the first copy of the direct repeat sequence and the first region of the second copy of the direct repeat sequence.
[0119] In certain embodiments, the second region of the first copy of the direct repeat is located 5' to the guide sequence, and the first region of the second copy of the direct repeat is located 3' to the guide sequence.
[0120] In certain embodiments, the second region of the first copy of the direct repeat sequence may or may not contain redundant nucleotides between the second region and the guide sequence.
[0121] In certain embodiments, the guide sequence may or may not contain superfluous nucleotides between the guide sequence and the first region of the second copy of the direct repeat sequence.
[0122] In certain embodiments, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in Section II above, the direct repeat sequence comprises the sequence shown in SEQ ID NO: 49 or consists of the sequence shown in SEQ ID NO: 49. In certain embodiments, the first region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 51 or consists of the sequence shown in SEQ ID NO: 51, and the second region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 52 or consists of the sequence shown in SEQ ID NO: 52.
[0123] In certain embodiments, when the Cas5a protein, Cas8a protein, Cas7 protein, Cas6 protein, and CSA5 protein are as defined in Sections I-II above, the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 53. In certain embodiments, the first region of the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 55, and the second region of the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 56.
[0124] In certain embodiments, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in Sections I-III above, the direct repeat sequence comprises the sequence shown in SEQ ID NO: 57 or consists of the sequence shown in SEQ ID NO: 57. In certain embodiments, the first region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 59 or consists of the sequence shown in SEQ ID NO: 59, and the second region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 60 or consists of the sequence shown in SEQ ID NO: 60.
[0125] In certain embodiments, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are as defined in Sections I-IV above, the direct repeat sequence comprises the sequence shown in SEQ ID NO: 61 or consists of the sequence shown in SEQ ID NO: 61. In certain embodiments, the first region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 63 or consists of the sequence shown in SEQ ID NO: 63, and the second region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 64 or consists of the sequence shown in SEQ ID NO: 64.
[0126] III. Type IA System - Protein and Dual-Targeting Guide RNA
[0127] In certain embodiments, the system further comprises one or more guide RNAs of a Type IA CRISPR-Cas system or a nucleotide sequence encoding the one or more guide RNAs; wherein the one or more guide RNAs comprise a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, and a second guide sequence capable of hybridizing to a second target sequence;
[0128] The first target sequence and the second target sequence are respectively located on the flanks of the region to be modified (eg, the region to be deleted) in the double-stranded target nucleic acid molecule.
[0129] In certain embodiments, the first target sequence and the second target sequence are respectively located on two single strands of the region to be modified; for example, the first target sequence and the second target sequence are respectively located at the 5' end of the region to be modified in their respective single strands.
[0130] In certain embodiments, the direct repeat sequence comprises a stem-loop structure.
[0131] In certain embodiments, the direct repeat sequence can bind to one or more cas proteins in the system; for example, the direct repeat sequence can bind to one or more proteins selected from cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein. In certain embodiments, the guide RNA can bind to the Cascade complex formed by cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein.
[0132] In certain embodiments, when the target sequence is DNA, the protospacer adjacent motif (PAM) recognized by the system has a sequence represented by 5'CCN-; in certain embodiments, the PAM has a sequence represented by 5'CCT- or 5'CCC-.
[0133] In certain embodiments, the direct repeat sequence comprises a first region and a second region, and the first region comprises a stem-loop structure.
[0134] In certain embodiments, the first region is located 5' to the second region.
[0135] In certain embodiments, there may or may not be extra nucleotides between the first region and the second region.
[0136] In certain embodiments, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in Section II above, the direct repeat sequence is as shown in SEQ ID NO: 69. In certain embodiments, the first region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 51 or consists of the sequence shown in SEQ ID NO: 51, and the second region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 52 or consists of the sequence shown in SEQ ID NO: 52.
[0137] In certain embodiments, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in Sections I-II above, the direct repeat sequence is as shown in SEQ ID NO: 53. In certain embodiments, the first region of the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 55, and the second region of the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 56.
[0138] In certain embodiments, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in Sections I-III above, the direct repeat sequence is as shown in SEQ ID NO: 57. In certain embodiments, the first region of the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 59, and the second region of the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 60.
[0139] In certain embodiments, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in Sections I-IV above, the direct repeat sequence is as shown in SEQ ID NO: 61. In certain embodiments, the first region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 63 or consists of the sequence shown in SEQ ID NO: 63, and the second region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 64 or consists of the sequence shown in SEQ ID NO: 64.
[0140] In certain embodiments, the one guide RNA comprises:
[0141] (i) a first copy of a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, a second copy of a direct repeat sequence, a second guide sequence capable of hybridizing to a second target sequence, and a third copy of a direct repeat sequence; or
[0142] (ii) a second region of the first copy of the direct repeat sequence, a first guide sequence capable of hybridizing to the first target sequence, a second copy of the direct repeat sequence, a second guide sequence capable of hybridizing to the second target sequence, and a first region of the third copy of the direct repeat sequence;
[0143] In certain embodiments, the guide RNA in (i) comprises, from 5' to 3' direction: a first copy of the direct repeat sequence, the first guide sequence, a second copy of the direct repeat sequence, the second guide sequence, and a third copy of the direct repeat sequence. In certain embodiments, the guide RNA in (ii) comprises, from 5' to 3' direction: a second region of the first copy of the direct repeat sequence, the first guide sequence, the second copy of the direct repeat sequence, the second guide sequence, and a first region of the third copy of the direct repeat sequence.
[0144] In certain embodiments, the plurality of guide RNAs comprises:
[0145] a first guide RNA comprising a direct repeat sequence and a first guide sequence capable of hybridizing to a first target sequence; and
[0146] A second guide RNA comprises a direct repeat sequence and a second guide sequence capable of hybridizing to a second target sequence.
[0147] In certain embodiments, the first guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a first guide sequence located between the two copies of the repeat sequence; or, the first guide RNA comprises, from a 5' to 3' direction, a second region of the first copy of the direct repeat sequence, a first guide sequence, and a first region of the second copy of the direct repeat sequence.
[0148] In certain embodiments, the second guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a second guide sequence located between the two copies of the repeat sequence; or, the second guide RNA comprises, from a 5' to 3' direction, a second region of the first copy of the direct repeat sequence, a second guide sequence, and a first region of the second copy of the direct repeat sequence.
[0149] IV. Effector proteins of the Type IA system
[0150] On the other hand, the present application provides a cas protein of a Type IA CRISPR-Cas system selected from:
[0151] (1) Cas5a protein, wherein the Cas5a protein has an amino acid sequence as shown in any one of SEQ ID NOs: 2, 8, 14, 20 or an ortholog, homolog, variant or functional fragment thereof;
[0152] (2) Cas8a protein, wherein the Cas8a protein has an amino acid sequence as shown in any one of SEQ ID NOs: 3, 9, 15, 21 or an ortholog, homolog, variant or functional fragment thereof;
[0153] (3) Cas7 protein, wherein the Cas7 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 4, 10, 16, 22 or an ortholog, homolog, variant or functional fragment thereof;
[0154] (4) Cas6 protein, wherein the Cas6 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 5, 11, 17, 23 or an ortholog, homolog, variant or functional fragment thereof;
[0155] (5) csa5 protein, wherein the csa5 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 6, 12, 18, 24, or an ortholog, homolog, variant, or functional fragment thereof;
[0156] (6) Cas3 protein, wherein the Cas3 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 1, 7, 13, 19 or an ortholog, homolog, variant or functional fragment thereof;
[0157] Wherein, in any one of (1) to (6), the ortholog, homolog, variant or functional fragment substantially retains the biological function of the sequence from which it is derived.
[0158] In certain embodiments, the orthologs, homologs, variants have one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence from which they are derived, or have at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity and substantially retain the biological function of the sequence from which they are derived.
[0159] In the present invention, the biological function of the above sequence refers to the activity of the Cas effector protein, including but not limited to the activity of binding to the guide RNA, the endonuclease activity, and the activity of binding to and cutting a specific site of the target sequence or its complementary sequence under the guidance of the guide RNA.
[0160] The protein of the present invention can be derivatized, for example, by being linked to another molecule (e.g., another polypeptide or protein). Typically, the derivatization (e.g., labeling) of the protein will not adversely affect the desired activity of the protein (e.g., activity of binding to a guide RNA, endonuclease activity, activity of binding to and cutting a specific site of a target sequence or its complementary sequence under the guidance of a guide RNA). Therefore, the protein of the present invention is also intended to include such derivatized forms. For example, the protein of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage or other means) to one or more other molecular groups, such as another protein or polypeptide, a detection reagent, a pharmaceutical agent, etc.
[0161] In particular, the protein of the present invention can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to improve the ability of the protein of the present invention to enter the cell nucleus. For example, it can be linked to a targeting moiety to make the protein of the present invention targeted. For example, it can be linked to a detectable label to facilitate detection of the protein of the present invention. For example, it can be linked to an epitope tag to facilitate expression, detection, tracing and / or purification of the protein of the present invention.
[0162] In certain embodiments, the protein described in any one of (1) to (6) optionally comprises an additional protein or polypeptide, wherein the additional protein or polypeptide is selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), an adenosine deaminase (e.g., TadA8e), a cytosine deaminase (e.g., APOBEC3), a domain having an activity selected from the following: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity (e.g., single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity) and nucleic acid binding activity; and any combination thereof.
[0163] In certain embodiments, at least one (e.g., at least two, at least three, at least four, or all five) of the proteins described in any one of (1)-(6) comprises the additional protein or polypeptide; for example, the protein described in each of (1)-(6) comprises the additional protein or polypeptide.
[0164] In certain embodiments, the additional protein or polypeptide is an NLS sequence; for example, the protein described in each of (1)-(6) comprises an NLS sequence.
[0165] In certain embodiments, the NLS sequence is as shown in SEQ ID NO:65.
[0166] In certain embodiments, the additional protein or polypeptide is linked to the protein with or without a linker.
[0167] In certain embodiments, the linker is a peptide linker or a non-peptide linker.
[0168] In certain embodiments, the peptide linker sequence is shown in SEQ ID NO: 66, 67 or 95.
[0169] In certain embodiments, the NLS sequence is located at or near a terminus (eg, the N-terminus or the C-terminus) of the protein.
[0170] In certain embodiments, the additional protein or polypeptide is an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3). In certain embodiments, one of the proteins described in any one of (1) to (5) comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3).
[0171] In certain embodiments, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (eg, N-terminus or C-terminus) of the protein (eg, Cas8a protein).
[0172] For example, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
[0173] In certain embodiments, the cas5a protein comprises an NLS sequence and comprises, or consists of, a sequence selected from the group consisting of: (i) a sequence shown in any one of SEQ ID NOs: 69, 75, 81, and 87; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to the sequence shown in any one of SEQ ID NOs: 69, 75, 81, and 87; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence shown in any one of SEQ ID NOs: 69, 75, 81, and 87.
[0174] In certain embodiments, the cas8a protein comprises an NLS sequence and comprises, or consists of, a sequence selected from the group consisting of: (i) a sequence shown in any one of SEQ ID NOs: 70, 76, 82, and 88; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to the sequence shown in any one of SEQ ID NOs: 70, 76, 82, and 88; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence shown in any one of SEQ ID NOs: 70, 76, 82, and 88.
[0175] In certain embodiments, the cas7 protein comprises an NLS sequence and comprises, or consists of, a sequence selected from the group consisting of: (i) a sequence shown in any one of SEQ ID NOs: 71, 77, 83, and 89; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to the sequence shown in any one of SEQ ID NOs: 71, 77, 83, and 89; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence shown in any one of SEQ ID NOs: 71, 77, 83, and 89.
[0176] In certain embodiments, the cas6 protein comprises an NLS sequence and comprises, or consists of, a sequence selected from the group consisting of: (i) a sequence shown in any one of SEQ ID NOs: 72, 78, 84, and 90; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to the sequence shown in any one of SEQ ID NOs: 72, 78, 84, and 90; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence shown in any one of SEQ ID NOs: 72, 78, 84, and 90.
[0177] In certain embodiments, the csa5 protein comprises an NLS sequence and comprises, or consists of, a sequence selected from the group consisting of: (i) a sequence as set forth in any one of SEQ ID NOs: 73, 79, 85, and 91; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to the sequence as set forth in any one of SEQ ID NOs: 73, 79, 85, and 91; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence as set forth in any one of SEQ ID NOs: 73, 79, 85, and 91.
[0178] In certain embodiments, the Cas3 protein comprises an NLS sequence and comprises, or consists of, a sequence selected from the group consisting of: (i) a sequence shown in any one of SEQ ID NOs: 68, 74, 80, and 86; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to the sequence shown in any one of SEQ ID NOs: 68, 74, 80, and 86; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence shown in any one of SEQ ID NOs: 68, 74, 80, and 86.
[0179] V. Directly Repeated Sequence-Related Nucleic Acid Molecules
[0180] In another aspect, the present application provides an isolated nucleic acid molecule comprising or consisting of a sequence selected from the following:
[0181] (i) a sequence shown in any one of SEQ ID NOs: 49, 53, 57, and 61;
[0182] (ii) comprising the sequences shown in SEQ ID NOs: 51 and 52, comprising the sequences shown in SEQ ID NOs: 55 and 56, comprising the sequences shown in SEQ ID NOs: 59 and 60, or comprising the sequences shown in SEQ ID NOs: 63 and 64;
[0183] (iii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in (i) or (ii);
[0184] (iv) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in (i) or (ii);
[0185] (v) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iv); or
[0186] (vi) a complementary sequence of the sequence described in any one of (i) to (iv);
[0187] Furthermore, the sequence of any one of (iii) to (vi) substantially retains the biological function of the sequence from which it is derived.
[0188] In certain embodiments, the nucleic acid molecule is capable of binding to one or more of the cas proteins described in Section IV above. In certain embodiments, the nucleic acid molecule is capable of binding to one or more proteins selected from the group consisting of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein.
[0189] In certain embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:
[0190] (a) the nucleotide sequence shown in any one of SEQ ID NOs: 49, 53, 57, and 61;
[0191] (b) comprising the sequences shown in SEQ ID NOs: 51 and 52, comprising the sequences shown in SEQ ID NOs: 55 and 56, comprising the sequences shown in SEQ ID NOs: 59 and 60, or comprising the sequences shown in SEQ ID NOs: 63 and 64;
[0192] (c) a sequence that hybridizes under stringent conditions to the sequence described in (a) or (b); or
[0193] (d) The complementary sequence of the sequence described in (a) or (b).
[0194] In certain embodiments, the isolated nucleic acid molecule is RNA.
[0195] In certain embodiments, the isolated nucleic acid molecule is a direct repeat sequence in a CRISPR / Cas system or a fragment thereof.
[0196] VI. Protein Expression-Related Nucleic Acid Molecules / Vectors / Host Cells
[0197] In another aspect, the present application provides an isolated nucleic acid molecule encoding a protein as described in Section IV above.
[0198] In another aspect, the present application provides a vector comprising the isolated nucleic acid molecule as described in Section VI.
[0199] In another aspect, the present application provides a host cell comprising the isolated nucleic acid molecule or vector as described in Section VI.
[0200] Such host cells include, but are not limited to, prokaryotic cells such as bacterial cells (e.g., E. coli cells), and eukaryotic cells such as fungal cells (e.g., yeast cells), insect cells, plant cells, and animal cells (e.g., mammalian cells, e.g., mouse cells, human cells, etc.).
[0201] In certain embodiments, the cell or its progeny is incapable of developing into a whole animal or plant.
[0202] In certain embodiments, the host cell is a microorganism.
[0203] VII. Type IA vector system - protein part
[0204] On the other hand, the present application provides a Type IA CRISPR-Cas vector system, comprising one or more vectors, wherein the one or more vectors comprise: a nucleotide sequence encoding a cas protein in a Type IA CRISPR-Cas system, wherein the cas protein comprises: cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein;
[0205] Wherein, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are as defined in any one of the above sections I to IV.
[0206] In certain embodiments, the nucleotide sequence encoding the cas protein is located in one or more expression cassettes.
[0207] In certain embodiments, the nucleotide sequences encoding the cas proteins located in the same expression cassette are arranged in any order.
[0208] In certain embodiments, the nucleotide sequences encoding the cas protein located in the same expression cassette are linked to each other by a nucleotide sequence encoding a self-cleaving peptide (eg, T2A).
[0209] In certain embodiments, the expression cassettes each independently comprise a promoter, such as an inducible promoter.
[0210] In certain embodiments, the one or more vectors further comprise a nucleotide sequence encoding a cas3 protein;
[0211] Wherein, the Cas3 protein is as defined in any one of the above sections I-IV.
[0212] In certain embodiments, the nucleotide sequences encoding Cas5a protein, Cas8a protein, Cas7 protein, Cas6 protein, Csa5 protein and Cas3 protein are located in the same expression cassette.
[0213] In certain embodiments, the one or more vectors comprise:
[0214] A first expression cassette comprising a nucleotide sequence encoding a cas3 protein and a csa5 protein; and
[0215] The second expression cassette comprises a nucleotide sequence encoding a cas7 protein, a cas5a protein, a cas6 protein and a cas8a protein.
[0216] In certain embodiments, the one or more vectors do not comprise a nucleotide sequence encoding a cas3 protein;
[0217] Wherein, the cas protein in the system is as defined in any one of the above sections I-IV.
[0218] In certain embodiments, the one or more vectors comprise:
[0219] A first expression cassette comprising a nucleotide sequence encoding a cas8a protein; and
[0220] The second expression cassette comprises a nucleotide sequence encoding Cas7 protein, Cas5a protein, Cas6 protein and Csa5 protein.
[0221] In certain embodiments, the cas8a protein is as defined in any one of Sections I-IV above.
[0222] VIII. Type IA Vector System - Protein and Guide RNA
[0223] In certain embodiments, the one or more vectors further comprise: a nucleotide sequence encoding a guide RNA in a Type IA CRISPR-Cas system, wherein the guide RNA is as defined in Section II above.
[0224] In certain embodiments, the nucleotide sequence encoding the guide RNA in the Type IA CRISPR-Cas system is located in an additional expression cassette. In certain embodiments, the additional expression cassette comprises a promoter, such as an inducible promoter.
[0225] In certain embodiments, the nucleotide sequences encoding the cas proteins are all located on the same vector.
[0226] In certain embodiments, the nucleotide sequence encoding the cas protein and the nucleotide sequence encoding the guide RNA are both located on the same vector.
[0227] IX.Type IA Vector System - Protein and Dual-Targeting Guide RNA
[0228] In certain embodiments, the one or more vectors further comprise: a nucleotide sequence encoding one or more guide RNAs in a Type IA CRISPR-Cas system, wherein the one or more guide RNAs are as defined in Section III above.
[0229] In certain embodiments, the nucleotide sequence encoding one or more guide RNAs in the Type IA CRISPR-Cas system is located in an additional expression cassette. In certain embodiments, the additional expression cassette comprises a promoter, such as an inducible promoter.
[0230] In certain embodiments, the nucleotide sequences encoding the cas proteins are all located on the same vector.
[0231] In certain embodiments, the nucleotide sequence encoding the cas protein and the nucleotide sequence encoding the guide RNA are both located on the same vector.
[0232] X.Type IA System - Dual-targeting guide RNA portion
[0233] On the other hand, the present application provides a Type IA CRISPR-Cas system, comprising: one or more guide RNAs or nucleotide sequences encoding the one or more guide RNAs; wherein the one or more guide RNAs comprise a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, and a second guide sequence capable of hybridizing to a second target sequence;
[0234] The first target sequence and the second target sequence are respectively located on the flanks of the region to be modified (eg, the region to be deleted) in the double-stranded target nucleic acid molecule.
[0235] In certain embodiments, the first target sequence and the second target sequence are respectively located on two single strands of the region to be modified. In certain embodiments, the first target sequence and the second target sequence are respectively located at the 5' end of the region to be modified in their respective single strands.
[0236] In certain embodiments, the direct repeat sequence comprises a stem-loop structure.
[0237] In certain embodiments, the direct repeat sequence is capable of binding to one or more cas proteins in a Type IA CRISPR-Cas system.
[0238] In certain embodiments, when the target sequence is DNA, the protospacer adjacent motif (PAM) recognized by the system has a sequence represented by 5'CCN-; in certain embodiments, the PAM has a sequence represented by 5'CCT- or 5'CCC-.
[0239] In certain embodiments, the direct repeat sequence comprises a first region and a second region, and the first region comprises a stem-loop structure.
[0240] In certain embodiments, the first region is located 5' to the second region.
[0241] In certain embodiments, there may or may not be extra nucleotides between the first region and the second region.
[0242] In certain embodiments, the one guide RNA comprises:
[0243] (i) a first copy of a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, a second copy of a direct repeat sequence, a second guide sequence capable of hybridizing to a second target sequence, and a third copy of a direct repeat sequence; or
[0244] (ii) a second region of the first copy of the direct repeat sequence, a first guide sequence capable of hybridizing to the first target sequence, a second copy of the direct repeat sequence, a second guide sequence capable of hybridizing to the second target sequence, and a first region of the third copy of the direct repeat sequence.
[0245] In certain embodiments, the one guide RNA in (i) comprises, from 5' to 3' direction: a first copy of the direct repeat sequence, the first guide sequence, a second copy of the direct repeat sequence, the second guide sequence, and a third copy of the direct repeat sequence;
[0246] In certain embodiments, the one guide RNA in (ii) comprises, from 5' to 3' direction: the second region of the first copy of the direct repeat sequence, the first guide sequence, the second copy of the direct repeat sequence, the second guide sequence, and the first region of the third copy of the direct repeat sequence.
[0247] In certain embodiments, the plurality of guide RNAs comprises:
[0248] a first guide RNA comprising a direct repeat sequence and a first guide sequence capable of hybridizing to a first target sequence; and
[0249] A second guide RNA comprises a direct repeat sequence and a second guide sequence capable of hybridizing to a second target sequence.
[0250] In certain embodiments, the first guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a first guide sequence located between the two copies of the repeat sequence; or, the first guide RNA comprises, from a 5' to 3' direction, a second region of the first copy of the direct repeat sequence, a first guide sequence, and a first region of the second copy of the direct repeat sequence.
[0251] In certain embodiments, the second guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a second guide sequence located between the two copies of the repeat sequence; or, the second guide RNA comprises, from a 5' to 3' direction, a second region of the first copy of the direct repeat sequence, a second guide sequence, and a first region of the second copy of the direct repeat sequence.
[0252] XI.Type IA System - Dual-targeting guide RNA and protein
[0253] In certain embodiments, the system further comprises: a cas protein in a Type IA CRISPR-Cas system or a nucleotide sequence encoding the cas protein.
[0254] For example, each of the cas proteins further comprises an additional protein or polypeptide selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), an adenosine deaminase (e.g., TadA8e), a cytosine deaminase (e.g., APOBEC3), a domain having an activity selected from the following: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity (e.g., single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity) and nucleic acid binding activity; and any combination thereof.
[0255] In certain embodiments, the additional protein or polypeptide is an NLS sequence.
[0256] In certain embodiments, the additional protein or polypeptide is an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
[0257] In certain embodiments, the Cas protein comprises Cas3 protein, Cas5a protein, Cas8a protein, Cas6 protein, Csa5 protein and Cas7 protein.
[0258] In certain embodiments, the Cas3 protein, Cas5a protein, Cas8a protein, Cas6 protein, Csa5 protein and Cas7 protein are as defined in any one of Sections I-IV above.
[0259] XII. Type IA vector system - dual-targeting guide RNA portion
[0260] On the other hand, the present application provides a Type IA CRISPR-Cas vector system, which comprises one or more vectors, wherein the one or more vectors comprise: a nucleotide sequence encoding one or more guide RNAs in the Type IA CRISPR-Cas system, wherein the one or more guide RNAs are as defined in Part III above.
[0261] XIII. Type IA Vector System - Dual-Targeting Guide RNA and Protein
[0262] In certain embodiments, the one or more vectors further comprise: a nucleotide sequence encoding a cas protein in a Type IA CRISPR-Cas system.
[0263] In certain embodiments, each of the cas proteins further comprises an additional protein or polypeptide selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), an adenosine deaminase (e.g., TadA8e), a cytosine deaminase (e.g., APOBEC3), a domain having an activity selected from the following: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity (e.g., single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity) and nucleic acid binding activity; and any combination thereof.
[0264] In certain embodiments, the additional protein or polypeptide is an NLS sequence.
[0265] In certain embodiments, the additional protein or polypeptide is an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
[0266] In certain embodiments, the Cas protein comprises Cas3 protein, Cas5a protein, Cas8a protein, Cas6 protein, Csa5 protein and Cas7 protein.
[0267] In certain embodiments, the Cas3 protein, Cas5a protein, Cas8a protein, Cas6 protein, Csa5 protein and Cas7 protein are as defined in any one of Sections I-IV above.
[0268] In certain embodiments, the nucleotide sequence encoding one or more guide RNAs in the Type IA CRISPR-Cas system and the nucleotide sequence encoding the cas protein in the Type IA CRISPR-Cas system are located in different expression cassettes.
[0269] In certain embodiments, the nucleotide sequences encoding the cas proteins are all located on the same vector.
[0270] In certain embodiments, the nucleotide sequence encoding the cas protein and the nucleotide sequence encoding one or more guide RNAs are both located on the same vector.
[0271] XIV. Host Cells
[0272] On the other hand, the present application provides a host cell comprising the vector system as described in any one of Sections VII-IX and XII-XIII above.
[0273] Such host cells include, but are not limited to, prokaryotic cells such as bacterial cells (e.g., E. coli cells), and eukaryotic cells such as fungal cells (e.g., yeast cells), insect cells, plant cells, and animal cells (e.g., mammalian cells, e.g., mouse cells, human cells, etc.).
[0274] In certain embodiments, the cell or its progeny is incapable of developing into a whole animal or plant.
[0275] XV. Application
[0276] On the other hand, the present application provides a kit comprising a system as described in any one of Parts I-III above, a protein as described in Part IV above, an isolated nucleic acid molecule as described in Part V or Part VI above, a vector as described in Part VI above, a host cell as described in Part VI above, a vector system as described in any one of Parts VII-IX above, a system as described in any one of Parts X-XI above, a vector system as described in any one of Parts XII-XIII above, or a host cell as described in Part XIV above; and instructions for using the system for nucleic acid editing (e.g., gene or genome editing, gene or genome large fragment deletion, gene or genome base modification, genomic structural variation).
[0277] In certain embodiments, the kit comprises a system as described in any of Sections II-III above.
[0278] In certain embodiments, the kit comprises a vector system as described in any one of Sections VIII-IX above.
[0279] In certain embodiments, the kit comprises a system as described in Section XI above.
[0280] In certain embodiments, the kit comprises a vector system as described in Section XIII above.
[0281] On the other hand, the present application also provides a delivery composition comprising a system as described in any one of Sections I-III above, a carrier system as described in any one of Sections VII-IX above, a system as described in any one of Sections X-XI above, or a carrier system as described in any one of Sections XII-XIII above, and a delivery system.
[0282] In certain embodiments, the delivery system is selected from a particle, a vesicle, or a viral vector.
[0283] In certain embodiments, the particle comprises a lipid, a sugar, a metal, or a protein.
[0284] In certain embodiments, the vesicle comprises an exosome or a liposome.
[0285] In certain embodiments, the viral vector comprises an adenovirus, a lentivirus, or an adeno-associated virus.
[0286] In certain embodiments, the delivery composition comprises a system as described in any of Sections II-III above.
[0287] In certain embodiments, the delivery composition comprises a carrier system as described in any of Sections VIII-IX above.
[0288] In certain embodiments, the delivery composition comprises a system as described in Section XI above.
[0289] In certain embodiments, the delivery composition comprises a carrier system as described in Section XIII above.
[0290] On the other hand, the present application provides a method for inducing a deletion in a target genome, wherein the target genome comprises a complementary first nucleic acid chain and a second nucleic acid chain, the method comprising: contacting the system as described in any one of Parts II-III above or the vector system as described in any one of Parts VIII-IX above, the system as described in Part XI above, or the vector system as described in Part XIII above with the target genome, or delivering it to a cell comprising the target genome.
[0291] In certain embodiments, the one or more cas proteins contained in the system or vector system are capable of forming a complex with a guide RNA and, upon binding of the complex to a target sequence and / or its complementary sequence, inducing deletion of a region comprising the target sequence and / or its complementary sequence.
[0292] In certain embodiments, the method comprises contacting the system described in Section III above, or the vector system described in Section IX above, the system described in Section XI above, or the vector system described in Section XIII above with the target genome, or delivering it into a cell comprising the target genome.
[0293] In certain embodiments, the deletion is a large deletion, e.g., greater than 0.1 kb, greater than 0.2 kb, greater than 0.5 kb, greater than 1 kb, greater than 1.5 kb, greater than 2 kb, greater than 10 kb, greater than 50 kb, greater than 100 kb, e.g., less than 500 kb, less than 400 kb, less than 300 kb, less than 200 kb.
[0294] In certain embodiments, the one or more guide RNAs comprised by the system or vector system comprise a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, and a second guide sequence capable of hybridizing to a second target sequence; wherein the first target sequence and the second target sequence are respectively located on either side of the region to be deleted in the target genome.
[0295] In certain embodiments, the first target sequence is located on the first nucleic acid strand of the target genome, and the second target sequence is located on the second nucleic acid strand of the target genome; for example, in the first nucleic acid strand, the first target sequence is located on the 5' end of the region to be deleted, and, in the second nucleic acid strand, the second target sequence is located on the 5' end of the region to be deleted.
[0296] In certain embodiments, the length of the region to be deleted is greater than 0.1 kb, for example, greater than 0.2 kb, greater than 0.3 kb, greater than 0.4 kb, greater than 0.5 kb; for example, the length of the region to be deleted is less than 500 kb, for example, less than 400 kb, less than 300 kb, less than 200 kb; for example, the length of the region to be deleted is 0.2 kb-200 kb (for example, 0.2 kb-2 kb, 0.2 kb-5 kb, 0.2 kb-10 kb, 0.2 kb-100 kb, 0.2 kb-200 kb; for example, 0.5 kb-1.5 kb, 0.5 kb-2 kb, 0.5 kb-10 kb).
[0297] In certain embodiments, the target genome is present within a cell, or alternatively, the target genome is present in a nucleic acid molecule (eg, a plasmid) in vitro.
[0298] In certain embodiments, the cell is a prokaryotic cell.
[0299] In certain embodiments, the cell is a eukaryotic cell.
[0300] In certain embodiments, the cell is selected from an animal cell (eg, a mammalian cell, such as a human cell), a plant cell (eg, a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast).
[0301] In certain embodiments, the methods are used for chromosome elimination.
[0302] On the other hand, the present application provides a method for inducing structural variation in a genome, wherein the genome comprises a complementary first nucleic acid chain and a second nucleic acid chain, and the method comprises: contacting the system as described in any one of Parts II-III above or the vector system as described in any one of Parts VIII-IX above, the system as described in Part XI above, or the vector system as described in Part XIII above with a target genome, or delivering it to a cell comprising the target genome.
[0303] In certain embodiments, one or more cas proteins contained in the system or vector system are capable of forming a complex with a guide RNA, and after the complex binds to the target sequence and / or its complementary sequence, it induces the deletion of the region containing the target sequence and / or its complementary sequence, thereby inducing genomic structural variation.
[0304] In certain embodiments, the genome comprises a complementary first nucleic acid chain and a second nucleic acid chain, and the method comprises: contacting the system as described in Part III above or the vector system as described in Part IX above, the system as described in Part XI above, or the vector system as described in Part XIII above with the target genome, or delivering it to a cell comprising the target genome.
[0305] In certain embodiments, the deletion is a large deletion, e.g., greater than 0.1 kb, greater than 0.2 kb, greater than 0.5 kb, greater than 1 kb, greater than 1.5 kb, greater than 2 kb, greater than 10 kb, greater than 50 kb, greater than 100 kb, e.g., less than 500 kb, less than 400 kb, less than 300 kb, less than 200 kb.
[0306] In certain embodiments, the one or more guide RNAs comprised by the system or vector system comprise a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, and a second guide sequence capable of hybridizing to a second target sequence; wherein the first target sequence and the second target sequence are respectively located on either side of the region to be deleted in the target genome.
[0307] In certain embodiments, the first target sequence is located on the first nucleic acid strand of the target genome, and the second target sequence is located on the second nucleic acid strand of the target genome; for example, in the first nucleic acid strand, the first target sequence is located on the 5' end of the region to be deleted, and, in the second nucleic acid strand, the second target sequence is located on the 5' end of the region to be deleted.
[0308] In certain embodiments, the length of the region to be deleted is greater than 0.1 kb, for example, greater than 0.2 kb, greater than 0.3 kb, greater than 0.4 kb, greater than 0.5 kb; for example, the length of the region to be deleted is less than 500 kb, for example, less than 400 kb, less than 300 kb, less than 200 kb; for example, the length of the region to be deleted is 0.2 kb-200 kb (for example, 0.2 kb-2 kb, 0.2 kb-5 kb, 0.2 kb-10 kb, 0.2 kb-100 kb, 0.2 kb-200 kb; for example, 0.5 kb-1.5 kb, 0.5 kb-2 kb, 0.5 kb-10 kb).
[0309] In certain embodiments, the target genome is present within a cell, or alternatively, the target genome is present in a nucleic acid molecule (eg, a plasmid) in vitro.
[0310] In certain embodiments, the cell is a prokaryotic cell.
[0311] In certain embodiments, the cell is a eukaryotic cell.
[0312] In certain embodiments, the cell is selected from an animal cell (eg, a mammalian cell, such as a human cell), a plant cell (eg, a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast).
[0313] On the other hand, the present application provides a method for modifying a target nucleic acid molecule, comprising: contacting the system as described in any one of Sections II-III above, the vector system as described in any one of Sections VIII-IX above, the system as described in Section XI above, or the vector system as described in Section XIII above with the target nucleic acid molecule, or delivering it to a cell containing the target nucleic acid molecule.
[0314] In certain embodiments, the one or more cas proteins contained in the system or vector system are capable of forming a complex with a guide RNA and, after the complex binds to the target sequence and / or its complementary sequence, inducing modification of the target nucleic acid molecule comprising the target sequence and / or its complementary sequence.
[0315] In certain embodiments, the target nucleic acid molecule is RNA or DNA.
[0316] In certain embodiments, the target nucleic acid molecule is double-stranded DNA.
[0317] In certain embodiments, the target nucleic acid molecule is a gene or genome.
[0318] In certain embodiments, the target nucleic acid molecule is present within a cell, or alternatively, the target nucleic acid molecule is present in a nucleic acid molecule (eg, a plasmid) in vitro.
[0319] In certain embodiments, the cell is a prokaryotic cell.
[0320] In certain embodiments, the cell is a eukaryotic cell.
[0321] In certain embodiments, the cell is selected from an animal cell (eg, a mammalian cell, such as a human cell), a plant cell (eg, a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast).
[0322] In certain embodiments, the modification is a large deletion of the target nucleic acid molecule.
[0323] In certain embodiments, the modification refers to a break in the target nucleic acid molecule, such as a double-strand break in DNA; for example, the modification further includes inserting an exogenous nucleic acid into the break.
[0324] In certain embodiments, the modification refers to a change in a base (eg, cytosine, adenine) in the target nucleic acid molecule.
[0325] On the other hand, the present application provides a method for inducing base mutations in a target nucleic acid molecule, comprising: contacting the system described in any one of Parts II-III above, the vector system described in any one of Parts VIII-IX above, the system described in Part XI above, or the vector system described in Part XIII above with the target nucleic acid molecule, or delivering it to a cell containing the target nucleic acid molecule.
[0326] In certain embodiments, the system or vector system does not comprise a Cas3 protein or a nucleotide sequence encoding a Cas3 protein.
[0327] In certain embodiments, the one or more cas proteins contained in the system or vector system are capable of forming a complex with a guide RNA, and after the complex binds to the target sequence and / or its complementary sequence, it induces modification of the bases in the target nucleic acid molecule comprising the target sequence and / or its complementary sequence, and generates base mutations during nucleic acid repair or replication;
[0328] In certain embodiments, the modification of the base refers to a modification that changes the base pairing pattern of the base to be modified. In certain embodiments, before modification, the base to be modified is complementary to a first base, and after modification, the modified base is complementary to a second base.
[0329] In certain embodiments, the one or more cas proteins comprised in the system or vector system further comprise an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
[0330] In certain embodiments, the one or more cas proteins (e.g., cas8a protein) contained in the system or vector system further comprises an adenosine deaminase (e.g., TadA8e), and the base to be modified is adenine. Before modification, adenine is complementary to thymine, and after modification, adenine is modified to hypoxanthine, and hypoxanthine is complementary to cytosine.
[0331] In certain embodiments, the one or more cas proteins (e.g., cas8a protein) contained in the system or vector system further comprises a cytosine deaminase (e.g., APOBEC3), and the base to be modified is cytosine. Before modification, cytosine is complementary to guanine. After modification, cytosine is modified to uracil, and uracil is complementary to thymine.
[0332] In certain embodiments, the target nucleic acid molecule is RNA or DNA.
[0333] In certain embodiments, the target nucleic acid molecule is double-stranded DNA.
[0334] In certain embodiments, the target nucleic acid molecule is a gene or genome.
[0335] In certain embodiments, the target nucleic acid molecule is present within a cell, or alternatively, the target nucleic acid molecule is present in a nucleic acid molecule (eg, a plasmid) in vitro.
[0336] In certain embodiments, the cell is a prokaryotic cell.
[0337] In certain embodiments, the cell is a eukaryotic cell.
[0338] In certain embodiments, the cell is selected from an animal cell (eg, a mammalian cell, such as a human cell), a plant cell (eg, a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast).
[0339] On the other hand, the present application provides a method for altering the expression of a gene product, comprising: contacting the system described in any one of Sections II-III above, the vector system described in any one of Sections VIII-IX above, the system described in Section XI above, or the vector system described in Section XIII above with a target nucleic acid molecule encoding the gene product, or delivering it to a cell containing the target nucleic acid molecule.
[0340] In certain embodiments, the one or more cas proteins contained in the system or vector system are capable of forming a complex with a guide RNA, and after the complex binds to the target sequence and / or its complementary sequence, it induces modification of the target nucleic acid molecule comprising the target sequence and / or its complementary sequence, thereby changing the expression of the gene product.
[0341] In certain embodiments, the target nucleic acid molecule is present within a cell, or the target nucleic acid molecule is present in a nucleic acid molecule (eg, a plasmid) in vitro.
[0342] In certain embodiments, the cell is a prokaryotic cell.
[0343] In certain embodiments, the cell is a eukaryotic cell.
[0344] In certain embodiments, the cell is selected from an animal cell (eg, a mammalian cell, such as a human cell), a plant cell (eg, a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast).
[0345] In certain embodiments, expression of the gene product is altered (eg, increased or decreased).
[0346] In certain embodiments, the gene product is a protein.
[0347] On the other hand, the present application provides a method for producing a plant with a modified trait, the method comprising contacting a plant cell with the system described in any one of Sections II-III above, the vector system described in any one of Sections VIII-IX above, the system described in Section XI above, or the vector system described in Section XIII above, or subjecting the plant cell to the method described in any one of the above, thereby modifying or editing the target gene or target nucleic acid molecule in the genome of the plant cell, and regenerating a plant from the plant cell.
[0348] In certain embodiments, the method comprises contacting a plant cell with a system as described in Section III above, or a vector system as described in Section IX above, a system as described in Section XI above, or a vector system as described in Section XIII above.
[0349] In certain embodiments, the plant is an agricultural plant, such as corn, barley, cotton, rice, soybean, wheat, or rice.
[0350] In certain embodiments, in the method described in any one of the above items, the cas protein or the nucleotide sequence encoding the cas protein, the guide RNA or the nucleotide sequence encoding the guide RNA contained in the system or vector system is present in the delivery system.
[0351] In certain embodiments, the delivery system is selected from a particle, a vesicle, or a viral vector.
[0352] In certain embodiments, the particle comprises a lipid, a sugar, a metal, or a protein.
[0353] In certain embodiments, the vesicle comprises an exosome or a liposome.
[0354] In certain embodiments, the viral vector comprises an adenovirus, a lentivirus, or an adeno-associated virus.
[0355] On the other hand, the present application provides a system as described in any of Sections I-III above, a protein as described in Section IV above, an isolated nucleic acid molecule as described in Section V or Section VI above, a vector as described in Section VI above, a host cell as described in Section VI above, a vector system as described in any of Sections VII-IX above, a system as described in any of Sections X-XI above, a vector system as described in any of Sections XII-XIII above, a host cell as described in Section XIV above, a kit as described in Section XV above, or a delivery composition as described in Section XV above, for use in nucleic acid editing, or for use in preparing a preparation for nucleic acid editing.
[0356] In certain embodiments, the nucleic acid editing comprises gene or genome editing.
[0357] In certain embodiments, the gene or genome editing includes deletion of large nucleic acid fragments, modification of genes, knockout of genes, alteration of gene product expression, repair of mutations, and / or insertion of polynucleotides or base mutations.
[0358] In certain embodiments, the nucleic acid editing comprises inducing genomic structural variation or chromosome elimination.
[0359] On the other hand, the present application provides a system as described in any of Sections I-III above, a protein as described in Section IV above, an isolated nucleic acid molecule as described in Section V or Section VI above, a vector as described in Section VI above, a host cell as described in Section VI above, a vector system as described in any of Sections VII-IX above, a system as described in any of Sections X-XI above, a vector system as described in any of Sections XII-XIII above, a host cell as described in Section XIV above, a kit as described in Section XV above, or a delivery composition as described in Section XV above, for use in preparing a preparation for editing a target nucleotide sequence in a target locus to modify an organism or a non-human organism (e.g., a plant).
[0360] In another aspect, the present application provides a cell or progeny thereof obtained by any of the methods described above, wherein the cell comprises a modification that is not present in its wild type.
[0361] In certain embodiments, the cell or its progeny is incapable of developing into a whole animal or plant.
[0362] In another aspect, the present application also provides a cell product of the cell or its progeny as described above.
[0363] Definition of terms
[0364] Unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the virology, biochemistry, and immunology laboratory procedures used herein are conventional procedures widely used in the respective fields. To facilitate a better understanding of the present invention, definitions and explanations of relevant terms are provided below.
[0365] When the terms "for example," "such as," "including," "including," "comprising," or variations thereof are used herein, these terms will not be considered as limiting terms, but will be interpreted to mean "but not limited to" or "not limited to."
[0366] The terms "a" and "an" and "the" and similar referents in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context.
[0367] As used herein, the term "Type IA CRISPR-CAS system" refers to a Class 1 CRISPR-CAS system comprising a multi-subunit crRNA-effector complex, more particularly to a Type I system, and even more particularly to a subtype IA system. The subtype IA system can include multiple different CAS components, for example, CAS components including Cas3, Cas5 (e.g., cas5a), Cas6, Csa5, Cas7, and Cas8 (e.g., Cas8a), and optionally other CAS components (see, for example, Makarova et al. 2020. Nature Reviews Microbiology 18(2):67-83. https: / / doi.org / 10.1038 / s41579-019-0299-x., Koonin, Makarova, and Zhang 2017. Current Opinion in Microbiology 37:67-78. https: / / doi.org / 10.1016 / j.mib.2017.05.008., Koonin and Makarova 2019. Russian Veterinary Journal 2019(2):29-36. http: / / dx.doi.org / 10.1098 / rstb.2018.0087, the entire text of which is incorporated herein by reference). In certain embodiments, the CAS protein used in this application is derived from or derived from a prokaryotic organism having a natural IA system. However, it should be understood that CAS proteins from any source (e.g., Cas3, Cas5 (e.g., Cas5a), Cas7, Cas6, Cas8 (e.g., Cas8a), Csa5) or derivatives thereof can be used. In certain embodiments, the different CAS components used in this application can be derived from or derived from the same organism or different organisms.
[0368] In certain embodiments, the amino acid sequence of the Cas3 protein can be found in SEQ ID NOs: 1, 7, 13, and 19. However, those skilled in the art understand that mutations or variations (including but not limited to substitutions, deletions, and / or additions, such as Cas3 proteins in IA CRISPR-CAS systems from different sources) may be naturally or artificially introduced into the amino acid sequence of the Cas3 protein without affecting its biological function. Therefore, in the present invention, the term "Cas3 protein" should include all such sequences, including, for example, the sequences shown in SEQ ID NOs: 1, 7, 13, and 19, as well as natural or artificial variants thereof.
[0369] In certain embodiments, the amino acid sequence of the cas5a protein can be found in SEQ ID NOs: 2, 8, 14, and 20. However, it will be appreciated by those skilled in the art that mutations or variations (including but not limited to substitutions, deletions, and / or additions, such as cas5a proteins in IA CRISPR-CAS systems from different sources) may be naturally or artificially introduced into the amino acid sequence of the cas5a protein without affecting its biological function. Therefore, in the present invention, the term "cas5a protein" shall include all such sequences, including, for example, sequences shown in SEQ ID NOs: 2, 8, 14, and 20, as well as natural or artificial variants thereof.
[0370] In certain embodiments, the amino acid sequence of the cas8a protein can be found in SEQ ID NO: 3, 9, 15, 21. However, it is understood by those skilled in the art that mutations or variations (including but not limited to, substitutions, deletions and / or additions, such as cas8a proteins in IA CRISPR-CAS systems from different sources) may be naturally or artificially introduced into the amino acid sequence of the cas8a protein without affecting its biological function. Therefore, in the present invention, the term "cas8a protein" shall include all such sequences, including, for example, sequences shown in SEQ ID NO: 3, 9, 15, 21 and natural or artificial variants thereof.
[0371] In certain embodiments, the amino acid sequence of the cas7 protein can be found in SEQ ID NO: 4, 10, 16, 22. However, it is understood by those skilled in the art that mutations or variations (including but not limited to, substitutions, deletions and / or additions, such as cas7 proteins in IA CRISPR-CAS systems from different sources) may be naturally or artificially introduced into the amino acid sequence of the cas7 protein without affecting its biological function. Therefore, in the present invention, the term "cas7 protein" shall include all such sequences, including, for example, the sequences shown in SEQ ID NO: 4, 10, 16, 22 and their natural or artificial variants.
[0372] In certain embodiments, the amino acid sequence of the cas6 protein can be found in SEQ ID NO: 5, 11, 17, 23. However, it is understood by those skilled in the art that mutations or variations (including but not limited to, substitutions, deletions and / or additions, such as cas6 proteins in IA CRISPR-CAS systems from different sources) may be naturally or artificially introduced into the amino acid sequence of the cas6 protein without affecting its biological function. Therefore, in the present invention, the term "cas6 protein" shall include all such sequences, including, for example, sequences shown in SEQ ID NO: 5, 11, 17, 23 and natural or artificial variants thereof.
[0373] In certain embodiments, the amino acid sequence of the csa5 (cas11) protein can be found in SEQ ID NOs: 6, 12, 18, and 24. However, those skilled in the art will appreciate that mutations or variations (including but not limited to substitutions, deletions, and / or additions, such as csa5 proteins in IA CRISPR-Cas systems from different sources) may occur naturally or be artificially introduced into the amino acid sequence of the csa5 protein without affecting its biological function. Therefore, in the present invention, the term "csa5 protein" should include all such sequences, including, for example, the sequences shown in SEQ ID NOs: 6, 12, 18, and 24, as well as natural or artificial variants thereof.
[0374] In addition, the sequences shown in SEQ ID NO: 1-24 of the present application do not include amino acids (such as methionine (Met)) encoded by start codons (such as ATG) at their N-termini. It is understood by those skilled in the art that in the process of preparing proteins by genetic engineering, due to the effect of the start codon, the first position of the produced polypeptide chain is often an amino acid (such as Met) encoded by the start codon. The cas protein of the present invention not only encompasses amino acid sequences that do not include amino acids (such as Met) encoded by start codons at its N-termini, but also encompasses amino acid sequences that include amino acids (such as Met) encoded by start codons at its N-termini. Therefore, sequences that further include amino acids (such as Met) encoded by start codons at the N-termini of the above-mentioned amino acid sequences are also within the scope of protection of the present invention.
[0375] As used herein, the terms "guide RNA (guide RNA)", "mature crRNA" are used interchangeably and have the meanings generally understood by those skilled in the art. In general, a guide RNA may comprise a direct repeat sequence and a guide sequence (guide sequence), or may consist essentially of or consist of a direct repeat sequence and a guide sequence (also referred to as a spacer in the context of an endogenous CRISPR system). In some cases, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target sequence to hybridize with the target sequence and guide the specific binding of a CRISPR / Cas complex to the target sequence or its complementary sequence. In certain embodiments, when optimally aligned, the degree of complementarity between a guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. It is within the capabilities of those of ordinary skill in the art to determine optimal alignment. For example, there are publicly and commercially available alignment algorithms and programs such as, but not limited to, ClustalW, Smith-Waterman in matlab, Bowtie, Geneious, Biopython, and SeqMan.
[0376] In some cases, the guide sequence is at least 5, at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 45, or at least 50 nucleotides in length. In some cases, the guide sequence is no more than 50, 45, 40, 35, 30, 25, 24, 23, 22, 21, 20, 15, 10, or fewer nucleotides in length. In certain embodiments, the guide sequence is 10-50, or 15-40, or 20-40 nucleotides in length.
[0377] In some cases, the direct repeat sequence is at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, or at least 70 nucleotides in length. In some cases, the same direction repeat sequence is no more than 70,65,64,63,62,61,60,59,58,57,56,55,50,45,40,35,30,29,28,27,26,25,24,23,22,21,20,15,10 or less nucleotides in length. In certain embodiments, the same direction repeat sequence is 55-70 nucleotides in length, such as 55-65 nucleotides, such as 60-65 nucleotides, such as 62-65 nucleotides, such as 63-64 nucleotides. In certain embodiments, the same direction repeat sequence is 15-40 nucleotides in length, such as 15-38 nucleotides, such as 20-40 nucleotides, such as 22-38 nucleotides, such as 32 nucleotides. In certain embodiments, the direct repeat sequence is no less than 30 nt in length, such as 30 nt-40 nt, such as 37 nt.
[0378] As used herein, the term "CRISPR / Cas complex" refers to a ribonucleoprotein complex formed by the binding of a guide RNA or mature crRNA to a Cas protein, comprising a guide sequence that can hybridize to a target sequence and bind to the Cas protein. The ribonucleoprotein complex is capable of recognizing and / or cleaving a polynucleotide and / or its complementary strand that can hybridize to the guide RNA or mature crRNA.
[0379] Therefore, in the case of forming a CRISPR / Cas complex, a "target sequence" refers to a polynucleotide targeted by a guide sequence designed to be targeted, such as a sequence complementary to the guide sequence, wherein the hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex. Complete complementarity is not required, as long as there is enough complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex. The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located in the nucleus or cytoplasm of the cell. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast.
[0380] In the present invention, the expression "target sequence" or "target polynucleotide" can be any endogenous or exogenous polynucleotide to a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, it is believed that the target sequence should be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for PAM vary depending on the Cas effector enzyme used, but PAM is typically a 2-5 base pair sequence adjacent to the protospacer sequence (i.e., target sequence). Those skilled in the art will be able to identify PAM sequences for use with a given Cas effector protein.
[0381] As used herein, the term "adenosine deaminase" refers to a protein that catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine to inosine in deoxyribonucleic acid (DNA). In some embodiments, the adenosine deaminase is TadA8e. In certain embodiments, the amino acid sequence of the adenosine deaminase can be found in NCBI Genbank ID: UNJ19119.1 or NCBI Genbank ID: QHD44350.1. However, it is understood by those skilled in the art that mutations or variations (including but not limited to substitutions, deletions and / or additions, such as adenosine deaminases from different sources) can be naturally or artificially introduced into the amino acid sequence of the adenosine deaminase without affecting its biological function. Therefore, in the present invention, the term "adenosine deaminase" should include all such sequences, including, for example, the sequences shown in NCBI Genbank ID: UNJ19119.1 or NCBI Genbank ID: QHD44350.1 and their natural or artificial variants.
[0382] As used herein, the term "cytosine deaminase" refers to a protein that catalyzes the hydrolytic deamination of cytidine or cytosine. In certain embodiments, the cytosine deaminase is APOBEC3. In certain embodiments, the amino acid sequence of the cytosine deaminase can be found in NCBI Genbank ID: 76096346 or NCBI Genbank ID: 176865758. However, it will be appreciated by those skilled in the art that mutations or variations (including but not limited to substitutions, deletions and / or additions, such as cytosine deaminases from different sources) may be naturally occurring or artificially introduced into the amino acid sequence of the cytosine deaminase without affecting its biological function. Therefore, in the present invention, the term "cytosine deaminase" shall include all such sequences, including, for example, the sequences shown in NCBI Genbank ID: 76096346 or NCBI Genbank ID: 176865758 and their natural or artificial variants.
[0383] As used herein, the term "identity" is used to refer to the matching of sequences between two polypeptides or between two nucleic acids. In order to determine the percent identity of two amino acid sequences or two nucleic acid sequences, the sequences are aligned for optimal comparison purposes (e.g., a gap can be introduced in the first amino acid sequence or nucleic acid sequence to optimally align with the second amino acid or nucleic acid sequence). The amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., percent identity = number of identical overlapping positions / total number of positions × 100%). In certain embodiments, the two sequences are the same length.
[0384] The determination of percent identity between two sequences can also be achieved using a mathematical algorithm. A non-limiting example of a mathematical algorithm for the comparison of two sequences is the algorithm of Karlin and Altschul, 1990, Proc. Natl. Acad. Sci. USA 87: 2264-2268, as modified in Karlin and Altschul, 1993, Proc. Natl. Acad. Sci. USA 90: 5873-5877. Such an algorithm is incorporated into the NBLAST and XBLAST programs of Altschul et al., 1990, J. Mol. Biol. 215: 403.
[0385] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which a polynucleotide can be inserted. When a vector is capable of expressing a protein encoded by the inserted polynucleotide, it is referred to as an expression vector. A vector can be introduced into a host cell via transformation, transduction, or transfection, allowing the genetic material elements it carries to be expressed in the host cell. Vectors are well known to those skilled in the art and include, but are not limited to, plasmids; phagemids; cosmids; artificial chromosomes, such as yeast artificial chromosomes (YACs), bacterial artificial chromosomes (BACs), or P1-derived artificial chromosomes (PACs); bacteriophages such as lambda phage or M13 phage; and animal viruses. Animal viruses that can be used as vectors include, but are not limited to, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpes viruses (such as herpes simplex virus), poxviruses, baculoviruses, papillomaviruses, and papillomas (such as SV40). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. In addition, a vector may also contain an initiation of replication site.
[0386] Advantageous Effects of the Invention
[0387] The IA CRISPR-Cas effector protein and system provided by the present invention have significant application value.
[0388] For example, the IA CRISPR-Cas system provided by the present invention can be used to achieve precise large-scale deletion of target genes or genomes (e.g., knockout of gene coding regions, knockout of long lncRNAs or enhancers, chromosome elimination) and / or other target nucleic acid editing (e.g., modifying genes, knocking out genes, changing the expression of gene products, repairing mutations, inserting polynucleotides, and / or single-base mutations, etc.).
[0389] For example, the IA CRISPR-Cas system provided by the present invention has pre-crRNA processing activity. Compared with the Cas9 system, it does not require tracrRNA and can be more easily applied to multi-target gene editing.
[0390] For example, the IA CRISPR-Cas system provided by the present invention recognizes a PAM motif having a structure represented by 5'CCN- (eg, 5'CCT- or 5'CCC-).
[0391] For example, the guide RNA provided by the present invention comprising two reverse target sites can achieve precise genomic fragment deletion compared to gene editing using a single target site.
[0392] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples, but it will be understood by those skilled in the art that the following drawings and examples are intended only to illustrate the present invention and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art based on the following detailed description of the accompanying drawings and preferred embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0393] FIG1 is a simplified flow chart of the experimental design for PAM identification in Example 2.
[0394] Figure 2 is a diagram of the carrier design in Example 3.
[0395] FIG3 shows the editing sites of Type 1-A-2 and Type 1-A-3 on the ROS1 gene in the maize genome in Example 4.
[0396] Figure 4 shows the detection results of maize endogenous gene editing activity in Example 4. Figure 4A shows the PCR detection results of the editing products of the type IA-2 system, Figure 4B shows the PCR detection results of the editing products of the type 1-A-3 system, Figure 4C shows the first-generation sequencing alignment results of the editing sites of the type IA-2 system editing products on the ROS1 gene, and Figure 4D shows the first-generation sequencing alignment results of the editing sites of the type IA-3 system editing products on the ROS1 gene.
[0397] Figure 5 shows the dual-targeted editing sites of Type 1-A-1, Type 1-A-2, and Type 1-A-3 on the ROS1 gene in the maize genome in Example 5.
[0398] Figure 6 shows the detection results of dual-targeted editing activity of maize endogenous genes in Example 5. Figure 6A shows the detection results of PCR detection of the editing products of the type IA-1 system, Figure 6B shows the detection results of PCR detection of the editing products of the type 1-A-2 system, Figure C shows the detection results of PCR detection of the editing products of the type 1-A-3 system, Figure 6D shows the first-generation sequencing alignment results of the editing sites of the type IA-1 system editing products on the ROS1 gene, Figure 6E shows the first-generation sequencing alignment results of the editing sites of the type IA-2 system editing products on the ROS1 gene, and Figure 6F shows the first-generation sequencing alignment results of the editing sites of the type IA-3 system editing products on the ROS1 gene.
[0399] Figure 7 is a design map of the adenine single-base editing vector (IA TadA8e) in implementation case 6.
[0400] Figure 8 shows the gene editing detection results of the type IA system in stable transgenic maize plants in Example 7. Figure 8A shows the dual-target design targeting the GA2 gene, with #g1 and #g2 as the two targets; Figure 8B shows the first-generation sequencing alignment of the editing sites of the type IA-2 system editing products on the GA2 gene in transgenic plants.
[0401] Figure 9 shows the gene editing detection results of the type IA system in the HEK293T fluorescent reporter cell line stably expressing Tdtomato in Example 8. Figure 9A is a schematic diagram of the animal cell expression vector; Figure 9B shows the target design for the Tdtomato red fluorescent gene, with G1 and G2 being two targets; Figure 9C shows the detection results of the editing efficiency of the type IA system and the CRISPR / Cas9 system using the red fluorescence system, where the ordinate represents the reduction ratio of the fluorescence value of each system, and on the abscissa, "Cas9" corresponds to the editing efficiency of the CRISPR / Cas9 system, "A3-CCT" corresponds to the editing efficiency of the type IA-3 system for the target G1 with a 5'-CCT sequence feature, and "A3-CCC" corresponds to the editing efficiency of the type IA-3 system for the target G2 with a 5'-CCC sequence feature.
[0402] Figure 10 shows the gene editing detection results of the Type IA-A system in the HEK293T cell line in Example 9. Figure 10A is a schematic diagram of the animal cell expression vector, Figure 10B shows the target design for the HPRT1 gene, with g1 and g2 being the two targets for dual targeting, and Figure 10C shows the first-generation sequencing alignment results of the editing products of the Type IA-2 system at the editing sites on the HPRT1 gene.
[0403] Sequence information
[0404] A description of the sequences involved in this application is provided in the table below.
[0405] Table 1: Sequence information DETAILED DESCRIPTION
[0406] The invention will now be described with reference to the following examples which are intended to illustrate the invention but not to limit it.
[0407] Unless otherwise indicated, the experiments and procedures described in the examples were performed essentially according to conventional methods well known in the art and described in various references. For example, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA used in the present invention can be found in Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd ed. (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (FM Ausubel et al., eds., (1987)); METHODS IN ENZYMOLOGY series (Academic Press): PCR 2: A PRACTICAL APPROACH (MJ MacPherson, BD Hames, and GR Taylor, eds. (1995)); and ANIMALCELL TECHNOLOGY (Animal Cell Culture, Inc.). Culture) (RI Freshney, ed. (1987)). In addition, if the specific conditions are not specified in the examples, the experiments were carried out according to conventional conditions or the conditions recommended by the manufacturer. The reagents or instruments used, if the manufacturers are not specified, are all conventional products that can be obtained commercially.
[0408] Those skilled in the art will appreciate that the embodiments describe the present invention by way of example and are not intended to limit the scope of the invention as claimed. All publications and other references mentioned herein are incorporated herein by reference in their entirety.
[0409] The formulas or sources of some of the reagents involved in the following examples are as follows:
[0410] LB liquid medium: 10 g tryptone, 5 g yeast extract, 10 g NaCl, dilute to 1 L, and sterilize.
[0411] CTAB solution: 16.7 g of CTAB (cetyltrimethylammonium bromide), 234 mL of 5 M NaCl, 83.5 mL of 1 M Tris-HCl (pH 8.0), and 33.4 mL of 0.5 M EDTA (pH 8.0). Add distilled water to a volume of 1 L and add β-mercaptoethanol in a ratio of 100:1 when using.
[0412] W5 solution: 154 mM NaCl, 125 mM CaCl2, 5 mM KCl, 4 mM MES, dilute to 500 mL, and adjust the pH to 5.7 with NaOH.
[0413] 20MMG solution: 0.4mM mannitol, 15mM MgCl2, 4mM MES, dilute to 10mL.
[0414] The large-scale plasmid extraction kit was purchased from QIAGEN, catalog number: 12963.
[0415] Blunt-smiple vector was purchased from Shanghai Yisheng Biotechnology Co., Ltd., catalog number: CB111-02.
[0416] Escherichia coli competent DH5α was purchased from Beijing Qingke Biotechnology Co., Ltd., product number: TSV-A07.
[0417] The prokaryotic expression vectors pACYC-Duet-1 and pUC19 were purchased from Beijing Quanshijin Biotechnology Co., Ltd.
[0418] Escherichia coli competent EC100 was purchased from Epicentre.
[0419] Unless otherwise specified, the sequence synthesis involved in the following examples was completed by Beijing Qingke Biotechnology Co., Ltd., and the sequencing involved was completed by Beijing Ruibo Xingke Biotechnology Co., Ltd., Sangon Bioengineering Co., Ltd. and Liuhe BGI.
[0420] Example 1: Acquisition of Type IA Gene and Type IA Guide RNA
[0421] 1. CRISPR and gene annotation: Prodigal was used to annotate the metagenomic and viral genome data from the JGI database to obtain all proteins. Piler-CR and Minced were used to annotate CRISPR loci. All parameters were default.
[0422] 2. Acquisition of CRISPR-associated proteins: Each CRISPR locus will be extended 10Kb upstream and downstream to identify non-redundant macromolecular proteins in the CRISPR adjacent region.
[0423] 3. Obtaining Type IA family Cas3 effector proteins: Since all Type IA family Cas3 effector proteins discovered so far are greater than 200 amino acids in length, to reduce computational complexity, we filtered the aforementioned CRISPR-associated proteins by protein length before mining. We collected Cas3 proteins from known Type IA families to build a library, and used the filtered CRISPR-associated proteins for Psi-blast comparisons, outputting comparison results with an Evalue < 1E-8.
[0424] 4. Cas3 effector protein domain annotation: Use the Pfam database to annotate the domains of the mapped Cas3 effector proteins. Screen for Type IA Cas3 proteins containing both HD and DEAD domains.
[0425] 5. Identification of Cas8a and Csa5 Marker Proteins in the Type IA Family: Cas8a and Csa5 proteins from known Type IA families were collected and separately constructed. The CRISPR locus containing the identified Cas3 protein was extended 10 kb upstream and downstream. All protein sequences within this range were collected and Psi-blasted against the proteins in the library, with an E-value < 1E-8 output. After protein redundancy was removed using cd-hit software, systems were screened for Cas3, Cas8a, and Csa5 proteins in the same CRISPR locus.
[0426] 6. Identification of Type IA family Cas5, Cas6, and Cas7 protein components: Cas5, Cas6, and Cas7 proteins from known Type IA families were collected and compiled into libraries. Psi-blast alignments were performed on all protein sequences from the same CRISPR locus in step 5 against the proteins in the libraries, with an E-value < 1E-8. Systems containing Cas3, Cas8a, Csa5, Cas5, Cas6, and Cas7 proteins coexisting on a single CRISPR locus were considered complete candidate Type IA systems.
[0427] On this basis, the inventors obtained a new Cas effector protein, namely Type IA, and its four active homolog sequences, respectively named Type IA-1 (the amino acid sequences of the proteins Cas3, Cas5a, Cas8a, Cas7, Cas6, and Csa5 contained therein are shown in SEQ ID NOs: 1-6, respectively), Type IA-2 (the amino acid sequences of the proteins Cas3, Cas5a, Cas8a, Cas7, Cas6, and Csa5 contained therein are shown in SEQ ID NOs: 7-12, respectively), Type IA-3 (the amino acid sequences of the proteins Cas3, Cas5a, Cas8a, Cas7, Cas6, and Csa5 contained therein are shown in SEQ ID NOs: 13-18, respectively), and Type IA-4 (the amino acid sequences of the proteins Cas3, Cas5a, Cas8a, Cas7, Cas6, and Csa5 contained therein are shown in SEQ ID NOs: 19-24, respectively). The encoding DNAs of the four homologs are shown in SEQ ID NOs: 25-48, respectively. The prototype direct repeat sequences corresponding to Type IA-1, Type IA-2, Type IA-3, and Type IA-4 are shown in SEQ ID NOs: 49, 53, 57, and 61, respectively.
[0428] Example 2. Identification of the PAM domain of Type IA gene
[0429] 1. Construction and sequencing of the recombinant plasmid pACYC-Duet-1+CRISPR / type IA. Taking Type IA-1 as an example, the structure of the recombinant plasmid pACYC-Duet-1+CRISPR / type IA is described as follows: the small fragment between the restriction endonuclease EcoN I and EC00109I recognition sequences of the vector pACYC-Duet-1 is replaced with the double-stranded DNA molecule shown in positions 1 to 7368 from the 5' end of the sequence shown in SEQ ID NO: 100. The recombinant plasmid pACYC-Duet-1+CRISPR / type IA expresses the type I-A-1 cas3, cas5a, cas8a, cas7, cas6, and csa5 proteins shown in SEQ ID NO: 1-6, as well as the type IA guide RNA targeting the PAM library sequence.
[0430] 2. The recombinant plasmid pACYC-Duet-1+CRISPR / type IA contains an expression cassette, the nucleotide sequence of which is shown in SEQ ID NO: 100. From the 5' end, positions 1 to 7252 contain the nucleotide sequence of the type IA-1 gene, and positions 7253 to 7338 contain the nucleotide sequence of the terminator (used to terminate transcription). From the 5' end, positions 7339 to 7373 contain the nucleotide sequence of the J23119 promoter, positions 7374 to 7410 and 7447 to 7483 contain the nucleotide sequence of the CRISPR array, and positions 7484 to 7511 contain the nucleotide sequence of the rrnB-T1 terminator (used to terminate transcription).
[0431] 3. Construction of the PAM library: The sequence shown in SEQ ID NO:104, including eight random bases at the 5' end and the target sequence, was artificially synthesized and ligated into the pUC19 vector. Eight random bases were designed in front of the 5' end of the target sequence in the PAM library to construct the plasmid library.
[0432] 4. Obtaining Recombinant E. coli: The recombinant plasmid pACYC-Duet-1 + CRISPR / type IA and the PAM library plasmid were co-introduced into E. coli EC100 (a partial schematic diagram of the recombinant plasmid pACYC-Duet-1 + CRISPR / type IA and the PAM library plasmid is shown in Figure 1). The cells were incubated at 37°C for 12-14 hours. The plasmids were extracted, and the PAM region sequence was amplified and sequenced by PCR.
[0433] 5. Obtaining PAM library domains: The number of occurrences of 65,536 combinations of PAM sequences in the experimental and control groups was counted and normalized by the total number of PAM sequences in each group to analyze the PAM sequences recognized by each type IA system.
[0434] Example 3. Experimental related vector design
[0435] 1. As described in Example 1, the present invention obtained the protein amino acid sequence information (SEQ ID NOs: 1-24) of Cas3, Cas5a, Cas8a, Cas7, Cas6, and Csa5 of type-IA-1, type-IA-2, type-IA-3, and type-IA-4 respectively through metagenome and phage database mining, and performed codon optimization for eukaryotic corn. The optimized protein coding sequences are shown in SEQ ID NOs: 25-48.
[0436] 2. Using the corn UBI promoter and T2A cleavage peptide, a monocistronic expression vector for expressing Cas8a, Cas7, Cas5a, and Cas6 was designed. The Cas3 protein and Csa5 in the vector were expressed using the CMV35S promoter, and the two were connected with a T2A cleavage peptide. A nuclear localization signal was added to the N-terminus of each protein (the amino acid sequence of the nuclear localization signal is shown in SEQ ID NO: 65). The guide RNA was expressed through the OsU3 promoter, and the above proteins and RNA components were constructed into the P3301 vector (purchased from Youbao Bio, catalog number: VT1386) for subsequent experimental detection. The vector structure is shown in Figure 2.
[0437] Example 4. Detection of endogenous gene editing activity in corn
[0438] 1. To detect the editing activity of the I-A system in endogenous genes in eukaryotes, we selected maize ROS1 as the target gene, as shown in Figure 3. The target site for the detection site of the ROS1 gene is designed as shown in Figure 3 g.
[0439] 2. Following the target site design method in step 1, we selected a 37-nt DNA sequence with a 5'-CCT signature as the target site and used it as a spacer sequence to construct U3-RNA vectors. Each constructed U3-RNA vector was then ligated into the p3301 vector.
[0440] 3. Protoplast Extraction
[0441] The middle part of the leaf was selected to separate the protoplasts, which were cut into strips of about 0.5 mm in width with a sharp blade. 20 to 30 strips could be cut together; the strips were transferred to a prepared enzymolysis solution, protected from light, and vacuumed at -15 to -20 inHg for 30 minutes; the enzymolysis was then performed in the dark for 5 to 6 hours while slowly shaking (decolorization shaker, speed 10 rpm); after the enzymolysis was completed, an equal amount of W5 solution was added and the protoplasts were shaken horizontally for 10 seconds with a slight force by hand to release the protoplasts; the protoplasts were filtered through a 40 μm nylon membrane into a 50 mL round-bottom centrifuge tube, and the protoplasts were precipitated by horizontal centrifugation at 100 g for 3 minutes, and the supernatant was removed; the protoplasts were resuspended in W5 and ice-bathed for 30 minutes to allow the protoplasts to settle naturally, and the supernatant was discarded as much as possible; an appropriate amount of MMG solution was added to resuspend the protoplasts to a protoplast concentration of 2 × 10 6 / ml, counted on a hemacytometer.
[0442] 4. The vectors constructed in step 2 were used to transform protoplasts, and after culturing at 28°C for 48 hours, maize genomic DNA was extracted. Primers were designed to amplify the region of about 2 kb upstream and downstream of the target site. The amplified products were ligated to the Blunt-simple vector. 96 recombinant clones were randomly selected and tested for colony PCR using M13F / M13R primers (sequences shown in SEQ ID NOs: 114 and 115), and gel electrophoresis analysis was performed. The electrophoresis results are shown in Figures 4A and 4B. The results show that the type I-A-2 and type I-A-3 system editing products have PCR bands at the ROS1 gene locus that are smaller than the wild-type genome amplification length (4.3 kb) (as shown in the lanes marked by arrows in Figures 4A and 4B). The PCR products marked by arrows were sequenced using M13F, and the sequence results were aligned with the B73 reference genome. The results showed that the sequences marked by red arrows contained large fragment deletions, and the deleted fragments are shown in Figures 4C and 4D.
[0443] Example 5. Detection of dual-targeted endogenous gene editing activity in maize
[0444] 1. To test the editing activity of the I-A system on endogenous eukaryotic genes, we selected maize ROS1 as the target gene, as shown in Figure 5. To improve the accuracy of gene deletion length, we designed two opposing dual-target sites for the detection site. The dual-target sites for the detection site of the ROS1 gene are shown as g1 and g2 in Figure 5.
[0445] 2. Following the target site design method in Step 1, we selected a 37-nt DNA sequence with a 5'-CCT signature as the targeting site and two 37-nt DNA sequences approximately 1 kb apart in opposite directions as spacer sequences for the construction of U3-RNA vectors. Each constructed U3-RNA vector was then ligated into the p3301 vector.
[0446] 3. Protoplast Extraction
[0447] The middle part of the leaf was selected to separate the protoplasts, which were cut into strips of about 0.5 mm in width with a sharp blade. 20 to 30 strips could be cut together; the strips were transferred to a prepared enzymolysis solution, protected from light, and vacuumed at -15 to -20 inHg for 30 minutes; the enzymolysis was then performed in the dark for 5 to 6 hours while slowly shaking (decolorization shaker, speed 10 rpm); after the enzymolysis was completed, an equal amount of W5 solution was added and the protoplasts were shaken horizontally for 10 seconds with a slight force by hand to release the protoplasts; the protoplasts were filtered through a 40 μm nylon membrane into a 50 mL round-bottom centrifuge tube, and the protoplasts were precipitated by horizontal centrifugation at 100 g for 3 minutes, and the supernatant was removed; the protoplasts were resuspended in W5 and ice-bathed for 30 minutes to allow the protoplasts to settle naturally, and the supernatant was discarded as much as possible; an appropriate amount of MMG solution was added to resuspend the protoplasts to a protoplast concentration of 2 × 10 6 / ml, counted on a hemacytometer.
[0448] 4. The vectors constructed in step 2 were transformed into protoplasts respectively, and the maize genomic DNA was extracted after culturing at 28°C for 48 hours. Primers were designed to amplify the region of about 1 kb upstream and downstream of the target site. The amplified products were ligated to the Blunt-simple vector. 96 recombinant clones were randomly selected and subjected to colony PCR detection using M13F / M13R primers (sequences as shown in SEQ ID NO: 114 and 115), and gel electrophoresis analysis was performed. The electrophoresis results are shown in Figures 6A, 6B, and 6C. The results show that the type I-A-1, type I-A-2, and type I-A-3 system editing products have PCR bands at the ROS1 gene site that are smaller than the wild-type genome amplification length (2 Kb) (as shown in the lanes marked with arrows in Figures 6A, 6B, and 6C). The PCR products marked by arrows were sequenced using M13F, and the sequence of the first-generation sequencing results was compared with the B73 reference genome. The results showed that the sequences marked by red arrows contained large deletions, and the deleted fragments were mainly between the two target sites (as shown in Figures 6D, 6E, and 6F).
[0449] Example 6. Type I-A system for adenine base editing
[0450] 1. Design of adenine single-base editing vector (IA TadA8e)
[0451] Using the maize UBI promoter and T2A cleavage peptide, monocistronic expression vectors for expressing Cas7, Cas5a, Cas6, and Cas5 were designed. The TadA8e-Cas8a fusion protein was expressed using the CMV35S promoter, and a nuclear localization signal was added to the N-terminus of each protein (the amino acid sequence of the nuclear localization signal is shown in SEQ ID NO: 65). Guide RNA was expressed via the OsU3 promoter, and the above protein and RNA components were constructed into the P3301 vector (purchased from Youbao Bio, catalog number: VT1386) for subsequent experiments. The vector design map is shown in Figure 7.
[0452] 2. The DNA sequence containing type IA recognition PAM (CCT) on the corn genome was selected as the target sequence to construct an adenine single-base editing vector (IA TadA8e).
[0453] 3. The constructed vector was used to transform corn protoplasts, and the transformed DNA was extracted. PCR amplification was performed upstream and downstream of the target site. The DNA product after PCR amplification was connected to the B vector and sequenced. The sequencing results were used to determine whether there was an A to G base substitution near the target sequence.
[0454] Example 7. Detection of Gene Editing in Stable Transgenic Maize Plants Using the Type Ⅰ-A System
[0455] 1. To test the editing efficiency of the Type IA system in stable transgenic maize plants, we selected the maize GA2 gene (GRMZM2G368411) as the target gene and designed two opposing target sites on the gene, as shown in Figure 8A. For example, for the two detection sites of the GA2 gene, the dual target sites were designed as #g1 and #g2 in Figure 8A.
[0456] 2. Following the target site design method in Step 1, we selected a 37-nt DNA sequence with a 5'-CCT signature as the targeting site and two 37-nt DNA sequences approximately 1 kb apart in opposite directions as spacer sequences for the construction of U3-RNA vectors. Each constructed U3-RNA vector was then ligated into the p3301 vector.
[0457] 3. The vector constructed in step 2 was used for Agrobacterium transformation and callus regeneration. DNA from leaves of T0 transgenic plants was extracted and amplified by PCR. The detection method was the same as that of step 4 in Example 4. Genome-specific primers were designed 2 kb upstream and downstream of the target site for PCR amplification. The amplified PCR product was ligated to the Blunt-simple vector. Twenty-four recombinant clones were randomly selected for colony PCR detection and first-generation sequencing using the M13F / M13R primer pair. The first-generation sequencing results were compared with the genomic sequence of the reference gene. The alignment results are shown in Figure 8B (GA2 gene).
[0458] 4. Based on the first-generation sequencing results of step 3, each transgenic event containing one or more deletion clones was considered a gene-editing-positive plant. The proportion of gene-editing-positive plants in the Type Ⅰ-A-2 system on the GA2 gene was statistically analyzed to be 66.67%.
[0459] Example 8. Gene Editing Detection of the Type I-A System in HEK293T Fluorescent Reporter Cell Lines Stably Expressing Tdtomato
[0460] 1. To test the gene editing efficiency of the Type IA system in a HEK293T fluorescent reporter cell line stably expressing Tdtomato, we selected the Tdtomato gene as the target gene and designed two target sites on this gene with 5'-CCT and 5'-CCC sequence characteristics, respectively, as shown in Figure 9A, #G1 and #G2.
[0461] 2. Following the target site design method in Step 1, we selected DNA sequences (37 nt) with 5'-CCT and 5'-CCC features as targeting sites and used them as spacer sequences to construct U6-RNA vectors. Each constructed U3-RNA vector was then ligated into the PX458 vector. The vector construction is shown in Figure 9B.
[0462] 3. The vector constructed in Step 2 was transfected into a HEK293T fluorescent reporter cell line stably expressing Tdtomato. After 5 days of culture, flow cytometric analysis was performed. The mean fluorescence intensity (MFI) changes between the experimental and control groups were compared by flow cytometry. The editing efficiency was measured by the reduction in MFI. The data were processed and plotted using Graphpad Prism 8.0. The editing efficiency of the Type I-A-3 system in the HEK293T fluorescent reporter cell line stably expressing Tdtomato was statistically analyzed. The results are shown in Figure 9C.
[0463] Example 9. Gene Editing Detection of Type I-A System in HEK293T Cell Line
[0464] 1. To test the editing efficiency of the Type IA system in HEK293T cells, we selected the human genome HPRT1 gene as the target gene and designed two opposite target sites on this gene. The design of the dual target sites is shown in Figure 10A, #g1 and #g2.
[0465] 2. Following the target site design method in Step 1, we selected a 37-nt DNA sequence with a 5'-CCT signature as the targeting site and two 37-nt DNA sequences approximately 1 kb apart in opposite directions as spacer sequences for the construction of U6-RNA vectors. Each constructed U3-RNA vector was then ligated into the PX458 vector. The vector construction is shown in Figure 10B.
[0466] 3. The vector constructed in step 2 was transfected into HEK293T cells. After culturing for 2 days, DNA was extracted and PCR amplified. The detection method was the same as that of step 4 in Example 5. Genome-specific primers were designed 1 kb upstream and downstream of the target site for PCR amplification. The amplified PCR product was ligated to the Blunt-simple vector. 96 recombinant clones were randomly selected for colony PCR detection and first-generation sequencing using the M13F / M13R primer pair. The first-generation sequencing results were compared with the genomic sequence of the reference gene. The results of the Type IA-2 system editing in HEK293T cells are shown in Figure 10C.
[0467] Although the specific embodiments of the present invention have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details based on all the teachings published, and these changes are all within the scope of protection of the present invention. The entire invention is given by the appended claims and any equivalents thereof.
Claims
1. A Type IA CRISPR-Cas system comprising: (1) a cas5a protein or a nucleotide sequence encoding a cas5a protein, wherein the cas5a protein has an amino acid sequence as shown in any one of SEQ ID NOs: 2, 8, 14, 20 or an ortholog, homolog, variant or functional fragment thereof; (2) a Cas8a protein or a nucleotide sequence encoding a Cas8a protein, wherein the Cas8a protein has an amino acid sequence as shown in any one of SEQ ID NOs: 3, 9, 15, 21 or an ortholog, homolog, variant or functional fragment thereof; (3) a Cas7 protein or a nucleotide sequence encoding a Cas7 protein, wherein the Cas7 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 4, 10, 16, 22 or an ortholog, homolog, variant or functional fragment thereof; (4) a Cas6 protein or a nucleotide sequence encoding a Cas6 protein, wherein the Cas6 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 5, 11, 17, 23 or an ortholog, homolog, variant or functional fragment thereof; and, (5) a csa5 protein or a nucleotide sequence encoding a csa5 protein, wherein the csa5 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 6, 12, 18, 24 or an ortholog, homolog, variant or functional fragment thereof; in, In any one of (1) to (5), the ortholog, homolog, variant or functional fragment substantially retains the biological function of the sequence from which it is derived; Preferably, the ortholog, homolog, variant has one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence from which it is derived, or, has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, and substantially retains the biological function of the sequence from which it is derived.
2. The system of claim 1, wherein: The system further comprises: (6) a cas3 protein or a nucleotide sequence encoding a cas3 protein, wherein the cas3 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 1, 7, 13, 19 or an ortholog, homolog, variant or functional fragment thereof; wherein the ortholog, homolog, variant or functional fragment substantially retains the biological function of the sequence from which it is derived; Preferably, the orthologues, homologues, variants have one or more amino acids different from the sequence from which they are derived. or at least 99% sequence identity, and substantially retains the biological function of the sequence from which it is derived.
3. The system of claim 1 or 2, wherein: Any cas protein in the system optionally comprises an additional protein or polypeptide selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), an adenosine deaminase (e.g., TadA8e), a cytosine deaminase (e.g., APOBEC3), a domain having an activity selected from the following: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity (e.g., single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity) and nucleic acid binding activity; and any combination thereof; For example, at least one cas protein in the system comprises the additional protein or polypeptide; for example, the protein described in each of (1)-(6) comprises the additional protein or polypeptide.
4. The system of claim 3, wherein: The additional protein or polypeptide is an NLS sequence; For example, each cas protein in the system comprises an NLS sequence; For example, the NLS sequence is shown in SEQ ID NO:65; For example, the NLS sequence is located at or near a terminus (eg, the N-terminus or the C-terminus) of the protein.
5. The system of claim 3 or 4, wherein: The additional protein or polypeptide is an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); For example, one of the proteins described in any one of (1) to (5) in the system comprises an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3); For example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the protein (e.g., cas8a protein); For example, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
6. The system of any one of claims 3 to 5, wherein: The additional protein or polypeptide is connected to the protein via a linker or not; For example, the linker is a peptide linker or a non-peptide linker; For example, the peptide linker sequence is shown in SEQ ID NO: 66, 67 or 95.
7. The system according to any one of claims 1, 3 to 6, wherein: The system does not contain cas3 protein or nucleotide sequence encoding cas3 protein; Preferably, a cas protein (e.g., cas5a protein, cas8a protein, cas7 protein, cas6 protein or csa5 protein) in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the cas protein; For example, the cas8a protein in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); For example, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein; For example, the adenosine deaminase or cytosine deaminase is connected to the protein via a linker or without a linker; For example, the linker is a peptide linker or a non-peptide linker; For example, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; preferably, the peptide linker sequence is as shown in SEQ ID NO: 95; For example, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises a sequence as shown in any one of SEQ ID NOs: 96-99.
8. The system of any one of claims 1 to 7, wherein: The cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein respectively comprise the amino acid sequences shown in SEQ ID NOs: 2-6; For example, one or more (e.g., all 5) of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are connected to an NLS sequence (e.g., a sequence shown in SEQ ID NO: 65) via a linker or not; For example, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein connected to the NLS sequence respectively contain the amino acid sequences shown in SEQ ID NOs: 69-73.
9. The system of claim 8, wherein: The system further comprises: (6) cas3 protein or a nucleotide sequence encoding cas3 protein; Wherein, the cas3 protein comprises the amino acid sequence shown in SEQ ID NO:1; For example, the cas3 protein is connected to an NLS sequence (e.g., the sequence shown in SEQ ID NO: 65) through a linker or without a linker; For example, the cas3 protein connected to the NLS sequence comprises the amino acid sequence shown in SEQ ID NO:
68.
10. The system of claim 8, wherein: The system does not contain cas3 protein or nucleotide sequence encoding cas3 protein; For example, a cas protein (e.g., cas5a protein, cas8a protein, cas7 protein, cas6 protein or csa5 protein) in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the cas protein; For example, the cas8a protein in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); For example, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein; For example, the adenosine deaminase or cytosine deaminase is connected to the protein via a linker or without a linker; For example, the linker is a peptide linker or a non-peptide linker; For example, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; preferably, the peptide linker sequence is as shown in SEQ ID NO: 95; For example, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in SEQ ID NO:
96.
11. The system of claims 1-7, wherein: The cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein respectively contain the amino acid sequences shown in SEQ ID NOs: 8-12; For example, one or more (e.g., all 5) of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are connected to an NLS sequence (e.g., a sequence shown in SEQ ID NO: 65) via a linker or not; For example, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein connected to the NLS sequence respectively contain the amino acid sequences shown in SEQ ID NOs: 75-79.
12. The system of claim 11, wherein: The system further comprises: (6) cas3 protein or a nucleotide sequence encoding cas3 protein; Wherein, the cas3 protein comprises the amino acid sequence shown in SEQ ID NO:7; For example, the cas3 protein is connected to an NLS sequence (e.g., the sequence shown in SEQ ID NO: 65) through a linker or without a linker; For example, the cas3 protein connected to the NLS sequence comprises the amino acid sequence shown in SEQ ID NO:
74.
13. The system of claim 11, wherein: The system does not contain cas3 protein or nucleotide sequence encoding cas3 protein; For example, a cas protein (e.g., cas5a protein, cas8a protein, cas7 protein, cas6 protein or csa5 protein) in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the cas protein; For example, the cas8a protein in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); For example, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein; For example, the adenosine deaminase or cytosine deaminase is connected to the protein via a linker or without a linker; For example, the linker is a peptide linker or a non-peptide linker; For example, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; preferably, the peptide linker sequence is as shown in SEQ ID NO: 95; For example, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in SEQ ID NO:
97.
14. The system of any one of claims 1 to 7, wherein: The cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein respectively contain the amino acid sequences shown in SEQ ID NOs: 14-18; For example, one or more (e.g., all 5) of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are connected to an NLS sequence (e.g., a sequence shown in SEQ ID NO: 65) via a linker or not; For example, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein connected to the NLS sequence respectively contain the amino acid sequences shown in SEQ ID NOs: 81-85.
15. The system of claim 14, wherein: The system further comprises: (6) cas3 protein or a nucleotide sequence encoding cas3 protein; Wherein, the cas3 protein comprises the amino acid sequence shown in SEQ ID NO:13; For example, the cas3 protein is connected to an NLS sequence (e.g., the sequence shown in SEQ ID NO: 65) through a linker or without a linker; For example, the cas3 protein connected to the NLS sequence comprises the amino acid sequence shown in SEQ ID NO:
80.
16. The system of claim 14, wherein: The system does not contain cas3 protein or nucleotide sequence encoding cas3 protein; For example, a cas protein (e.g., cas5a protein, cas8a protein, cas7 protein, cas6 protein or csa5 protein) in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the cas protein; For example, the cas8a protein in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); For example, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein; For example, the adenosine deaminase or cytosine deaminase is connected to the protein via a linker or without a linker; For example, the linker is a peptide linker or a non-peptide linker; For example, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; preferably, the peptide linker sequence is as shown in SEQ ID NO: 95; For example, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in SEQ ID NO:
98.
17. The system of any one of claims 1 to 7, wherein: The cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein respectively comprise the amino acid sequences shown in SEQ ID NOs: 20-24; For example, one or more (e.g., all 5) of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are connected to an NLS sequence (e.g., a sequence shown in SEQ ID NO: 65) via a linker or not; For example, the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein connected to the NLS sequence respectively contain the amino acid sequences shown in SEQ ID NOs: 87-91.
18. The system of claim 17, wherein: The system further comprises: (6) cas3 protein or a nucleotide sequence encoding cas3 protein; Wherein, the cas3 protein comprises the amino acid sequence shown in SEQ ID NO:19; For example, the cas3 protein is connected to an NLS sequence (e.g., the sequence shown in SEQ ID NO: 65) through a linker or without a linker; For example, the cas3 protein connected to the NLS sequence comprises the amino acid sequence shown in SEQ ID NO:
86.
19. The system of claim 17, wherein: The system does not contain cas3 protein or nucleotide sequence encoding cas3 protein; For example, a cas protein (e.g., cas5a protein, cas8a protein, cas7 protein, cas6 protein or csa5 protein) in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the cas protein; For example, the cas8a protein in the system comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); For example, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein; For example, the adenosine deaminase or cytosine deaminase is connected to the protein via a linker or without a linker; For example, the linker is a peptide linker or a non-peptide linker; For example, the peptide linker sequence is as shown in SEQ ID NO: 66, 67 or 95; preferably, the peptide linker sequence is as shown in SEQ ID NO: 95; For example, the cas8a protein in the system comprises TadA8e, and the cas8a protein comprises the sequence shown in SEQ ID NO:
99.
20. The system of any one of claims 1 to 19, further comprising a guide RNA of a Type IA CRISPR-Cas system or a nucleotide sequence encoding the guide RNA; wherein, The guide RNA comprises a direct repeat sequence and a guide sequence capable of hybridizing with a target sequence; For example, the direct repeat sequence comprises a stem-loop structure; For example, the direct repeat sequence can bind to one or more cas proteins in the system; for example, the direct repeat sequence can bind to one or more proteins selected from cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein; for example, the guide RNA can bind to the Cascade complex formed by cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein; For example, when the target sequence is DNA, the protospacer adjacent motif (PAM) recognized by the system has a sequence represented by 5'CCN-; preferably, the PAM has a sequence represented by 5'CCT- or 5'CCC-.
21. The system of claim 20, wherein: The directly repeated sequence comprises a first region and a second region, wherein the first region comprises a stem-loop structure; For example, the first region is located at the 5' end of the second region; For example, there may or may not be extra nucleotides between the first region and the second region.
22. The system of claim 20 or 21, wherein: The guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a guide sequence located between the first copy of the direct repeat sequence and the second copy of the direct repeat sequence.
23. The system of claim 20 or 21, wherein: The guide RNA comprises a second region of a first copy of a direct repeat sequence, a guide sequence, and a first region of a second copy of a direct repeat sequence; Preferably, the guide sequence is located between the second region of the first copy of the direct repeat sequence and the first region of the second copy of the direct repeat sequence; Preferably, the second region of the first copy of the direct repeat sequence is located at the 5' end of the guide sequence, and the first region of the second copy of the direct repeat sequence is located at the 3' end of the guide sequence; Preferably, the second region of the first copy of the direct repeat sequence contains or does not contain redundant nucleotides between the guide sequence; Preferably, there may or may not be extra nucleotides between the guide sequence and the first region of the second copy of the direct repeat sequence.
24. The system of any one of claims 20 to 23, wherein: When the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in any one of claims 8 to 10, the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 49; Preferably, the first region of the directly repeated sequence comprises or consists of the sequence shown in SEQ ID NO:51, and the second region of the directly repeated sequence comprises or consists of the sequence shown in SEQ ID NO:
52.
25. The system of any one of claims 20-23, wherein: When the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in any one of claims 11 to 13, the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 53; Preferably, the first region of the directly repeated sequence comprises or consists of the sequence shown in SEQ ID NO:55, and the second region of the directly repeated sequence comprises or consists of the sequence shown in SEQ ID NO:
56.
26. The system of any one of claims 20-23, wherein: When the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in any one of claims 14 to 16, the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 57; Preferably, the first region of the directly repeated sequence comprises or consists of the sequence shown in SEQ ID NO:59, and the second region of the directly repeated sequence comprises the sequence shown in SEQ ID NO:60 Or consisting of the sequence shown in SEQ ID NO:
60.
27. The system of any one of claims 20-23, wherein: When the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are as defined in any one of claims 17 to 19, the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 61; Preferably, the first region of the directly repeated sequence comprises or consists of the sequence shown in SEQ ID NO:63, and the second region of the directly repeated sequence comprises or consists of the sequence shown in SEQ ID NO:
64.
28. The system of any one of claims 1-19, further comprising one or more guide RNAs of a Type IA CRISPR-Cas system or a nucleotide sequence encoding the one or more guide RNAs; wherein, The one or more guide RNAs comprise a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, and a second guide sequence capable of hybridizing to a second target sequence; Wherein, the first target sequence and the second target sequence are respectively located on the flanks of the region to be modified (e.g., the region to be deleted) in the double-stranded target nucleic acid molecule; For example, the first target sequence and the second target sequence are respectively located on two single strands of the region to be modified; for example, the first target sequence and the second target sequence are respectively located at the 5' end of the region to be modified in their respective single strands; For example, the direct repeat sequence comprises a stem-loop structure; For example, the direct repeat sequence can bind to one or more cas proteins in the system; for example, the direct repeat sequence can bind to one or more proteins selected from cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein; for example, the guide RNA can bind to the Cascade complex formed by cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein; For example, when the target sequence is DNA, the protospacer adjacent motif (PAM) recognized by the system has a sequence represented by 5'CCN-; preferably, the PAM has a sequence represented by 5'CCT- or 5'CCC-.
29. The system of claim 28, wherein: The directly repeated sequence comprises a first region and a second region, wherein the first region comprises a stem-loop structure; For example, the first region is located at the 5' end of the second region; For example, there may or may not be extra nucleotides between the first region and the second region; For example, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are defined as in any one of claims 8 to 10, the direct repeat sequence is as shown in SEQ ID NO: 69; preferably, the first region of the direct repeat sequence comprises or consists of the sequence shown in SEQ ID NO: 51, The second region of the directly repeated sequence comprises or consists of the sequence shown in SEQ ID NO:52; For example, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are defined as in any one of claims 11 to 13, the direct repeat sequence is as shown in SEQ ID NO: 53; preferably, the first region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 55 or consists of the sequence shown in SEQ ID NO: 55, and the second region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 56 or consists of the sequence shown in SEQ ID NO: 56; For example, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are defined as in any one of claims 14 to 16, the direct repeat sequence is as shown in SEQ ID NO: 57; preferably, the first region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 59 or consists of the sequence shown in SEQ ID NO: 59, and the second region of the direct repeat sequence comprises the sequence shown in SEQ ID NO: 60 or consists of the sequence shown in SEQ ID NO: 60; For example, when the cas5a protein, cas8a protein, cas7 protein, cas6 protein and csa5 protein are defined in any one of claims 17 to 19, the direct repeat sequence is as shown in SEQ ID NO: 61; preferably, the first region of the direct repeat sequence contains the sequence shown in SEQ ID NO: 63 or consists of the sequence shown in SEQ ID NO: 63, and the second region of the direct repeat sequence contains the sequence shown in SEQ ID NO: 64 or consists of the sequence shown in SEQ ID NO:
64.
30. The system of claim 28 or 29, wherein: The one guide RNA comprises: (i) a first copy of a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, a second copy of a direct repeat sequence, a second guide sequence capable of hybridizing to a second target sequence, and a third copy of a direct repeat sequence; or, (ii) a second region of a first copy of a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, a second copy of a direct repeat sequence, a second guide sequence capable of hybridizing to a second target sequence, and a first region of a third copy of a direct repeat sequence; Preferably, in (i), the one guide RNA comprises from 5' to 3' direction: the first copy of the direct repeat sequence, the first guide sequence, the second copy of the direct repeat sequence, the second guide sequence, and the third copy of the direct repeat sequence; Preferably, in (ii), the one guide RNA comprises from 5' to 3' direction: the second region of the first copy of the direct repeat sequence, the first guide sequence, the second copy of the direct repeat sequence, the second guide sequence, and the first region of the third copy of the direct repeat sequence.
31. The system of claim 28 or 29, wherein: The plurality of guide RNAs comprises: a first guide RNA comprising a direct repeat sequence and a first guide sequence capable of hybridizing to a first target sequence; and A second guide RNA comprises a direct repeat sequence and a second guide sequence capable of hybridizing to a second target sequence.
32. The system of claim 31, wherein: The first guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a first guide sequence located between the two copies of the repeat sequence; or, the first guide RNA comprises, from 5' to 3' direction, a second region of the first copy of the direct repeat sequence, a first guide sequence, and a first region of the second copy of the direct repeat sequence; The second guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a second guide sequence located between the two copies of the repeat sequence; or, the second guide RNA comprises, from the 5' to the 3' direction, a second region of the first copy of the direct repeat sequence, a second guide sequence, and a first region of the second copy of the direct repeat sequence.
33. A cas protein of a Type IA CRISPR-Cas system selected from: (1) Cas5a protein, wherein the Cas5a protein has an amino acid sequence shown in any one of SEQ ID NOs: 2, 8, 14, 20 or an ortholog, homolog, variant or functional fragment thereof; (2) a Cas8a protein having an amino acid sequence as shown in any one of SEQ ID NOs: 3, 9, 15, 21 or an ortholog, homolog, variant or functional fragment thereof; (3) a Cas7 protein having an amino acid sequence as shown in any one of SEQ ID NOs: 4, 10, 16, 22 or an ortholog, homolog, variant or functional fragment thereof; (4) a Cas6 protein having an amino acid sequence as shown in any one of SEQ ID NOs: 5, 11, 17, 23 or an ortholog, homolog, variant or functional fragment thereof; (5) a csa5 protein having an amino acid sequence as shown in any one of SEQ ID NOs: 6, 12, 18, 24 or an ortholog, homolog, variant or functional fragment thereof; (6) Cas3 protein, wherein the Cas3 protein has an amino acid sequence as shown in any one of SEQ ID NOs: 1, 7, 13, 19 or an ortholog, homolog, variant or functional fragment thereof; in, In any one of (1) to (6), the ortholog, homolog, variant or functional fragment substantially retains the biological function of the sequence from which it is derived; Preferably, the ortholog, homolog, variant has one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence from which it is derived, or has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least 135%, at least 136%, at least 137%, at least 138%, at least 139 or at least 99% sequence identity and substantially retains the biological function of the sequence from which it is derived.
34. The protein of claim 33, wherein The protein described in any one of (1) to (6) optionally comprises an additional protein or polypeptide, wherein the additional protein or polypeptide is selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcription activation domain (e.g., VP64), a transcription repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), an adenosine deaminase (e.g., TadA8e), a cytosine deaminase (e.g., APOBEC3), a domain having an activity selected from the following: methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, nuclease activity (e.g., single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity) and nucleic acid binding activity; and any combination thereof; For example, at least one (e.g., at least two, at least three, at least four or all five) of the proteins described in any one of (1) to (6) comprises the additional protein or polypeptide; for example, the protein described in each of (1) to (6) comprises the additional protein or polypeptide; For example, the additional protein or polypeptide is an NLS sequence; for example, the protein described in each of (1)-(6) comprises an NLS sequence; For example, the NLS sequence is shown in SEQ ID NO:65; For example, the additional protein or polypeptide is connected to the protein via a linker or not via a linker; For example, the linker is a peptide linker or a non-peptide linker; For example, the peptide linker sequence is shown in SEQ ID NO: 66, 67 or 95; For example, the NLS sequence is located at or near a terminus (e.g., the N-terminus or the C-terminus) of the protein; For example, the additional protein or polypeptide is an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); for example, one of the proteins described in any one of (1) to (5) comprises an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); For example, the amino acid sequence of the adenosine deaminase or cytosine deaminase is located at or near the end (e.g., N-terminus or C-terminus) of the protein (e.g., cas8a protein); For example, the adenosine deaminase or cytosine deaminase amino acid sequence is located at or near the N-terminus of the cas8a protein.
35. The protein of claim 33 or 34, wherein: (1) The cas5a protein comprises an NLS sequence and comprises a sequence selected from the following, or consists of a sequence selected from the following: (i) a sequence shown in any one of SEQ ID NOs: 69, 75, 81, and 87; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to the sequence shown in any one of SEQ ID NOs: 69, 75, 81, and 87. or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 69, 75, 81, 87; (2) The cas8a protein comprises an NLS sequence and comprises a sequence selected from the following, or consists of a sequence selected from the following: (i) a sequence shown in any one of SEQ ID NOs: 70, 76, 82, and 88; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to a sequence shown in any one of SEQ ID NOs: 70, 76, 82, and 88; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a sequence shown in any one of SEQ ID NOs: 70, 76, 82, and 88; (3) The cas7 protein comprises an NLS sequence and comprises a sequence selected from the following, or consists of a sequence selected from the following: (i) a sequence shown in any one of SEQ ID NOs: 71, 77, 83, and 89; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to a sequence shown in any one of SEQ ID NOs: 71, 77, 83, and 89; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a sequence shown in any one of SEQ ID NOs: 71, 77, 83, and 89; (4) The cas6 protein comprises an NLS sequence and comprises a sequence selected from the following, or consists of a sequence selected from the following: (i) a sequence shown in any one of SEQ ID NOs: 72, 78, 84, and 90; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to a sequence shown in any one of SEQ ID NOs: 72, 78, 84, and 90; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a sequence shown in any one of SEQ ID NOs: 72, 78, 84, and 90; (5) The csa5 protein comprises an NLS sequence and comprises a sequence selected from the following, or consists of a sequence selected from the following: (i) a sequence as shown in any one of SEQ ID NOs: 73, 79, 85, and 91; (ii) a sequence having one or more amino acid substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions) compared to the sequence as shown in any one of SEQ ID NOs: 73, 79, 85, and 91; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence as shown in any one of SEQ ID NOs: 73, 79, 85, and 91; (6) The cas3 protein comprises an NLS sequence and comprises a sequence selected from the following, or is composed of a sequence selected from the following Composition: (i) a sequence as shown in any one of SEQ ID NOs: 68, 74, 80, 86; (ii) a sequence having one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 68, 74, 80, 86; or (iii) a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the sequence shown in any one of SEQ ID NOs: 68, 74, 80, 86.
36. An isolated nucleic acid molecule comprising or consisting of a sequence selected from the group consisting of: (i) a sequence shown in any one of SEQ ID NOs: 49, 53, 57, and 61; (ii) comprising the sequence shown in SEQ ID NOs: 51 and 52, comprising the sequence shown in SEQ ID NOs: 55 and 56, comprising the sequence shown in SEQ ID NOs: 59 and 60, or comprising the sequence shown in SEQ ID NOs: 63 and 64; (iii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in (i) or (ii); (iv) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to the sequence shown in (i) or (ii); (v) a sequence that hybridizes under stringent conditions to a sequence described in any one of (i) to (iv); or (vi) a complementary sequence of the sequence described in any one of (i) to (iv); Furthermore, the sequence described in any one of (iii) to (vi) substantially retains the biological function of the sequence from which it is derived; For example, the nucleic acid molecule can bind to one or more of the cas proteins described in any one of claims 33-35; for example, the nucleic acid molecule can bind to one or more proteins selected from the group consisting of the cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein; For example, the nucleic acid molecule comprises a sequence selected from the following, or consists of a sequence selected from the following: (a) the nucleotide sequence shown in any one of SEQ ID NOs: 49, 53, 57, and 61; (b) comprising the sequence shown in SEQ ID NOs: 51 and 52, comprising the sequence shown in SEQ ID NOs: 55 and 56, comprising the sequence shown in SEQ ID NOs: 59 and 60, or comprising the sequence shown in SEQ ID NOs: 63 and 64; (c) a sequence that hybridizes under stringent conditions to the sequence described in (a) or (b); or (d) The complementary sequence of the sequence described in (a) or (b); For example, the isolated nucleic acid molecule is RNA; For example, the isolated nucleic acid molecule is a direct repeat sequence or a fragment thereof in the CRISPR / Cas system.
37. An isolated nucleic acid molecule encoding the protein of any one of claims 33-35.
38. A vector comprising the isolated nucleic acid molecule of claim 37.
39. A host cell comprising the isolated nucleic acid molecule of claim 37 or the vector of claim 38.
40. A Type IA CRISPR-Cas vector system, comprising one or more vectors, wherein the one or more vectors comprise: a nucleotide sequence encoding a cas protein in a Type IA CRISPR-Cas system, wherein the cas protein comprises: a cas5a protein, a cas8a protein, a cas7 protein, a cas6 protein, and a csa5 protein; in, The cas5a protein, cas8a protein, cas7 protein, cas6 protein, and csa5 protein are defined in any one of claims 1-19 and 33-35.
41. The vector system of claim 40, wherein The nucleotide sequence encoding the cas protein is located in one or more expression cassettes; Preferably, the nucleotide sequences encoding the cas proteins located in the same expression cassette are arranged in any order; Preferably, the nucleotide sequences encoding the cas protein located in the same expression cassette are connected to each other by a nucleotide sequence encoding a self-cleaving peptide (eg, T2A); Preferably, the expression cassettes each independently comprise a promoter, such as an inducible promoter.
42. The vector system of claim 40 or 41, wherein The one or more vectors further comprise a nucleotide sequence encoding a cas3 protein; Wherein, the cas3 protein is defined in any one of claims 1-6, 8-9, 11-12, 14-15, 17-18, 33-35; For example, the nucleotide sequences encoding cas5a protein, cas8a protein, cas7 protein, cas6 protein, csa5 protein and cas3 protein are located in the same expression cassette; For example, the one or more vectors comprise: A first expression cassette comprising a nucleotide sequence encoding a cas3 protein and a csa5 protein; and The second expression cassette comprises a nucleotide sequence encoding a cas7 protein, a cas5a protein, a cas6 protein and a cas8a protein.
43. The vector system of claim 40 or 41, wherein The one or more vectors do not contain a nucleotide sequence encoding a cas3 protein; Wherein, the cas protein in the system is any one of claims 1, 3-8, 10-11, 13-14, 16-17, 19. definition; For example, the one or more vectors comprise: A first expression cassette comprising a nucleotide sequence encoding a cas8a protein; and A second expression cassette comprising a nucleotide sequence encoding a cas7 protein, a cas5a protein, a cas6 protein, and a csa5 protein; For example, the cas8a protein is defined in any one of claims 1, 3-8, 10-11, 13-14, 16-17, and 19.
44. The vector system of any one of claims 40 to 43, wherein The one or more vectors further comprise: a nucleotide sequence encoding a guide RNA in a Type IA CRISPR-Cas system, the guide RNA being defined as in any one of claims 20-27; For example, the nucleotide sequence encoding the guide RNA in the Type IA CRISPR-Cas system is located in an additional expression cassette; for example, the additional expression cassette comprises a promoter, such as an inducible promoter.
45. The vector system of any one of claims 40 to 43, wherein The one or more vectors further comprise: a nucleotide sequence encoding one or more guide RNAs in a Type IA CRISPR-Cas system, the one or more guide RNAs being defined as in any one of claims 28 to 32; For example, the nucleotide sequence encoding one or more guide RNAs in the Type IA CRISPR-Cas system is located in an additional expression cassette; for example, the additional expression cassette comprises a promoter, such as an inducible promoter.
46. The vector system of any one of claims 40 to 45, wherein: The nucleotide sequences encoding the cas proteins are all located on the same vector; For example, the nucleotide sequence encoding the cas protein and the nucleotide sequence encoding the guide RNA are both located on the same vector.
47. A Type IA CRISPR-Cas system comprising: one or more guide RNAs or nucleotide sequences encoding the one or more guide RNAs; wherein: The one or more guide RNAs comprise a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, and a second guide sequence capable of hybridizing to a second target sequence; Wherein, the first target sequence and the second target sequence are respectively located on the flanks of the region to be modified (e.g., the region to be deleted) in the double-stranded target nucleic acid molecule; For example, the first target sequence and the second target sequence are respectively located on two single strands of the region to be modified; for example, the first target sequence and the second target sequence are respectively located at the 5' end of the region to be modified in their respective single strands; For example, the direct repeat sequence comprises a stem-loop structure; For example, the direct repeat sequence can bind to one or more cas proteins in the Type IA CRISPR-Cas system; For example, when the target sequence is DNA, the protospacer adjacent motif (PAM) recognized by the system has a sequence represented by 5'CCN-; preferably, the PAM has a sequence represented by 5'CCT- or 5'CCC-.
48. The system of claim 47, wherein: The directly repeated sequence comprises a first region and a second region, wherein the first region comprises a stem-loop structure; Preferably, the first region is located at the 5' end of the second region; For example, there may or may not be extra nucleotides between the first region and the second region.
49. The system of claim 47 or 48, wherein: The one guide RNA comprises: (i) a first copy of a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, a second copy of a direct repeat sequence, a second guide sequence capable of hybridizing to a second target sequence, and a third copy of a direct repeat sequence; or, (ii) a second region of a first copy of a direct repeat sequence, a first guide sequence capable of hybridizing to a first target sequence, a second copy of a direct repeat sequence, a second guide sequence capable of hybridizing to a second target sequence, and a first region of a third copy of a direct repeat sequence; Preferably, in (i), the one guide RNA comprises from 5' to 3' direction: a first copy of the direct repeat sequence, the first guide sequence, a second copy of the direct repeat sequence, the second guide sequence, and a third copy of the direct repeat sequence; Preferably, in (ii), the one guide RNA comprises from 5' to 3' direction: the second region of the first copy of the direct repeat sequence, the first guide sequence, the second copy of the direct repeat sequence, the second guide sequence, and the first region of the third copy of the direct repeat sequence.
50. The system of claim 47 or 48, wherein: The plurality of guide RNAs comprises: a first guide RNA comprising a direct repeat sequence and a first guide sequence capable of hybridizing to a first target sequence; and A second guide RNA comprises a direct repeat sequence and a second guide sequence capable of hybridizing to a second target sequence.
51. The system of claim 50, wherein: The first guide RNA comprises two copies of the direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a first guide sequence located between the two copies of the repeat sequence; or, the first guide RNA comprises, from the 5' to the 3' direction, a second region of the first copy of the direct repeat sequence, a first guide sequence, and a first region of a second copy of a direct repeat sequence; The second guide RNA comprises two copies of a direct repeat sequence, i.e., a first copy of the direct repeat sequence and a second copy of the direct repeat sequence, and a second guide sequence located between the two copies of the repeat sequence; or, the second guide RNA comprises, from the 5' to the 3' direction, a second region of the first copy of the direct repeat sequence, a second guide sequence, and a first region of the second copy of the direct repeat sequence.
52. The system of any one of claims 47-51, wherein: The system further comprises: a cas protein in a Type IA CRISPR-Cas system or a nucleotide sequence encoding the cas protein; For example, each of the cas proteins further comprises an additional protein or polypeptide selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), an adenosine deaminase (e.g., TadA8e), a cytosine deaminase (e.g., APOBEC3), a domain having an activity selected from the following: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity (e.g., single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity) and nucleic acid binding activity; And any combination thereof; For example, the additional protein or polypeptide is an NLS sequence; For example, the additional protein or polypeptide is an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
53. The system of claim 52, wherein: The cas protein comprises cas3 protein, cas5a protein, cas8a protein, cas6 protein, csa5 protein and cas7 protein; Preferably, the cas3 protein, cas5a protein, cas8a protein, cas6 protein, csa5 protein and cas7 protein are defined in any one of claims 1-6, 8-9, 11-12, 14-15, 17-18, 33-35.
54. A Type IA CRISPR-Cas vector system, comprising one or more vectors, wherein the one or more vectors comprise: a nucleotide sequence encoding one or more guide RNAs in a Type IA CRISPR-Cas system, wherein the one or more guide RNAs are defined in any one of claims 47-51.
55. The vector system of claim 54, wherein The one or more vectors further comprise: a nucleotide sequence encoding a cas protein in a Type IA CRISPR-Cas system; For example, each of the cas proteins further comprises an additional protein or polypeptide selected from the epitope Tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting moieties, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB domains or SID domains), nuclease domains (e.g., Fok1), adenosine deaminase (e.g., TadA8e), cytosine deaminase (e.g., APOBEC3), domains having an activity selected from the following: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity (e.g., single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity) and nucleic acid binding activity; and any combination thereof; For example, the additional protein or polypeptide is an NLS sequence; For example, the additional protein or polypeptide is an adenosine deaminase (eg, TadA8e) or a cytosine deaminase (eg, APOBEC3).
56. The vector system of claim 55, wherein The cas protein comprises cas3 protein, cas5a protein, cas8a protein, cas6 protein, csa5 protein and cas7 protein; Preferably, the cas3 protein, cas5a protein, cas8a protein, cas6 protein, csa5 protein and cas7 protein are defined in any one of claims 1-6, 8-9, 11-12, 14-15, 17-18, 33-35.
57. The vector system of any one of claims 54 to 56, wherein: The nucleotide sequence encoding one or more guide RNAs in the Type IA CRISPR-Cas system and the nucleotide sequence encoding the cas protein in the Type IA CRISPR-Cas system are located in different expression cassettes.
58. The vector system of any one of claims 54 to 57, wherein: The nucleotide sequences encoding the cas proteins are all located on the same vector; For example, the nucleotide sequence encoding the cas protein and the nucleotide sequence encoding one or more guide RNAs are both located on the same vector.
59. A host cell comprising the vector system of any one of claims 40-46, 54-58.
60. A kit comprising the system of any one of claims 1-32, the protein of any one of claims 33-35, the isolated nucleic acid molecule of claim 36 or 37, the vector of claim 38, the host cell of claim 39, the vector system of any one of claims 40-46, the system of any one of claims 47-53, the vector system of any one of claims 54-58, or the host cell of claim 59; and instructions for using the system for nucleic acid editing (e.g., gene or genome editing, gene or genome large fragment deletion, gene or genome base modification, genome structural variation); For example, the kit comprises the system of any one of claims 20-32; For example, the kit comprises the vector system of any one of claims 44-46; For example, the kit comprises the system of claim 52 or 53; For example, the kit comprises the vector system of any one of claims 55-58.
61. A delivery composition comprising the system of any one of claims 1-32, the carrier system of any one of claims 40-46, the system of any one of claims 47-53, or the carrier system of any one of claims 54-58, and a delivery system; For example, the delivery system is selected from particles, vesicles or viral vectors; For example, the particles comprise lipids, sugars, metals, or proteins; For example, the vesicle comprises an exosome or a liposome; For example, the viral vector comprises an adenovirus, a lentivirus or an adeno-associated virus; For example, the delivery composition comprises the system of any one of claims 20-32; For example, the delivery composition comprises the carrier system of any one of claims 44-46; For example, the delivery composition comprises the system of claim 52 or 53; For example, the delivery composition comprises the carrier system of any one of claims 55-58.
62. A method for inducing a deletion in a target genome, the target genome comprising a complementary first nucleic acid strand and a second nucleic acid strand, the method comprising: Contacting the system of any one of claims 20 to 32, or the vector system of any one of claims 44 to 47, the system of claim 52 or 53, or the vector system of any one of claims 53 to 58 with the target genome, or delivering it to a cell comprising the target genome; For example, one or more cas proteins contained in the system or vector system are capable of forming a complex with the guide RNA, and after the complex binds to the target sequence and / or its complementary sequence, inducing the deletion of the region containing the target sequence and / or its complementary sequence; For example, the method comprises: contacting the system of any one of claims 28 to 32 or the vector system of claim 45 or 46, the system of claim 52 or 53, or the vector system of any one of claims 55 to 58 with the target genome, or delivering it into a cell comprising the target genome; For example, the deletion is a large deletion, such as a deletion of greater than 0.1 kb, greater than 0.2 kb, greater than 0.5 kb, greater than 1 kb, greater than 1.5 kb, greater than 2 kb, greater than 10 kb, greater than 50 kb, greater than 100 kb, such as less than 500 kb, less than 400 kb, less than 300 kb, less than 200 kb; For example, the one or more guide RNAs included in the system or vector system include direct repeat sequences, A first guide sequence that hybridizes with a first target sequence and a second guide sequence that can hybridize with a second target sequence; wherein the first target sequence and the second target sequence are respectively located on the flanks of the region to be deleted in the target genome; For example, the first target sequence is located in the first nucleic acid strand of the target genome, and the second target sequence is located in the second nucleic acid strand of the target genome; for example, in the first nucleic acid strand, the first target sequence is located at the 5' end of the region to be deleted, and, in the second nucleic acid strand, the second target sequence is located at the 5' end of the region to be deleted; For example, the length of the region to be deleted is greater than 0.1 kb, for example, greater than 0.2 kb, greater than 0.3 kb, greater than 0.4 kb, greater than 0.5 kb; for example, the length of the region to be deleted is less than 500 kb, for example, less than 400 kb, less than 300 kb, less than 200 kb; for example, the length of the region to be deleted is 0.2 kb-200 kb (for example, 0.2 kb-2 kb, 0.2 kb-5 kb, 0.2 kb-10 kb, 0.2 kb-100 kb, 0.2 kb-200 kb; for example, 0.5 kb-1.5 kb, 0.5 kb-2 kb, 0.5 kb-10 kb); For example, the target genome is present in a cell, or the target genome is present in a nucleic acid molecule (e.g., a plasmid) in vitro; For example, the cell is a prokaryotic cell; For example, the cell is a eukaryotic cell; For example, the cell is selected from an animal cell (e.g., a mammalian cell, such as a human cell), a plant cell (e.g., a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast); For example, the method is used for chromosome elimination.
63. A method for inducing structural variation in a genome, the genome comprising a first nucleic acid strand and a second nucleic acid strand that are complementary, the method comprising: Contacting the system of any one of claims 20 to 32, or the vector system of any one of claims 44 to 46, the system of claim 52 or 53, or the vector system of any one of claims 55 to 58 with a target genome, or delivering it into a cell comprising the target genome; For example, one or more cas proteins contained in the system or vector system can form a complex with the guide RNA, and after the complex binds to the target sequence and / or its complementary sequence, it induces the deletion of the region containing the target sequence and / or its complementary sequence, thereby inducing genomic structural variation; For example, the genome comprises a complementary first nucleic acid strand and a second nucleic acid strand, and the method comprises: contacting the system of any one of claims 28 to 32 or the vector system of claim 45 or 46, the system of claim 52 or 53, or the vector system of any one of claims 55 to 58 with a target genome, or delivering it into a cell comprising the target genome; For example, the deletion is a large deletion, such as a deletion of greater than 0.1 kb, greater than 0.2 kb, greater than 0.5 kb, greater than 1 kb, greater than 1.5 kb, greater than 2 kb, greater than 10 kb, greater than 50 kb, greater than 100 kb, such as less than 500 kb, less than 400 kb, less than 300 kb, less than 200 kb; For example, the one or more guide RNAs included in the system or vector system include direct repeat sequences, A first guide sequence that hybridizes with a first target sequence and a second guide sequence that can hybridize with a second target sequence; wherein the first target sequence and the second target sequence are respectively located on the flanks of the region to be deleted in the target genome; For example, the first target sequence is located in the first nucleic acid strand of the target genome, and the second target sequence is located in the second nucleic acid strand of the target genome; for example, in the first nucleic acid strand, the first target sequence is located at the 5' end of the region to be deleted, and, in the second nucleic acid strand, the second target sequence is located at the 5' end of the region to be deleted; For example, the length of the region to be deleted is greater than 0.1 kb, for example, greater than 0.2 kb, greater than 0.3 kb, greater than 0.4 kb, greater than 0.5 kb; for example, the length of the region to be deleted is less than 500 kb, for example, less than 400 kb, less than 300 kb, less than 200 kb; for example, the length of the region to be deleted is 0.2 kb-200 kb (for example, 0.2 kb-2 kb, 0.2 kb-5 kb, 0.2 kb-10 kb, 0.2 kb-100 kb, 0.2 kb-200 kb; for example, 0.5 kb-1.5 kb, 0.5 kb-2 kb, 0.5 kb-10 kb); For example, the target genome is present in a cell, or the target genome is present in a nucleic acid molecule (e.g., a plasmid) in vitro; For example, the cell is a prokaryotic cell; For example, the cell is a eukaryotic cell; For example, the cell is selected from an animal cell (eg, a mammalian cell, such as a human cell), a plant cell (eg, a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast).
64. A method for modifying a target nucleic acid molecule, comprising: contacting the system of any one of claims 20 to 32, the vector system of any one of claims 44 to 46, the system of claims 52 or 53, or the vector system of any one of claims 55 to 58 with the target nucleic acid molecule, or delivering it to a cell comprising the target nucleic acid molecule; For example, one or more cas proteins contained in the system or vector system are capable of forming a complex with a guide RNA, and after the complex binds to a target sequence and / or its complementary sequence, induces modification of a target nucleic acid molecule comprising the target sequence and / or its complementary sequence; For example, the target nucleic acid molecule is RNA or DNA; For example, the target nucleic acid molecule is double-stranded DNA; For example, the target nucleic acid molecule is a gene or genome; For example, the target nucleic acid molecule is present in a cell, or the target nucleic acid molecule is present in a nucleic acid molecule (e.g., a plasmid) in vitro; For example, the cell is a prokaryotic cell; For example, the cell is a eukaryotic cell; For example, the cell is selected from an animal cell (e.g., a mammalian cell, such as a human cell), a plant cell (e.g., a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast); For example, the modification refers to the deletion of a large fragment of the target nucleic acid molecule; For example, the modification refers to a break in the target nucleic acid molecule, such as a double-strand break in DNA; for example, the modification also includes inserting an exogenous nucleic acid into the break; For example, the modification refers to a change in a base (eg, cytosine, adenine) in the target nucleic acid molecule.
65. A method for inducing base mutation in a target nucleic acid molecule, comprising: contacting the system of any one of claims 20 to 32, the vector system of any one of claims 44 to 46, the system of claims 52 or 53, or the vector system of any one of claims 55 to 58 with the target nucleic acid molecule, or delivering it to a cell comprising the target nucleic acid molecule; Wherein, the cas protein contained in the system or vector system is defined in any one of claims 5-7, 10, 13, 16, and 19; For example, one or more cas proteins contained in the system or vector system can form a complex with the guide RNA, and after the complex binds to the target sequence and / or its complementary sequence, it induces modification of the bases in the target nucleic acid molecule containing the target sequence and / or its complementary sequence, and generates base mutations during nucleic acid repair or replication; For example, the modification of the base refers to a modification that can change the base complementary pairing mode of the base to be modified; for example, before modification, the base to be modified is complementary paired with the first base, and after modification, the modified base is complementary paired with the second base; For example, the one or more cas proteins contained in the system or vector system further comprise an adenosine deaminase (e.g., TadA8e) or a cytosine deaminase (e.g., APOBEC3); For example, the one or more cas proteins (such as cas8a proteins) contained in the system or vector system further comprise an adenosine deaminase (such as TadA8e), the base to be modified is adenine, and before modification, adenine is complementary to thymine, and after modification, adenine is modified to hypoxanthine, and hypoxanthine is complementary to cytosine; For example, the one or more cas proteins (such as cas8a protein) contained in the system or vector system further comprises a cytosine deaminase (such as APOBEC3), the base to be modified is cytosine, before modification, cytosine is complementary to guanine, and after modification, cytosine is modified to uracil, and uracil is complementary to thymine; For example, the target nucleic acid molecule is RNA or DNA; For example, the target nucleic acid molecule is double-stranded DNA; For example, the target nucleic acid molecule is a gene or genome; For example, the target nucleic acid molecule is present in a cell, or the target nucleic acid molecule is present in a nucleic acid molecule (e.g., a plasmid) in vitro; For example, the cell is a prokaryotic cell; For example, the cell is a eukaryotic cell; For example, the cell is selected from an animal cell (e.g., a mammalian cell, such as a human cell), a plant cell (e.g., corn cells, corn protoplasts, rice cells, Arabidopsis cells, Arabidopsis protoplasts).
66. A method for altering the expression of a gene product, comprising: contacting the system of any one of claims 20 to 32, the vector system of any one of claims 44 to 46, the system of claims 52 or 53, or the vector system of any one of claims 55 to 58 with a target nucleic acid molecule encoding the gene product, or delivering it to a cell comprising the target nucleic acid molecule; For example, one or more cas proteins contained in the system or vector system are capable of forming a complex with a guide RNA, and after the complex binds to a target sequence and / or its complementary sequence, it induces modification of a target nucleic acid molecule comprising the target sequence and / or its complementary sequence, thereby changing the expression of a gene product; For example, the target nucleic acid molecule is present in a cell, or the target nucleic acid molecule is present in a nucleic acid molecule (e.g., a plasmid) in vitro; For example, the cell is a prokaryotic cell; For example, the cell is a eukaryotic cell; For example, the cell is selected from an animal cell (e.g., a mammalian cell, such as a human cell), a plant cell (e.g., a corn cell, a corn protoplast, a rice cell, an Arabidopsis cell, an Arabidopsis protoplast); For example, expression of the gene product is altered (e.g., increased or decreased); For example, the gene product is a protein.
67. A method for producing a plant with a modified trait, the method comprising contacting a plant cell with the system of any one of claims 20-32, the vector system of any one of claims 44-46, the system of claims 52 or 53, or the vector system of any one of claims 55-58, or subjecting the plant cell to the method of any one of claims 62-66, thereby modifying or editing a target gene or target nucleic acid molecule in the genome of the plant cell, and regenerating a plant from the plant cell; For example, the method comprises contacting a plant cell with the system of any one of claims 28 to 32 or the vector system of claim 45 or 46, the system of claim 52 or 53, or the vector system of any one of claims 55 to 58; For example, the plant is an agricultural plant, such as corn, barley, cotton, rice, soybean, wheat, or paddy rice.
68. The method of any one of claims 62-67, wherein the cas protein or nucleotide sequence encoding the cas protein, the guide RNA or nucleotide sequence encoding the guide RNA contained in the system or vector system is present in the delivery system; For example, the delivery system is selected from particles, vesicles or viral vectors; For example, the particles comprise lipids, sugars, metals, or proteins; For example, the vesicle comprises an exosome or a liposome; For example, the viral vector comprises an adenovirus, a lentivirus or an adeno-associated virus.
69. The system of any one of claims 1-32, the protein of any one of claims 33-35, the isolated nucleic acid molecule of claim 36 or 37, the vector of claim 38, the host cell of claim 39, the vector system of any one of claims 40-46, the system of any one of claims 47-53, the vector system of any one of claims 54-58, the host cell of claim 59, the kit of claim 60 or the delivery composition of claim 61 for use in nucleic acid editing, or in the preparation of a formulation for nucleic acid editing; For example, the nucleic acid editing includes gene or genome editing; For example, the gene or genome editing includes deletion of large nucleic acid fragments, modification of genes, knockout of genes, alteration of expression of gene products, repair of mutations, and / or insertion of polynucleotides, base mutations; For example, the nucleic acid editing includes inducing genomic structural variation or chromosome elimination.
70. Use of the system of any one of claims 1-32, the protein of any one of claims 33-35, the isolated nucleic acid molecule of claim 36 or 37, the vector of claim 38, the host cell of claim 39, the vector system of any one of claims 40-46, the system of any one of claims 47-53, the vector system of any one of claims 54-58, the host cell of claim 59, the kit of claim 60 or the delivery composition of claim 61 in the preparation of a formulation for editing a target nucleotide sequence in a target locus to modify an organism or a non-human organism (e.g., a plant).
71. A cell or progeny thereof obtained by the method of any one of claims 62-68, wherein the cell comprises a modification not present in its wild type.
72. A cell product of the cell of claim 71 or its progeny.