CRISPR-Cas σ enzymes and systems

By developing new RNA-guided endonuclease and CRISPR/Cas systems, the shortcomings of the existing system are solved and a more robust gene editing effect is achieved.

CN119193541BActive Publication Date: 2025-08-12CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411237286.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-09-04
Filing Date
2024-09-04
Publication Date
2025-08-12
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

The existing CRISPR/Cas systems have their own advantages and disadvantages, and lack a robust gene editing technology with multiple aspects and good performance.

Method used

A novel RNA-guided endonuclease was developed based on the enzyme's CRISPR/Cas system and its gene editing methods, including proteins, conjugates and complexes, capable of binding to guide RNA and cleaving DNA at specific sites.

Benefits of technology

It provides a more robust CRISPR/Cas system that can effectively perform gene editing and has many good performance, including activity of binding to guide RNA and cleavage of target sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119193541B_ABST
    Figure CN119193541B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of nucleic acid editing, in particular to the technical field of regularly clustered interspaced short palindromic repeats (CRISPR). Specifically, the present invention relates to Cas effector proteins, fusion proteins comprising such proteins, and nucleic acid molecules encoding them. The present invention also relates to complexes and compositions for nucleic acid editing (e.g., gene or genome editing), comprising proteins or fusion proteins of the present invention, or nucleic acid molecules encoding them. The present invention also relates to methods for nucleic acid editing (e.g., gene or genome editing), which use proteins or fusion proteins comprising the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of nucleic acid editing, in particular to the technical field of regularly clustered interspaced short palindromic repeats (CRISPR). Specifically, the present invention relates to Cas effector proteins, fusion proteins comprising such proteins, and nucleic acid molecules encoding them. The present invention also relates to complexes and compositions for nucleic acid editing (e.g., gene or genome editing), comprising proteins or fusion proteins of the present invention, or nucleic acid molecules encoding them. The present invention also relates to methods for nucleic acid editing (e.g., gene or genome editing), which use proteins or fusion proteins comprising the present invention. Background Art

[0002] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to guide specific binding to target sequences on the genome and cut DNA to produce double-strand breaks, and uses biological non-homologous end joining or homologous recombination to perform site-directed gene editing.

[0003] The CRISPR / Cas9 system is the most commonly used Type II CRISPR system, which recognizes a 3'-NGG PAM motif and performs blunt-end cleavage on target sequences. The CRISPR / Cas Type V system, a newly discovered CRISPR system in the past two years, utilizes a 5'-TTN motif and performs sticky-end cleavage on target sequences. Examples include Cpf1, C2c1, CasX, and CasY. However, the various CRISPR / Cas systems currently available have varying advantages and disadvantages. For example, Cas9, C2c1, and CasX all require two guide RNAs, while Cpf1 requires only a single guide RNA and can be used for multiplexed gene editing. CasX is 980 amino acids long, while the more common Cas9, C2c1, CasY, and Cpf1 are typically around 1300 amino acids. Furthermore, the PAM sequences of Cas9, Cpf1, CasX, and CasY are relatively complex and diverse, while C2c1 recognizes a strict 5'-TTN motif, making its target site more predictable than other systems, thereby reducing potential off-target effects.

[0004] In summary, given that the currently available CRISPR / Cas systems are limited by some defects, the development of a new CRISPR / Cas system that is more robust and has good performance in many aspects is of great significance to the development of biotechnology. Summary of the Invention

[0005] After extensive experimentation and repeated exploration, the inventors of this application unexpectedly discovered a new type of RNA-guided nuclease. Based on this discovery, the inventors developed a new CRISPR / Cas system and a gene editing method based on this system.

[0006] Cas effector proteins

[0007] Therefore, in a first aspect, the present invention provides a protein having an amino acid sequence as shown in any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 and 13, or a direct homologue, homologue, variant or functional fragment thereof; wherein the direct homologue, homologue, variant or functional fragment substantially retains the biological function of the sequence from which it is derived.

[0008] In the present invention, the biological functions of the above sequences include, but are not limited to, the activity of binding to guide RNA, endonuclease activity, and the activity of binding to and cutting specific sites of the target sequence under the guidance of the guide RNA.

[0009] In certain embodiments, the orthologs, homologs, variants have at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity compared to the sequence from which they are derived.

[0010] In certain embodiments, the orthologs, homologs, and variants have at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity compared to the sequence shown in any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, and 13, and substantially retain the biological function of the sequence from which it is derived (e.g., activity of binding to a guide RNA, endonuclease activity, activity of binding to and cleaving a specific site of a target sequence under the guidance of a guide RNA).

[0011] In certain embodiments, the protein is an effector protein in a CRISPR / Cas system.

[0012] In certain embodiments, the protein of the present invention comprises or consists of a sequence selected from the group consisting of:

[0013] (i) the sequence shown in any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, and 13;

[0014] (ii) a sequence having one or more amino acid substitutions, deletions or additions compared to the sequence shown in any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, and 13 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 and 40 amino acid substitutions, deletions or additions); or

[0015] (iii) a sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence set forth in any one of SEQ ID NOs: 1-13.

[0016] Derivatized proteins

[0017] The protein of the present invention can be derivatized, for example, by being linked to another molecule (e.g., another polypeptide or protein). Generally, the derivatization (e.g., labeling) of the protein will not adversely affect the desired activity of the protein (e.g., activity binding to a guide RNA, endonuclease activity, activity binding to and cutting a specific site of a target sequence under the guidance of a guide RNA). Therefore, the protein of the present invention is also intended to include such derivatized forms. For example, the protein of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage or other means) to one or more other molecular groups, such as another protein or polypeptide, a detection reagent, a pharmaceutical agent, etc.

[0018] In particular, the protein of the present invention can be linked to other functional units. For example, it can be linked to a nuclear localization signal (NLS) sequence to improve the ability of the protein of the present invention to enter the cell nucleus. For example, it can be linked to a targeting moiety to make the protein of the present invention targeted. For example, it can be linked to a detectable label to facilitate detection of the protein of the present invention. For example, it can be linked to an epitope tag to facilitate expression, detection, tracing and / or purification of the protein of the present invention.

[0019] Conjugate

[0020] Thus, in a second aspect, the present invention provides a conjugate comprising a protein as described above and a modifying moiety.

[0021] In certain embodiments, the modifying moiety is selected from another protein or polypeptide, a detectable label, or any combination thereof.

[0022] In certain embodiments, the additional protein or polypeptide is selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), a domain having an activity selected from the group consisting of nucleotide deaminase, methylase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof.

[0023] In certain embodiments, the conjugates of the present invention comprise one or more NLS sequences, such as the NLS of the SV40 virus large T antigen. In certain exemplary embodiments, the NLS sequence is as shown in SEQ ID NO: 53. In certain embodiments, the NLS sequence is located at, near, or close to a terminus (e.g., N-terminus or C-terminus) of the protein of the present invention. In certain exemplary embodiments, the NLS sequence is located at, near, or close to the C-terminus of the protein of the present invention.

[0024] In certain embodiments, the conjugates of the present invention comprise an epitope tag. Such epitope tags are well known to those skilled in the art, and examples thereof include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag according to the desired purpose (e.g., purification, detection, or tracing).

[0025] In certain embodiments, the conjugates of the present invention comprise a reporter gene sequence. Such reporter genes are well known to those skilled in the art, and examples thereof include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, and the like.

[0026] In certain embodiments, the conjugates of the present invention comprise a domain capable of binding to a DNA molecule or an intracellular molecule, such as maltose binding protein (MBP), the DNA binding domain (DBD) of Lex A, the DBD of GAL4, and the like.

[0027] In certain embodiments, the conjugates of the invention comprise a detectable label, such as a fluorescent dye, eg, FITC or DAPI.

[0028] In certain embodiments, the protein of the present invention is coupled, conjugated or fused to the modifying moiety, optionally via a linker.

[0029] In certain embodiments, the modifying moiety is directly linked to the N-terminus or C-terminus of a protein of the invention.

[0030] In certain embodiments, the modified portion is connected to the N-terminus or C-terminus of the protein of the present invention via a linker. Such linkers are well known in the art, and examples include, but are not limited to, linkers comprising one or more (e.g., 1, 2, 3, 4, or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.

[0031] Fusion protein

[0032] In a third aspect, the present invention provides a fusion protein comprising the protein of the present invention and another protein or polypeptide.

[0033] In certain embodiments, the additional protein or polypeptide is selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), a domain having an activity selected from the group consisting of nucleotide deaminase, methylase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof.

[0034] In certain embodiments, the fusion proteins of the present invention comprise one or more NLS sequences, such as the NLS of the SV40 virus large T antigen. In certain embodiments, the NLS sequence is located at, near, or close to the terminus (e.g., N-terminus or C-terminus) of the protein of the present invention. For example, the NLS sequence is as shown in SEQ ID NO: 53. In certain exemplary embodiments, the NLS sequence is located at, near, or close to the C-terminus of the protein of the present invention.

[0035] In certain embodiments, the fusion proteins of the invention comprise an epitope tag.

[0036] In certain embodiments, the fusion proteins of the present invention comprise a reporter gene sequence.

[0037] In certain embodiments, the fusion proteins of the present invention comprise a domain capable of binding to a DNA molecule or an intracellular molecule.

[0038] In certain embodiments, the protein of the present invention is fused to the additional protein or polypeptide, optionally via a linker.

[0039] In certain embodiments, the additional protein or polypeptide is directly linked to the N-terminus or C-terminus of the protein of the invention.

[0040] In certain embodiments, the additional protein or polypeptide is linked to the N-terminus or C-terminus of the protein of the invention via a linker.

[0041] In certain exemplary embodiments, the fusion protein of the present invention has an amino acid sequence as shown in any one of SEQ ID NOs: 54-66.

[0042] The protein, conjugate or fusion protein of the present invention is not limited by the method of production. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.

[0043] Directly repeated sequences

[0044] In a fourth aspect, the present invention provides an isolated nucleic acid molecule comprising or consisting of a sequence selected from the group consisting of:

[0045] (i) the sequence shown in any one of SEQ ID NOs: 27-39;

[0046] (ii) a sequence having one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 base substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID NOs: 27-39;

[0047] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 27-39;

[0048] (iv) a sequence that hybridizes under stringent conditions to the sequence described in any one of (i) to (iii); or

[0049] (v) a complementary sequence of the sequence described in any one of (i) to (iii);

[0050] Furthermore, the sequence described in any one of (ii) to (v) substantially retains the biological function of the sequence from which it is derived, wherein the biological function of the sequence refers to the activity as a direct repeat sequence in the CRISPR-Cas system.

[0051] In certain embodiments, the isolated nucleic acid molecule is a direct repeat sequence in a CRISPR-Cas system.

[0052] In certain embodiments, the nucleic acid molecule comprises or consists of a sequence selected from the group consisting of:

[0053] (a) the nucleotide sequence shown in any one of SEQ ID NOs: 27-39;

[0054] (b) a sequence that hybridizes under stringent conditions to the sequence described in (a); or

[0055] (c) The complementary sequence of the sequence described in (a).

[0056] In certain embodiments, the isolated nucleic acid molecule is RNA.

[0057] CRISPR / Cas complex

[0058] In a fifth aspect, the present invention provides a composite comprising:

[0059] (i) a protein component selected from the group consisting of a protein, a conjugate or a fusion protein of the present invention, and any combination thereof; and

[0060] (ii) a nucleic acid component comprising, from 5' to 3' direction, an isolated nucleic acid molecule as described above and a guide sequence capable of hybridizing to a target sequence,

[0061] Wherein, the protein component and the nucleic acid component combine with each other to form a complex.

[0062] In certain embodiments, the targeting sequence is linked to the 3' end of the nucleic acid molecule.

[0063] In certain embodiments, the guide sequence comprises the complement of the target sequence.

[0064] In certain embodiments, the nucleic acid component is a guide RNA in a CRISPR-Cas system.

[0065] In certain embodiments, the nucleic acid molecule is RNA.

[0066] In certain embodiments, the complex does not comprise a trans-acting crRNA (tracrRNA).

[0067] In certain embodiments, the guide sequence is at least 5, at least 10, at least 15, at least 20, at least 25, or at least 30 nucleotides in length. In certain embodiments, the guide sequence is 10-30, or 15-25, or 15-22, or 19-25, or 19-22 nucleotides in length.

[0068] In certain embodiments, the nucleic acid molecules of the separation are 55-70 nucleotides in length, for example 55-65 nucleotides, for example 60-65 nucleotides, for example 62-65 nucleotides, for example 63-64 nucleotides. In certain embodiments, the nucleic acid molecules of the separation are 15-30 nucleotides in length, for example 15-25 nucleotides, for example 20-25 nucleotides, for example 22-24 nucleotides, for example 23 nucleotides.

[0069] Encoding nucleic acid, vector and host cell

[0070] In a sixth aspect, the present invention provides an isolated nucleic acid molecule comprising:

[0071] (i) a nucleotide sequence encoding a protein or fusion protein of the present invention;

[0072] (ii) a nucleotide sequence encoding the isolated nucleic acid molecule according to the fourth aspect; or

[0073] (iii) a nucleotide sequence comprising (i) and (ii).

[0074] In certain embodiments, the nucleotide sequence described in any one of (i)-(iii) is codon-optimized for expression in prokaryotes. In certain embodiments, the nucleotide sequence described in any one of (i)-(iii) is codon-optimized for expression in eukaryotic cells.

[0075] In a seventh aspect, the present invention further provides a vector comprising the isolated nucleic acid molecule described in the sixth aspect. The vector of the present invention can be a cloning vector or an expression vector. In certain embodiments, the vector of the present invention is, for example, a plasmid, a cosmid, a phage, a cosmid, or the like. In certain embodiments, the vector is capable of expressing a protein of the present invention, a fusion protein, the isolated nucleic acid molecule described in the fourth aspect, or the complex described in the fifth aspect in a subject (e.g., a mammal, such as a human).

[0076] In an eighth aspect, the present invention further provides a host cell comprising the isolated nucleic acid molecule or vector as described above. Such host cells include, but are not limited to, prokaryotic cells such as Escherichia coli cells, and eukaryotic cells such as yeast cells, insect cells, plant cells, and animal cells (e.g., mammalian cells, such as mouse cells, human cells, etc.). The cells of the present invention can also be cell lines, such as 293T cells.

[0077] Compositions and carrier compositions

[0078] In a ninth aspect, the present invention further provides a composition comprising:

[0079] (i) a first component selected from the group consisting of a protein, a conjugate, a fusion protein, a nucleotide sequence encoding the protein or fusion protein of the present invention, and any combination thereof; and

[0080] (ii) a second component, which is a nucleotide sequence comprising a guide RNA, or a nucleotide sequence encoding the nucleotide sequence comprising a guide RNA;

[0081] Wherein, the guide RNA comprises a direct repeat sequence and a guide sequence from the 5' to the 3' direction, and the guide sequence is capable of hybridizing with the target sequence;

[0082] The guide RNA is capable of forming a complex with the protein, conjugate or fusion protein described in (i).

[0083] In certain embodiments, the direct repeat sequence is an isolated nucleic acid molecule as defined in the fourth aspect.

[0084] In certain embodiments, the guide sequence is linked to the 3' end of the direct repeat sequence.In certain embodiments, the guide sequence comprises a complementary sequence to the target sequence.

[0085] In certain embodiments, the composition does not comprise crRNA (tracrRNA).

[0086] In certain embodiments, the composition is non-naturally occurring or modified. In certain embodiments, at least one component of the composition is non-naturally occurring or modified. In certain embodiments, the first component is non-naturally occurring or modified; and / or, the second component is non-naturally occurring or modified.

[0087] In certain embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has a sequence represented by 5'-NTN, wherein each N is independently selected from A, G, T or C; for example, the sequence of the PAM is ATG, ATG, GTG, ATA, ATA, GTA, GTA and / or GTG.

[0088] In certain embodiments, when the target sequence is RNA, the target sequence is not restricted by a PAM domain.

[0089] In certain embodiments, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell. In certain embodiments, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0090] In certain embodiments, the target sequence is present in a cell. In certain embodiments, the target sequence is present in the nucleus or in the cytoplasm (e.g., an organelle). In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a prokaryotic cell.

[0091] In certain embodiments, the protein is linked to one or more NLS sequences. In certain embodiments, the conjugate or fusion protein comprises one or more NLS sequences. In certain embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In certain embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the protein.

[0092] In a tenth aspect, the present invention further provides a composition comprising one or more carriers, wherein the one or more carriers comprise:

[0093] (i) a first nucleic acid comprising a nucleotide sequence encoding a protein or fusion protein of the present invention; optionally, the first nucleic acid is operably linked to a first regulatory element; and

[0094] (ii) a second nucleic acid comprising a nucleotide sequence encoding a guide RNA; optionally, the second nucleic acid is operably linked to a second regulatory element;

[0095] in:

[0096] The first nucleic acid and the second nucleic acid are present on the same or different vectors;

[0097] The guide RNA comprises a direct repeat sequence and a guide sequence from the 5' to the 3' direction, and the guide sequence is capable of hybridizing with the target sequence;

[0098] The guide RNA is capable of forming a complex with the effector protein or fusion protein described in (i).

[0099] In certain embodiments, the direct repeat sequence is an isolated nucleic acid molecule as defined in the fourth aspect.

[0100] In certain embodiments, the guide sequence is linked to the 3' end of the direct repeat sequence.In certain embodiments, the guide sequence comprises a complementary sequence to the target sequence.

[0101] In certain embodiments, the composition does not comprise trans-acting crRNA (tracrRNA).

[0102] In certain embodiments, the composition is non-naturally occurring or modified. In certain embodiments, at least one component in the composition is non-naturally occurring or modified.

[0103] In certain embodiments, the first regulatory element is a promoter, such as an inducible promoter.

[0104] In certain embodiments, the second regulatory element is a promoter, such as an inducible promoter.

[0105] In certain embodiments, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has a sequence represented by 5'-NTN, wherein each N is independently selected from A, G, T or C; for example, the sequence of the PAM is ATG, ATG, GTG, ATA, ATA, GTA, GTA and / or GTG.

[0106] In certain embodiments, when the target sequence is RNA, the target sequence is not restricted by a PAM domain.

[0107] In certain embodiments, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell. In certain embodiments, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0108] In certain embodiments, the target sequence is present in a cell. In certain embodiments, the target sequence is present in the nucleus or in the cytoplasm (e.g., an organelle). In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a prokaryotic cell.

[0109] In certain embodiments, the protein is linked to one or more NLS sequences. In certain embodiments, the conjugate or fusion protein comprises one or more NLS sequences. In certain embodiments, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In certain embodiments, the NLS sequence is fused to the N-terminus or C-terminus of the protein.

[0110] In certain embodiments, a type of vector is a plasmid, which refers to a circular double-stranded DNA loop in which other DNA fragments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which virally derived DNA or RNA sequences are present in a vector for packaging viruses (for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by viruses for transfection into a host cell. Some vectors (for example, bacterial vectors and additional mammalian vectors with bacterial replication origins) can replicate autonomously in the host cell into which they are introduced. Other vectors (for example, non-additional mammalian vectors) are integrated into the genome of the host cell after introducing the host cell, and thus replicate together with the host genome. Moreover, some vectors can instruct the expression of their operably connected genes. Such a vector is referred to as an "expression vector" at this. The common expression vector used in recombinant DNA technology is typically a plasmid form.

[0111] Recombinant expression vectors may contain the nucleic acid molecules of the present invention in a form suitable for nucleic acid expression in a host cell, which means that these recombinant expression vectors contain one or more regulatory elements selected based on the host cell to be used for expression, which are operably linked to the nucleic acid sequence to be expressed.

[0112] Delivery and delivery compositions

[0113] The proteins, conjugates, fusion proteins of the present invention, the isolated nucleic acid molecules of the fourth aspect, the complexes of the present invention, the isolated nucleic acid molecules of the sixth aspect, the vectors of the seventh aspect, and the compositions of the ninth and tenth aspects can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipofection, nucleofection, microinjection, sonoporation, gene gun, calcium phosphate-mediated transfection, cationic transfection, lipofection, dendritic transfection, heat shock transfection, nucleofection, magnetofection, lipofection, puncture transfection, optical transfection, agent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial virions, and the like.

[0114] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: a protein, a conjugate, a fusion protein of the present invention, an isolated nucleic acid molecule as described in the fourth aspect, a complex of the present invention, an isolated nucleic acid molecule as described in the sixth aspect, a vector as described in the seventh aspect, a composition as described in the ninth aspect and the tenth aspect.

[0115] In certain embodiments, the delivery vehicle is a particle.

[0116] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, a gene gun, or a viral vector (e.g., a replication-defective retrovirus, a lentivirus, an adenovirus, or an adeno-associated virus).

[0117] Reagent test kit

[0118] In another aspect, the present invention provides a kit comprising one or more of the components described above. In certain embodiments, the kit comprises one or more components selected from the group consisting of a protein, conjugate, fusion protein of the present invention, an isolated nucleic acid molecule as described in the fourth aspect, a complex of the present invention, an isolated nucleic acid molecule as described in the sixth aspect, a vector as described in the seventh aspect, and a composition as described in the ninth and tenth aspects.

[0119] In certain embodiments, the kit of the present invention comprises the composition of the ninth aspect. In certain embodiments, the kit further comprises instructions for using the composition.

[0120] In certain embodiments, the kit of the present invention comprises the composition of the tenth aspect. In certain embodiments, the kit further comprises instructions for using the composition.

[0121] In certain embodiments, the components included in the kits of the present invention may be provided in any suitable container.

[0122] In certain embodiments, the kit further comprises one or more buffers. The buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer and combinations thereof. In certain embodiments, the buffer is alkaline. In certain embodiments, the buffer has a pH from about 7 to about 10.

[0123] In certain embodiments, the kit further comprises one or more oligonucleotides corresponding to a targeting sequence for insertion into a vector so as to operably link the targeting sequence and regulatory elements. In certain embodiments, the kit comprises a homologous recombination template polynucleotide.

[0124] Methods and uses

[0125] In another aspect, the present invention provides a method for modifying a target gene, comprising: contacting the complex described in the fifth aspect, the composition described in the ninth aspect, or the composition described in the tenth aspect with the target gene, or delivering it to a cell containing the target gene; the target sequence is present in the target gene.

[0126] In certain embodiments, the method is used to modify a target gene in vitro or ex vivo. In certain embodiments, the method is not a method for treating a human or animal by therapy. In certain embodiments, the method does not include a step of modifying human germline genetic characteristics.

[0127] In certain embodiments, the target gene is present in a cell. In certain embodiments, the cell is a prokaryotic cell. In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is selected from non-human primate, cattle, pig or rodent cells. In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as poultry or fish. In certain embodiments, the cell is a plant cell, such as a cell of a cultivated plant (such as cassava, corn, sorghum, wheat or rice), algae, tree or vegetable.

[0128] In certain embodiments, the target gene is present in a nucleic acid molecule (e.g., a plasmid) in vitro. In certain embodiments, the target gene is present in a plasmid.

[0129] In certain embodiments, the method results in a target sequence break (e.g., a DNA double-strand break or an RNA single-strand break). In certain embodiments, the break results in a reduction in transcription of the target gene.

[0130] In certain embodiments, the method further comprises: contacting an editing template (e.g., an exogenous nucleic acid) with the target gene, or delivering it to a cell comprising the target gene. In such embodiments, the method repairs the broken target gene by homologous recombination with an editing template (e.g., an exogenous nucleic acid), wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides of the target gene. In certain embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.

[0131] Thus, in certain embodiments, the modification further comprises inserting an editing template (eg, an exogenous nucleic acid) into the break.

[0132] In certain embodiments, the protein, protein truncation, conjugate, fusion protein, isolated nucleic acid molecule, complex, vector or composition is contained in a delivery vehicle.

[0133] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0134] In certain embodiments, the methods are used to modify a cell, cell line, or organism by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

[0135] In another aspect, the present invention provides a method for altering the expression of a gene product, comprising: contacting the complex described in the fifth aspect, the composition described in the ninth aspect, or the composition described in the tenth aspect with a nucleic acid molecule encoding the gene product, or delivering it to a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

[0136] In certain embodiments, the method is used to modify the expression of a gene product in vitro or in vitro. In certain embodiments, the method is not a method for treating a human or animal by therapy. In certain embodiments, the method does not include a step of modifying human germline genetic characteristics.

[0137] In certain embodiments, the nucleic acid molecule is present in a cell. In certain embodiments, the cell is a prokaryotic cell. In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is selected from non-human primate, cattle, pig or rodent cells. In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as poultry or fish. In certain embodiments, the cell is a plant cell, such as a cell of a cultivated plant (such as cassava, corn, sorghum, wheat or rice), algae, tree or vegetable.

[0138] In certain embodiments, the nucleic acid molecule is present in a nucleic acid molecule (e.g., a plasmid) in vitro. In certain embodiments, the nucleic acid molecule is present in a plasmid.

[0139] In certain embodiments, the expression of the gene product is altered (e.g., enhanced or reduced). In certain embodiments, the expression of the gene product is enhanced. In certain embodiments, the expression of the gene product is reduced.

[0140] In certain embodiments, the gene product is a protein.

[0141] In certain embodiments, the protein, protein truncation, conjugate, fusion protein, isolated nucleic acid molecule, complex, vector or composition is contained in a delivery vehicle.

[0142] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, viral vectors (such as replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0143] In certain embodiments, the methods are used to modify a cell, cell line, or organism by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

[0144] In another aspect, the present invention relates to the protein as described in the first aspect, the conjugate as described in the second aspect, the fusion protein as described in the third aspect, the isolated nucleic acid molecule as described in the fourth aspect, the complex as described in the fifth aspect, the isolated nucleic acid molecule as described in the sixth aspect, the vector as described in the seventh aspect, the composition as described in the ninth aspect, the composition as described in the tenth aspect, and the kit of the present invention, and their use in preparing a preparation for nucleic acid editing (e.g., in vitro or ex vivo nucleic acid editing).

[0145] In certain embodiments, the nucleic acid to be edited is present in a cell. In certain embodiments, the cell is a prokaryotic cell or a eukaryotic cell. In certain embodiments, the nucleic acid to be edited is present in an in vitro nucleic acid molecule (e.g., a plasmid).

[0146] In certain embodiments, the nucleic acid editing includes gene or genome editing, such as modifying a gene, knocking out a gene, changing the expression of a gene product, repairing a mutation, and / or inserting a polynucleotide. In certain embodiments, the gene or genome editing does not include a step of modifying human germline genetic characteristics. In certain embodiments, the use is not a method for treating a human or animal by therapy.

[0147] In certain embodiments, the use further comprises repairing the edited target sequence by homologous recombination with an exogenous template polynucleotide, wherein the repair can produce a mutation of the target sequence, including insertion, deletion or substitution of one or more nucleotides.

[0148] In another aspect, the present invention relates to the protein as described in the first aspect, the conjugate as described in the second aspect, the fusion protein as described in the third aspect, the isolated nucleic acid molecule as described in the fourth aspect, the complex as described in the fifth aspect, the isolated nucleic acid molecule as described in the sixth aspect, the vector as described in the seventh aspect, the composition as described in the ninth aspect, the composition as described in the tenth aspect, and the kit of the present invention, and their use in preparing a preparation for: (i) in vitro or ex vivo DNA detection; (ii) editing a target sequence in a target locus to modify an organism or non-human organism (e.g., a prokaryotic organism).

[0149] In certain embodiments, the preparation is used for the detection of single-stranded DNA or double-stranded DNA (eg, the detection of single-stranded or double-stranded DNA in prokaryotic cells).

[0150] In certain embodiments, the DNA detection is used to detect tumors, viruses, or bacteria. Without being limited by theory, it is believed that due to the non-specific cleavage properties of Casσ on single-stranded DNA after target DNA recognition, when target DNA (e.g., tumor-specific marker, virus- or bacteria-specific marker) is present, by adding detectable single-stranded DNA and detecting the non-specific cleavage of the single-stranded DNA, it is possible to detect tumors, Ebola, avian influenza, African swine fever, and other viruses or bacteria.

[0151] In another aspect, the present invention also provides a method for detecting the presence of a target nucleic acid in a sample, comprising the following steps:

[0152] (1) contacting the sample with a labeled DNA probe and any one of the following components: the complex of the present invention, the composition of the ninth aspect and the tenth aspect, or the kit of the present invention;

[0153] wherein the guide sequence contained in the complex, composition or kit is capable of hybridizing to the target nucleic acid, and the DNA probe does not hybridize to the guide sequence;

[0154] In certain embodiments, the DNA probe emits a detectable signal upon cleavage;

[0155] (2) detecting a detectable signal generated by the protein or protein truncation contained in the complex, composition or kit cleaving the DNA probe, thereby determining whether the target nucleic acid is present in the sample.

[0156] In certain embodiments, one end (eg, the 5' end) of the DNA probe is labeled with a fluorescent group, and the other end (eg, the 3' end) is labeled with a quencher group.

[0157] In certain embodiments, the sequence of the target nucleic acid is a sequence obtained from a pathogen. In certain embodiments, the pathogen is selected from a virus, a bacterium, a fungus, a protozoan, a parasite, or any combination thereof.

[0158] In certain embodiments, the sequence of the target nucleic acid is obtained from the genome of a tumor cell.

[0159] The target nucleic acid detected by the present application can be DNA or RNA. Therefore, in certain embodiments, the method further comprises the step of contacting the sample with a reagent for reverse transcription. In certain embodiments, the reagent for reverse transcription is selected from reverse transcriptase, oligonucleotide primers, dNTPs or any combination thereof.

[0160] In certain embodiments, the target nucleic acid is single-stranded or double-stranded. In certain embodiments, the sequence of the target nucleic acid is a DNA or RNA sequence from a prokaryotic or eukaryotic cell; or, the sequence of the target nucleic acid is a non-naturally occurring DNA or RNA sequence.

[0161] In certain embodiments, the detectable signal is determined by one or more methods selected from the group consisting of imaging-based detection, sensor-based detection, color detection, gold nanoparticle-based detection, fluorescence polarization, colloidal phase transition / dispersion, electrochemical detection, and semiconductor-based sensing.

[0162] In certain embodiments, the method further comprises the step of amplifying the target nucleic acid in the sample.

[0163] Cells and cell progeny

[0164] In some cases, the modifications introduced into the cell by the methods of the present invention can cause the cell and its progeny to be altered to improve the production of its biological product (such as an antibody, starch, ethanol or other desired cellular output). In some cases, the modifications introduced into the cell by the methods of the present invention can cause the cell and its progeny to include changes that cause the produced biological product to change.

[0165] Therefore, in another aspect, the present invention also relates to a cell obtained by the method as described above, or a progeny thereof, wherein the cell contains a modification that is not present in its wild type form.

[0166] The present invention also relates to cell products of the cells described above or their progeny.

[0167] The present invention also relates to an in vitro, ex vivo or in vivo cell or cell line or their progeny, which comprises: the protein as described in the first aspect, the conjugate as described in the second aspect, the fusion protein as described in the third aspect, the isolated nucleic acid molecule as described in the fourth aspect, the complex as described in the fifth aspect, the isolated nucleic acid molecule as described in the sixth aspect, the vector as described in the seventh aspect, the composition as described in the ninth aspect, the composition as described in the tenth aspect, the kit or delivery composition of the present invention.

[0168] In certain embodiments, the cell is a prokaryotic cell.

[0169] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, such as a cell of a non-human primate, cattle, sheep, pig, dog, monkey, rabbit, rodent (such as rat or mouse). In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as a cell of poultry (such as chicken), fish or crustacean (such as clams, shrimp). In certain embodiments, the cell is a plant cell, such as a cell or cultivated plant or food crop such as cassava, corn, sorghum, soybean, wheat, oat or rice that a monocot or dicot has, such as an algae, tree or production plant, fruit or vegetable (for example, trees such as citrus trees, nut trees; Solanaceae, cotton, tobacco, tomato, grape, coffee, cocoa, etc.).

[0170] In certain embodiments, the cell is a stem cell or a stem cell line.

[0171] Definition of terms

[0172] Unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, procedures in molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA used herein are conventional procedures widely used in the relevant fields. To facilitate a better understanding of the present invention, definitions and explanations of relevant terms are provided below.

[0173] In the present invention, the expression “Casσ” refers to a Cas effector protein first discovered and identified by the present inventors, which has an amino acid sequence selected from the following:

[0174] (i) the sequence shown in any one of SEQ ID NOs: 1-13;

[0175] (ii) a sequence having one or more amino acid substitutions, deletions or additions compared to the sequence shown in any one of SEQ ID NOs: 1-13 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 and 40 amino acid substitutions, deletions or additions); or

[0176] (iii) a sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the sequence set forth in any one of SEQ ID NOs: 1-13.

[0177] The Casσ of the present invention is a nuclease that binds to and cuts a specific site of a target sequence under the guidance of a guide RNA.

[0178] As used herein, the terms "clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meaning generally understood by those skilled in the art, which generally include transcripts or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements that can guide the activity of the Cas genes. Such transcripts or other elements can include sequences encoding Cas effector proteins and guide RNAs comprising CRISPR RNA (crRNA), as well as trans-acting crRNA (tracrRNA) sequences contained in the CRISPR-Cas9 system, or other sequences or transcripts from the CRISPR locus. In the Casσ-based CRISPR system of the present invention, a tracrRNA sequence is not required.

[0179] As used herein, the terms "Cas effector protein" and "Cas effector enzyme" are used interchangeably and refer to any protein greater than 800 amino acids in length that is present in the CRISPR-Cas system. In some cases, such proteins refer to proteins identified from the Cas locus.

[0180] As used herein, the terms "guide RNA (guide RNA)", "mature crRNA" are used interchangeably and have the meanings generally understood by those skilled in the art. In general, guide RNA can include a direct repeat sequence and a guide sequence (guide sequence), or essentially consist of or consist of a direct repeat sequence and a guide sequence (also referred to as a spacer sequence in the context of an endogenous CRISPR system). In some cases, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target sequence to hybridize with the target sequence and guide the specific binding of the CRISPR / Cas complex to the target sequence. In certain embodiments, when optimally aligned, the degree of complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. It is within the capabilities of those of ordinary skill in the art to determine the optimal alignment. For example, there are publicly available and commercially available alignment algorithms and programs, such as, but not limited to, ClustalW, Smith-Waterman algorithm (Smith-Waterman), Bowtie, Geneious, Biopython, and SeqMan in matlab.

[0181] In some cases, the guide sequence is at least 5, at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 45, or at least 50 nucleotides in length. In some cases, the guide sequence is no more than 50, 45, 40, 35, 30, 25, 24, 23, 22, 21, 20, 15, 10, or fewer nucleotides in length. In certain embodiments, the guide sequence is 10-30, or 15-25, or 15-22, or 19-25, or 19-22 nucleotides in length.

[0182] In some cases, the direct repeat sequence is at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, or at least 70 nucleotides in length. In some cases, the repeat sequence in the same direction is no more than 70, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 15, 10 or less nucleotides in length. In certain embodiments, the repeat sequence in the same direction is 55-70 nucleotides in length, such as 55-65 nucleotides, such as 60-65 nucleotides, such as 62-65 nucleotides, such as 63-64 nucleotides. In certain embodiments, the repeat sequence in the same direction is 15-30 nucleotides in length, such as 15-25 nucleotides, such as 20-25 nucleotides, such as 22-24 nucleotides, such as 23 nucleotides.

[0183] As used herein, the term "CRISPR / Cas complex" refers to a ribonucleoprotein complex formed by the binding of a guide RNA or mature crRNA to a Cas protein, which comprises a guide sequence that hybridizes to a target sequence and binds to the Cas protein. The ribonucleoprotein complex is capable of recognizing and cleaving a polynucleotide that can hybridize to the guide RNA or mature crRNA.

[0184] Therefore, in the case of forming a CRISPR / Cas complex, a "target sequence" refers to a polynucleotide targeted by a guide sequence designed to be targeted, such as a sequence having complementarity with the guide sequence, wherein hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex. Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex. The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located in the nucleus or cytoplasm of the cell. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast. A sequence or template that can be used to recombine into a target locus comprising the target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence". In certain embodiments, the editing template is an exogenous nucleic acid. In certain embodiments, the recombination is homologous recombination.

[0185] In the present invention, the expression "target sequence" or "target polynucleotide" can be any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, it is believed that the target sequence should be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for PAM vary depending on the Cas effector enzyme used, but PAM is typically a 2-5 base pair sequence adjacent to the protospacer sequence (i.e., target sequence). Those skilled in the art will be able to identify PAM sequences for use with a given Cas effector protein. Herein, "specific motif sequence recognized by Cas protein" or "motif sequence" refers to a PAM sequence.

[0186] In some cases, the target sequence or target polynucleotide may include a plurality of disease-related genes and polynucleotides and signal transduction biochemical pathway-related genes and polynucleotides. Non-limiting examples of such target sequences or target polynucleotides include those listed in U.S. Provisional Patent Applications 61 / 736,527 and 61 / 748,427, filed on December 12, 2012 and January 2, 2013, respectively, and International Application No. PCT / US2013 / 074667, filed on December 12, 2013, all of which are incorporated herein by reference.

[0187] In some cases, examples of target sequences or target polynucleotides include sequences associated with signal transduction biochemical pathways, such as signal transduction biochemical pathway-related genes or polynucleotides. Examples of target polynucleotides include disease-associated genes or polynucleotides. "Disease-associated" genes or polynucleotides refer to any genes or polynucleotides that produce transcriptional or translational products at abnormal levels or in abnormal forms in cells derived from disease-affected tissues compared to tissues or cells of non-disease controls. In cases where the altered expression is associated with the emergence and / or progression of the disease, it can be a gene that is expressed at an abnormally high level; alternatively, it can be a gene that is expressed at an abnormally low level. Disease-associated genes also refer to genes with one or more mutations or genetic variations that are directly responsible for or are disequilibrium with one or more genes responsible for the etiology of the disease. The transcribed or translated products can be known or unknown and can be at normal or abnormal levels.

[0188] As used herein, the term "wild type" has the meaning generally understood by those skilled in the art to refer to the typical form of an organism, strain, gene, or characteristic as it exists in nature, as distinguished from mutant or variant forms, which can be isolated from a source in nature and has not been intentionally modified by man.

[0189] As used herein, the terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of human effort. When these terms are used to describe a nucleic acid molecule or polypeptide, they indicate that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is associated in nature or as found in nature.

[0190] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As a further guide, an "orthologue" of a protein as described herein refers to a protein belonging to a different species that performs the same or similar function as the protein to which it is an orthologue.

[0191] As used herein, the term "identity" refers to the match between two polypeptides or between two nucleic acids. When a position in both sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of the two DNA molecules is occupied by adenine, or a position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared x 100. For example, if 6 out of 10 positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (3 out of 6 positions match). Typically, two sequences are compared when they are aligned for maximum identity. Such an alignment can be achieved, for example, by using the method of Needleman et al. (1970) J. Mol. Biol. 48:443-453, which can be conveniently performed using a computer program such as the Align program (DNAstar, Inc.). The percent identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)), which has been incorporated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. In addition, the percent identity between two amino acid sequences can be determined using the Needleman and Wunsch (J Mol. Biol. 48:444-453 (1970)) algorithm, which has been incorporated into the GAP program in the GCG software package (available at www.gcg.com), using either a Blossum 62 matrix or a PAM250 matrix and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6.

[0192] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which a polynucleotide can be inserted. When a vector is capable of expressing a protein encoded by the inserted polynucleotide, it is referred to as an expression vector. A vector can be introduced into a host cell via transformation, transduction, or transfection, allowing the genetic material elements it carries to be expressed in the host cell. Vectors are well known to those skilled in the art and include, but are not limited to, plasmids; phagemids; cosmids; artificial chromosomes, such as yeast artificial chromosomes (YACs), bacterial artificial chromosomes (BACs), or P1-derived artificial chromosomes (PACs); bacteriophages such as lambda phage or M13 phage; and animal viruses. Animal viruses that can be used as vectors include, but are not limited to, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpes viruses (such as herpes simplex virus), poxviruses, baculoviruses, papillomaviruses, and papillomas (such as SV40). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. In addition, a vector may also contain an initiation of replication site.

[0193] As used herein, the term "host cell" refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, fungal cells such as yeast cells or Aspergillus, insect cells such as S2 Drosophila cells or Sf9, or animal cells such as fibroblasts, CHO cells, COS cells, NSO cells, HeLa cells, BHK cells, HEK 293 cells or human cells.

[0194] Those skilled in the art will appreciate that the design of the expression vector may depend on factors such as the choice of the host cell to be transformed, the desired expression level, etc. A vector can be introduced into a host cell to thereby produce a transcript, protein, or peptide, including a protein, fusion protein, isolated nucleic acid molecule, etc. as described herein (e.g., a CRISPR transcript, such as a nucleic acid transcript, protein, or enzyme).

[0195] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, California (1990). In some cases, regulatory elements include those that direct the constitutive expression of a nucleotide sequence in many types of host cells and those that direct the nucleotide sequence to be expressed only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or special cell types (e.g., lymphocytes). In some cases, regulatory elements can also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell-type-specific. In some cases, the term "regulatory element" encompasses enhancer elements such as WPRE; the CMV enhancer; the R-U 5' segment in the LTR of HTLV-I ((Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).

[0196] As used herein, the term "promoter" has a meaning well known to those skilled in the art and refers to a non-coding nucleotide sequence located upstream of a gene that can initiate expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to the promoter.

[0197] As used herein, the term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to the one or more regulatory elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0198] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence by means of traditional Watson-Crick or other non-traditional types. Percent complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues of a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, "substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.

[0199] As used herein, " stringent conditions " for hybridization refer to a nucleic acid with complementarity to a target sequence that primarily hybridizes with the target sequence and does not substantially hybridize to the conditions on a non-target sequence. Stringent conditions are typically sequence-dependent and vary depending on many factors. Generally speaking, the longer the sequence, the higher the temperature at which the sequence-specific hybridization occurs on its target sequence. Non-limiting examples of stringent conditions are described in " Laboratory Techniques In Biochemistry And Molecular Biology - Hybridization With Nucleic Acid Probes " by Tijssen (1993), Part I, Chapter II, " Overview of principles of hybridization and the strategy of nucleic acid probe assay ", Elsevier, New York.

[0200] As used herein, the term "hybridization" refers to a reaction in which one or more polynucleotide reactions form a complex that is stabilized via hydrogen bonding of the bases between the nucleotide residues. Hydrogen bonding can occur by means of Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. The complex can comprise two chains forming a duplex, three or more chains forming a multi-chain complex, a single self-hybridizing chain, or any combination thereof. A hybridization reaction can constitute a step in a broader process (such as the beginning of PCR or the cutting of a polynucleotide via an enzyme). A sequence that can hybridize with a given sequence is referred to as the "complement" of the given sequence.

[0201] As used herein, the term "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.

[0202] As used herein, the term "linker" refers to a linear polypeptide formed by connecting multiple amino acid residues via peptide bonds. The linker of the present invention can be an artificially synthesized amino acid sequence, or a naturally occurring polypeptide sequence, such as a polypeptide having a hinge region function. Such linker polypeptides are well known in the art (see, for example, Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90: 6444-6448; Poljak, RJ et al. (1994) Structure 2: 1121-1123).

[0203] As used herein, the term "treat" refers to treating or curing a disorder, delaying the onset of symptoms of a disorder, and / or delaying the progression of a disorder.

[0204] As used herein, the term "subject" includes, but is not limited to, various animals, such as mammals, such as bovines, equines, ovines, porcines, canines, felines, lagomorphs, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., human) suffers from a disorder (e.g., a disorder caused by a disease-associated gene defect).

[0205] Advantageous Effects of the Invention

[0206] Compared with the prior art, the Cas protein and system of the present invention have significant advantages. For example, the Cas effector protein of the present invention is smaller than Cas9, C2c1, CasY and Cpf1 proteins in molecular size, so the transfection efficiency is better than Cas9, C2c1, CasY and Cpf1 proteins, which can improve the delivery efficiency in eukaryotic cells. For example, when using viral vectors (such as AAV vectors, etc.), it can be used for delivery to eukaryotic cells (such as mammalian cells, human cells, mouse cells, etc.), and can be applied to research and / or clinical applications. In addition, the Cas effector protein of the present invention can perform DNA cutting in eukaryotic organisms, and compared to the FnCpf1 whose PAM domain is 5'-TTN, the Cas protein of the present invention also has a wider PAM recognition site, which is 4 times larger than that of Cas9 or Cas12a.

[0207] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples, but it will be understood by those skilled in the art that the following drawings and examples are intended only to illustrate the present invention and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art based on the following detailed description of the accompanying drawings and preferred embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0208] Figure 1 The PAM structure and analysis results in Example 3.

[0209] Figure 2 This is the result of verifying the in vitro cleavage activity of PAM in Example 3.

[0210] Figure 3 This is the in vivo verification result of the PAM domain in Escherichia coli in Example 3.

[0211] Figure 4 These are the editing activity detection results in human cells in Example 4.

[0212] Sequence information

[0213] Information on the partial sequences involved in the present invention is provided in Table 1 below.

[0214] Table 1: Description of sequences

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226]

[0227]

[0228]

[0229]

[0230]

[0231]

[0232]

[0233]

[0234]

[0235]

[0236]

[0237]

[0238]

[0239]

[0240]

[0241]

[0242]

[0243]

[0244]

[0245]

[0246]

[0247]

[0248]

[0249]

[0250]

[0251]

[0252]

[0253]

[0254]

[0255]

[0256]

[0257]

[0258]

[0259]

[0260] DETAILED DESCRIPTION

[0261] The invention will now be described with reference to the following examples which are intended to illustrate the invention but not to limit it.

[0262] Unless otherwise specified, the experiments and methods described in the embodiments are carried out substantially according to conventional methods well known in the art and described in various references. For example, conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA used in the present invention can be found in Sambrook, Fritsch and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd edition (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (editor F M Ausubel et al., (1987)); METHODS IN ENZYMOLOGY series (Academic Publishing Company): PCR 2: A PRACTICAL METHOD. APPROACH) (MJ MacPherson, BD Hames and GR Taylor, eds. (1995)), Harlow and Lane, eds. (1988) ANTIBODIES, A LABORATORY MANUAL, and ANIMAL CELL CULTURE (RI Freshney, ed. (1987)).

[0263] In addition, if specific conditions are not specified in the examples, the experiments were performed under conventional conditions or the conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments is not specified, they are all conventional products that can be obtained commercially. It is understood that the examples describe the present invention by way of example and are not intended to limit the scope of the present invention. All publications and other references mentioned herein are incorporated herein by reference in their entirety.

[0264] The sources of some of the reagents involved in the following examples are as follows:

[0265] LB liquid medium: 10g tryptone, 5g yeast extract, 10g NaCl, dilute to 1L, and sterilize. If antibiotics are needed, add them after the medium has cooled, to a final concentration of 50μg / mL.

[0266] Chloroform / isoamyl alcohol: Add 10 mL of isoamyl alcohol to 240 mL of chloroform and mix well.

[0267] RNP buffer: 100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg / mL BSA, pH 7.9.

[0268] The prokaryotic expression vectors pET-30a, pUC19, and pACYCDuet-1 were purchased from Beijing Quanshijin Biotechnology Co., Ltd.

[0269] Escherichia coli competent TSC-E03 was purchased from Beijing Qingke Biotechnology Co., Ltd.

[0270] Example 1. Acquisition of Casσ gene and Casσ guide RNA

[0271] 1. CRISPR and gene annotation: Prodigal was used to annotate the microbial genome and metagenomic data from the NCBI and JGI databases to obtain all proteins. Piler-CR was used to annotate the CRISPR loci. All parameters were default.

[0272] 2. Protein filtering: De-redundancy of annotated proteins is performed through sequence consistency, and proteins with completely identical sequences are removed.

[0273] 3. Acquisition of CRISPR-associated proteins: Each CRISPR locus will be extended 10Kb upstream and downstream to identify non-redundant proteins in the CRISPR adjacent region.

[0274] 4. Clustering of CRISPR-associated proteins: Use BLASTP to perform internal pairwise alignments of non-redundant CRISPR-associated proteins, outputting alignments with an Evalue < 1E-10. Use MCL to perform cluster analysis of the BLASTP output to identify CRISPR-associated protein families.

[0275] 5. Identification of CRISPR-enriched protein families: Use BLASTP to align proteins from CRISPR-associated protein families to a non-redundant protein database excluding CRISPR-associated proteins, outputting alignments with an E-value < 1E-10. If the number of homologous proteins found in a non-CRISPR-associated protein database is less than 100%, it indicates that the proteins in this family are enriched in the CRISPR region. This method allows us to identify CRISPR-enriched protein families.

[0276] 6. Protein Function and Domain Annotation: We annotated CRISPR-enriched protein families using the Pfam database, the NR database, and Cas proteins collected from NCBI to identify new CRISPR / Cas protein families. We performed multiple sequence alignments of each CRISPR / Cas family protein using Mafft, followed by conserved domain analysis using JPred and HHpred, identifying protein families containing the RuvC domain.

[0277] On this basis, the inventors obtained a new Cas effector protein, which was named Casσ-1 to Casσ-13, respectively. The protein sequences are shown in SEQ ID NOs: 1-13, and the nucleotide sequences encoding the proteins are shown in SEQ ID NOs: 14-26. The prototype direct repeat sequences corresponding to Casσ-1 to Casσ-13 (the repeat sequences contained in the pre-crRNA) are shown in SEQ ID NOs: 27-39.

[0278] Example 2. Description of the sequence structure of the Casσ gene

[0279] 1. The CRISPR / Casσ sequence fragment was synthesized by Beijing Qingke Biotechnology Co., Ltd. and constructed into the protein expression vector pET-30a(+), and confirmed by first-generation sequencing. Based on the sequencing results, the recombinant plasmid pET-30a+CRISPR / Casσ is described as follows:

[0280] (1) The recombinant plasmid pET-30a+CRISPR / Casσ-1 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 67. In the sequence shown in SEQ ID NO: 67, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2742 are the nucleotide sequence of Casσ-1, and positions 2743 to 2802 are the nucleoplasmin NLS signal peptide.

[0281] (2) The recombinant plasmid pET-30a+CRISPR / Casσ-2 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 68. In the sequence shown in SEQ ID NO: 68, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2901 are the nucleotide sequence of Casσ-2, and positions 2902 to 2961 are the nucleoplasmin NLS signal peptide.

[0282] (3) The recombinant plasmid pET-30a+CRISPR / Casσ-3 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 69. In the sequence shown in SEQ ID NO: 69, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2700 are the nucleotide sequence of Casσ-3, and positions 2701 to 2856 are the nucleoplasmin NLS signal peptide.

[0283] (4) The recombinant plasmid pET-30a+CRISPR / Casσ-4 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 70. In the sequence shown in SEQ ID NO: 70, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 1977 are the nucleotide sequence of Casσ-4, and positions 1978 to 2037 are the nucleoplasmin NLS signal peptide.

[0284] (5) The recombinant plasmid pET-30a+CRISPR / Casσ-5 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 71. In the sequence shown in SEQ ID NO: 71, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2877 are the nucleotide sequence of Casσ-5, and positions 2878 to 2937 are the nucleoplasmin NLS signal peptide.

[0285] (6) The recombinant plasmid pET-30a+CRISPR / Casσ-6 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 72. In the sequence shown in SEQ ID NO: 72, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2796 are the nucleotide sequence of Casσ-6, and positions 2797 to 2856 are the nucleoplasmin NLS signal peptide.

[0286] (7) The recombinant plasmid pET-30a+CRISPR / Casσ-7 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 73. In the sequence shown in SEQ ID NO: 73, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2901 are the nucleotide sequence of Casσ-7, and positions 2902 to 2961 are the nucleoplasmin NLS signal peptide.

[0287] (8) The recombinant plasmid pET-30a+CRISPR / Casσ-8 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 74. In the sequence shown in SEQ ID NO: 74, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2784 are the nucleotide sequence of Casσ-8, and positions 2785 to 2844 are the nucleoplasmin NLS signal peptide.

[0288] (9) The recombinant plasmid pET-30a+CRISPR / Casσ-9 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 75. In the sequence shown in SEQ ID NO: 75, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2757 are the nucleotide sequence of Casσ-9, and positions 2758 to 2817 are the nucleoplasmin NLS signal peptide.

[0289] (10) The recombinant plasmid pET-30a+CRISPR / Casσ-10 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 76. In the sequence shown in SEQ ID NO: 76, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2559 are the nucleotide sequence of Casσ-10, and positions 2560 to 2619 are the nucleoplasmin NLS signal peptide.

[0290] (11) The recombinant plasmid pET-30a+CRISPR / Casσ-11 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 77. In the sequence shown in SEQ ID NO: 77, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2958 are the nucleotide sequence of Casσ-11, and positions 2959 to 3018 are the nucleoplasmin NLS signal peptide.

[0291] (12) The recombinant plasmid pET-30a+CRISPR / Casσ-12 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 78. In the sequence shown in SEQ ID NO: 78, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 3099 are the nucleotide sequence of Casσ-12, and positions 3100 to 3159 are the nucleoplasmin NLS signal peptide.

[0292] (13) The recombinant plasmid pET-30a+CRISPR / Casσ-13 contains an expression cassette, and the expression cassette sequence is shown in SEQ ID NO: 79. In the sequence shown in SEQ ID NO: 79, positions 1 to 27 from the 5' end are the nucleotide sequence of SV40-NLS, positions 28 to 96 are the nucleotide sequence of 3×FLAG, positions 97 to 2559 are the nucleotide sequence of Casσ-13, and positions 2560 to 2619 are the nucleoplasmin NLS signal peptide.

[0293] Example 3. Identification of PAM and DNA cleavage patterns in the CRISPR / Casσ system

[0294] 1. In vitro expression and purification of Casσ protein

[0295] The steps for in vitro expression and purification of Casσ protein are as follows:

[0296] 1. Artificially synthesize the nucleotide sequences shown in SEQ ID NOs: 67-79.

[0297] 2. Introduce the recombinant plasmid pET-30a-CRISPR / Casσ-1 to 13 into Escherichia coli TSC-E03 to obtain recombinant bacteria, which are named TSC-E03-CRISPR / Casσ-1 to 13. Pick a single clone of TSC-E03-CRISPR / Casσ-1 to 13 and inoculate it into 100 mL of LB liquid medium (containing 50 μg / mL kanamycin). Incubate with shaking at 37°C and 200 rpm for 12 h to obtain a culture solution.

[0298] 3. Take the culture solution and inoculate it into 50mL LB liquid medium (containing 50μg / mL kanamycin) at a volume ratio of 1:100. Culture at 37℃ and 200rpm with shaking until the OD 600nm The value was 0.6, and then IPTG was added to a concentration of 1 mM. The culture was shaken at 18°C ​​and 220 rpm for 14 h, and the cells were centrifuged at 4°C and 7000 rpm for 10 min to collect the bacterial precipitate.

[0299] 5. Take the bacterial pellet, add 100 mL of pH 8.0, 100 mM Tris-HCl buffer, resuspend and ultrasonically disrupt (ultrasonic power 600 W, cycle program: disruption 4 s, pause 6 s, total 20 min), then centrifuge at 4°C, 10000 rpm for 10 min, and collect supernatant A.

[0300] 6. Take supernatant A, centrifuge at 4°C, 12000 rpm for 10 min, and collect supernatant B.

[0301] 7. Supernatant B was purified using a nickel column produced by GE (refer to the instructions of the nickel column for the specific purification steps), and then Casσ-1 to Casσ-13 proteins were quantified using a protein quantification kit produced by Thermo Fisher Scientific.

[0302] 2. Transcription and purification of Casσ protein guide RNA:

[0303] 1. Design the templates for guide RNA transcription. The structure of the transcription template is as follows: (1) T7 promoter + prototype direct repeats of Casσ-1 to Casσ-13 (SEQ ID NO: 27-39) + guide sequence (SEQ ID NO: 81). Primers were designed using Primer 5.0 software, ensuring that the forward primer and reward primer had at least 18 bp of overlapping sequence.

[0304] 2. Prepare the following reaction system, pipette gently to mix, centrifuge briefly, and place in a PCR instrument for slow annealing. The PCR system is as follows:

[0305]

[0306] 3. Use MinElute PCR Purification Kit to purify the template. The steps are as follows:

[0307] 1) Add 5 volumes of PB to the PCR product, place a MinElute column on a 2 ml collection tube, let stand at room temperature for 2 minutes, and centrifuge at 12,000 g for 1 minute;

[0308] 2) Discard the waste solution and add 750 μL Buffer PE (remember to add ethanol before use) and incubate at 12,000 g / 2 min.

[0309] 3) Discard the waste solution, add 350 μL Buffer PE, centrifuge at 12000 g for 1 min, discard the waste solution, and centrifuge at 12000 g for 2 min;

[0310] 4) Transfer the MinElute column to a new 1.5 ml centrifuge tube, open the tube lid, and incubate at 65°C for 2 minutes.

[0311] 5) Add 20 μL of preheated EB solution, let it stand for 2 minutes, and then centrifuge at 12,000 g for 1 minute. To improve the recovery rate, the contents of the centrifuge tube can be passed through a MinElute centrifuge column 2-3 times;

[0312] 6) Measure the concentration using Nanodrop and store at -20°C until use.

[0313] 4. Purification of guide RNA: Phenol:chloroform:isoamyl alcohol (25:24:1) extraction to remove DNase I in the system;

[0314] 1) Add 80 μL RNA-free HO to the post-transcription reaction system to adjust the volume to 100 μL;

[0315] 2) Remove 2 ml of Phase Lock Gel (PLG) Heavy and centrifuge at 15,000 g for 2 minutes. Add 100 μL of phenol:chloroform:isoamyl alcohol (25:24:1) and 100 μL of DNAseI-digested RNA. Gently flick the Phase-Lock tube 5-10 times to mix thoroughly. Centrifuge at 15°C / 16,000 g for 12 minutes.

[0316] 3) Take a new RNA-free 1.5ml centrifuge tube and aspirate the supernatant from the previous step into the tube, being careful not to aspirate the gel. Add an equal volume of isopropanol and one-tenth the volume of sodium acetate solution to the supernatant. Mix thoroughly by pipetting with a pipette and place in a -20°C refrigerator for 1 hour or overnight.

[0317] 4) Centrifuge at 4°C / 16,000 g for 30 min, discard the supernatant, add 75% pre-chilled ethanol, pipette and mix the precipitate, centrifuge at 4°C / 16,000 g for 12 min, discard the supernatant, let stand in a fume hood for 2-3 min to dry the ethanol on the RNA surface, add 100 μL of RNA-free HO, and pipette and mix.

[0318] 5. Use Nanodrop to determine the concentration of purified crRNA, and uniformly dilute to 250 ng / μL, dispense into 200 μL PCR centrifuge tubes, and freeze at -80°C for later use.

[0319] 3. In vitro cleavage of Casσ protein and PAM depletion:

[0320] 1. Establishment of double-stranded DNA enzyme digestion system:

[0321] (1) Prepare the following reaction system, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 15 minutes. The DNA cleavage reaction system is as follows:

[0322]

[0323] (2) Add 300 ng of substrate DNA (100 ng / μL), 3 μL of the tube, gently pipette to mix, and briefly centrifuge. Incubate at 37°C for 8 hours.

[0324] (3) Add RNAse and incubate at 37°C for 15 min to fully digest RNA impurities in the system;

[0325] (4) Add proteinase K and incubate at 55°C for 15 min to digest Casσ-1 to Casσ-13 proteins;

[0326] (5) Agarose gel testing.

[0327] The gel running results showed that Casσ-1 can effectively cut double-stranded DNA.

[0328] 2. PAM site identification:

[0329] (1) Prepare the reaction system as in step 1 above, replace the substrate DNA with a plasmid library containing 8 random bases before the target, and place at 37°C for 8 hours. The secondary control sample is a sample with Casσ added but no crRNA added. Three replicates are performed for each protein.

[0330] (2) After the reaction is completed, the reaction sample is column purified and the purified product is used as a template for the construction of a second-generation library. The system and method of library construction are the same as the library construction method in step 2 of the in vivo PAM library consumption of Escherichia coli. The specific operation process is as follows:

[0331] (Each sample corresponds to one R-direction primer and multiple F-direction primers), prepare the following reagents:

[0332]

[0333] Place the prepared reaction system in the PCR instrument and follow the procedure as follows:

[0334] temperature time 98℃ 3min 98℃ 15s 60℃ 30s 72℃ 20s Go to step 2 20 cycles 72℃ 5min 10℃ forever

[0335] 1G sequencing per sample;

[0336] (3) The number of occurrences of the combined PAM sequences in the experimental and control groups was counted separately, and the results were normalized by the number of all PAM sequences in each group. For any PAM sequence, when log2 (normalized value of the control group / normalized value of the experimental group) was greater than 3.5, we considered that this PAM was significantly consumed. We obtained the significantly consumed PAM sequences from all PAM sequences. Furthermore, we used Weblogo to predict the significantly consumed PAM sequences, and finally obtained the PAM domain of Casσ ( Figure 1 ).

[0337] (4) Verification of the PAM library domain: Through the PAM library depletion experiment, we obtained the PAM domain of Casσ-1. In order to verify the rigor of this domain, we set up TTT PAM for in vivo experiments to test the editing activity of Casσ-1 on this PAM. First, we integrated the T7 promoter with the 26nt target of the corresponding PAM site and the sequence of the T7 terminator into the pET30a-Casσ-1 vector, and then co-transformed it with the pACYCDuet-1 plasmid and coated it on the kanamycin and chloramphenicol resistance plates for screening. The double-resistant monoclonal plaques were selected for shaking, and the OD value was 1.0 for IPTG induction for 12 hours. Then, the bacteria before and after induction were gradiently diluted and spotted for observation. If the chloramphenicol gene was edited, the growth on the chloramphenicol resistance plate was poor. Through the experimental results ( Figure 2 and Figure 3 ), we can see that the CRISPR / Casσ system can only effectively edit target sequences with specific PAM domains (for example, TTT), while it has no editing activity on other target sequences (for example, CCC), thus verifying the accuracy of Casσ-1's PAM domain recognition. Through the above experimental results, it is confirmed that Casσ-1 has a broad PAM recognition mode (i.e., NTN; where the two Ns can independently be A, G, T, or C), so Casσ-1 is easier to select targets.

[0338] Example 4. Analysis of Casσ cleavage activity in human cell lines

[0339] The eukaryotic expression vector containing the Casσ-1 gene and the PCR product containing the U6 promoter and guide RNA (containing the prototype direct repeat sequence shown in SEQ ID NO: 27 and the eukaryotic editing guide sequence shown in SEQ ID NO: 82) were transferred into human HELA cells by lipofectamine transfection and cultured at 37 degrees Celsius and 5% carbon dioxide for 72 hours. DNA from all cells was extracted, and a 700bp sequence containing the target site was amplified. The PCR product was connected to the B-simple vector for first-generation sequencing. The sequencing was completed by Thermo Fisher Scientific. The sequencing results were aligned to the AAVS1 gene in the human genome, and it was identified that Casσ-1 can perform double-stranded DNA editing on the target site, thereby causing base deletion ( Figure 4 ).

[0340] Although the specific embodiments of the present invention have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details based on all the teachings published, and these changes are all within the scope of protection of the present invention. The entire invention is given by the appended claims and any equivalents thereof.

Claims

1. An effector protein in a CRISPR / Cas system, the amino acid sequence of which is shown in SEQ ID NO:

1.

2. A conjugate comprising the protein of claim 1 and a modified portion; in, The modified portion is selected from another protein or polypeptide, a detectable label, and any combination thereof; wherein the other protein or polypeptide is selected from an epitope tag, a reporter gene sequence, a nuclear localization signal sequence, a transcription activation domain, a transcription repression domain, a nuclease domain, and any combination thereof.

3. The conjugate according to claim 2, wherein The modification portion is connected to the N-terminus or C-terminus of the protein through a linker, or the modification portion is fused to the N-terminus or C-terminus of the protein.

4. The conjugate of claim 2, which has one or more of the following features: (i) the transcriptional activation domain is VP64; (ii) the transcriptional repression domain is a KRAB domain or a SID domain; (iii) the nuclease domain is Fok1.

5. The conjugate according to claim 2, wherein The conjugate comprises an epitope tag.

6. The conjugate according to claim 2, wherein The conjugate comprises an NLS sequence.

7. The conjugate according to claim 6, wherein The NLS sequence is shown in SEQ ID NO:

53.

8. The conjugate according to claim 7, wherein The NLS sequence is located at the N-terminus or C-terminus of the protein.

9. A fusion protein comprising the protein of claim 1 and another protein or polypeptide; wherein, The additional protein or polypeptide is selected from the group consisting of an epitope tag, a reporter gene sequence, a nuclear localization signal sequence, a transcriptional activation domain, a transcriptional repression domain, a nuclease domain, and any combination thereof.

10. The fusion protein according to claim 9, wherein The additional protein or polypeptide is connected to the N-terminus or C-terminus of the protein via a linker.

11. The fusion protein of claim 9, which has one or more of the following features: (i) the transcriptional activation domain is VP64; (ii) the transcriptional repression domain is a KRAB domain or a SID domain; (iii) the nuclease domain is Fok1.

12. The fusion protein according to claim 11, wherein The fusion protein comprises an epitope tag.

13. The fusion protein according to claim 11, wherein The fusion protein comprises an NLS sequence.

14. The fusion protein according to claim 13, wherein The NLS sequence is shown in SEQ ID NO:

53.

15. The fusion protein according to claim 14, wherein The NLS sequence is located at the N-terminus or C-terminus of the protein.

16. The fusion protein according to claim 9, wherein The fusion protein has the amino acid sequence shown in SEQ ID NO:

54.

17. An isolated nucleic acid molecule, the sequence of which is shown in SEQ ID NO:

27.

18. A composite comprising: (i) a protein component selected from the group consisting of: the protein of claim 1, the conjugate of any one of claims 2 to 8, and the fusion protein of any one of claims 9 to 16; and (ii) a nucleic acid component comprising, from 5' to 3' direction, the isolated nucleic acid molecule of claim 17 and a guide sequence capable of hybridizing to a target sequence, in, The protein component and the nucleic acid component combine with each other to form a complex.

19. The complex according to claim 18, wherein The targeting sequence is linked to the 3' end of the nucleic acid molecule.

20. The complex according to claim 18, wherein The guide sequence comprises a complementary sequence to the target sequence.

21. The complex according to claim 18, wherein The nucleic acid component is the guide RNA in the CRISPR / Cas system.

22. The complex according to claim 18, wherein The nucleic acid molecule is RNA.

23. The complex according to claim 18, wherein The complex does not contain a trans-acting crRNA.

24. An isolated nucleic acid molecule comprising: (i) a nucleotide sequence encoding the protein of claim 1, or the fusion protein of any one of claims 9 to 16; (ii) the nucleotide sequence of the isolated nucleic acid molecule of claim 17; and / or, (iii) a nucleotide sequence comprising (i) and (ii).

25. The isolated nucleic acid molecule of claim 24, wherein The nucleotide sequence described in any one of (i) to (iii) is codon-optimized for expression in prokaryotes or eukaryotic cells.

26. A vector comprising the isolated nucleic acid molecule of claim 24 or 25.

27. A host cell comprising the isolated nucleic acid molecule of claim 24 or 25 or the vector of claim 26.

28. A composition comprising: (i) a first component selected from the group consisting of: the protein of claim 1, the conjugate of any one of claims 2 to 8, the fusion protein of any one of claims 9 to 16, a nucleotide sequence encoding the protein or fusion protein, and any combination thereof; and (ii) a second component, which is the nucleotide sequence of the guide RNA; in, The guide RNA comprises a direct repeat sequence and a guide sequence from the 5' to the 3' direction, and the guide sequence is capable of hybridizing with the target sequence; The guide RNA is capable of forming a complex with the protein, conjugate or fusion protein described in (i).

29. The composition of claim 28, wherein The direct repeat sequence is an isolated nucleic acid molecule as defined in claim 17.

30. The composition of claim 28, wherein The targeting sequence is linked to the 3' end of the direct repeat sequence.

31. The composition of claim 28, wherein The guide sequence comprises a complementary sequence to the target sequence.

32. The composition of claim 28, wherein The composition does not comprise trans-acting crRNA.

33. A composition comprising one or more carriers comprising: (i) a first nucleic acid comprising a nucleotide sequence encoding the protein of claim 1 or the fusion protein of any one of claims 9 to 16; optionally, the first nucleic acid is operably linked to a first regulatory element; and (ii) a second nucleic acid comprising a nucleotide sequence encoding a guide RNA; optionally, the second nucleic acid is operably linked to a second regulatory element; in: The first nucleic acid and the second nucleic acid are present on the same or different vectors; The guide RNA comprises a direct repeat sequence and a guide sequence from the 5' to the 3' direction, and the guide sequence is capable of hybridizing with the target sequence; The guide RNA is capable of forming a complex with the protein or fusion protein described in (i).

34. The composition of claim 33, wherein The direct repeat sequence is an isolated nucleic acid molecule as defined in claim 17.

35. The composition of claim 33, wherein The targeting sequence is linked to the 3' end of the direct repeat sequence.

36. The composition of claim 33, wherein The guide sequence comprises a complementary sequence to the target sequence.

37. The composition of claim 33, wherein The composition does not comprise trans-acting crRNA.

38. The composition of claim 33, wherein The first regulatory element and / or the second regulatory element is a promoter.

39. The composition of claim 38, wherein The promoter is an inducible promoter.

40. The composition of any one of claims 28 to 38, wherein When the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has a sequence represented by 5'-NTN, wherein each N is independently selected from A, G, T or C.

41. The composition of claim 40, wherein The target sequence is a DNA or RNA sequence from a prokaryotic cell or a eukaryotic cell; alternatively, the target sequence is a non-naturally occurring DNA or RNA sequence.

42. The composition according to any one of claims 28 to 39, 41, wherein The protein is linked to one or more NLS sequences, or the conjugate or fusion protein comprises one or more NLS sequences.

43. A kit comprising one or more components selected from the group consisting of: the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the host cell of claim 27, and the composition of any one of claims 28-42.

44. A delivery composition comprising a delivery vector and one or more selected from the following: the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the host cell of claim 27, and the composition of any one of claims 28-42.

45. The delivery composition of claim 44, wherein the delivery vehicle is selected from lipid particles, metal particles, protein particles, exosomes, or viral vectors.

46. The delivery composition of claim 44, wherein the delivery vector is selected from a replication-defective retrovirus, a lentivirus, an adenovirus, or an adeno-associated virus.

47. A method for modifying a target gene for purposes other than disease diagnosis and treatment, comprising: The complex according to any one of claims 18 to 23 or the composition according to any one of claims 28 to 42 is contacted with the target gene, or delivered to a cell comprising the target gene; the target sequence is present in the target gene.

48. The method of claim 47, wherein the target gene exists in a cell, or the target gene exists in a nucleic acid molecule in vitro.

49. The method of claim 47, wherein the cell is a prokaryotic cell.

50. The method of claim 47, wherein the cell is a eukaryotic cell.

51. The method of claim 47, wherein the cell is selected from the group consisting of an animal cell and a plant cell.

52. The method of claim 47, wherein the modification is cleavage of the target sequence.

53. The method of claim 47, which results in DNA double-strand breaks.

54. A method for altering the expression of a gene product for purposes other than disease diagnosis and treatment, comprising: The complex of any one of claims 18 to 23 or the composition of any one of claims 28 to 42 is contacted with a nucleic acid molecule encoding the gene product, or delivered to a cell comprising the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

55. The method of claim 54, wherein the nucleic acid molecule is present in a cell.

56. The method of claim 54, wherein the cell is a prokaryotic cell.

57. The method of claim 54, wherein the cell is a eukaryotic cell.

58. The method of claim 54, wherein the cell is selected from the group consisting of an animal cell and a plant cell.

59. The method of claim 54, wherein the nucleic acid molecule is present in a nucleic acid molecule in vitro.

60. The method of claim 54, wherein expression of the gene product is altered.

61. The method of claim 54, wherein the gene product is a protein.

62. The method of any one of claims 47-61, wherein the protein, conjugate, fusion protein, complex, vector or composition is contained in a delivery vehicle.

63. The method of claim 62, wherein the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, and viral vectors.

64. The method of any one of claims 47-61, which is used to modify a cell, cell line or organism by altering one or more target sequences in a target gene or a nucleic acid molecule encoding a target gene product.

65. An isolated cell or cell line or progeny thereof, comprising: the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, or the composition of any one of claims 28-42.

66. The cell or cell line or progeny thereof of claim 65, wherein the cell is a prokaryotic cell or a eukaryotic cell.

67. Use of the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the composition of any one of claims 28-42, or the kit of claim 43 in the preparation of a formulation for nucleic acid editing.

68. The use of claim 67, wherein the nucleic acid editing comprises gene editing.

69. The use of claim 68, wherein the gene editing comprises modifying a gene, knocking out a gene, altering the expression of a gene product, repairing a mutation, and / or inserting a polynucleotide.

70. Use of the protein of claim 1, the conjugate of any one of claims 2-8, the fusion protein of any one of claims 9-16, the isolated nucleic acid molecule of claim 17, the complex of any one of claims 18-23, the isolated nucleic acid molecule of claim 24 or 25, the vector of claim 26, the composition of any one of claims 28-42, or the kit of claim 43 in the preparation of a formulation for: (i) in vitro or ex vivo DNA detection; and / or, (ii) editing a target sequence in a target locus to modify an organism.

71. A method for detecting the presence of a target nucleic acid in a sample, comprising the steps of: (1) contacting the sample with a labeled DNA probe and any of the following components: The complex according to any one of claims 18 to 23, the composition according to any one of claims 28 to 42, or the kit according to claim 43; wherein the guide sequence contained in the complex, composition or kit is capable of hybridizing to the target nucleic acid, and the DNA probe does not hybridize to the guide sequence; and the DNA probe emits a detectable signal after being cleaved; (2) detecting a detectable signal generated by cleavage of the DNA probe by the protein or protein truncation contained in the complex, composition or kit, thereby determining whether the target nucleic acid is present in the sample. The method according to claim 71 , wherein one end of the DNA probe is labeled with a fluorescent group, and the other end is labeled with a quencher group.

73. The method of claim 71, wherein The method further comprises the step of contacting the sample with a reagent for reverse transcription.

74. The method of claim 71, wherein The target nucleic acid is single-stranded or double-stranded.

75. The method of claim 71, wherein The target nucleic acid sequence is a DNA or RNA sequence from a prokaryotic cell or a eukaryotic cell; or, the target nucleic acid sequence is a non-naturally occurring DNA or RNA sequence.

76. The method of claim 71, wherein The detectable signal is determined by one or more methods selected from the group consisting of imaging-based detection, sensor-based detection, color detection, gold nanoparticle-based detection, fluorescence polarization, colloidal phase transition / dispersion, electrochemical detection, and semiconductor-based sensing.

77. The method of claim 71, wherein The method further comprises the step of amplifying the target nucleic acid in the sample.

Citation Information

Patent Citations

  • Novel CRISPR-Cas12L enzyme and system

    CN113930410A

  • Crispr-cas12j enzyme and system

    WO2020098772A1