Novel Cas enzyme, system and application

By developing a new Cas enzyme and corresponding CRISPR/Cas system, the existing system's shortcomings in robustness and multi-faceted performance are solved, and more efficient and precise gene editing effects are achieved.

CN120098967APending Publication Date: 2025-06-06SHANDONG SHUNFENG BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332668.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-15
Filing Date
2025-03-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing CRISPR/Cas system has various shortcomings, which limits its application in the field of gene editing, especially in terms of robustness and various good performance.

Method used

A new Cas enzyme was developed and a new CRISPR/Cas system and corresponding gene editing tools were constructed based on the novel Cas enzyme. The novel Cas enzyme has a variety of amino acid sequence variations, is able to recognize different PAM motifs, and combines guide RNA to cleave target sequences in a specific manner.

Benefits of technology

A more robust and efficient gene editing capability is achieved, improving performance in many aspects, including the accuracy of target sites and the reduction of off-target effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention belongs to the field of nucleic acid editing, and particularly relates to the technical field of regularly clustered interval short palindromic repeat (CRISPR). Specifically, the invention provides a novel CRISPR (clustered regularly interspaced short palindromic repeats) enzyme (Cas enzyme), and the Cas enzyme belongs to a class of novel Cas proteins and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application claims the priority of Chinese patent application 2024104452730, entitled “A novel Cas enzyme, system and application”, filed on April 15, 2024. The entire contents of the above application are incorporated herein by reference. Technical Field

[0002] The present invention relates to the field of gene editing, in particular to the field of clustered regularly interspaced short palindromic repeats (CRISPR) technology. Specifically, the present invention screened a new type of Cas enzyme, and developed corresponding gene editing tools and applications based on the new Cas enzyme. Background Art

[0003] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to guide specific binding to target sequences on the genome and cut DNA to produce double-strand breaks, and uses biological non-homologous end joining or homologous recombination for site-specific gene editing.

[0004] The CRISPR / Cas9 system is the most commonly used Type II CRISPR system, which recognizes the PAM motif of 3'-NGG and performs blunt-end cleavage on the target sequence. The CRISPR / Cas Type V system is a newly discovered CRISPR system that has a 5'-TTN motif and performs sticky-end cleavage on the target sequence, such as Cpf1, C2c1, CasX, and CasY. However, the different CRISPR / Cas systems currently available have different advantages and disadvantages. For example, Cas9, C2c1, and CasX all require two RNAs for guide RNA, while Cpf1 only requires one guide RNA and can be used for multiple gene editing. CasX has a size of 980 amino acids, while the common Cas9, C2c1, CasY, and Cpf1 are usually around 1,300 amino acids in size. In addition, the PAM sequences of Cas9, Cpf1, CasX, and CasY are relatively complex and diverse, while C2c1 recognizes the rigorous 5'-TTN, so its target site is easier to predict than other systems, thereby reducing potential off-target effects.

[0005] In summary, given that the currently available CRISPR / Cas systems are limited by some defects, the development of a more robust new CRISPR / Cas system with multi-faceted good performance is of great significance to the development of biotechnology. Summary of the invention

[0006] After a lot of experiments and repeated exploration, the inventors of this application unexpectedly discovered a new type of nuclease (Cas enzyme). Based on this discovery, the inventors developed a new CRISPR / Cas system and a gene editing method and nucleic acid detection method based on the system.

[0007] Cas effector proteins

[0008] On the one hand, the present invention provides a Cas protein, which is an effector protein in the CRISPR / Cas system. In the present invention, it is referred to as Cas-sf6727, Cas-sf6729, Cas-sf6733, Cas-sf6734, Cas-sf6801, Cas-sf6802, Cas-sf6803, Cas-sf6805, Cas-sf6806, Cas-sf6811, Cas-sf6812, Cas-sf6813, Cas-sf6814, Cas-sf6818, Cas-sf6820, Cas-sf6824 and Cas-sf6825, and the amino acid sequences of the above proteins are shown in SEQ ID No.1-17, respectively.

[0009] In one embodiment, the Cas protein amino acid sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to any one of SEQ ID No.1-17, and substantially retains the biological function of the sequence from which it is derived. Preferably, the Cas protein is derived from the same species as Cas-sf6727, Cas-sf6729, Cas-sf6733, Cas-sf6734, Cas-sf6801, Cas-sf6802, Cas-sf6803, Cas-sf6805, Cas-sf6806, Cas-sf6811, Cas-sf6812, Cas-sf6813, Cas-sf6814, Cas-sf6818, Cas-sf6820, Cas-sf6824 or Cas-sf6825.

[0010] In one embodiment, the Cas protein amino acid sequence has one or more amino acid substitutions, deletions or additions compared to any one of SEQ ID No.1-17; and substantially retains the biological function of the sequence from which it is derived; the one or more amino acid substitutions, deletions or additions include 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions. Preferably, the Cas protein is derived from the same species as Cas-sf6727, Cas-sf6729, Cas-sf6733, Cas-sf6734, Cas-sf6801, Cas-sf6802, Cas-sf6803, Cas-sf6805, Cas-sf6806, Cas-sf6811, Cas-sf6812, Cas-sf6813, Cas-sf6814, Cas-sf6818, Cas-sf6820, Cas-sf6824 or Cas-sf6825.

[0011] In some embodiments, the Cas protein of the invention is capable of recognizing a protospacer adjacent motif (PAM), and the target nucleic acid comprises or consists of the PAM.

[0012] In one embodiment, the preferred PAM sequence of Cas-sf6727 is WTAWDK, the preferred PAM sequence of Cas-sf6729 is DNHWWR, the preferred PAM sequence of Cas-sf6733 is TWD, the preferred PAM sequence of Cas-sf6734 is TWWD, the preferred PAM sequence of Cas-sf6801 is WWD, the preferred PAM sequence of Cas-sf6802 is WWA, the preferred PAM sequence of Cas-sf6803 is TTC, and the preferred PAM sequence of Cas-sf680 The preferred PAM sequence of Cas-5 is NVCWYT, the preferred PAM sequence of Cas-sf6811 is DVTWAG, the preferred PAM sequence of Cas-sf6812 is DWTC, the preferred PAM sequence of Cas-sf6813 is TTTWC, the preferred PAM sequence of Cas-sf6814 is TYTTC, the preferred PAM sequence of Cas-sf6818 is TWY, the preferred PAM sequence of Cas-sf6820 is TWWR, and the preferred PAM sequence of Cas-sf6824 is WWR. In the above preferred PAM sequences: R represents A or G; Y represents C or T; M represents A or C; K represents G or T; S represents C or G; W represents A or T; H represents A or C or T; B represents C or G or T; V represents A or C or G; D represents A or G or T; N represents A or C or G or T.

[0013] It is clear to those skilled in the art that the structure of a protein can be changed without adversely affecting its activity and functionality, for example, one or more conservative amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples and embodiments of conservative amino acid substitutions are clear to those skilled in the art. Specifically, the amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted, that is, a non-polar amino acid residue can be substituted for another non-polar amino acid residue, a polar uncharged amino acid residue can be substituted for another polar uncharged amino acid residue, a basic amino acid residue can be substituted for another basic amino acid residue, and an acidic amino acid residue can be substituted for another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. As long as the substitution does not result in the inactivation of the biological activity of the protein, a conservative substitution in which an amino acid is replaced by other amino acids belonging to the same group falls within the scope of the present invention. Therefore, the protein of the present invention may contain one or more conservative substitutions in the amino acid sequence, and these conservative substitutions are preferably generated by substitution according to the following table. In addition, the present invention also covers proteins that also contain one or more other non-conservative substitutions, as long as the non-conservative substitutions do not significantly affect the desired function and biological activity of the protein of the present invention.

[0014] Conservative amino acid substitutions can be performed at one or more predicted non-essential amino acid residues. "Non-essential" amino acid residues are amino acid residues that can be changed (deleted, substituted or replaced) without changing the biological activity, while "essential" amino acid residues are required for biological activity. "Conservative amino acid substitutions" are substitutions in which amino acid residues are replaced by amino acid residues with similar side chains. Amino acid substitutions can be performed in the non-conserved regions of the above-mentioned Cas proteins. In general, such substitutions are not performed on conserved amino acid residues, or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, it will be appreciated by those skilled in the art that functional variants may have fewer conservative or non-conservative changes in conserved regions.

[0015]

[0016]

[0017] It is well known in the art that one or more amino acid residues can be changed (replaced, deleted, truncated or inserted) from the N and / or C terminus of a protein while still retaining its functional activity. Therefore, proteins in which one or more amino acid residues are changed from the N and / or C terminus of a Cas protein while retaining its desired functional activity are also within the scope of the present invention. These changes may include changes introduced by modern molecular methods such as PCR, which includes PCR amplification of a protein coding sequence by means of including an amino acid coding sequence in an oligonucleotide used in PCR amplification to change or extend the protein coding sequence.

[0018] It will be appreciated that proteins can be altered in a variety of ways, including amino acid substitutions, deletions, truncations and insertions, and methods for such manipulations are generally known in the art. For example, amino acid sequence variants of the above-mentioned proteins can be prepared by mutations in the DNA. Other forms of mutagenesis and / or directed evolution can also be accomplished, for example, using known mutagenesis, recombination and / or shuffling methods, in combination with related screening methods, to perform single or multiple amino acid substitutions, deletions and / or insertions.

[0019] Those skilled in the art will appreciate that these minor amino acid changes in the Cas proteins of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domain of the protein, the properties of the polypeptide may be changed, but the polypeptide may retain its activity. If the mutations present are not close to the catalytic domain, active site, or other functional domain, lesser effects may be expected.

[0020] Those skilled in the art can identify the essential amino acids of the Cas protein of the present invention according to methods known in the art, such as site-directed mutagenesis or protein evolution or analysis of a bioinformatics system. The catalytic domain, active site or other functional domain of the protein can also be determined by physical analysis of the structure, such as by the following techniques: such as nuclear magnetic resonance, crystallography, electron diffraction or photoaffinity labeling, combined with mutations of amino acids at putative key sites.

[0021] In one embodiment, the Cas protein contains an amino acid sequence shown in any one of SEQ ID No.1-17.

[0022] In one embodiment, the Cas protein is an amino acid sequence shown in any one of SEQ ID No.1-17.

[0023] In one embodiment, the Cas protein is a derivatized protein having the same biological function as a protein having a sequence shown in any one of SEQ ID No. 1-17.

[0024] The biological functions include, but are not limited to, the activity of binding to the guide RNA, the endonuclease activity, the activity of binding to and cutting a specific site of the target sequence under the guidance of the guide RNA, including but not limited to Cis cutting activity and Trans cutting activity.

[0025] The present invention also provides a fusion protein, which includes the Cas protein as described above and other modified parts.

[0026] In one embodiment, the modifying moiety is selected from another protein or polypeptide, a detectable label, or any combination thereof.

[0027] In one embodiment, the modified portion is selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting portion, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), and a domain having an activity selected from the following: nucleotide deaminase, methylase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity and nucleic acid binding activity; and any combination thereof. The NLS sequence is well known to those skilled in the art, and examples thereof include, but are not limited to, the SV40 large T antigen, EGL-13, c-Myc and TUS protein.

[0028] In one embodiment, the NLS sequence is located at, near or close to a terminus (e.g., the N-terminus, the C-terminus, or both) of the Cas protein of the invention.

[0029] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can select other suitable epitope tags (e.g., purification, detection or tracing).

[0030] The reporter gene sequence is well known to those skilled in the art, and examples thereof include but are not limited to GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.

[0031] In one embodiment, the fusion protein of the present invention comprises a domain capable of binding to a DNA molecule or an intracellular molecule, such as maltose binding protein (MBP), the DNA binding domain (DBD) of Lex A, the DBD of GAL4, and the like.

[0032] In one embodiment, the fusion protein of the invention comprises a detectable label, such as a fluorescent dye, such as FITC or DAPI.

[0033] In one embodiment, the Cas protein of the present invention is optionally coupled, conjugated or fused to the modified portion via a linker.

[0034] In one embodiment, the modification portion is directly linked to the N-terminus or C-terminus of the Cas protein of the present invention.

[0035] In one embodiment, the modified portion is connected to the N-terminus or C-terminus of the Cas protein of the present invention via a linker. Such linkers are well known in the art, and examples thereof include but are not limited to linkers comprising one or more (e.g., 1, 2, 3, 4 or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA or Ava), or PEG, etc.

[0036] The Cas protein, protein derivative or fusion protein of the present invention is not limited by the way it is produced. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.

[0037] Cas protein nucleic acid

[0038] In another aspect, the present invention provides an isolated polynucleotide comprising:

[0039] (a) a polynucleotide sequence encoding the Cas protein or fusion protein of the present invention;

[0040] or,

[0041] (b) a polynucleotide complementary to the polynucleotide described in (a).

[0042] In one embodiment, the nucleotide sequence described in any one of (a)-(b) is codon optimized for expression in prokaryotes. In one embodiment, the nucleotide sequence described in any one of (a)-(b) is codon optimized for expression in eukaryotic cells.

[0043] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.

[0044] Guide RNA (gRNA)

[0045] On the other hand, the present invention provides a gRNA, which includes tracrRNA and crRNA; the crRNA includes a sequence that can pair with the tracrRNA, and the tracrRNA can pair with the pairing region of the crRNA to form a duplex; the crRNA also includes a region that hybridizes with the target sequence (i.e., a targeting sequence of the targeting nucleic acid).

[0046] In one embodiment, the pairing region sequence of the crRNA and tracrRNA is shown in any one of SEQ ID No.18-34.

[0047] In one embodiment, the sequence of the tracrRNA is shown in any one of SEQ ID No.35-53.

[0048] In one embodiment, the gRNA of the present invention (also known as guide RNA or guide RNA) is composed of crRNA and tracrRNA molecules that are partially complementary to form a complex, wherein the crRNA comprises a sequence that is sufficiently complementary to the target sequence to hybridize with the complementary sequence of the target sequence and guide the Cas enzyme to bind to the target sequence in a sequence-specific manner. The gRNA of the present invention includes tracrRNA and crRNA.

[0049] The targeting sequence of the targeting nucleic acid of the present invention or the targeting section of the targeting nucleic acid comprises a nucleotide sequence complementary to the sequence in the target nucleic acid. In other words, the targeting sequence of the targeting nucleic acid of the present invention or the targeting section of the targeting nucleic acid interacts with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the targeting sequence of the targeting nucleic acid or the targeting section of the targeting nucleic acid can be changed, or can be modified to hybridize any desired sequence in the target nucleic acid. The nucleic acid is selected from DNA or RNA.

[0050] The percent complementarity between the targeting sequence of a targeting nucleic acid or the targeting segment of a targeting nucleic acid and the target sequence of a target nucleic acid can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).

[0051] The gRNA of the present invention can form a complex with the Cas protein.

[0052] Carrier

[0053] The present invention also provides a vector, which comprises the Cas protein, isolated nucleic acid molecule or polynucleotide as described above; preferably, it also includes a regulatory element operably linked thereto.

[0054] In one embodiment, the regulatory element is selected from one or more of the following groups: enhancer, transposon, promoter, terminator, leader sequence, polyadenylation sequence, marker gene.

[0055] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.

[0056] In some embodiments, the vector included in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a phage, etc., which are well known to those skilled in the art.

[0057] CRISPR system

[0058] The present invention provides an engineered non-naturally occurring vector system, or a CRISPR-Cas system, which includes a Cas protein or a nucleic acid sequence encoding the Cas protein and a nucleic acid encoding one or more guide RNAs, wherein the guide RNAs include tracrRNA and crRNA.

[0059] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more guide RNAs are artificially synthesized.

[0060] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more guide RNAs do not naturally co-exist.

[0061] The one or more guide RNAs target one or more target sequences in the cell. The one or more target sequences hybridize with the genomic loci of the DNA molecule encoding the one or more gene products, and guide the Cas protein to the genomic loci of the DNA molecule encoding the one or more gene products. After the Cas protein reaches the target sequence position, it modifies, edits or cuts the target sequence, thereby changing or modifying the expression of the one or more gene products.

[0062] The cells of the present invention include one or more of animals, plants or microorganisms.

[0063] In some embodiments, the Cas protein is codon optimized for expression in a cell.

[0064] In some embodiments, the Cas protein directs cleavage of one or both strands at the location of the target sequence.

[0065] The present invention also provides an engineered non-naturally occurring vector system, which may include one or more vectors, wherein the one or more vectors include:

[0066] a) a first regulatory element, which is operably linked to the gRNA,

[0067] b) a second regulatory element, which is operably linked to the Cas protein;

[0068] Components (a) and (b) are located on the same or different carriers of the system.

[0069] The first and second regulatory elements include a promoter (e.g., a constitutive promoter or an inducible promoter), an enhancer (e.g., a 35S promoter or a 35S enhanced promoter), an internal ribosome entry site (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences).

[0070] In some embodiments, the vector in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a phage, etc., which are well known to those skilled in the art.

[0071] In some embodiments, the systems provided herein are in a delivery system. In some embodiments, the delivery system is a nanoparticle, a liposome, an exosome, a microbubble, and a gene gun.

[0072] In one embodiment, the target sequence is a DNA or RNA sequence from a prokaryotic cell or a eukaryotic cell. In one embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0073] In one embodiment, the target sequence is present in a cell. In one embodiment, the target sequence is present in the nucleus or in the cytoplasm (e.g., an organelle). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.

[0074] In one embodiment, the Cas protein is connected to one or more NLS sequences. In one embodiment, the fusion protein comprises one or more NLS sequences. In one embodiment, the NLS sequence is connected to the N-terminus or C-terminus of the protein. In one embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the protein.

[0075] On the other hand, the present invention relates to an engineered CRISPR system, comprising the above-mentioned Cas protein and one or more guide RNAs, wherein the guide RNA includes tracrRNA and crRNA, and the Cas protein is capable of binding to the guide RNA and targeting a target nucleic acid sequence complementary to the targeting sequence.

[0076] Protein-nucleic acid complexes / compositions

[0077] In another aspect, the present invention provides a compound or composition comprising:

[0078] (i) a protein component selected from: the above-mentioned Cas protein, derivatized protein or fusion protein, and any combination thereof; and

[0079] (ii) a nucleic acid component selected from: gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA of the gRNA, or a precursor RNA nucleic acid encoding the gRNA; the gRNA includes tracrRNA and crRNA.

[0080] The protein component and the nucleic acid component are combined with each other to form a complex.

[0081] In one embodiment, the nucleic acid component is a guide RNA in a CRISPR-Cas system.

[0082] In one embodiment, the complex or composition is non-naturally occurring or modified. In one embodiment, at least one component of the complex or composition is non-naturally occurring or modified. In one embodiment, the first component is non-naturally occurring or modified; and / or, the second component is non-naturally occurring or modified.

[0083] Activated CRISPR complex

[0084] On the other hand, the present invention also provides an activated CRISPR complex, the activated CRISPR complex comprising: (1) a protein component selected from: a Cas protein, a derivatized protein or a fusion protein of the present invention, and any combination thereof; (2) a nucleic acid component selected from: a gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA of the gRNA, or a precursor RNA nucleic acid encoding the gRNA; the gRNA includes tracrRNA and crRNA; and (3) a target sequence bound to the gRNA. Preferably, the binding is carried out by binding the target nucleic acid to the targeting sequence of the targeting nucleic acid on the gRNA.

[0085] The terms "activated CRISPR complex", "activated complex" or "ternary complex" used in this article refer to the complex after the Cas protein, gRNA and target nucleic acid in the CRISPR system are bound or modified.

[0086] The Cas protein and gRNA of the present invention can form a binary complex, which is activated when bound to a nucleic acid substrate to form an activated CRISPR complex. The nucleic acid substrate is complementary to the targeting sequence in the gRNA (or referred to as a guide sequence hybridized with the target nucleic acid). In some embodiments, the targeting sequence of the gRNA completely matches the target substrate. In other embodiments, the targeting sequence of the gRNA matches a portion (continuous or discontinuous) of the target substrate.

[0087] In a preferred embodiment, the activated CRISPR complex can exhibit side branch nuclease cleavage activity, which refers to the non-specific cleavage activity or random cleavage activity of the activated CRISPR complex on single-stranded nucleic acids, also known as trans cleavage activity in the art.

[0088] Delivery and delivery compositions

[0089] The Cas protein, gRNA, fusion protein, nucleic acid molecule, vector, system, complex and composition of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipofection, nuclear transfection, microinjection, sonoporation, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetofection, lipofection, puncture transfection, optical transfection, agent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial virions, etc.

[0090] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: the Cas protein, fusion protein, nucleic acid molecule, vector, system, complex and composition of the present invention.

[0091] In one embodiment, the delivery vehicle is a particle.

[0092] In one embodiment, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).

[0093] Host cells

[0094] The present invention also relates to an in vitro, ex vivo or in vivo cell or cell line or their progeny, which comprises: the Cas protein, fusion protein, nucleic acid molecule, protein-nucleic acid complex, activated CRISPR complex, vector, and delivery composition of the present invention.

[0095] In certain embodiments, the cell is a prokaryotic cell.

[0096] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, such as a cell of a non-human primate, a cow, a sheep, a pig, a dog, a monkey, a rabbit, a rodent (such as a rat or a mouse). In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as a cell of a poultry bird (such as a chicken), a fish or a crustacean (such as a clam, a shrimp). In certain embodiments, the cell is a plant cell, such as a cell or a cultivated plant or a food crop such as cassava, corn, sorghum, soybean, wheat, oat or rice, such as algae, a tree or a production plant, a fruit or a vegetable (for example, a tree such as a citrus tree, a nut tree; Solanum, cotton, tobacco, tomato, grape, coffee, cocoa, etc.).

[0097] In certain embodiments, the cell is a stem cell or a stem cell line.

[0098] In certain cases, the host cells of the invention contain genetic or genomic modifications that are not present in their wild type.

[0099] Gene Editing Methods and Applications

[0100] The Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition, activated CRISPR complex or host cell of the present invention can be used for any one or more of the following purposes: targeting and / or editing target nucleic acid; cutting double-stranded DNA, single-stranded DNA or single-stranded RNA; non-specific cutting and / or degradation of side branch nucleic acid; non-specific cutting of single-stranded nucleic acid; nucleic acid detection; detection of nucleic acid in target sample; specific editing of double-stranded nucleic acid; base editing of double-stranded nucleic acid; base editing of single-stranded nucleic acid. In other embodiments, it can also be used to prepare reagents or kits for any one or more of the above purposes.

[0101] The present invention also provides the use of the above-mentioned Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition or activated CRISPR complex in gene editing, gene targeting or gene cutting; or, use in the preparation of reagents or kits for gene editing, gene targeting or gene cutting.

[0102] In one embodiment, the gene editing, gene targeting or gene cleavage is performed inside and / or outside the cell.

[0103] The present invention also provides a method for editing a target nucleic acid, targeting a target nucleic acid, or cutting a target nucleic acid, the method comprising contacting the target nucleic acid with the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition, or the above-mentioned activated CRISPR complex. In one embodiment, the method is to edit the target nucleic acid, target the target nucleic acid, or cut the target nucleic acid in a cell or outside the cell.

[0104] The gene editing or editing of target nucleic acid includes modifying genes, knocking out genes, changing the expression of gene products, repairing mutations, and / or inserting polynucleotides, gene mutations.

[0105] The editing can be performed in prokaryotic cells and / or eukaryotic cells.

[0106] On the other hand, the present invention also provides the use of the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition or the above-mentioned activated CRISPR complex in nucleic acid detection, or in the preparation of a reagent or kit for nucleic acid detection.

[0107] On the other hand, the present invention also provides a method for cutting single-stranded nucleic acids, the method comprising contacting a nucleic acid population with the above-mentioned Cas protein and gRNA, wherein the nucleic acid population comprises a target nucleic acid and a plurality of non-target single-stranded nucleic acids, and the Cas protein cuts the plurality of non-target single-stranded nucleic acids.

[0108] The gRNA is capable of binding to the Cas protein.

[0109] The gRNA is capable of targeting the target nucleic acid.

[0110] The contacting can be inside a cell in vitro, ex vivo or in vivo.

[0111] Preferably, the cleavage of the single-stranded nucleic acid is non-specific cleavage.

[0112] On the other hand, the present invention also provides the use of the above-mentioned Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition or activated CRISPR complex in non-specific cleavage of single-stranded nucleic acid, or in the preparation of a reagent or kit for non-specific cleavage of single-stranded nucleic acid.

[0113] On the other hand, the present invention also provides a kit for gene editing, gene targeting or gene cutting, which comprises the above-mentioned Cas protein, gRNA, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition, the above-mentioned activated CRISPR complex or the above-mentioned host cell.

[0114] It is known in the art that the precursor RNA can be cleaved or processed into the mature guide RNA described above.

[0115] On the other hand, the invention provides the use of the above-mentioned Cas protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition, activated CRISPR complex or host cell in the preparation of a preparation or a kit, wherein the preparation or kit is used for:

[0116] (i) gene or genome editing;

[0117] (ii) target nucleic acid detection and / or diagnosis;

[0118] (iii) editing a target sequence in a target locus to modify an organism or non-human organism;

[0119] (iv) treatment of disease;

[0120] (v) targeting target genes;

[0121] (vi) Cutting the target gene.

[0122] Preferably, the above-mentioned gene or genome editing is performed inside or outside the cell.

[0123] Preferably, the target nucleic acid detection and / or diagnosis is performed in vitro.

[0124] Preferably, the treatment of the disease is the treatment of a condition caused by a defect in the target sequence in the target locus.

[0125] Method for specifically modifying target nucleic acid

[0126] On the other hand, the present invention also provides a method for specifically modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with the above-mentioned Cas protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition or the above-mentioned activated CRISPR complex.

[0127] The specific modification can occur in vivo or in vitro.

[0128] The specific modification can occur inside or outside the cell.

[0129] In some cases, the cell is selected from a prokaryotic cell or a eukaryotic cell, for example, an animal cell, a plant cell, or a microbial cell.

[0130] In one embodiment, the modification refers to a break in the target sequence, such as a single-strand / double-strand break in DNA, or a single-strand break in RNA.

[0131] In some cases, the method further comprises contacting the target nucleic acid with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target nucleic acid.

[0132] In one embodiment, the modification further comprises inserting an editing template (eg, an exogenous nucleic acid) into the break.

[0133] In one embodiment, the method further comprises: contacting the editing template with the target nucleic acid, or delivering it to a cell comprising the target nucleic acid. In this embodiment, the method repairs the broken target gene by homologous recombination with an exogenous template polynucleotide; in some embodiments, the repair results in a mutation, including insertion, deletion, or substitution of one or more nucleotides of the target gene, and in other embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.

[0134] Definition of terms

[0135] In the present invention, unless otherwise specified, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. In addition, the molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics and recombinant DNA procedures used herein are conventional procedures widely used in the corresponding fields. At the same time, in order to better understand the present invention, the definitions and explanations of the relevant terms are provided below.

[0136] In the present invention, amino acid residues can be represented by single letters or three letters, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamine (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine ​​(Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), arginine (Arg, R).

[0137] Cas proteins

[0138] In the present invention, Cas protein, Cas enzyme, and Cas effector protein can be used interchangeably; the inventors have discovered and identified for the first time a Cas effector protein having an amino acid sequence selected from the following:

[0139] (i) any one of SEQ ID No. 1-17;

[0140] (ii) a sequence having one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence shown in any one of SEQ ID No. 1-17; or

[0141] (iii) a sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID Nos. 1-17.

[0142] Nucleic acid cleavage or nucleic acid cleavage herein includes: DNA or RNA breakage (Cis cleavage) in the target nucleic acid produced by the Cas enzyme described herein, DNA or RNA breakage in the side branch nucleic acid substrate (single-stranded nucleic acid substrate) (i.e., non-specific or non-targeted, Trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.

[0143] CRISPR system

[0144] As used herein, the terms "Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meaning generally understood by those skilled in the art, which generally includes transcription products or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcription products or other elements capable of directing the activity of the Cas genes.

[0145] CRISPR / Cas complex

[0146] As used herein, the term "CRISPR / Cas complex" refers to a complex formed by the binding of guide RNA to Cas protein, which includes tracrRNA and crRNA, and the complex is capable of recognizing and cleaving a polynucleotide that can hybridize with the guide RNA.

[0147] Guide RNA (gRNA)

[0148] As used herein, the terms "guide RNA (gRNA)", "guide sequence" are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally speaking, the guide RNA includes tracrRNA and crRNA, or consists essentially of or consists of tracrRNA and crRNA.

[0149] In some cases, the targeting sequence is any polynucleotide sequence that has sufficient complementarity with the target sequence to hybridize with the target sequence and guide the specific binding of the CRISPR / Cas complex to the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between the targeting sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining the optimal alignment is within the capabilities of ordinary technicians in the field. For example, there are publicly available and commercially available alignment algorithms and programs, such as, but not limited to, ClustalW, Smith-Waterman algorithm (Smith-Waterman) in matlab, Bowtie, Geneious, Biopython, and SeqMan.

[0150] Target sequence

[0151] "Target sequence" refers to a polynucleotide targeted by a targeting sequence in a gRNA, such as a sequence having complementarity with the targeting sequence, wherein hybridization between the target sequence and the targeting sequence will promote the formation of a CRISPR / Cas complex (including Cas protein and gRNA). Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex.

[0152] The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the nucleus or cytoplasm of the cell. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast. A sequence or template that can be used to recombine into a target locus comprising the target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence". In one embodiment, the editing template is an exogenous nucleic acid. In one embodiment, the recombination is homologous recombination.

[0153] In the present invention, "target sequence" or "target polynucleotide" or "target nucleic acid" can be any endogenous or exogenous polynucleotide to a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).

[0154] wild type

[0155] As used herein, the term "wild type" has the meaning generally understood by those skilled in the art, which refers to the typical form of an organism, strain, gene, or the characteristics that distinguish it from mutant or variant forms when it exists in nature, which can be isolated from a source in nature and has not been intentionally modified by man.

[0156] Derivatization

[0157] As used herein, the term "derivatization" refers to a chemical modification of an amino acid, polypeptide or protein wherein one or more substituents have been covalently attached to the amino acid, polypeptide or protein. The substituents may also be referred to as side chains.

[0158] A derivatized protein is a derivative of the protein. Generally, derivatization of the protein does not adversely affect the desired activity of the protein (e.g., the activity of binding to the guide RNA, the endonuclease activity, the activity of binding to and cutting a specific site of the target sequence under the guidance of the guide RNA), that is, the derivative of the protein has the same activity as the protein.

[0159] Derivatized Protein

[0160] Also known as "protein derivatives", refers to modified forms of proteins, for example, wherein one or more amino acids of the protein may be deleted, inserted, modified and / or substituted.

[0161] Non-natural

[0162] As used herein, the terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of human effort. When these terms are used to describe a nucleic acid molecule or polypeptide, it means that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is associated in nature or as found in nature.

[0163] Orthologue (ortholog)

[0164] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As a further guide, an "orthologue" of a protein as described herein refers to a protein belonging to a different species that performs the same or similar function as its orthologue.

[0165] Identity

[0166] As used herein, the term "identity" is used to refer to the matching of sequences between two polypeptides or between two nucleic acids. When a position in both sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of the two DNA molecules is occupied by adenine, or a position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared x 100. For example, if 6 out of 10 positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (3 out of a total of 6 positions match). Typically, the comparison is made when the two sequences are aligned to produce maximum identity. Such an alignment can be achieved by using, for example, the method of Needleman et al. (1970) J. Mol. Biol. 48: 443-453, which can be conveniently performed by a computer program such as the Align program (DNAstar, Inc.). The percent identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4: 11-17 (1988)), which has been incorporated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. In addition, the percent identity between two amino acid sequences can be determined using the Needleman and Wunsch (J Mol. Biol. 48: 444-453 (1970)) algorithm, which has been incorporated into the GAP program in the GCG software package (available at www.gcg.com), using a Blossum 62 matrix or a PAM250 matrix and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6.

[0167] Carrier

[0168] The term "vector" refers to a nucleic acid molecule that is capable of transporting another nucleic acid molecule to which it is attached. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, no free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and other various polynucleotides known in the art. The vector can be introduced into a host cell by transformation, transduction, or transfection so that the genetic material elements it carries are expressed in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription start sequences, enhancer sequences, selection elements, and reporter genes. In addition, the vector may also contain a replication initiation site.

[0169] One type of vector is a "plasmid," which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques.

[0170] Another type of vector is a viral vector, wherein a virally derived DNA or RNA sequence is present in a vector for packaging a virus (e.g., a retrovirus, a replication-defective retrovirus, an adenovirus, a replication-defective adenovirus, and an adeno-associated virus). The viral vector also comprises a polynucleotide carried by a virus for transfection into a host cell. Some vectors (e.g., bacterial vectors and episomal mammalian vectors with a bacterial origin of replication) can replicate autonomously in the host cell into which they are introduced.

[0171] Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell after introduction into the host cell, and are replicated together with the host genome. Moreover, some vectors are capable of directing the expression of genes to which they are operably linked. Such vectors are referred to as "expression vectors" herein.

[0172] Host cells

[0173] As used herein, the term "host cell" refers to cells that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, eukaryotic cells such as microbial cells, fungal cells, animal cells and plant cells.

[0174] Those skilled in the art will appreciate that the design of the expression vector may depend on factors such as the choice of the host cell to be transformed, the level of expression desired, and the like.

[0175] Regulatory elements

[0176] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, California (1990). In some cases, regulatory elements include those sequences that direct the constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct the nucleotide sequence to be expressed only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in desired tissues of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or special cell types (e.g., lymphocytes). In some cases, the regulatory element can also direct expression in a timing-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue or cell type specific. In some cases, the term "regulatory element" encompasses enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I ((Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequences between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).

[0177] Promoter

[0178] As used herein, the term "promoter" has a meaning well known to those skilled in the art, and refers to a non-coding nucleotide sequence located upstream of a gene that can initiate expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of a gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell essentially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of a gene product in the cell essentially only when the cell is a cell of the tissue type corresponding to the promoter.

[0179] NLS

[0180] A "nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" a protein for import into the nucleus by nuclear transport, i.e., a protein with an NLS is transported to the nucleus. Typically, an NLS comprises a positively charged Lys or Arg residue exposed on the surface of the protein. Exemplary nuclear localization sequences include, but are not limited to, NLSs from: SV40 large T antigen, EGL-13, c-Myc, and TUS proteins. In some embodiments, the NLS comprises a PKKKRKV sequence. In some embodiments, the NLS comprises an AVKRPAATKKAGQAKKKKLD sequence. In some embodiments, the NLS comprises a PAAKRVKLD sequence. In some embodiments, the NLS comprises an MSRRRKANPTKLSENAKKLAKEVEN sequence. In some embodiments, the NLS comprises a KLKIKRPVK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the sequence KIPIK in the yeast transcription repressor Matα2, and PY-NLS.

[0181] Operatively connected

[0182] As used herein, the term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to the one or more regulatory elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0183] Complementarity

[0184] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence by means of traditional Watson-Crick or other non-traditional types. The percentage of complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues of a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, "substantially complementary" refers to a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.

[0185] Strict conditions

[0186] As used herein, "stringent conditions" for hybridization refer to conditions under which a nucleic acid having complementarity with a target sequence predominantly hybridizes to the target sequence and substantially does not hybridize to non-target sequences. Stringent conditions are typically sequence-dependent and vary depending on many factors. In general, the longer the sequence, the higher the temperature at which the sequence specifically hybridizes to its target sequence.

[0187] Hybridization

[0188] The terms "hybridize" or "complementary" or "substantially complementary" refer to a nucleic acid (e.g., RNA, DNA) comprising a nucleotide sequence that enables it to non-covalently bind, i.e., form base pairs and / or G / U base pairs, "anneal" or "hybridize" with another nucleic acid in a sequence-specific, anti-parallel manner (i.e., nucleic acids specifically bind to complementary nucleic acids).

[0189] Hybridization requires that the two nucleic acids contain complementary sequences, although there may be mismatches between the bases. Suitable conditions for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, which are variables well known in the art. Typically, the length of a hybridizable nucleic acid is 8 nucleotides or more (e.g., 10 nucleotides or more, 12 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 22 nucleotides or more, 25 nucleotides or more, or 30 nucleotides or more).

[0190] It should be understood that the sequence of a polynucleotide need not be 100% complementary to the sequence of its target nucleic acid to specifically hybridize. A polynucleotide may comprise 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% sequence complementarity to the target region in the target nucleic acid sequence with which it hybridizes.

[0191] The hybridization of the target sequence and the gRNA represents that at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or represents that at least 12, 15, 16, 17, 18, 19, 20, 21, 22 or more bases of the nucleic acid sequences of the target sequence and the gRNA can complementarily pair and hybridize to form a complex.

[0192] Express

[0193] As used herein, the term "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. Transcripts and encoded polypeptides may be collectively referred to as "gene products." If the polynucleotide is derived from genomic DNA, expression may include splicing of mRNA in eukaryotic cells.

[0194] Connectors

[0195] As used herein, the term "joint" refers to a linear polypeptide formed by connecting a plurality of amino acid residues via peptide bonds. The joint of the present invention can be an artificially synthesized amino acid sequence, or a naturally occurring polypeptide sequence, such as a polypeptide having a hinge region function. Such joint polypeptides are well known in the art (see, for example, Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90: 6444-6448; Poljak, RJ et al. (1994) Structure 2: 1121-1123).

[0196] treat

[0197] As used herein, the term "treat" refers to treating or curing a disorder, delaying the onset of symptoms of a disorder, and / or delaying the progression of a disorder.

[0198] Subjects

[0199] As used herein, the term "subject" includes, but is not limited to, various animals, plants, and microorganisms.

[0200] animal

[0201] For example, mammals, such as bovines, equines, ovines, porcines, canines, felines, lagomorphs, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., a human) suffers from a disorder (e.g., a disorder caused by a disease-related gene defect).

[0202] plant

[0203] The term "plant" is to be understood as any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development, in particular monocotyledonous or dicotyledonous plants, vegetable crops, including artichokes, Brussels sprouts, rocket, leeks, asparagus, lettuce (e.g., head lettuce, leaf lettuce, romaine lettuce), bok choy, yellow taro, melons (e.g., melon, watermelon, Crenshaw melon, honeydew melon, cantaloupe), rapeseed crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, kale, kale, Chinese cabbage, pak choi), cardoon, carrot, napa, okra, onion, celery, parsley, chickpeas, parsnips, endive, peppers, potatoes, cucurbits (e.g., zucchini, cucumber, courgette, squash, pumpkin), radish, bulb onions, rutabagas, eggplant (also called eggplant), salsify, lettuce, shallots, endive, garlic, spinach, green onions, squash, greens, beets (sugar beets and fodder beets), sweet potatoes, Swiss chard, horseradish, tomatoes, turnips, and spices; fruits and / or vines such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, quince, almonds, chestnuts, hazelnuts, pecans, pistachios, walnuts, citrus, blueberries, boysenberries, erry), cranberries, currants, loganberries, raspberries, strawberries, blackberries, grapes, avocados, bananas, kiwis, persimmons, pomegranates, pineapples, tropical fruits, pomegranates, melons, mangoes, papayas, and lychees; field crops such as clover, alfalfa, evening primrose, silver grass, corn / maize (fodder corn, sweet corn, popcorn), hops, jojoba, peanuts, rice, safflower, small grain cereals (barley, oats, rye, wheat, etc.), sorghum, tobacco, kapok, legumes (beans, lentils, peas beans, soybeans), oil plants (rapeseed, mustard, olives, sunflowers, coconuts, castor oil plants, cocoa beans, peanuts), Arabidopsis, fiber plants (cotton, flax, jute), Lauraceae (cinnamon, camphor), or a plant such as coffee, sugar cane, tea, and natural rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants, and trees such as forests (broadleaf trees and evergreen trees, such as conifers), fruit trees, ornamental trees, and nut-bearing trees, as well as shrubs and other seedlings.

[0204] Advantageous Effects of the Invention

[0205] The present invention discovered a new type of Cas enzyme. The Blast results showed that the Cas enzyme of the present application had low consistency with the reported Cas enzymes and was a new type of Cas protein with broad application prospects.

[0206] Embodiments of the present invention will be described in detail below in conjunction with examples, but it will be appreciated by those skilled in the art that the following examples are only used to illustrate the present invention, rather than to limit the scope of the present invention. According to the following detailed description of preferred embodiments, various objects and advantages of the present invention will become apparent to those skilled in the art. DETAILED DESCRIPTION

[0207] The following examples are only used to describe the present invention, but not to limit the present invention. Unless otherwise specified, the experiment and method described in the embodiment are carried out basically according to the conventional methods well known in the art and described in various references. For example, the conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA used in the present invention can be found in Sambrook, Fritsch and Maniatis, "Molecular cloning: laboratory manual" (MOLECULAR CLONING: A LABORATORY MANUAL), 2nd edition (1989); "Contemporary molecular biology experimental manual" (CURRENT PROTOCOLS IN MOLECULAR BIOLOGY) (FM Ausubel (FM Ausubel) et al. edited, (1987)); "Enzymology" (METHODS IN ENZYMOLOGY) series (Academic Publishing Company): "PCR 2: Practical method" (PCR 2: A PRACTICAL METHOD) APPROACH) (MJ MacPherson, BD Hames, and GR Taylor, eds. (1995)), ANTIBODIES, A LABORATORY MANUAL, eds. Harlow and Lane, eds. (1988), and ANIMAL CELL CULTURE (RI Freshney, ed. (1987)).

[0208] In addition, if the specific conditions are not specified in the examples, they are carried out according to the conventional conditions or the conditions recommended by the manufacturer. If the manufacturer is not specified in the reagents or instruments used, they are all conventional products that can be obtained commercially. It is known to those skilled in the art that the embodiments describe the present invention by way of example and are not intended to limit the scope of the present invention. All public cases and other references mentioned herein are incorporated herein by reference in their entirety.

[0209] Example 1. Acquisition of Cas protein

[0210] The inventors analyzed the metagenome of uncultured organisms and identified 17 new Cas enzymes through redundancy removal and protein clustering analysis. The Blast results showed that the sequence identity of the Cas protein was low with the reported Cas protein, and they were named Cas-sf6727, Cas-sf6729, Cas-sf6733, Cas-sf6734, Cas-sf6801, Cas-sf6802, Cas-sf6803, Cas-sf6805, Cas-sf6806, Cas-sf6811, Cas-sf6812, Cas-sf6813, Cas-sf6814, Cas-sf6818, Cas-sf6820, Cas-sf6824 and Cas-sf6825 respectively in the present invention; the amino acid sequences of the above proteins are shown in Table 1, the pairing region sequences of crRNA and tracrRNA corresponding to the Cas protein are shown in Table 2, and the tracrRNA sequences predicted according to different proteins are shown in Table 3.

[0211] Table 1. Amino acid sequences of Cas proteins

[0212]

[0213]

[0214]

[0215] Table 2. Pairing region sequences of crRNA and tracrRNA corresponding to Cas proteins

[0216]

[0217]

[0218] Table 3. TracrRNA sequences corresponding to Cas proteins

[0219]

[0220]

[0221] Example 2. PAM identification of Cas proteins

[0222] The expression plasmid of the protein in Example 1 was constructed: the nucleic acid sequence was optimized for E. coli codons and then gene synthesis was performed, and the vector PeT28 (a) + was connected to the E. coli expression vector, and the JM23119 promoter was added to start the transcription of the gRNA of the Cas protein to form the vector: PeT28 (a) + -Cas-JM23119-gRNA. The gRNA sequence is shown in Table 4 below, and the underlined target sequence is the target sequence.

[0223] Table 4. gRNA sequences corresponding to Cas proteins

[0224]

[0225]

[0226] Construction of PAM library: Synthetic sequence: CGTGTTTCGTAAAGTCTGGAAACGCGGAA GCCCCCAGCGCTTCAG CGTTC NNNNNNT CCCCTACGTGCTGCTGAAGTTGCCCGCAA, N is a random deoxynucleotide, and the underline is the target sequence. After being filled with Klenow enzyme, it was connected to the pacyc184 vector. After transforming E. coli, the plasmid was extracted to form a PAM library.

[0227] PAM library subtraction experiment: The expression vector PeT28(a)+-Cas-JM23119-gRNA and the PAM library plasmid were co-transformed into competent BL21 (DE3), spread on LB plates containing kanamycin and chloramphenicol, and the cells were collected after overnight culture at 37°C. The bacterial solution concentration was adjusted to OD 600 0.6-0.8, add IPTG 0.2mM, induce at 37℃ for 4h. FastPure EndoFree PlasmidMaxi Kit (vazyme) was used for plasmid extraction to obtain the subtracted PAM library. Primers: PAM-F: GGTCTTCGGTTTCCGTGTT; PAM-R: TGGCGTTGACTCTCAGTCAT. PCR reaction was performed with 30ng / μL plasmid (PAM library) as template to obtain control group samples, and PCR reaction was performed with 30ng / μL plasmid (PAM library after subtraction) as template to obtain experimental group samples. The control group samples and experimental group samples were sent to the second generation sequencing for data analysis. For the 4096 PAM sequences, the number of occurrences in the experimental group and the control group were counted respectively, and the total number of all PAM sequences in each group was used for standardization, and then the significantly consumed PAM sequences were analyzed using Weblogo. The PAM preferences of 15 Cas proteins were obtained, as shown in Table 5, where R represents A or G; Y represents C or T; M represents A or C; K represents G or T; S represents C or G; W represents A or T; H represents A or C or T; B represents C or G or T; V represents A or C or G; D represents A or G or T; N represents A or C or G or T.

[0228] Table 5. PAM preferences corresponding to Cas proteins

[0229]

[0230]

[0231] Example 3. Editing efficiency of Cas proteins in animal cells

[0232] The protein of PAM measured in Example 2 was used to verify the gene editing activity of different Cas proteins in animal cells. The vector pcDNA3.3 was modified to carry ECFP fluorescent protein, and the SV40 NLS-Cas-NLS fusion protein was inserted through the restriction site BsmB1, and the fusion protein was started by the CMV promoter. The Cas protein and the protein CFP were connected with the connecting peptide T2A; the U6 promoter and gRNA sequence were inserted through the restriction site Mfe1. The gRNA sequence is the same as in Table 4 above, and the target sequence 5spacer1: GCAACTTCAGCAGCACGTAGGGGA ; Obtain the intracellular editing vector: pcDNA3.3 CS2.0-Flag-Cas-NLS-T2A-ECFP-gRNA. After the pUC19 vector was modified, the promoter EF-1α initiated the expression of the tdTomato-T2A-GF(5spacer1 24bp)FP gene. After the Cas protein recognized the target site 5spacer1 24bp and edited it, the proportion of GFP-positive cells in CFP and tdTomato double-positive cells was analyzed as the editing efficiency of the Cas protein.

[0233] Plating: 293T cells were plated when the confluence reached 70-80%, and the number of cells seeded in a 12-well plate was 8*10^4 cells / well.

[0234] Transfection: 12-24 hours after plating, add 2ug plasmid (pcDNA3.3:pUC19=1:1) to 100μL opti-MEM; add 4μL diluted plasmid EL Transfection Reagent (TRAN), incubate at room temperature for 15-20 minutes. Add the incubated mixture to the culture medium with cells for transfection. Change to normal culture medium 24 hours after transfection, and perform flow cytometry analysis 48 hours after transfection. The editing efficiency is shown in Table 6 below.

[0235] Table 6. Editing efficiency of Cas proteins in animal cells

[0236] Cas proteins 293T three-color analysis (%) Cas-sf6727 3.69 Cas-sf6729 3.89 Cas-sf6733 7.84 Cas-sf6801 10.41 Cas-sf6802 7.86 Cas-sf6803 3.44 Cas-sf6805 4.21 Cas-sf6811 8.27 Cas-sf6812 1.61 Cas-sf6813 3.55 Cas-sf6814 3.56 Cas-sf6818 2.98 Cas-sf6820 4.92 Cas-sf6824 0.45

[0237] Although the specific embodiments of the present invention have been described in detail, it will be understood by those skilled in the art that various modifications and changes may be made to the details according to all the teachings that have been published, and these changes are within the scope of protection of the present invention. The entire invention is given by the attached claims and any equivalents thereof.

Claims

1. A Cas protein, characterized in that The Cas protein is any one of the following I-III: I. The amino acid sequence of the Cas protein is the same as SEQ ID No.5, SEQ ID No.10, SEQ ID No.6, SEQ ID No.3, SEQ ID No.15, SEQ ID No.8, SEQ ID No.2, SEQ ID No.1, SEQ ID No.7, SEQ ID No.12, SEQ ID No.13, SEQ ID No.14, SEQ ID No.11, SEQ ID No.16, SEQ ID No.4, SEQ ID No.9 or SEQ ID No. 17, having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to any sequence of No. 17, and substantially retains the biological function of the sequence from which it is derived; II. The amino acid sequence of the Cas protein has one or more amino acid substitutions, deletions or additions compared to any of SEQ ID No.5, SEQ ID No.10, SEQ ID No.6, SEQ ID No.3, SEQ ID No.15, SEQ ID No.8, SEQ ID No.2, SEQ ID No.1, SEQ ID No.7, SEQ ID No.12, SEQ ID No.13, SEQ ID No.14, SEQ ID No.11, SEQ ID No.16, SEQ ID No.4, SEQ ID No.9 or SEQ ID No.17, and substantially retains the biological function of the sequence from which it is derived; III. The Cas protein comprises the amino acid sequence shown in any one of SEQ ID No.5, SEQ ID No.10, SEQ ID No.6, SEQ ID No.3, SEQ ID No.15, SEQ ID No.8, SEQ ID No.2, SEQ ID No.1, SEQ ID No.7, SEQ ID No.12, SEQ ID No.13, SEQ ID No.14, SEQ ID No.11, SEQ ID No.16, SEQ ID No.4, SEQ ID No.9 or SEQ ID No.

17.

2. A fusion protein comprising the Cas protein according to claim 1 and other modified parts.

3. An isolated polynucleotide, characterized in that The polynucleotide is a polynucleotide sequence encoding the Cas protein according to claim 1, or a polynucleotide sequence encoding the fusion protein according to claim 2.

4. A carrier, characterized in that The vector comprises the polynucleotide according to claim 3 and a regulatory element operably linked thereto.

5. A CRISPR-Cas system, characterized in that: The system comprises the Cas protein of claim 1 and gRNA, and the gRNA is capable of binding to the Cas protein of claim 1.

6. A composition, characterized in that The composition comprises: (i) a protein component selected from: the Cas protein according to claim 1 or the fusion protein according to claim 2; (ii) a nucleic acid component selected from: gRNA, or a nucleic acid encoding a gRNA, or a precursor RNA of a gRNA, or a precursor RNA nucleic acid encoding a gRNA, wherein the gRNA is capable of binding to the Cas protein of claim 1; The protein component and the nucleic acid component are combined with each other to form a complex.

7. An engineered host cell, characterized in that The host cell comprises the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6.

8. Use of the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7 in any one or more of the following: gene editing, gene targeting, gene cutting, cutting of double-stranded DNA, single-stranded DNA or single-stranded RNA, specifically editing double-stranded nucleic acid, base editing double-stranded nucleic acid, base editing single-stranded nucleic acid.

9. Use of the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7 in the preparation of a preparation or a kit for: gene editing, gene targeting, gene cutting, cutting double-stranded DNA, single-stranded DNA or single-stranded RNA, specifically editing double-stranded nucleic acids, base editing double-stranded nucleic acids, base editing single-stranded nucleic acids.

10. A method for editing, targeting or cutting a target nucleic acid, the method comprising contacting the target nucleic acid with the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the composition of claim 6, or the host cell of claim 7.