A cas protein, a corresponding gene editing system and application thereof

By providing novel Cas proteins and their variant peptides, as well as the CRISPR-Cas system, the differences in target recognition and cleavage between existing systems have been addressed, enabling broader gene editing applications and more efficient target recognition capabilities.

CN119095956BActive Publication Date: 2026-01-23YOLTECH THERAPEUTICS CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202480002205.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-11-10
Filing Date
2024-05-11
Publication Date
2026-01-23
Estimated Expiration
2044-05-11

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems vary in size, guide RNA, and PAM structure, making it difficult to meet diverse application needs.

Method used

A novel Cas protein and its variant peptides, including peptides with specific amino acid sequences, fusion proteins, isolated polynucleotides, and guide RNAs, are provided to construct a CRISPR-Cas system for target recognition and cleavage.

Benefits of technology

It enables broader target identification and cutting capabilities, adapts to diverse gene editing needs, and improves the efficiency and accuracy of gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GDA0005088044750000621
    Figure GDA0005088044750000621
  • Figure GDA0005088044750000631
    Figure GDA0005088044750000631
  • Figure GDA0005088044750000633
    Figure GDA0005088044750000633
Patent Text Reader

Abstract

The present disclosure provides a Cas protein. The present disclosure also provides a composition, a CRISPR-Cas system. Further, the present disclosure also provides methods and uses of using the composition, the CRISPR-Cas system. The present disclosure also provides a cell comprising the Cas protein, the composition, the CRISPR-Cas system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of gene editing, in particular, to a Cas protein and a variant polypeptide thereof, a corresponding gene editing system and applications thereof. BACKGROUND

[0002] Clustered regularly interspaced short palindromic repeats (CRISPR) system is formed by bacteria and archaea to defend against invading phage DNA. The most common one is CRISPR / Cas9 system, Cas9 protein can process pre-crRNA into mature crRNA combined with tracrRNA with the help of trans-encoded small RNA (tracrRNA). Then, it is found that the guide RNA (gRNA) which is a single-stranded chimeric body simulating the crRNA-tracrRNA complex can effectively mediate the recognition and cutting of Cas9 protein to the target. The three bases adjacent to the 3' end of the target must be in the form of 5'-NGG-3', thereby forming the PAM (protospacer adjacent motif) structure required for Cas / crRNA complex to recognize the target.

[0003] However, the different CRISPR / Cas currently exist have different advantages and defects, for example, the size of different Cas proteins, guide RNA and PAM are different.

[0004] Therefore, there is still a need to develop new Cas proteins and CRISPR-Cas systems to meet the diversified application requirements. SUMMARY

[0005] The main purpose of the present application is to provide a new Cas protein and a variant polypeptide thereof, and a CRISPR-Cas system comprising the same, to meet the above application requirements.

[0006] In one aspect, the present disclosure provides a Cas protein selected from the group consisting of:

[0007] (a) a polypeptide having the amino acid sequence shown in SEQ ID NO. 1;

[0008] (b) a polypeptide having at least about 60% (e.g., 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5%) identity to the amino acid sequence set forth in SEQ ID NO. 1;

[0009] (c) a substitution, deletion, or addition of one or more amino acid residues to the amino acid sequence set forth in SEQ ID NO. 1.

[0010] In yet another aspect, the present disclosure provides a fusion protein comprising a Cas protein of the present disclosure; and one or more functional domains.

[0011] In yet another aspect, the present disclosure provides an isolated polynucleotide encoding a Cas protein of the present disclosure or a fusion protein of the present disclosure;

[0012] Preferably, the polynucleotide has been codon-optimized for expression in a eukaryotic cell;

[0013] Preferably, the polynucleotide comprises a nucleotide sequence having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity to the nucleotide sequence set forth in any one of SEQ ID NOs. 2, 37, 83, 85, 87.

[0014] Preferably, the polynucleotide is a nucleotide sequence set forth in any one of SEQ ID NOs. 2, 37, 83, 85, 87.

[0015] In yet another aspect, the present disclosure provides a guide RNA (gRNA) comprising

[0016] (1) a Direct Repeat (DR) sequence capable of forming a complex with the Cas protein of the first aspect, or the fusion protein of the second aspect;

[0017] (2) a spacer sequence capable of hybridizing to a target sequence of a target DNA, thereby directing the complex to the target DNA.

[0018] Optionally, the spacer sequence is at least 15 nt in length, for example, the spacer sequence is 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nt in length.

[0019] Optionally, the spacer sequence is 15-50 nt in length.

[0020] Optionally, the spacer sequence is 18-41 nt in length.

[0021] Optionally, the spacer sequence is 18-27 nt in length.

[0022] Optionally, the spacer sequence is 18-24 nt in length.

[0023] Optionally, the spacer sequence is 18-22 nt in length.

[0024] Optionally, the 5' end of the direct repeat sequence is linked to the spacer sequence.

[0025] Optionally, the spacer sequence comprises at least 15 contiguous nucleotides of the nucleotide sequence set forth in any one of SEQ ID NOs. 6, 11, 53, 55, 57, 59, 61, 101.

[0026] In yet another aspect, the present disclosure provides an isolated nucleic acid molecule comprising or consisting of a sequence selected from the group consisting of:

[0027] (i) the sequence set forth in SEQ ID NO: 5 or 71;

[0028] (ii) a sequence comprising one or more substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 substitutions, deletions, or additions) of bases as compared to the sequence set forth in SEQ ID NO: 5 or 71;

[0029] (iii) a sequence having at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% sequence identity to the sequence set forth in SEQ ID NO: 5 or 71;

[0030] (iv) a sequence that hybridizes to the sequence recited in any one of (i)-(iii) under stringent conditions; or

[0031] (v) a complement of the sequence recited in any one of (i)-(iii);

[0032] and the sequence recited in any one of (ii)-(v) substantially retains the biological function of the sequence from which it is derived;

[0033] For example, the isolated nucleic acid molecule is RNA.

[0034] For example, the isolated nucleic acid molecule comprises a direct repeat (DR) sequence in a CRISPR / Cas system.

[0035] In yet another aspect, the present disclosure provides a delivery composition comprising a delivery vehicle, and one or more selected from the group consisting of a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a complex of the present disclosure, a vector of the present disclosure, a CRISPR-Cas composition of the present disclosure, or a system of the present disclosure.

[0036] In yet another aspect, the present disclosure provides a complex comprising:

[0037] (i) a protein component selected from the group consisting of a Cas protein of the present disclosure, a fusion protein of the present disclosure, or a combination thereof; and

[0038] (ii) a guide RNA of the present disclosure;

[0039] wherein the protein component and the nucleic acid component are associated with each other to form the complex.

[0040] In yet another aspect, the present disclosure provides a CRISPR-Cas system comprising:

[0041] (i) a Cas protein of the present disclosure or a fusion protein of the present disclosure, or a nucleotide encoding the Cas protein or the fusion protein; and

[0042] (ii) a guide RNA, or a nucleotide encoding the guide RNA;

[0043] the guide RNA comprises:

[0044] (a) a direct repeat (DR) sequence capable of forming a complex with the Cas protein or the fusion protein, and

[0045] (b) a spacer sequence capable of hybridizing to a target sequence of a target DNA, thereby directing the complex to the target DNA.

[0046] In yet another aspect, the present disclosure provides a CRISPR-Cas composition comprising:

[0047] (i) a first component selected from the group consisting of a Cas protein of the present disclosure, a fusion protein of the present disclosure, a nucleotide sequence encoding a Cas protein of the present disclosure or a fusion protein of the present disclosure, and any combination thereof; and

[0048] (ii) a second component, which is one or more guide RNAs of the present disclosure, or a nucleotide sequence encoding the one or more guide RNAs of the present disclosure;

[0049] the guide RNA is capable of forming a complex with the protein or fusion protein described in (i).

[0050] In yet another aspect, the present disclosure provides a CRISPR-Cas system comprising one or more vectors, which comprises:

[0051] (i) a first nucleic acid, which is a nucleotide sequence encoding a Cas protein of the present disclosure or a fusion protein of the present disclosure; optionally the first nucleic acid is operably linked to a first regulatory element; and

[0052] (ii) a second nucleic acid, which encodes a nucleotide sequence comprising a guide RNA of the present disclosure; optionally the second nucleic acid is operably linked to a second regulatory element;

[0053] wherein:

[0054] the first nucleic acid and the second nucleic acid are present on the same or different vectors;

[0055] the guide RNA is capable of forming a complex with the protein or fusion protein described in (i).

[0056] The present disclosure provides a kit comprising one or more components selected from the group consisting of a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a complex of the present disclosure, a vector of the present disclosure, a CRISPR-Cas composition of the present disclosure, or a system of the present disclosure.

[0057] In certain embodiments, the kit further comprises a label or an instruction.

[0058] In certain embodiments, the kit is used for gene or genome editing, disease treatment, targeting a target gene, cleaving one or more of a gene of interest or a non-gene of interest.

[0059] In yet another aspect, the present disclosure provides a host cell comprising a Cas protein of the present disclosure, a fusion protein of the present disclosure, a polynucleotide of the present disclosure, a complex of the present disclosure, a vector of the present disclosure, a composition of the present disclosure, a system of the present disclosure, or a delivery composition of the present disclosure.

[0060] In yet another aspect, the present disclosure provides an enzyme preparation comprising a Cas protein of the present disclosure, a fusion protein of the present disclosure, a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure, or a system of the present disclosure, or a delivery composition of the present disclosure.

[0061] In certain embodiments, the enzyme preparation comprises an injection, and / or a lyophilized preparation.

[0062] In yet another aspect, the present disclosure provides a kit comprising:

[0063] a first container, and a complex of the present disclosure or a composition of the present disclosure or a system of the present disclosure, or a medicament containing the complex of the present disclosure or the composition of the present disclosure or the system of the present disclosure, located in the first container.

[0064] In yet another aspect, the present disclosure provides a kit comprising:

[0065] (a1) a first container, and a Cas protein of the present disclosure or a fusion protein of the present disclosure or a gene encoding the same or an expression vector thereof, or a medicament containing the Cas protein of the present disclosure or the fusion protein of the present disclosure or a gene encoding the same or an expression vector thereof, located in the first container.

[0066] (b1) an optional second container, and a guide RNA of the present disclosure or an expression vector thereof, or a medicament containing the guide RNA of the present disclosure or an expression vector thereof, located in the second container.

[0067] In yet another aspect, the present disclosure provides a method of targeting and editing a target gene or cleaving a target gene, comprising: contacting a Cas protein of the present disclosure or a fusion protein of the present disclosure or a complex of the present disclosure or a composition of the present disclosure or a system of the present disclosure or a delivery composition of the present disclosure or an enzyme preparation of the present disclosure or a kit of the present disclosure with the target gene, or delivering into a cell comprising the target gene, a target sequence being present in the target gene.

[0068] In yet another aspect, the present disclosure provides a method of inducing a change in a cell state, the method comprising contacting a Cas protein of the present disclosure or a fusion protein of the present disclosure or a complex of the present disclosure or a composition of the present disclosure or a system of the present disclosure or a delivery composition of the present disclosure or an enzyme preparation of the present disclosure or a kit of the present disclosure with a target gene in a cell.

[0069] In yet another aspect, the disclosure provides a method of altering expression of a gene product, comprising: contacting a Cas protein of the disclosure, or a fusion protein of the disclosure, or a complex of the disclosure, or a composition of the disclosure, or a system of the disclosure, or a delivery composition of the disclosure, or an enzyme preparation of the disclosure, or the disclosure, or a kit of the disclosure, with a nucleic acid molecule encoding the gene product, or delivering into a cell comprising the nucleic acid molecule, the target sequence being present in the nucleic acid molecule.

[0070] In yet another aspect the disclosure provides a cell obtained from a method of the disclosure, or a progeny thereof, wherein the cell comprises a modification that is not present in its wild type.

[0071] In yet another aspect, the disclosure provides a cell product of a cell of the disclosure, or a progeny thereof.

[0072] In yet another aspect, the disclosure provides an in vitro, ex vivo, or in vivo cell or cell line, or a progeny thereof, comprising: a Cas protein of the disclosure, a fusion protein of the disclosure, a polynucleotide of the disclosure, a complex of the disclosure, a vector of the disclosure, a CRISPR-Cas composition of the disclosure, or a system of the disclosure, or a delivery composition of the disclosure.

[0073] In yet another aspect, the disclosure provides use of a Cas protein of the disclosure, a fusion protein of the disclosure, a polynucleotide of the disclosure, a complex of the disclosure, a vector of the disclosure, a CRISPR-Cas composition of the disclosure, or a system of the disclosure, or a kit of the disclosure, or a delivery composition of the disclosure, or an enzyme preparation of the disclosure, or the disclosure, or a kit of the disclosure, for the manufacture of a medicament or a preparation for nucleic acid editing (e.g., gene or genome editing).

[0074] In yet another aspect, the disclosure provides use of a Cas protein of the disclosure, a fusion protein of the disclosure, a polynucleotide of the disclosure, a complex of the disclosure, a vector of the disclosure, a CRISPR-Cas composition of the disclosure, or a system of the disclosure, or a kit of the disclosure, or a delivery composition of the disclosure, or an enzyme preparation of the disclosure, or the disclosure, or a kit of the disclosure, for the manufacture of a medicament or a preparation for one or more selected from the group consisting of:

[0075] (i) ex vivo gene or genome editing;

[0076] (ii) ex vivo detection of single-stranded DNA;

[0077] (iii) editing a target sequence in a target locus to modify an organism or a non-human organism;

[0078] (iv) treating a disorder caused by a defect in a target sequence in a target locus;

[0079] (v) treating a disorder or disease in a subject in need thereof.

[0080] In yet another aspect, the present disclosure provides a method of detecting the presence or absence of a target nucleic acid molecule in a sample, the method comprising contacting the sample with a Cas protein of the present disclosure, a fusion protein of the present disclosure, or a complex of the present disclosure, a CRISPR-Cas composition of the present disclosure or a system of the present disclosure, a kit of the present disclosure or a delivery composition of the present disclosure or an enzyme preparation of the present disclosure and a non-target sequence, detecting a detectable signal produced by cleavage of the non-target sequence, which does not hybridize to the guide RNA, thereby detecting the target nucleic acid molecule.

[0081] In certain embodiments, the cleavage of the non-target sequence by the protein in the complex or CRISPR-Cas composition or system or delivery composition indicates the presence of the target nucleic acid molecule in the sample; and the non-cleavage of the non-target sequence by the protein in the complex or CRISPR-Cas composition or system or delivery composition indicates the absence of the target nucleic acid molecule in the sample.

[0082] In yet another aspect, the present disclosure provides a method for diagnosing, preventing or treating a disease in a subject in need thereof, the method comprising administering to the subject a vector of the present disclosure, a complex of the present disclosure, a CRISPR-Cas system of the present disclosure, a CRISPR-Cas composition of the present disclosure or a system of the present disclosure or a kit of the present disclosure or a delivery composition of the present disclosure or an enzyme preparation of the present disclosure, wherein the disease is associated with a target DNA, wherein the spacer sequence is capable of hybridizing to a target sequence of the target DNA, the target DNA is modified, thereby diagnosing, preventing or treating the disease.

[0083] In yet another aspect, the present disclosure provides a pharmaceutical composition comprising a vector of the present disclosure, a complex of the present disclosure, a CRISPR-Cas system of the present disclosure, a CRISPR-Cas composition of the present disclosure or a system of the present disclosure or a kit of the present disclosure or a delivery composition of the present disclosure or a host cell of the present disclosure or an enzyme preparation of the present disclosure, and a pharmaceutically acceptable carrier or excipient.

[0084] According to WIPO Standard ST.26, the symbol "t" is used to represent both T in DNA and U in RNA. Thus in the Sequence Listing prepared according to ST.26, a T in the sequence should be read as a U when the sequence is RNA.

[0085] It should be understood that, in the scope of the present application, all combinations between the above-mentioned technical features of the present application and the technical features specifically described hereinafter (e.g. in the examples) can be made and can form new or preferred technical solutions. Due to the limited space, they are not listed one by one here. BRIEF DESCRIPTION OF DRAWINGS

[0086] An understanding of certain features and advantages of the present disclosure will be obtained by reference to the following detailed description and drawings, which sets forth illustrative embodiments in which the principles of the disclosure can be utilized, and in which:

[0087] Figure 1 shows the map of CasY7 recombinant expression plasmid (pET28a-CasY7) and LbCpf1 recombinant expression plasmid (pET28a-LbCpf1). Figure 1A ) and LbCpf1 recombinant expression plasmid map (pET28a-LbCpf1). Figure 1B ) and LbCpf1 recombinant expression plasmid map (pET28a-LbCpf1).

[0088] Figure 2 Figure 4 shows the predicted secondary structure of the RNA transcript of the Direct Repeat (DR) sequence corresponding to CasY7.

[0089] Figure 3 Figure 5 shows the map of Target plasmid used to evaluate the cleavage activity of CasY7 in E. coli.

[0090] Figure 4 Figure 6 shows the comparison of editing efficiency of CasY7 and LbCpf1 in E. coli.

[0091] Figure 5 Figure 7 shows the map of PHK09T plasmid used to express gRNA.

[0092] Figure 6 Figure 8 shows the comparison of editing efficiency of CasY7 and LbCpf1 in HEK293T cells.

[0093] Figure 7 Figure 9 shows four mutants of CasY7 with different point mutations, which eliminate the catalytic activity (cleavage activity) compared to WT CasY7 protein.

[0094] Figure 8 shows the activity test of base editor containing dCasY7 in base editing (A>G), wherein, Figure 8A Figure 10 shows the single base editing (A>G) efficiency at the A2, A4, A13, A15-A17 sites of the TTR gene target sequence. Figure 8B Figure 11 shows the single base editing (A>G) efficiency at the A7, A13, A20 sites of another target sequence of the TTR gene.

[0095] Figure 9 Figure 12 shows the comparison of cleavage activity of different mutants of CasY7 at the target sequence of the TTR gene.

[0096] Figure 10 Figure 13 shows the comparison of cleavage activity of different mutants of CasY7 at the target sequence of the PD-1 gene, wherein WT represents the cleavage activity of wild type CasY7.

[0097] Figure 11 The cleavage activity comparison of CasY7 different mutants targeting the Trac-1 target sequence of TRAC gene is shown.

[0098] Figure 12 The cleavage activity comparison of CasY7 different mutants targeting the Trac-2 target sequence of TRAC gene is shown.

[0099] Figure 13 The cleavage activity comparison of CasY7 different mutants targeting the Trac-3 target sequence of TRAC gene is shown.

[0100] Figure 14 The cleavage activity comparison of CasY7 different mutants targeting the Trac-4 target sequence of TRAC gene is shown.

[0101] Figure 15 The schematic diagram of the predicted secondary structure of the RNA transcript of DR sequence variant DR-1 is shown.

[0102] Figure 16 The cleavage activity comparison of CasY7 targeting the TTR gene target sequence mediated by gRNA containing different DR sequences is shown.

[0103] Figure 17 shows a dual luciferase reporter system, wherein, Figure 17A The composition of the LUxUC sequence is shown, Figure 17B The luciferase reporter plasmid map is shown.

[0104] Figure 18 The editing activity of the fusion protein of CasY7 variant C05440 and T5 targeting the hHAOl gene target sequence is shown, and the results show that the editing activity of C05440-T5 is optimal, much higher than that of T5-C05440 and C05440.

[0105] Figure 19 The on-target cleavage activity of CasY7 mediated Trac gene target, and the off-target cleavage of 28 OT (off-target) targets are shown. DETAILED DESCRIPTION

[0106] The following examples are intended to describe embodiments of the present application and are not intended to limit the present application. Unless otherwise indicated, the experiments and methods described in the examples were performed essentially according to conventional methods well known in the art and described in various references.

[0107] In addition, unless otherwise indicated, the examples are performed under conventional conditions or manufacturer's recommended conditions. The reagents or instruments used are not specified by the manufacturer unless otherwise noted. Those skilled in the art know that the examples describe the present application by way of example and are not intended to limit the scope of the application as claimed. All publications and other references mentioned herein are incorporated by reference in their entirety.

[0108] To enable a better understanding of the present disclosure, certain terms are defined first. As used in this application, unless specifically identified otherwise, each of the following terms shall have the meaning given below. Additional definitions are set forth throughout the application.

[0109] The term

[0110] Unless otherwise described, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. For the purposes of interpreting this specification, the following descriptions will apply, and where appropriate, the singular forms also will include the plural and vice versa. The disclosures of all patents and other publications cited herein are incorporated by reference herein in their entirety. In the event that any definition set forth herein conflicts with the definition set forth in any document incorporated herein by reference, the definition set forth herein prevails.

[0111] As used herein, the singular forms "a," "an," and "the" include both singular and plural referents unless the context clearly dictates otherwise.

[0112] As used herein, the term "about" can mean a value or composition that is within an acceptable error range for the particular value or composition determined by one of ordinary skill in the art that is within a scope of practice in the art. For example, as used herein, the expression "about 100" includes all values between 99 and 101 and all values in between (e.g., 99.1, 99.2, 99.3, 99.4, etc.).

[0113] As used herein, the term "comprising" or "including" can be open, semi-closed or closed. In other words, the term also includes "consisting essentially of" or "consisting of."

[0114] As used herein, the term "optionally" means that the subsequently described event, condition, or substituent can or can not occur, and that the description includes instances where the event or condition occurs and instances where it does not.

[0115] The terms "substantially" and "essentially" as used herein refer to a degree, amount, level, value, number, frequency, percentage, dimension, size, quantity, weight, or length that is about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more, as compared to a reference amount, level, value, number, frequency, percentage, dimension, size, quantity, weight, or length. For example, as used herein, a "substantially identical sequence" can refer to a sequence, polynucleotide, or polypeptide that has a certain degree of identity to a reference sequence.

[0116] The term "and / or" as used herein refers to either or both of the alternatives.

[0117] As used herein, the terms "nucleic acid," "nucleic acid molecule," "polynucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides or their analogs, in either single- or double-stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can exist in a cell-free environment. A polynucleotide can be a gene or a fragment of a gene. A polynucleotide can be DNA or RNA. A polynucleotide can include one or more analogs of a natural nucleotide (e.g., altered backbone, sugar or nucleobase). If present, modifications to the nucleotide structure can be imparted before or after assembly into a polymer. Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acid, heteronomous nucleic acid, morpholino, locked nucleic acid, glycerol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, a number of loci defined as a contiguous stretch of DNA (one locus), exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. A nucleotide sequence can be interspersed with non-nucleotide components. A polynucleotide can include a mixture of nucleotides found in nature and nucleotide analogs (e.g., synthetic nucleotide analogs).

[0118] As used herein, the terms "peptide," "polypeptide," and "protein" are used interchangeably herein and generally refer to a polymer of at least two amino acid residues linked by peptide bonds. The terms do not denote a particular length of the polymer, nor are they intended to be limited by the particular amino acid content of the polymer. The terms also encompass amino acid polymers in which one or more amino acid residues are anabolic or catabolic modifications of a naturally occurring amino acid. In some instances, the polymer can be interspersed with non-amino acids. The terms include amino acid chains of any length including full-length proteins as well as proteins with or without secondary and / or tertiary structure (e.g., domains). The terms also encompass amino acid polymers that have been modified; for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation and any other manipulation, such as conjugation with a labeling component.

[0119] As used herein, the term "identity" refers to the overall relatedness between polymer molecules, e.g., between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. Sequence identity (or homology) is determined by comparing two aligned sequences over a predetermined comparison window, which can be 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the length of the reference nucleotide sequence or protein, and determining the number of positions at which the same residue occurs in both. Usually, this is expressed as a percentage. Using polypeptide sequences as an example, if the type of mutation is one or more of the following: substitution / replacement of one or more amino acids / nucleotides, insertion within the sequence, and deletion within the sequence, the total number of residues is calculated as the larger of the two molecules being compared. If the type of mutation also includes an insertion (extension) at either or both ends of the sequence or a deletion (truncation) at either or both ends of the sequence, the number of amino acids inserted or deleted at either or both ends (e.g., less than 20 inserted or deleted at both ends) is not counted in the total number of residues. In calculating the percentage identity, the sequences being compared are aligned in a way that maximizes the match between the sequences, and gaps in the alignment, if any, are addressed by a particular algorithm. Nucleotide identity is calculated similarly.

[0120] As used herein, the terms "amino acid" and "amino acids" generally refer to natural and unnatural amino acids, including but not limited to modified amino acids and amino acid analogs. Modified amino acids can include natural amino acids that have been chemically modified to include groups or chemical moieties that do not naturally occur on amino acids and unnatural amino acids. Amino acid analogs can refer to amino acid derivatives. The term "amino acid" includes D-amino acids and L-amino acids.

[0121] As used herein, the term "nuclease" refers to a polypeptide that is capable of cleaving the phosphodiester bond between nucleotide subunits of a nucleic acid; the term "endonuclease" refers to a polypeptide that is capable of catalyzing (e.g., cleaving) a phosphodiester bond within a polynucleotide (e.g., DNA or RNA) strand.

[0122] As used herein, the term "exonuclease" refers to a protein or polypeptide that is capable of digesting a nucleic acid (e.g., RNA or DNA) from a free end.

[0123] As used herein, the term“non-native” can generally refer to a nucleic acid or polypeptide sequence that is not found in a naturally occurring nucleic acid or protein. Non-native can refer to an affinity tag. Non-native can refer to a fusion. Non-native can refer to a naturally occurring nucleic acid or polypeptide sequence that includes a mutation, insertion, and / or deletion. A non-native sequence can exhibit and / or encode an activity (e.g., an enzymatic activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) that can also be exhibited by a nucleic acid and / or polypeptide sequence that is fused to the non-native sequence. A non-native nucleic acid or polypeptide sequence can be linked by genetic engineering to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to produce a chimeric nucleic acid and / or polypeptide sequence encoding a chimeric nucleic acid and / or polypeptide.

[0124] As used herein, the term“fusion protein” refers to a hybrid polypeptide that comprises protein domains from at least two different proteins. A protein can be positioned at an amino-terminal (N-terminal) portion or at a carboxy-terminal (C-terminal) protein of a fusion protein, thus forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively. A protein can comprise different domains, e.g., a nucleic acid binding domain (e.g., a gRNA binding domain of a Cas protein that directs binding of the protein to a target site) and a nucleic acid cleavage domain, or a catalytic domain of a nucleic acid editing protein. In some embodiments, there can be a linker between the proteins. In some embodiments, a protein comprises a protein portion (e.g., an amino acid sequence that constructs a nucleic acid binding domain) and an organic compound (e.g., a compound that can act as a nucleic acid cleavage agent). In some embodiments, a protein is complexed or associated with a nucleic acid (e.g., an RNA or DNA).

[0125] As used herein, the term "linker" refers to any means, entity, or moiety for joining two or more entities. In some embodiments, the linker is a covalent linker. In some embodiments, the linker is a non-covalent linker. Examples of covalent linkers include covalent bonds or linker moieties covalently attached to one or more proteins or domains to be connected. In some embodiments, the linker is a non-covalent bond, such as an organometallic bond through a metal center, such as a platinum atom. The joining can be permanent or reversible. For covalent attachment, various functional groups can be used, such as amide groups, including carbonic acid derivatives, ethers, esters (including organic and inorganic esters), amines, carbamates, ureas, and the like. To provide attachment, the domains can be modified to provide coupling sites by oxidation, hydroxylation, substitution, reduction, and the like. Conjugation methods are well known to those of skill in the art and are encompassed for use in the present application. Linker moieties include, but are not limited to, chemical linker moieties, or, for example, peptide linker moieties (linker sequences). The length and type of linker can be designed as desired. In some embodiments, the linker can be selected from an artificially synthesized amino acid sequence or a naturally occurring polypeptide sequence. It will be appreciated that modifications that do not significantly reduce the functionality of the RNA-binding domain and the effector domain are preferred.

[0126] As used herein, the term "complementary" refers to the ability of a nucleobase of a first polynucleotide sequence (e.g., a guide sequence) to base pair with a nucleobase of a second polynucleotide sequence (such as a target sequence) through traditional Watson-Crick base pairing. In some embodiments, a first polynucleotide can be substantially complementary to a second polynucleotide, i.e., have at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementarity to the second polynucleotide. In some embodiments, a first polynucleotide is fully complementary to a second polynucleotide, i.e., has 100% complementarity to the second polynucleotide. "Substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions. The term "stringent conditions" in relation to hybridization refers to conditions under which one nucleic acid having complementarity to a target sequence will hybridize primarily to that target sequence and not to non-target sequences. Stringent conditions are typically sequence dependent, and depend on many factors. Generally, the longer the sequence, the higher the temperature at which the sequence will specifically hybridize to its target sequence. "Hybridize" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues. The complex can comprise two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these. The hybridization reaction can constitute a step in a more extensive process, such as the initiation of PCR, or cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing to a given sequence is referred to as the "complement" of the given sequence.

[0127] As used herein, the term "regulatory element" refers to a DNA sequence that controls or influences one or more aspects of transcription and / or expression, and is intended to include promoters, enhancers, silencers, transcriptional termination signals, internal ribosome entry sites (IRES), protein degradation signals, and other expression control elements (e.g., such as polyadenylation signals and poly-U sequences, such regulatory elements are described in, for example, Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990)). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells, as well as those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Regulatory elements can also direct expression in a time-dependent manner, e.g., in a cell cycle-dependent or developmental stage-dependent manner, which can or can not be tissue- or cell type-specific.

[0128] As used herein, the term "domain" or "protein domain" refers to a portion of a protein sequence that can exist and function independently of the remainder of the protein chain.

[0129] As used herein, the term "operably linked" refers to functional linkage between a regulatory sequence and a target nucleic acid sequence, thereby inducing expression of the latter (e.g., in an in vitro transcription / translation system or in a host cell when the vector has been introduced into the host cell). For example, a first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence.

[0130] As used herein, the term "wild type" is a term understood by the skilled person and means the typical form of an organism, strain, gene or feature as it occurs in nature, as distinguished from a mutated or variant form. Thus, as used herein, when an amino acid or nucleotide sequence refers to a wild type sequence, a variant refers to a variant of the sequence, e.g., comprising a substitution, deletion, insertion. For example, a "wild type Cas protein" of the present disclosure refers to a Cas protein that occurs in nature, which has not been artificially modified, the nucleotide of which can be obtained by genetic engineering techniques, such as genome sequencing, polymerase chain reaction (PCR), etc., and the amino acid sequence of which can be deduced from the nucleotide sequence.

[0131] As used herein, the terms "parent," "parent polypeptide," and "parent sequence" refer to the original polypeptide to which alterations are made to produce the variant polypeptides of the present application (e.g., a reference polypeptide or starting polypeptide).

[0132] As used herein, the term "variant polypeptide" refers to a polypeptide comprising alterations (e.g., but not limited to, substitutions, insertions, deletions, additions, and / or fusions) at one or more residue positions compared to a parent polypeptide. In some embodiments, a variant polypeptide differs from its reference entity. For example, a variant polypeptide can differ from a reference polypeptide due to one or more differences in amino acid sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, etc.) covalently attached to the polypeptide backbone. In some embodiments, a variant polypeptide displays overall sequence identity to a reference polypeptide (e.g., a Cas protein of the disclosure) that is at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. Alternatively or additionally, in some embodiments, a variant polypeptide does not share at least one characteristic sequence element with a reference polypeptide. In some embodiments, a reference polypeptide has one or more biological activities. In some embodiments, a variant polypeptide shares one or more biological activities of a reference polypeptide, e.g., nuclease activity. In some embodiments, a variant polypeptide lacks one or more biological activities of a reference polypeptide. In some embodiments, a variant polypeptide displays reduced levels of one or more biological activities (e.g., nuclease activity, e.g., off-target nuclease activity) compared to a reference polypeptide. In some embodiments, a polypeptide of interest is considered a "variant" of a parent or reference polypeptide if it has the same amino acid sequence as the parent, but with a small number of sequence alterations at particular positions. Typically, less than 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the residues in a variant are substituted compared to the parent or reference polypeptide. In some embodiments, a variant has 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 substituted residues compared to the parent or reference polypeptide. Variants often have a small number (e.g., less than 5, 4, 3, 2, or 1) of substituted functional residues (i.e., residues involved in a particular biological activity). In some embodiments, a variant has no more than 5, 4, 3, 2, or 1 additions or deletions, and often no additions or deletions, compared to the parent or reference polypeptide. Furthermore, any additions or deletions are typically fewer than about 40, about 35, about 30, about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 12, about 11, about 10, about 9, about 8, about 7, about 6 residues, and often fewer than about 5, about 4, about 3, or about 2 residues. In some embodiments, the parent or reference polypeptide is wild type. A variant of a polynucleotide or polypeptide can be naturally occurring, such as an allelic variant, or it can be a variant that has been made by mutagenesis not known to occur naturally.Non-naturally occurring variants of polynucleotides and polypeptides can be prepared by mutagenesis techniques, by direct synthesis, and by other recombinant methods known to those of skill in the art.

[0133] As used herein, the term "conservative amino acid substitution" refers to the substitution of an amino acid with a chemically or functionally similar amino acid without affecting the normal function of the protein, e.g., the interchangeability of amino acid residues with similar side chains: a group of amino acids having aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains consists of serine and threonine; a group of amino acids having amide-containing side chains consists of asparagine and glutamine; a group of amino acids having aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains consists of lysine, arginine, and histidine; a group of amino acids having sulfur-containing side chains consists of cysteine and methionine. In some embodiments, considered conservative amino acid substitutions include: aspartate-glutamate, lysine-arginine-histidine, serine-threonine-asparagine-glutamine, glycine-alanine-valine-leucine-isoleucine, cysteine-methionine-proline, phenylalanine-tyrosine-tryptophan. In some embodiments, considered conservative amino acid substitutions include: valine-leucine-methionine-isoleucine, alanine-serine-threonine.

[0134] As used herein, the term "Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meaning generally understood by those of skill in the art, and generally includes a transcription product or other element associated with the expression of a CRISPR-associated ("Cas") gene, or a transcription product or other element capable of directing the activity of said Cas gene.

[0135] As used herein, the term "Cas protein" is used interchangeably with Cas polypeptide in the present disclosure, and is used in its broadest sense to include a parent or reference Cas protein (e.g., comprising the Cas protein set forth in SEQ ID NO. 1), a derivative or variant thereof, and a functional fragment, such as a nucleic acid binding fragment thereof, including an endonuclease-deficient (inactive) Cas polypeptide (dCas).

[0136] As used herein, the term "complex" refers to a combination of two or more molecules. In some embodiments, a complex includes a polypeptide and a nucleic acid molecule that interact with each other (e.g., bind, contact, adhere). For example, the term "complex" can refer to a combination of a guide RNA and a polypeptide (e.g., a Cas protein). Alternatively, the term "complex" can refer to a combination of a guide RNA, a Cas protein, and a complementary region of a target nucleic acid.

[0137] As used herein, the term "cleave" refers to the breakage of the covalent backbone of a DNA molecule. Cleavage can be carried out by various methods, including but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand cleavage and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two distinct single-strand cleavage events. DNA cleavage can result in the generation of either blunt or sticky ends.

[0138] The term "cleave" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond within the backbone of a target nucleotide sequence, which can result in a single- or double-stranded break within the target sequence. For example, a Cas protein of the disclosure, or a variant polypeptide thereof, can cleave nucleotides within a polynucleotide as an endonuclease, or can be an exonuclease (removing consecutive nucleotides from the end (5' and / or 3' end) of a polynucleotide). In some embodiments, cleavage of a target polynucleotide by a Cas protein of the disclosure, or a variant polypeptide thereof, can result in staggered breaks or blunt ends.

[0139] As used herein, the term "protospacer adjacent motif" or "PAM" refers to a short sequence (or motif) adjacent to a protospacer sequence on a non-target strand of a target nucleic acid recognized by a CRISPR complex. The target nucleic acid is double-stranded DNA (dsDNA), one strand comprising a target sequence adjacent to the PAM and is referred to as the "PAM strand" (e.g., non-target strand or non-spacer complement strand), while the other complementary strand is referred to as the "non-PAM strand" (e.g., target strand or spacer complement strand). As used herein, the term "adjacent" includes instances in which the RNA guide of the complex specifically binds to, interacts with, or associates with the target sequence immediately adjacent to the PAM. In such instances, there are no nucleotides between the target sequence and the PAM. The term "adjacent" also includes instances in which there are a small number (e.g., 1, 2, 3, 4, or 5) of nucleotides between the target sequence bound by the targeting moiety and the PAM.

[0140] As used herein, the terms “guide nucleic acid,” “RNA guide,” “RNA guide sequence,” “guide RNA (gRNA),” “single guide RNA (sgRNA)” are used interchangeably to refer to a nucleic acid-based molecule capable of forming a complex with a CRISPR-Cas protein (e.g., a Cas protein of the present disclosure) and comprises a sequence (e.g., a guide sequence) sufficiently complementary to a target nucleic acid to hybridize to the target nucleic acid and direct the complex to the target nucleic acid, including but not limited to RNA-based molecules, e.g., guide RNAs. A guide nucleic acid can comprise a segment that can be referred to as a “nucleic acid targeting segment” or “nucleic acid targeting sequence,” which can comprise a sub-segment that can be referred to as a “protein binding segment” or “protein binding sequence” or “Cas protein binding segment.” A guide nucleic acid can be a DNA molecule, an RNA molecule, or a DNA / RNA hybrid molecule. A “DNA / RNA hybrid molecule” refers to a nucleic acid comprising one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, whether contiguous or not. However, a “DNA molecule” or “RNA molecule” can also refer to a DNA molecule containing one or more modified or unmodified ribonucleotides, whether contiguous or not, or an RNA molecule containing one or more modified or unmodified deoxyribonucleotides, whether contiguous or not.

[0141] As used herein, the term “activity” refers to biological activity. In some embodiments, nuclease activity includes enzymatic activity, e.g., catalytic ability, of a nuclease. For example, nuclease activity can include nuclease activity. In some embodiments, nuclease activity includes binding activity, e.g., binding activity of a nuclease to an RNA guide and / or a target nucleic acid.

[0142] As used herein, the terms “upstream” and “downstream” refer to relative positions within individual nucleic acid (e.g., DNA) sequences in a nucleic acid molecule. “Upstream” and “downstream” relate to the 5’ to 3’ direction in which RNA transcription occurs. A first sequence is upstream of a second sequence when the 3’ end of the first sequence precedes the 5’ end of the second sequence. A first sequence is downstream of a second sequence when the 5’ end of the first sequence follows the 3’ end of the second sequence.

[0143] As used herein, "regulatory elements" include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals, poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some cases, regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells as well as those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired tissue of interest, such as muscle, neuronal, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). In other cases, regulatory elements can also direct expression in a temporal manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which can or can not be tissue- or cell type-specific. As used herein, the term "promoter" refers to a non-coding nucleotide sequence located upstream from a gene that initiates transcription of the downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or defining a gene product, will result in production of the gene product by a cell under most or all physiological conditions. An inducible promoter is one that selectively expresses a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, such as by a chemical compound (chemical inducer), or in response to an environmental, hormonal, chemical, and / or developmental signal. Inducible or regulated promoters include, for example, promoters that are induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounding, or chemicals such as ethanol, abscisic acid (ABA), jasmonate, salicylic acid, or safeners.

[0144] As used herein, the term "on-target" refers to a predetermined or intended region of DNA that is bound, cleaved, and / or edited by a Cas protein of the disclosure or a variant polypeptide thereof.

[0145] As used herein, the term "off-target" refers to a non-predetermined or non-intended region of DNA that is bound, cleaved, and / or edited, for example, by a Cas protein of the disclosure or a variant polypeptide thereof. In some embodiments, a region of DNA is an off-target region when it differs from a region of DNA that is predetermined or intended to be bound, cleaved, and / or edited by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. In some embodiments, such sites are detected using targeted sequencing of sites predicted via computer simulation or by other methods known in the art.

[0146] As used herein, the term "in vivo" refers to events that occur within a multicellular organism, such as a human or non-human animal. In the case of cell-based systems, it can be used to refer to events that occur within a living cell (as opposed to, for example, in vitro systems).

[0147] As used herein, the term "ex vivo / in vitro" refers to events that occur in a cell or tissue that is grown outside of a multicellular organism, rather than within a multicellular organism.

[0148] As used herein, the term "cell" is to be understood as referring not only to a particular individual cell, but also to progeny or potential progeny of the cell. Since certain modifications can occur in successive generations due to mutations or environmental influences, such progeny can not, in principle, be identical to the parent cell, but are still included in the scope of the term.

[0149] As used herein, the terms "subject," "individual," and "patient" are used interchangeably herein and refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Also encompassed are tissues, cells, and progeny of biological entities obtained in vivo or cultured in vitro.

[0150] As used herein, the term "disease" refers to any condition or disorder that impairs or interferes with the normal functioning of a cell, tissue, or organ, including, but not limited to, those specific diseases that have been medically or clinically defined.

[0151] As used herein, the term "treatment" refers to the administration of a therapeutic molecule (e.g., a CRISPR-Cas system described herein) that partially or completely alleviates, ameliorates, relieves, inhibits, delays onset of, reduces severity of, and / or reduces incidence of one or more symptoms or features of a disease, disorder, and / or condition. Such treatment can be of a subject who does not exhibit signs of the relevant disease, disorder, and / or condition and / or of a subject who exhibits only early signs of the disease, disorder, and / or condition. Alternatively or additionally, such treatment can be of a subject who exhibits one or more established signs of the relevant disease, disorder, and / or condition.

[0152] As used herein, a "biological sample" can contain whole cells and / or viable cells and / or cell fragments. A biological sample can comprise (or be derived from) a "body fluid." In some embodiments, a body fluid can be selected from amniotic fluid, aqueous humor, vitreous fluid, bile, blood serum, breast milk, cerebrospinal fluid, cerumen (earwax), chyle, chymus, endolymph, perilymph, exudate, fecal matter, female ejaculate, gastric acid, gastric juice, lymphatic fluid, mucus (including nasal drainage and sticky phlegm), pericardial fluid, peritoneal fluid, pleural fluid, pus, rheum, saliva, sebum (skin oil), semen, sputum, synovial fluid, sweat, tears, urine, vaginal secretion, vomit, and mixtures of one or more thereof. Biological samples include cell cultures, body fluids, cell cultures from body fluids. Body fluids can be obtained from a mammal, for example, by puncture or other collection or sampling procedures.

[0153] As used herein, the term "activity" refers to biological activity. In some embodiments, nuclease activity includes enzymatic activity, e.g., catalytic ability, of a nuclease. For example, nuclease activity can include nuclease activity. In some embodiments, nuclease activity includes binding activity, e.g., binding activity of a nuclease to an RNA guide and / or a target nucleic acid.

[0154] As used herein, "expression of a genomic locus" or "gene expression" is the process by which information from a gene is used in the synthesis of a functional gene product. Products of gene expression are usually proteins, but in non-protein coding genes such as rRNA genes or tRNA genes, the product is functional RNA. All known forms of life—eukaryotes (including multicellular organisms), prokaryotes (bacteria and archaea), and viruses—use processes of gene expression to produce functional products, thereby enabling life. As used herein, "expression" of a gene or nucleic acid encompasses not only the expression of a gene in a cell, but also the transcription and translation of a nucleic acid in a cloning or other context and in any other background. As used herein, "expression" also refers to the process by which a polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene product." If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell.

[0155] Any protein presented herein can be produced by any method known in the art. For example, the proteins presented herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described in Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.

[0156] Various embodiments are described below. It should be noted that specific embodiments are not intended as exhaustive descriptions or as limitations on the broader aspects discussed herein. An aspect described in connection with a particular embodiment is not necessarily limited to that embodiment and may be practiced in conjunction with any other embodiment. Throughout the specification, references to “one embodiment,” “implementation,” “some embodiments,” “implementation,” “example embodiment,” etc., refer to specific features, structures, or characteristics described in connection with an embodiment that are included in at least one embodiment of the invention. Therefore, the phrases “in one embodiment,” “in an embodiment,” or “an example embodiment” appearing throughout the specification do not necessarily all refer to the same embodiment, but may. Furthermore, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner, as will be apparent to those skilled in the art based on this disclosure. Moreover, although some embodiments described herein include some but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention. For example, in the appended claims, any claimed embodiment may be used in any combination.

[0157] Novel Cas protein

[0158] This disclosure discovers a novel Cas protein, which we name CasY7, possessing single-stranded or double-stranded DNA cleavage activity. The sequence of the protein has very low sequence identity with other known Cas proteins, and the protein molecule is small, enabling better delivery.

[0159] The disclosed CasY7 protein exhibits excellent cleavage activity against both exogenous and endogenous genes in vitro and at the cellular level, comparable to or even better than LbCpf1. At the cellular level, its cleavage activity against specific target sequences of exogenous or endogenous genes reaches greater than approximately 8%, 9%, 10%, 15%, 20%, 30%, 40%, 50%, or higher. Overall, its cleavage activity against specific target sequences of exogenous or endogenous genes at the cellular level is higher than that of LbCpf1. These novel Cas proteins may also contain amino acid mutations that have no substantial impact on catalytic activity (endonuclease cleavage activity) or nucleic acid binding function.

[0160] Parental or reference Cas protein

[0161] In some respects, the parental or reference Cas protein is selected from:

[0162] (1) The Cas protein (CasY7) represented by the amino acid sequence shown in SEQ ID NO.1 of this disclosure; (2) an ortholog, homolog, variant, or functional fragment of the Cas protein shown in (1); (3) a polypeptide having about 60% identity with any one of (1) or (2). Optionally, the parent or reference Cas protein may be wild-type or not.

[0163] In a preferred embodiment of the present invention, the wild-type Cas protein is CasY7, the sequence of which is shown in SEQ ID NO.1.

[0164] It should be understood that the amino acid numbering in the mutant proteins disclosed herein is based on the parental or reference Cas protein. When a specific mutant protein has 80% or more sequence homology with the parental or reference Cas protein, the amino acid numbering of the mutant protein may be misaligned relative to the amino acid numbering of the parental or reference Cas protein, such as misalignment 1-100 positions towards the N-terminus or C-terminus of the amino acid. Using conventional sequence alignment techniques in the art, those skilled in the art can generally understand that such misalignment is within a reasonable range, and mutant proteins with 80% (e.g., 90%, 95%, 98%) homology and having the same or similar gene editing activity (e.g., cleavage activity) comparable to or higher than the wild-type Cas protein (SEQ ID NO. 1) should not be excluded from the scope of the mutant proteins disclosed herein due to amino acid numbering misalignment.

[0165] Characterization of representative variant peptides (or mutant proteins) and variant peptides

[0166] In one aspect of this disclosure, the variant polypeptide of this disclosure has, or retains or has improved, endonuclease activity (mid-target) against target DNA for specific cleavage of the target DNA. In some embodiments, the variant polypeptide has mid-target cleavage activity against double-stranded DNA. In some embodiments, the variant polypeptide substantially retains the mid-target cleavage activity against double-stranded DNA of SEQ ID NO. 1.

[0167] In another aspect of this disclosure, the Cas protein variant peptides of this disclosure may substantially lack endonuclease activity (mid-target), but retain their ability to complex with guide RNA and be guided to target DNA, thereby realizing the ability of the functional domains associated with the Cas protein variant peptides to be guided to the target DNA and perform their functions. The characterization of the Cas protein variant peptides of this disclosure is not limited to their ability to cleave target DNA.

[0168] Increased mid-target cleavage activity

[0169] In some aspects, this disclosure provides variant peptides of the Cas protein, which are engineered based on a parental or reference Cas protein. Exemplarily, such variant peptides have mutations at multiple amino acid sites in the peptide shown in SEQ ID NO. 1. Furthermore, compared to the parental or reference Cas protein, it exhibits higher target cleavage activity in double-stranded DNA. For example, compared to the parental or reference Cas protein (using the same guide RNA-mediated approach), its target cleavage activity in double-stranded DNA is increased by at least approximately 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, or more.

[0170] In some embodiments, the variant polypeptide is a polypeptide having at least about 60%, for example 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity with the amino acid sequence shown in SEQ ID NO.1.

[0171] In some embodiments, the variant polypeptides of this disclosure have amino acid substitutions relative to the parental or reference Cas protein. In some embodiments, the substitutions are conserved amino acid substitutions.

[0172] In one aspect, this disclosure provides a variant polypeptide that, relative to SEQ ID NO.1, comprises any amino acid substitution at one or more of the following sites:

[0173] V7, R8, T11, S12, S110, I149, N150, H151, N152, L153, ​​Q160, E161, Y162, N163, C164, Y165, S166, S167, F168, K 195, S196, K197, S198, A199, C215, T216, A217, K219, S225, L227, L229, M231, D234, S240, S241, Q242, E243, I2 44. S250, F251, E252, K253, V254, K260, T261, E263, N276, Y282, D283, A285, N292, S303, I304, L307, Y308, S3 09. R315, E316, T317, I318, I319, V348, I349, E350, P351, L354, S355, N356, L357, K358, I373, G416, I417, E41 8. F419, D420, I455, R456, V458, H486, S488, L501, R509, P510, V511, L512, G513, N514, R515, V516, L525, I52 6. N527, K528, K529, C633, T634, T635, K636, N637, D638, R639, G640, E641, F642, E648, L650, A651, Y652, A661 , T681, N682, E683, S684, G727, K728, N729, E730, S743, E748, G749, S751, K752, K781, L782, G783, E784, C785 , S787, K865, P866, Y867, N868, I872, D896, N936, M957, F958, Q960, W961, P1018, S1019, R1020, N1021, S1022.

[0174] In some embodiments, the number of amino acid mutations occurring in the variant polypeptide relative to SEQ ID NO.1 includes about 1-50, about 1-40, about 1-30, about 2 to 25, preferably 2-20, for example 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20.

[0175] In some embodiments, the variant peptide, relative to SEQ ID NO.1, comprises the following mutations:

[0176] (a)M231+D234;

[0177] Preferably, it is selected from M231K+D234V;

[0178] (b)S240+S241+Q242+E243;

[0179] Preferably, the components are selected from S240A+S241R+Q242P+E243L and S240L+S241A+Q242S+E243H+I244L+L307R+Y308F+S309A;

[0180] (c)L227+L229;

[0181] Preferably, the components are selected from L227I+L229R and S225P+L227A+L229R;

[0182] (d)L307+Y308+S309;

[0183] Preferably, the components are selected from S240L+S241A+Q242S+E243H+I244L+L307R+Y308F+S309A, D283S+A285P+L307R+Y308F+S309A, L307R+Y308F+S309A+E648R+L650I+A651P+Y652F, and D283S+A285P+L30 7R+Y308F+S309A+I373V+E748G+S751G+K752R+S787F, S303T+I304M+L307R+Y308A+S 309I, S196N+K197F+S198N+A199T+I304A+L307R+Y308A+S309A, L307R+Y308F+S309A;

[0184] (e)S250+F251+E252+K253+V254;

[0185] Preferably, it is selected from S250W+F251A+E252A+K253R+V254R;

[0186] (f)K260+T261+E263;

[0187] Preferably, it is selected from K260R+T261P+E263H;

[0188] (g)R315+E316+T317+I318+I319;

[0189] Preferably, the components are selected from R315L+E316R+T317R+I318V+I319A, Y282F+D283Q+A285T+R315A+E316Q+T317R+I318L+I319Q, and Y282F+D283Q+A285T+R315S+E316R+T317K+I318K+I319W+E648S+A651L;

[0190] (h)L354+S355+N356+L357+K358;

[0191] Preferably, it is selected from L354F+S355G+N356D+L357I+K358R;

[0192] (i)E648+A651;

[0193] Preferably, the components are selected from E648S+A651L, E648S+A651T+Y652W, L307R+Y308F+S309A+E648R+L650I+A651P+Y652F, Y282F+D283Q+A285T+E648T+A651S+Y652F, Y282F+D283Q+A285T+E648L+A651G, N276S+D283N+E648S+A651L, Y282F+D283Q+A285T+R315S+E316R+T317K+I318K+I319W+E648S+A651L;

[0194] (j)T681+N682+E683;

[0195] Preferably, the following are selected from T681V+N682G+E683R+S684G, T681P+N682T+E683G+S684R, T681S+N682P+E683R+S684G, T681A+N682H+E683Y+S684R, Y282F+D283Q+A285T+T681P+N682P+E683Q+S684C, Y282F+D283Q+A285T+G416S+I417H+E418Q+F419T+D420V+L501P+T681N+N682P+E683G+S684A, Y165W+S166R+G41 6R+I417V+E418Q+F419V+D420T+T681D+N682A+E683R+S684D, Y165W+S1 66R+S198N+G416R+I417V+E418Q+F419V+D420T+T681S+N682R+E683R, Y 165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509N+P510G+V511G+ L512S+T681S+N682R+E683R, G416R+I417V+E418Q+F419V+D420T+T681S+ N682R+E683R+G727V+K728C+N729G+E730G, Y282F+D283Q+A285T+G416A +I417R+E418V+F419C+D420Q+R509V+P510R+V511L+L512A+T681S+N682 R+E683R, Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509Q+P51 0A+V511A+L512R+T681S+N682R+E683R+A661V, S110G+Y165W+S166R+G41 6R+I417V+E418Q+F419V+D420T+R509A+P510A+V511R+L512Q+T681S+N6 82R+E683R, Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509L+P 510V+V511S+L512F+T681S+N682R+E683R, Y282F+D283Q+A285T+G416A+ I417R+E418V+F419C+D420Q+R509G+P510K+L512C+T681S+N682R+E683R、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+T634V+T635Y+K636S+N637R+T681S+N682R+E683R;,

[0196] (k)Y165+S166;

[0197] Preferably, the Y165W+S166R, Y165W+S166V+S167I+F168I, Y165W+S166R+G416R+I417V+E418Q+F419V+D420T, Y165W+S166R+G416R+I417E+E418Q+F419V+D420A, Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+T681D+N682A+E683R+S684D, and Y165W+S166R+S198 N+G416R+I417V+E418Q+F419V+D420T+T681S+N682R+E683R, Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+N682G+E683A+ D896N、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509N+P510G+V511G+L512S+T681S+N682R+E683R、Y165W+S166R+G41 6R+I417V+E418Q+F419V+D420T+G513L+N514A+R515V+V516M+N682G+E683A+S684R+D896N, Y165W+S166R+G416R+I417V+E418Q +F419V+D420T+G513C+R515Y+V516M+N682G+E683A+S684R+D896N, Y165W+S166R+G416A+I417R+E418V+F419C+D420Q+H486S+S 488W+N682G+E683A+S684R+D896N, Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509Q+P510A+V511A+L512R+T681S+N68 2R+E683R+A661V, S110G+Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509A+P510A+V511R+L512Q+T681S+N682R+E683R,

[0198] Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509L+P510V+V511S+L512F+T681S+N682R+E683R, Y165W+S16 6R+G416R+I417V+E418Q+F419V+D420T+G513C+R515Y+V516M+T635S+K636G+N637L+N682G+E683A+S684R+D896N,

[0199] Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+T634V+T635Y+ K636S+N637R+T681S+N682R+E683R, Y165W+S166R+G416A+I417R+E418V+F419C+ D420Q+H486S+S488W+K781T+G783S+E784F+C785L+D896N, Y165W+S166R+G416A+ I417R+E418V+F419C+D420Q+H486S+S488W+D638F+R639P+G640N+E641Y+F642L、

[0200] Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+N682G+E683A+S684R+P1018L+S1019R+R1020P+N1021R+S1022V;

[0201] (l)Q160+E161+Y162+N163+C164;

[0202] Preferably, the components are selected from Q160V+E161I+Y162T+N163S+C164V, Q160S+E161S+Y162L+N163C+C164V, Q160P+E161V+Y162H+N163S+C164T, and Q160L+E161S+Y162F+N163S+C164A;

[0203] (m)V7+T11+S12;

[0204] Preferably, the components are selected from V7H+T11R+S12M, V7F+R8L+T11R+S12V, and V7I+T11V+S12M+Y282F+D283Q+A285T;

[0205] (n)I149+N150+H151+N152+L153;

[0206] Preferably, it is selected from I149M+N150S+H151C+N152T+L153Y;

[0207] (o)K195+K197+S198+A199;

[0208] Preferably, it is selected from K195R+K197A+S198A+A199G;

[0209] (p)T216+A217+K219;

[0210] Preferably, the material is selected from T216L+A217T+K219R and C215V+T216S+A217S+K219R;

[0211] (q)D283+A285;

[0212] Preferably, the molecule is selected from Y282F+D283Q+A285T, D283I+A285R+N292L, Y282F+D283Q+A285T+E648T+A651S+Y652F, Y282F+D283Q+A285T+E648L+A651G, D283S+A285P+L307R+Y308F+S309A, Y282F+D283Q+A285T+R315A+E316Q+T317R+I318L+I319Q, Y282F+D283Q+A285T+R315S+E316R+T317K+I318K+I319W+E6 48S+A651L, V7I+T11V+S12M+Y282F+D283Q+A285T, D283S+A285P+L307R +Y308F+S309A+I373V+E748G+S751G+K752R+S787F, Y282F+D283Q+A285 T+K865S+P866S+Y867H+N868F+I872L, Y282F+D283Q+A285T+M957V+F95 8L+Q960V+W961V, Y282F+D283Q+A285T+V348I+I349V+E350L+P351T, Y2 82F+D283Q+A285T+T681P+N682P+E683Q+S684C, Y282F+D283Q+A285T+E 748A+G749S+S751G, Y282F+D283Q+A285T+G416S+I417H+E418Q+F419T+ D420V, Y282F+D283Q+A285T+G416L+I417Q+E418M+F419R+D420A, Y282F +D283Q+A285T+G416A+I417R+E418V+F419C+D420Q、Y282F+D283Q+A285 T+G416S+I417H+E418Q+F419T+D420V+L501P+N936S, Y282F+D283Q+A28 5T+G416S+I417H+E418Q+F419T+D420V+L501P+T681N+N682P+E683G+S6 84A, Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q, Y282F+D 283Q+A285T+G416L+I417Q+E418M+F419R+D420A+I455G+R456E+V458E,Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+L525R+I526V+N527D+K528P+K529G+N682G+E683A+S684R+D896N、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509V+P510R+V511L+L512A+T681S+N682R+E683R、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509G+P510K+L512C+T681 S+N682R+E683R、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+N682G+E683A+S684R+G727E+K728H+N729Q+E730R、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+K781H+L782I+G783F+E784R+C785S、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+D638P+E641M+F642V、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+P1018S+S1019R+R1020V+N1021M+S1022P、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633N+T634E+T635E+K636G+N637L+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633F+T634P+T635F+K636P+N637C+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633S+T634C+T635G+K636H+N637F+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633R+T634Y+T635L+K636V+N637D+N682G+E683A+S684R+S743T;、

[0213] (r)G416+I417+E418+F419+D420;

[0214] Preferably, the following are selected from G416R+I417W+E418T+F419R+D420V, Y282F+D283Q+A285T+G416S+I417H+E418Q+F419T+D420V, Y165W+S166R+G416R+I417V+E418Q+F419V+D420T, Y282F+D283Q+A285T+G416L+I417Q+E418M+F419R+D420A, Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q, Y165W+S166R+G416R+I417E+ E418Q+F419V+D420A, Y282F+D283Q+A285T+G416S+I417H+E418Q+F419T+D 420V+L501P+N936S, Y282F+D283Q+A285T+G416S+I417H+E418Q+F419T+D42 0V+L501P+T681N+N682P+E683G+S684A, Y165W+S166R+G416R+I417V+E418 Q+F419V+D420T+T681D+N682A+E683R+S684D, Y165W+S166R+S198N+G416R+ I417V+E418Q+F419V+D420T+T681S+N682R+E683R, Y165W+S166R+G416R+I 417V+E418Q+F419V+D420T+N682G+E683A+D896N, Y282F+D283Q+A285T+G41 6A+I417R+E418V+F419C+D420Q, R509C+P510S+L512F+G416R+I417V+E418 Q+F419V+D420T+H486A+S488W+N682G+E683A+S684R+D896N, Y165W+S166R+ G416R+I417V+E418Q+F419V+D420T+R509N+P510G+V511G+L512S+T681S+N 682R+E683R, Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+G513L+N51 4A+R515V+V516M+N682G+E683A+S684R+D896N, Y165W+S166R+G416R+I417 V+E418Q+F419V+D420T+G513C+R515Y+V516M+N682G+E683A+S684R+D896N,G416R+I417V+E418Q+F419V+D420T+T681S+N682R+E683R+G727V+K728C+N729G+E730G、Y282F+D283Q+A285T+G416L+I417Q+E418M+F419R+D420A+I455G+R456E+V458E、Y165W+S166R+G416A+I417R+E418V+F419C+D420Q+H486S+S488W+N682G+E683A+S684R+D896N、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+L525R+I526V+N527D+K528P+K529G+N682G+E683A+S684R+D896N、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509V+P510R+V511L+L512A+T681S+N682R+E683R、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509Q+P510A+V511A+L512R+T681 S+N682R+E683R+A661V、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V、S110G+Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509A+P510A+V511R+L512Q+T681S+N682R+E683R、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509L+P510V+V511S+L512F+T681 S+N682R+E683R、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509G+P510K+L512C+T681 S+N682R+E683R、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+N682G+E683A+S684R+G727E+K728H+N729Q+E730R、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+G513C+R515Y+V516M+T635S+K636G+N637L+N682G+E683A+S684R+D896N、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+T634V+T635Y+K636S+N637R+T681 S+N682R+E683R、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+K781H+L782I+G783F+E784R+C785S、Y165W+S166R+G416A+I417R+E418V+F419C+D420Q+H486S+S488W+K781T+G783S+E784F+C785L+D896N、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+D638P+E641M+F642V、Y165W+S166R+G416A+I417R+E418V+F419C+D420Q+H486S+S488W+D638F+R639P+G640N+E641Y+F642L、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+N682G+E683A+S684R+P1018L+S1019R+R1020P+N1021R+S1022V、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+P1018S+S1019R+R1020V+N1021M+S1022P、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633N+T634E+T635E+K636G+N637L+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633F+T634P+T635F+K636P+ N637C+N682G+E683A+S684R+S743T, Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+ L512F+C633S+T634C+T635G+K636H+N637F+N682G+E683A+S684R+S743T, Y282F+D283Q+A285T+G416A+I417R+ E418V+F419C+D420Q+R509C+P510S+L512F+C633R+T634Y+T635L+K636V+N637D+N682G+E683A+S684R+S743T. ,

[0215] Endonuclease-deficient variant peptides (dCas)

[0216] In some aspects, this disclosure provides a nuclease-deficient variant polypeptide, which, exemplary, retains only a portion (e.g., less than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 4%, 3%, 2%, 1% or less) of spacer-specific nuclease cleavage activity against a target sequence of target DNA complementary to the guide RNA, relative to a parent or reference Cas protein.

[0217] In some implementations, the variant peptide essentially lacks spacer sequence-specific (targeted) double-stranded DNA cleavage activity.

[0218] In some embodiments, the variant peptide substantially lacks the spacer sequence-specific (targeted) double-stranded DNA cleavage activity of SEQ ID NO.1.

[0219] In some embodiments, when used in combination with the same guide RNA, the variant polypeptide has reduced spacer sequence-specific (targeted) double-stranded DNA cleavage activity compared to SEQ ID NO.1, for example, by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of spacer sequence-specific (targeted) double-stranded DNA cleavage activity.

[0220] In some embodiments, the variant polypeptide, relative to SEQ ID NO.1, comprises one or more amino acid substitutions at the following positions: D592, D643, E820, and D992.

[0221] In some embodiments, the variant polypeptide, relative to the amino acid sequence shown in SEQ ID NO.1, comprises one or more amino acid substitutions at the following positions: D592A, D643A, E820A, and D992A.

[0222] In some embodiments, the amino acid substitutions of the variant polypeptide relative to the amino acid sequence shown in SEQ ID NO.1 are selected from one of D592A, D643A, E820A, and D992A.

[0223] In some embodiments, the variant polypeptide comprises the amino acid sequence shown in any one of SEQ ID NO. 47-50, or an amino acid sequence having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with the amino acid sequence shown in any one of SEQ ID NO. 47-50.

[0224] Fusion protein

[0225] In one aspect, this disclosure provides a fusion protein comprising a Cas protein or variant polypeptide of this disclosure, and one or more functional domains associated with the Cas protein or variant polypeptide. In one embodiment, the functional domain is a heterologous functional domain.

[0226] As used herein, the term "functional domain" is used in its broadest sense, including proteins such as enzymes or factors themselves or fragments / domains that have a specific function. When a fusion protein includes more than one functional domain, the functional domains may be the same or different.

[0227] In one embodiment, the functional domain is selected from nuclear localization signals (NLS), nuclear output signals (NES), reporter proteins (e.g., fluorescent proteins), Cas protein targeting portions, DNA binding domains (e.g., Lex A DBD, Gal4 DBD, Sp1 DBD), epitope tags (e.g., His, myc, V5, FLAG, HA, VSV-G, etc.), transcriptional activation domains (e.g., VP64, VPR, p65, Rta), transcriptional repression domains (e.g., KRAB domain, SID domain, NuE domain, NcoR domain, or SID4X domain), nucleases, deaminases (e.g., adenosine deaminase or cytidine deaminase), methyltransferases (e.g., DNA methyltransferase DNMT), demethylases, transcription release factors, HDAC, cleavage active peptides, ligases, integrases, transposases, recombinases, polymerases, exonucleases (e.g., T5E), and base excision repair inhibitors (e.g., uracil-DNA glycosylation inhibitors (UGI)).

[0228] In some embodiments, the functional domain includes one or more of the following enzyme activities against the target sequence: methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, demyristylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), and deglycosylation activity.

[0229] In some implementations, the functional domain is selected from deaminases.

[0230] As used herein, "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that catalyze a hydrolytic deamination reaction that converts adenine (or the adenine portion of a molecule) to hypoxanthine (or the hypoxanthine portion of a molecule). In some embodiments, the functional domain is selected from adenosine deaminases. In some embodiments, the adenosine deaminase comprises the ability to deaminate adenine (A) to hypoxanthine (I). In some embodiments, the deamination of adenine to hypoxanthine converts adenosine (A) or deoxyadenosine (dA) containing adenine to guanosine (G) or deoxyguanosine (dG).

[0231] In some embodiments, the adenosine deaminase is, but is not limited to, a member of the enzyme family called an adenosine deaminase acting on RNA (ADAR), a member of the enzyme family called an adenosine deaminase acting on tRNA (ADAT), and other family members containing an adenosine deaminase domain (ADAD). In some embodiments, the adenosine deaminase is capable of targeting adenine in RNA / DNA and RNA duplexes. In some embodiments, the adenosine deaminase has been modified to increase its ability to edit DNA in RNA / DNA heteroduplexes of RNA duplexes.

[0232] In some embodiments, the adenosine deaminase is derived from one or more metazoan species, including but not limited to mammals, birds, frogs, squid, fish, flies, and worms. In some embodiments, the adenosine deaminase is human, squid, or fruit fly adenosine deaminase.

[0233] In some embodiments, adenosine deaminase is a human ADAR, including hADAR1, hADAR2, and hADAR3. In some embodiments, adenosine deaminase is a Caenorhabditis elegans ADAR protein, including ADR-1 and ADR-2. In some embodiments, adenosine deaminase is a Drosophila ADAR protein, including dADAR. In some embodiments, adenosine deaminase is a squid (Loligo pealeii) ADAR protein, including sqADAR2a and sqADAR2b. In some embodiments, adenosine deaminase is a human ADAT protein. In some embodiments, adenosine deaminase is a Drosophila ADAT protein. In some embodiments, adenosine deaminase is a human ADAD protein, including TENR (hADAD1) and TENRL (hADAD2).

[0234] In some embodiments, the adenosine deaminase is the TadA protein, such as E. coli TadA. See Kim et al., Biochemistry 45:6407-6416 (2006); Wolf et al., EMBO J.21:3841-3851 (2002). In some embodiments, the adenosine deaminase is mouse ADA. See Grunebaum et al., Curr. Opin. Allergy Clin. Immunol. 13:630-638 (2013). In some embodiments, the adenosine deaminase is human ADAT2. See Fukui et al., J. Nucleic Acids 2010:260512 (2010). In some embodiments, the deaminase (e.g., adenosine or cytidine deaminase) is one or more of those described in the following literature: Cox et al., Science. 2017 Nov 24; 358(6366):1019-1027; Komore et al., Nature. 2016 May 19; 533(7603):420-4; and Gaudelli et al., Nature. 2017 Nov 23; 551(7681):464-471.

[0235] In some embodiments, the adenosine deaminase protein recognizes one or more target adenosine residues in a double-stranded nucleic acid substrate and converts them into inosine residues. In some embodiments, the double-stranded nucleic acid substrate is an RNA-DNA hybrid double helix. In some embodiments, the adenosine deaminase protein recognizes a binding window on the double-stranded substrate. In some embodiments, the binding window contains at least one target adenosine residue. In some embodiments, the binding window is in the range of about 3 bp to about 100 bp. In some embodiments, the binding window is in the range of about 5 bp to about 50 bp. In some embodiments, the binding window is in the range of about 10 bp to about 30 bp. In some embodiments, the binding window is about 1 bp, 2 bp, 3 bp, 5 bp, 7 bp, 10 bp, 15 bp, 20 bp, 25 bp, 30 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, or 100 bp.

[0236] In some embodiments, the adenosine deaminase protein comprises one or more deaminase domains. Without being bound by any particular theory, it is anticipated that the deaminase domains will be used to recognize one or more target adenosine (A) residues contained in a double-stranded nucleic acid substrate and convert them to inosine (I) residues. In some embodiments, the deaminase domain contains an active site. In some embodiments, the active site contains a zinc ion. In some embodiments, during the AI ​​editing process, the base pairing at the target adenosine residue is disrupted, and the target adenosine residue is “flipped” out of the double helix to become accessible to the adenosine deaminase. In some embodiments, amino acid residues in or near the active site interact with one or more nucleotides at the 5' end of the target adenosine residue. In some embodiments, amino acid residues in or near the active site interact with one or more nucleotides at the 3' end of the target adenosine residue. In some embodiments, amino acid residues in or near the active site further interact with nucleotides complementary to the target adenosine residue on the opposite strand. In some embodiments, the amino acid residues form hydrogen bonds with the 2' hydroxyl group of the nucleotide.

[0237] In some embodiments, the adenosine deaminase comprises the human ADAR2 whole protein (hADAR2) or its deaminase domain (hADAR2-D). In some embodiments, the adenosine deaminase is an ADAR family member homologous to hADAR2 or hADAR2-D.

[0238] In some embodiments, the adenosine deaminase comprises the wild-type amino acid sequence of hADAR2-D. In some embodiments, the adenosine deaminase contains one or more mutations in the hADAR2-D sequence, such that the editing efficiency and / or substrate editing preference of hADAR2-D can be altered according to specific needs.

[0239] In some embodiments, adenosine deaminase is tRNA adenosine deaminase (TadA) or its deaminase domain or a functional variant or fragment thereof. For example, adenosine deaminase is selected from TadA8e (SEQ ID NO. 79), TadA8.17, TadA8.20, TadA9, TadA8EV106W, TadA8EV106W+D108Q, TadA-CDa, TadA-CDb, TadA-CDc, TadA-CDd, TadA-CDe, TadA-dual, TADAC-1.2, TADAC-1.14, TADAC-1.17, TADAC-1.19, TADAC-2.5, TADAC-2.6, TADAC-2.9, TADAC-2.19, TADAC-2.23, TadA8e-N46L, and TadA8e-N46P.

[0240] In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identity with the amino acid sequence shown in SEQ ID NO. 31 (selected from 005V1 deaminase in CN114634923A, with the amino acid sequence shown in SEQ ID NO. 1) and retains the deamination activity of the amino acid sequence shown in SEQ ID NO. 31.

[0241] In some embodiments, the amino acid sequence of the adenosine deaminase involves additions, insertions, deletions, and substitutions relative to the amino acid sequence shown in SEQ ID NO. 31. In some embodiments, the catalytic domain of the adenosine deaminase includes a mutant of the amino acid sequence shown in SEQ ID NO. 31. In some embodiments, the adenosine deaminase comprises the amino acid sequence shown in SEQ ID NO. 32.

[0242] In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identity with the amino acid sequence shown in SEQ ID NO. 32, and retains the deamination activity of the amino acid sequence shown in SEQ ID NO. 32.

[0243] In some embodiments, the amino acid sequence of the adenosine deaminase involves additions, insertions, deletions, and substitutions relative to the amino acid sequence shown in SEQ ID NO. 32. In some embodiments, the catalytic domain of the adenosine deaminase includes a mutant of the amino acid sequence shown in SEQ ID NO. 32. In some embodiments, the adenosine deaminase comprises the amino acid sequence shown in SEQ ID NO. 32.

[0244] In some embodiments, the amino acid sequence of the adenosine deaminase comprises an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identity with the amino acid sequence shown in SEQ ID NO. 70 (selected from 004V1 deaminase in CN114634923A, in which the amino acid sequence is SEQ ID NO. 1), and it retains the deamination activity of the amino acid sequence shown in SEQ ID NO. 70.

[0245] In some implementations, the functional domain is selected from cytidine deaminase.

[0246] As used herein, the term "cytidine deaminase" or "cytidine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that catalyzes a hydrolytic deamination reaction that converts cytosine (or the cytosine portion of a molecule) to uracil (or the uracil portion of a molecule). In some embodiments, the cytosine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0247] In some embodiments, the cytidine deaminase is selected from APOBEC (e.g., APOBEC3, such as APOBEC3A, APOBEC3B, APOBEC3C).

[0248] In some embodiments, the cytidine deaminase is selected from hAPOBEC3-W104A (SEQ ID NO.95).

[0249] In some embodiments, the cytidine deaminase comprises the wild-type amino acid sequence of cytosine deaminase. In some embodiments, the cytidine deaminase contains one or more mutations in the cytosine deaminase sequence, such that the editing efficiency and / or substrate editing preference of the cytosine deaminase can be altered according to specific needs.

[0250] In some embodiments, the functional domain is a base excision repair inhibitor; for example, a base excision repair inhibitor is a uracil-DNA glycosylation inhibitor (UGI).

[0251] In some embodiments, the uracil-DNA glycosylation inhibitor (UGI) is derived from the human UGI domain (exemplary of which is shown in SEQ ID NO. 96).

[0252] In one embodiment, the functional domain is a methyltransferase, such as HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3, ZMET2, CMT1, CMT2, etc.

[0253] In one embodiment, the functional domain is a demethylase, such as TET1 (ten-eleven translocation 1), ten-eleven translocation (TET) dioxygenase 1 (TET1CD), DME, DML1, DML2, ROS1, etc.

[0254] In one implementation, the functional domain is a transcriptional release factor, such as eukaryotic release factor 1 (ERF1) activity or eukaryotic release factor 3 (ERF3).

[0255] In one implementation, the nuclear output signal includes human protein tyrosine kinase 2.

[0256] In one embodiment, the reporter protein includes one or more of glutathione S-transferase, horseradish peroxidase, chloramphenicol acetyltransferase, β-galactosidase, β-glucuronidase, or autofluorescent protein.

[0257] In one embodiment, the autofluorescent protein includes one or more of green fluorescent protein, HcRed, DsRed, cyan fluorescent protein, yellow fluorescent protein, or blue fluorescent protein.

[0258] In one embodiment, the DNA binding domain includes one or more of methylation-binding proteins, LexADBD, Sp1DBD, or Gal4DBD.

[0259] In one embodiment, the epitope tag includes one or more of His (histidine tag), V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, or thioredoxin tag.

[0260] In one implementation, the transcriptional activation domain includes one or more of VP64, p65, Rta, and VPR.

[0261] In one implementation, the transcriptional repression domain includes KRAB and / or SID.

[0262] In one embodiment, the nuclease includes FokI.

[0263] In one embodiment, the cleavage-active polypeptide includes a polypeptide having single-stranded RNA cleavage activity, a polypeptide having double-stranded RNA cleavage activity, a polypeptide having single-stranded DNA cleavage activity, or a polypeptide having double-stranded DNA cleavage activity.

[0264] In one embodiment, the ligase includes DNA ligase and / or RNA ligase.

[0265] In some embodiments, the functional domain is an exonuclease. According to the disclosure herein, there are various eligible exonucleases or programmable exonucleases. Some examples of exonucleases suitable as part of the fusion protein in this application include MRE11, EXO1, EXO III, EXO VII, EXOT, DNA2, CtIP, TREX1, TREX2, Apollo, RecE, Red, T5, Lexo, RecBCD, and mung bean exonuclease. Other suitable exonucleases are also under consideration.

[0266] In some embodiments, the exonuclease is selected from T5 exonuclease or T7 exonuclease.

[0267] In some embodiments, the T5 exonuclease is an amino acid sequence containing at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with the amino acid sequence described in any one of SEQ ID NO. 81.

[0268] In some embodiments, the functional domain is a nuclear localization signal. As used herein, the terms "NLS," "nuclear localization sequence," or "nuclear localization signal" are used interchangeably and refer to the amino acid sequence that facilitates protein entry into the cell nucleus. Nuclear localization sequences are known in the art (e.g., described in Plank et al., International PCT Application PCT / EP2000 / 011690, filed November 23, 2000, and published as WO / 2001 / 038547 on May 31, 2001), which is incorporated herein by reference to its disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the following amino acid sequences: KRTADGSEFESPKKKRKV (SEQ ID NO.38), AVKRPAATKKAGQAKKKKLD (SEQ ID NO.39), KRPAATKKAGQAKKKK (SEQ ID NO.40), KKTELQTTNAENKTKKL (SEQ ID NO.41), KRGINDRNFWRGENGRKTR (SEQ ID NO.42), RKSGKIAAIVVKRPRK (SEQ ID NO.43), PKKKRKV (SEQ ID NO.44), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO.45).

[0269] Polynucleotides

[0270] In one aspect, this disclosure provides a polynucleotide that is a polynucleotide sequence encoding the Cas protein or a variant polypeptide of the disclosed protein, or a polynucleotide sequence encoding a fusion protein of the disclosed protein.

[0271] In one embodiment, the polynucleotide is a DNA molecule codon-optimized according to the codon preference of the host cell;

[0272] Optimization of this disclosure may require mutations in the nucleotide sequence of the encoded protein (e.g., the Cas protein of this disclosure or a variant polypeptide thereof) to mimic the codon bias of a intended host organism or cell simultaneously encoding the same protein. Therefore, codons can be changed, but the encoded protein remains unchanged. For example, if the intended target cell is a human cell, a human codon-optimized nucleotide sequence encoding the protein can be used. As another non-limiting embodiment, if the intended host cell is an animal cell (e.g., mouse cell, insect cell), an animal codon-optimized nucleotide sequence encoding the protein can be generated. As another non-limiting embodiment, if the intended host cell is a plant cell, a plant codon-optimized nucleotide sequence encoding the protein can be generated.

[0273] Lists of codon choices are readily available, for example, in the "Codon Usage Database" at www.kazusa.or.jp / codon. In some cases, the nucleic acids of this disclosure contain nucleotide sequences encoding CasY7 or its variant polypeptides or fusion proteins thereof, said nucleotide sequences being codon-optimized for expression in eukaryotic cells. In some cases, the nucleic acids of this disclosure contain nucleotide sequences encoding CasY7 or its variant polypeptides or fusion proteins thereof, said nucleotide sequences being codon-optimized for expression in animal cells. In some cases, the nucleic acids of this disclosure contain nucleotide sequences encoding CasY7 or its variant polypeptides or fusion proteins thereof, said nucleotide sequences being codon-optimized for expression in fungal cells. In some cases, the nucleic acids of this disclosure contain nucleotide sequences encoding CasY7 or its variant polypeptides or fusion proteins thereof, said nucleotide sequences being codon-optimized for expression in plant cells.

[0274] In some embodiments, the nucleic acid molecules of this disclosure comprise nucleotide sequences having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity.

[0275] In some embodiments, the polynucleotide is a nucleotide sequence such as that shown in any one of SEQ ID NO. 2, 37, 83, 85, 87.

[0276] In one implementation, the host cell includes a prokaryotic cell or a eukaryotic cell;

[0277] In some embodiments, this disclosure provides a polynucleotide encoding a guide RNA. In some embodiments, the polynucleotide encoding the guide RNA comprises DNA, RNA, or a DNA / RNA mixture. A “DNA / RNA mixture” refers to a nucleic acid comprising one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, whether continuous or not. However, “DNA” or “RNA” can also refer to DNA containing one or more modified or unmodified ribonucleotides, whether continuous or not, or RNA containing one or more modified or unmodified deoxyribonucleotides, whether continuous or not. In some embodiments, the polynucleotide encoding the guide RNA is operatively linked to a promoter or regulated by a promoter.

[0278] In some embodiments, the polynucleotide encoding the Cas protein is DNA, RNA, or a DNA / RNA mixture. A “DNA / RNA mixture” refers to a nucleic acid comprising one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, whether continuous or not. However, “DNA” or “RNA” can also refer to DNA containing one or more modified or unmodified ribonucleotides, whether continuous or not, or RNA containing one or more modified or unmodified deoxyribonucleotides, whether continuous or not.

[0279] In some embodiments, the polynucleotide encoding the Cas protein is operatively linked to or regulated by a promoter. In some embodiments, the promoter is a broad-spectrum promoter, a tissue-specific promoter, a cell-type-specific promoter, a constitutive promoter, or an inducible promoter.

[0280] In some embodiments, the polynucleotide encoding the guide RNA and / or Cas protein is functionally or operatively linked to the regulatory element and thus regulated. Optionally, the promoter may be selected from the group consisting of: pol I promoter, pol II promoter, pol III promoter, T7 promoter, U6 promoter, H1 promoter, retroviral Rouss sarcoma virus (RSV) LTR promoter, cytomegalovirus (CMV) promoter, SV40 promoter, dihydrofolate reductase promoter, β-actin promoter, glycerol phosphokinase (PGK) promoter, and EF1α promoter. In some embodiments, the promoter is the U6 promoter.

[0281] CRISPR system

[0282] The Cas protein and its variant peptides disclosed herein can be used in combination with guide RNA and guided by the guide RNA to target DNA to exert their effects on the target DNA.

[0283] In one aspect, this disclosure provides a CRISPR-Cas system comprising:

[0284] (i) The first component is selected from the group consisting of: the Cas protein or variant polypeptide of this disclosure, the fusion protein of this disclosure, the nucleotide sequence encoding the Cas protein or variant polypeptide of this disclosure, or the fusion protein of this disclosure, and any combination thereof; and

[0285] (ii) A second component comprising one or more of the guide RNAs disclosed herein, or encoding a nucleotide sequence comprising one or more of the guide RNAs disclosed herein.

[0286] In some embodiments, the guide RNA contains:

[0287] (i) a direct repeat (DR) sequence capable of forming a complex with the Cas protein or fusion protein; and

[0288] (ii) It is capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the spacer sequence of the target DNA.

[0289] In some implementations, the CRISPR-Cas system is either non-natural or engineered.

[0290] In some implementations, the system is a complex comprising a Cas protein compounded with guide RNA.

[0291] In some embodiments, the complex further comprises target DNA that hybridizes with the target sequence.

[0292] In one aspect, this disclosure provides a complex comprising:

[0293] (i) Protein components selected from the group consisting of: the Cas protein of this disclosure, the fusion protein of this disclosure, or combinations thereof; and

[0294] (ii) The guide RNA disclosed herein.

[0295] In one aspect, the present invention also provides a CRISPR-Cas composition comprising:

[0296] (1) Protein component: the Cas protein of this disclosure, or the fusion protein; or the nucleic acid molecule encoding the Cas protein of this disclosure or the fusion protein thereon;

[0297] (2) RNA component: the guide RNA disclosed herein, or one or more nucleic acids encoding the guide RNA, or the precursor RNA of the guide RNA, or the nucleic acid encoding the precursor RNA of the guide RNA; the protein component and the nucleic acid component combine to form a complex.

[0298] In one embodiment, the composition is an activated CRISPR complex, the activated CRISPR complex further comprising a target sequence of a target nucleic acid bound to the guide RNA.

[0299] Guide RNA

[0300] On the other hand, this disclosure provides a guide RNA (gRNA) comprising:

[0301] (1) A direct repeat (DR) sequence capable of forming a complex with the Cas protein of this disclosure, and

[0302] (2) It is capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the spacer sequence of the target DNA.

[0303] In some implementations, the guide RNA does not contain tracrRNA.

[0304] In some embodiments, the directional repeat (DR) sequence is located at the 5' end of the spacer sequence, which is capable of forming a complex with the Cas protein of this disclosure or its variant polypeptide, or the fusion protein of this disclosure.

[0305] In some embodiments, the guide RNA comprises, or consists essentially of, or is composed of a direct repeat (DR) sequence and a spacer sequence. In some embodiments, the guide RNA is a single nucleic acid molecule with the DR sequence linked to the spacer sequence.

[0306] In some implementations, the guide RNA comprises multiple tandemly arranged spacer sequences, optionally separated by nucleotide sequences such as directed repeat (DR) sequences as defined herein. The different spacer sequences are tandemly positioned without affecting activity.

[0307] In some embodiments, the spacer sequence comprises a series of 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 25, 25, 30, or more sequences, each capable of specifically hybridizing to a target sequence at a target genomic locus in the cell. In some embodiments, the functional CRISPR-Cas system or complex can edit multiple target sequences, for example, the target sequences may comprise genomic loci, and in some embodiments, alterations in gene expression may be present. In some embodiments, the functional CRISPR-Cas system or complex may comprise other functional domains. In some embodiments, this disclosure provides methods for altering or modifying the expression of multiple gene products. The methods may include introducing the target nucleic acid, such as a DNA molecule, into a cell containing and expressing the target nucleic acid, such as a DNA molecule; for example, the target nucleic acid may encode a gene product or provide expression of a gene product (e.g., a regulatory sequence). In some embodiments, the guide RNA comprises a direct repeat (DR) sequence, a spacer sequence, and a direct repeat (DR) sequence (DR-spacer-DR). This is a typical configuration of pre-crRNA. In some embodiments, the guide RNA comprises a DR-spacer-DR-spacer structure. In some embodiments, the guide RNA comprises two or more DRs and two or more spacer sequences. In some embodiments, the crRNA comprises truncated directional repeat (DR) sequences and spacer sequences. This is typical of processed or mature guide RNAs. In some embodiments, the Cas protein of this disclosure or a variant polypeptide thereof forms a complex with the guide RNA, and the spacer sequence guides the complex to a target nucleic acid complementary to the spacer sequence for sequence-specific binding.

[0308] Multiple guide RNAs

[0309] In one aspect, the CRISPR-Cas system disclosed herein includes a plurality of (exemplarily, 2, 3, 4, 5, or more) guide RNAs, for example, each of the plurality of guide RNAs specifically targets a DNA molecule in a cell encoding a gene product, such that each of the plurality of guide RNAs targets a specific DNA molecule encoding its gene product, and the Cas protein cleaves the target DNA molecule encoding the gene product, thereby altering the expression of the gene product; and wherein the CRISPR protein and the guide RNA are not naturally present together.

[0310] In some implementations, multiple guiding nucleic acids are operatively linked to the same regulatory element (e.g., a promoter) or to separate regulatory elements (e.g., promoters) or under their regulation.

[0311] In some embodiments, the guide RNA comprises non-naturally occurring nucleic acids and / or non-naturally occurring nucleotides and / or nucleotide analogs and / or chemical modifications. Preferably, these non-naturally occurring nucleic acids and non-naturally occurring nucleotides are located outside the guide RNA. The non-naturally occurring nucleic acids may include, for example, a mixture of natural and non-naturally occurring nucleotides. The non-naturally occurring nucleotides and / or nucleotide analogs may be modified at the ribose, phosphate, and / or base moieties. In one embodiment, the guide RNA comprises ribonucleotides and non-ribonucleotides.

[0312] In one embodiment, the guide RNA comprises one or more non-naturally occurring nucleotides or nucleotide analogs, such as nucleotides having phosphate-thioester bonds, locked nucleic acids (LNAs) or bridging nucleic acids (BNAs) containing a methylene bridge between the 2' and 4' carbon atoms of the ribose ring. Other examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. Other examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromouridine, pseudouridine, inosine, and 7-methylguanosine. Examples of chemical modifications to the guide RNA include, but are not limited to, incorporation of 2'-O-methyl (M), 2'-O-methyl 3'-phosphate-thioester (MS), S-restricted ethyl (cEt), or 2'-O-methyl 3'-thiophosphate (MSP) at one or more terminal nucleotides.

[0313] In one embodiment, the 5' and / or 3' ends of the guide RNA are modified with various functional moieties, including fluorescent dyes, polyethylene glycol, cholesterol, proteins, or detection tags (see Kelly et al., 2016, J. Biotech. 233:74-83). In one embodiment, the guide RNA contains a ribonucleotide in the region binding to the target sequence and one or more deoxyribonucleotides and / or nucleotide analogs in the region binding to a nucleic acid-directed nuclease. In one embodiment, 3-5 nucleotides at the 3' or 5' end of the guide RNA are chemically modified. In one embodiment, only minor modifications, such as 2'-F modifications, are introduced in the seed region. In one embodiment, the guide RNA is modified to include a chemical moieties at its 3' and / or 5' ends. Such moieties include, but are not limited to, amines, azides, alkynes, thiols, dibenzocyclooctyne (DBCO), or rhodamine. In some embodiments, the chemical moieties are conjugated to the guide RNA via linkers such as alkyl chains. In one embodiment, the modified chemical moieties of the guide RNA can be used to attach the guide RNA to another molecule, such as DNA, RNA, protein, or nanoparticles. This chemically modified guide RNA can be used to recognize or enrich cells edited by nucleases and related systems that are generally guided by nucleic acids (see Lee et al., eLife, 2017, 6:e25312, DOI:10.7554).

[0314] Directly Repeated (DR) Sequences

[0315] This disclosure provides DR sequences corresponding to Cas proteins and their variant peptides, which are capable of binding to Cas proteins or their variant peptides to form complexes. In some embodiments, the DR comprises the sequence shown in SEQ ID NO. 5 or SEQ ID NO. 71, or a nucleotide sequence having at least about 50% (e.g., at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identity with the nucleotide sequence shown in either SEQ ID NO. 5 or 71. In some embodiments, the DR is a “functional variant” (e.g., a “functionally truncated version,” a “functionally extended version,” or a “functionally replaced version”) of the RNA sequence shown in SEQ ID NO. 5 or 71, but still has DR function and retains at least partially (e.g., at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or higher) the function of the reference DR (parental DR).

[0316] In some implementations, the DR sequence includes a stem-loop structure (immediately adjacent spacer sequence) near the 3' end. A "stem-loop structure" refers to a nucleic acid having a secondary structure comprising nucleotide regions known or predicted to form a double-stranded (stem) portion, linked at one end by basic single-stranded nucleotides via a linker region (loop). The term "hairpin" structure is also used herein to refer to stem-loop structures. These structures are well known in the art, and these terms are used in the manner commonly known in the art. Stem-loop structures do not require precise base pairing. Therefore, the stem may contain one or more base mismatches. Alternatively, base pairing may be precise, i.e., not including any mismatches.

[0317] In one embodiment, the guide RNA of this disclosure comprises a stem-loop structure near the 3' end of the DR sequence. In some embodiments, the stem of the DR consists of 5 complementary base pairs that hybridize with each other, and the loop is 6, 7, 8, or 9 nucleotides in length. In some embodiments, the loop is 7 nucleotides in length. In some embodiments, the stem may contain at least 2, at least 3, at least 4, or at least 5 base pairs. In some embodiments, the DR comprises two complementary nucleotide segments of approximately 5 nucleotides in length, separated by approximately 7 nucleotides. In some embodiments, the stem-loop structure comprises a first stem nucleotide chain of 5 nucleotides in length; a second stem nucleotide chain of 5 nucleotides in length, wherein the first and second stem nucleotide chains can hybridize with each other; and a circular nucleotide chain arranged between the first and second stem nucleotide chains, wherein the circular nucleotide chain contains 6, 7, or 8 nucleotides.

[0318] As used herein, two or more guide RNAs having substantially identical or no substantial differences in secondary structure means that the stems and / or loops contained in these crRNAs differ in length by no more than 1, 2, or 3 nucleotides; and in terms of nucleotide type (A, U, G, or C), when comparing the nucleotide sequences of these guide RNAs by sequence alignment, the differences are no more than 1, 2, 3, 4, 5, 6, 7, or 8 nucleotides. In some embodiments, two or more guide RNAs having substantially identical or no substantial differences in secondary structure means that the stems contained in the crRNAs differ by at most one complementary base pair, and / or the loops differ by at most one nucleotide length, and / or contain stems of the same length but with mismatched bases.

[0319] In some embodiments, the stem-ring structure comprises 5'-X1X2X3X4X5NNNNNNNX6X7X8X9X 10 -3';X1,X2,X3,X4,X5,X6,X7,X8,X9,X 10X1, X2, X3, X4, X5 and X6, X7, X8, X9, X... are any bases containing A, T, C, or G; where X1, X2, X3, X4, X5 and X6, X7, X8, X9, X... are any bases containing A, T, C, or G; 10 They can hybridize to form stems and cause NNNNNNN to form loops; more preferably, the DR sequence includes any of the following stem-loop structures near the 3' end of the DR sequence:

[0320] 5'-CCGTCNNNNNNNGACGG-3' (SEQ ID NO.80); where N is any base containing A, T, C or G.

[0321] In some embodiments, the DR sequence that can guide any Cas protein or variant polypeptide of this disclosure to the target site contains one or more nucleotide changes selected from nucleotide addition, insertion, deletion and substitution, which do not result in a substantial difference in secondary structure compared to the DR sequence listed in SEQ ID NO. 5, 71 or a functionally truncated version thereof.

[0322] Spacer sequence

[0323] In some embodiments, the spacer sequence is at least about 15 nucleotides long, preferably about 15 to about 100 nucleotides, more preferably about 15 to about 50 nucleotides (e.g., any one of about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 nucleotides). In some embodiments, the spacer sequence is about 16 to about 27 nucleotides, for example, about 17 to about 24 nucleotides, about 18 to about 24 nucleotides, or about 18 to about 22 nucleotides.

[0324] In some embodiments, the complementarity between the spacer sequence and the target sequence is at least about 70% (e.g., at least about 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%). In some embodiments, there is at least about 15 nucleotide matches between the spacer sequence and the target sequence of the target nucleic acid (e.g., DNA) (e.g., at least about 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or more). For the spacer sequence, perfect complementarity is not required, as long as sufficient complementarity exists for the guide RNA to function (i.e., guide the Cas protein to the target site). In some implementations, Cas protein-mediated cleavage efficiency can be tuned by introducing one or more mismatches between the spacer and target sequences (e.g., one or two mismatches between the spacer and target sequences, including mismatches located along the spacer / target sequence). Mismatches (such as double mismatches) have a greater impact on cleavage efficiency when they are located more centrally within the spacer sequence (i.e., not at the 3' or 5' end of the spacer sequence). Therefore, the cleavage efficiency of the Cas protein can be modulated by selecting the location of the mismatch along the spacer sequence. For example, if a cleavage rate of less than 100% of the target sequence is desired (e.g., in a cell population), one or two mismatches between the spacer and target sequences can be introduced into the spacer sequence.

[0325] In some embodiments, the spacer sequence comprises at least 15 consecutive nucleotides from any one of the nucleotide sequences in SEQ ID NO. 6, 11, 53, 55, 57, 59, 61, 101.

[0326] In some embodiments, the spacer sequence comprises a nucleotide sequence as described in any one of SEQ ID NO. 6, 11, 53, 55, 57, 59, 61, 101.

[0327] PAM; target sequence

[0328] PAM identification assays can be performed according to any of the embodiments described in Karvelis et al. Methods. 2017 May 15; 121-122:3-8 (the entire contents of which are incorporated herein by reference) to identify PAM sequence specificity to allow optimal synthetic sequence targeting. In one embodiment, cells carrying plasmids encoding any of the enzymes described herein and a pre-intercalation region sequence targeting guide RNA are co-transformed with a plasmid library containing an antibiotic resistance gene and a pre-intercalation region sequence flanked by a randomized PAM sequence. Plasmids containing functional PAMs are cleaved by the enzyme, resulting in cell death. Deep sequencing of the pool of enzyme-resistant plasmids isolated from surviving cells reveals depleted plasmids containing functional, cleavable PAMs, thereby identifying PAM preference for Cas proteins.

[0329] In some embodiments, the Cas protein of this disclosure can recognize a PAM (protospacer adjacent motif) to act on target DNA. In some embodiments, the PAM comprises or consists of 5'-TTN-3' (where N is A, T, G, or C). In some embodiments, the PAM comprises or consists of 5'-TTC-3', 5'-TTA-3', 5'-TTT-3', or 5'-TTG-3'.

[0330] As used herein, the term "target nucleic acid" is used interchangeably with "target sequence," "target nucleic acid sequence," or "target nucleic acid molecule," referring to a specific nucleic acid containing a nucleic acid sequence that is wholly or partially complementary to a spacer sequence in a guide RNA. It is a polynucleotide targeted by the spacer sequence in the guide RNA, such as a sequence complementary to that spacer sequence, wherein hybridization between the target sequence and the spacer sequence will promote the formation of a CRISPR-Cas complex (including the Cas protein and the guide RNA), resulting in specific modification of the target DNA by the Cas protein. Complete complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR-Cas complex. In some embodiments, the target sequence contains a non-coding region (e.g., a promoter or terminator). In some embodiments, the target nucleic acid is single-stranded or double-stranded. The target sequence can contain any polynucleotide, such as DNA. In some cases, the target sequence is located intracellularly or extracellularly. In some cases, the target sequence is located within the cell nucleus, cytoplasm, or organelles (e.g., mitochondria or chloroplasts). The target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).

[0331] In some embodiments, the prototype spacer sequence is a segment of about or at least about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more consecutive nucleotides of the target DNA, or a segment of consecutive nucleotides within a numerical range between any two of the foregoing values, such as a segment of about 15 to about 50 or about 17 to about 22 consecutive nucleotides of the target DNA. In some implementations, the prototype spacer sequence is a segment of approximately 20 consecutive nucleotides of the target DNA.

[0332] In some embodiments, the prototype spacer sequence comprises about or at least about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more consecutive nucleotides of the target DNA, or consecutive nucleotides within a numerical range between any two of the foregoing values, for example, about 15 to about 50 or about 16 to about 23 consecutive nucleotides of the target DNA. In some embodiments, the prototype spacer sequence comprises about 20 consecutive nucleotides of the target DNA.

[0333] In some implementations, the target sequence is a continuous nucleotide segment identified from the target strand of the target DNA. The target sequence is the inverse complementary sequence of the prototype spacer sequence. In some embodiments, the target sequence is a segment of about or at least about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more consecutive nucleotides on the target strand of the target DNA, or a segment of consecutive nucleotides in a numerical range between any two of the foregoing values, for example, a segment of about 15 to about 50 or about 16 to about 22 consecutive nucleotides on the target strand of the target DNA. In some implementations, the target sequence is a segment of about 20 consecutive nucleotides on the target strand of the target DNA.

[0334] In some embodiments, the target sequence comprises about or at least about 15 consecutive nucleotides of the target DNA, such as about or at least about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more consecutive nucleotides of the target DNA, or a numerical range between any two of the foregoing values, such as about 16 to about 50 or about 17 to about 22 consecutive nucleotides of the target DNA. In some implementations, the target sequence comprises approximately 20 consecutive nucleotides of the target DNA.

[0335] In some embodiments, the target sequence comprises about or at least about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more consecutive nucleotides on the target strand of the target DNA, or consecutive nucleotides in a numerical range between any two of the foregoing values, for example, about 15 to about 50 or about 17 to about 23 consecutive nucleotides on the target strand of the target DNA. In some implementations, the target sequence comprises approximately 20 consecutive nucleotides on the target strand of the target DNA.

[0336] In some implementations, the inverse complementary sequence of the target sequence is adjacent to the 3' end of the prototype spacer neighboring motif (PAM).

[0337] In some embodiments, the non-target strand is the sense strand of the target DNA. In some embodiments, the non-target strand is the antisense strand of the target DNA. In some embodiments, the target strand is the sense strand of the target DNA. In some embodiments, the target strand is the antisense strand of the target DNA.

[0338] In some embodiments, the prototype spacer sequence or target sequence is located within an exon of the target DNA. In some embodiments, the prototype spacer sequence or target sequence is located within an intron of the target DNA. In some embodiments, the target DNA comprises DNA derived from eukaryotes or DNA derived from prokaryotes. In some embodiments, the target DNA is eukaryotic DNA. In some embodiments, the target nucleic acid includes non-human mammalian DNA, human DNA, insect DNA, avian DNA, reptile DNA, amphibian DNA, rodent DNA, fish DNA, worm DNA, nematode DNA, or yeast DNA;

[0339] In one embodiment, the non-human mammalian DNA includes non-human primate DNA. In some embodiments, the target DNA is associated with disease mutations.

[0340] In some embodiments, the length of the spacer sequence is about or at least about 15 nucleotides, for example, about or at least about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more nucleotides, or the length is a numerical range between any two of the foregoing values, for example, about 15 to about 50 nucleotides, about 17 to about 22 nucleotides, or about 18 to about 20 nucleotides. In some real modes, the spacer sequence is about 20 nucleotides in length.

[0341] In some embodiments, (1) the spacer sequence is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% (complete), optionally about 100% (complete) anticomplementary to the target sequence; (2) the spacer sequence contains no more than 5, 4, 3, 2, or 1 mismatch with the target sequence or contains no mismatch; or (3) the spacer sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, preceding the 5' end of the guide RNA sequence. Nucleotides 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 do not contain mismatches with the target sequence. In some embodiments, the spacer sequence is approximately 100% (completely) anticomplementary to the target sequence.

[0342] In some embodiments, a variety of strategies and methods may be employed to use the systems of this disclosure to modify target DNA in cells. As used herein, “modification” includes, but is not limited to, cutting, nicking, editing, deletion, knockout, knockdown, mutation, correction, exon skipping, etc. In some embodiments, the modification results in correction or compensation of mutations in the cell, thereby producing edited cells that alter the expression of the target DNA, such as decreasing or increasing the level of the target DNA transcript (mRNA) or the level of the target DNA expression product (e.g., protein). In some embodiments, the modification includes suppressing or eliminating the expression of gene products (e.g., reducing it by at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, or more) by knocking down or knocking out the gene. In some embodiments, the modification increases the level of the expression product (e.g., protein) of the target DNA (e.g., by at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, or more). In some embodiments, the modification decreases the level of the expression product (e.g., protein) of the target DNA (e.g., by at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, or more).

[0343] Reagent test kit

[0344] In one aspect, this disclosure provides a kit comprising one or more components selected from: the Cas protein of this disclosure, the fusion protein of this disclosure, the polynucleotide of this disclosure, the complex of this disclosure, the vector of this disclosure, the CRISPR-Cas composition of this disclosure, or the system of this disclosure.

[0345] In some embodiments, the kit further includes one or more labels or instructions, and / or instructions for use in combination with one or more additional components that may be available elsewhere or are required. In some embodiments, the kit further includes instructions for using the kit, such as instructions in more than one language.

[0346] In one aspect, the present invention also provides a container comprising the aforementioned reagent kit.

[0347] In one embodiment, the container includes a sterile container.

[0348] In one embodiment, the container includes a syringe.

[0349] In some embodiments, the kit further includes one or more buffer solutions that can be used to dissolve any one of the components contained therein, and / or to provide suitable reaction conditions for one or more of the components. Such buffer solutions may include one or more of the following: PBS, HEPES, Tris, MOPS, Na₂CO₃, NaHCO₃, NaB, or combinations thereof. In some embodiments, reaction conditions include an appropriate pH, such as an alkaline pH. In some embodiments, the pH is between 7 and 10. In some embodiments, any one or more of the kit components may be stored in a suitable container or at a suitable temperature, such as 4 degrees Celsius. In some embodiments, the kit is used for gene or genome editing, disease treatment, targeting a target gene, cutting a target gene or a non-target gene, or one or more other similar applications.

[0350] deliver

[0351] The components of the CRISPR-Cas system / composition disclosed herein can be delivered in various forms; for example, the Cas protein can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA, or as a protein. The guide can be delivered as a DNA-encoded polynucleotide or RNA. All possible combinations are contemplated, including mixed delivery methods.

[0352] In one aspect, this disclosure provides a delivery composition comprising a delivery medium and one or more of the following: the Cas protein of this disclosure, the fusion protein of this disclosure, the polynucleotide of this disclosure, the vector of this disclosure, the complex of this disclosure, the CRISPR-Cas system of this disclosure, the CRISPR-Cas composition of this disclosure, or the system of this disclosure.

[0353] As used herein, methods for introducing nucleic acids into host cells are known in the art, and any convenient method may be used to introduce polynucleotides (e.g., expression constructs / vectors of Cas proteins) into target cells (e.g., prokaryotic cells, eukaryotic cells, plant cells, animal cells, mammalian cells, human cells, etc.). Suitable methods include, for example, viral infection, transfection, conjugation, protoplast fusion, liposome transfection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (see, for example, Panyam et al. Adv Drug Deliv Rev. Sep 13, 2012. pii:S0169-409X(12)00283-9. doi:10.1016 / j.addr.2012.09.023), etc.

[0354] In some embodiments, the Cas protein of this disclosure may be provided as a polynucleotide encoding the Cas protein (e.g., mRNA, DNA, plasmid, expression vector, viral vector, etc.). In some embodiments, the Cas protein of this disclosure may be provided directly as a protein (e.g., without or with an associated guide RNA, i.e., as a ribonucleoprotein complex (RNP)). The Cas protein of this disclosure may be introduced into cells (provided to cells) by any convenient method; such methods are known to those skilled in the art. Exemplarily, for example, the Cas protein of this disclosure or a fusion protein comprising it may be injected directly into cells (e.g., with or without the Cas protein guide RNA or the nucleic acid encoding the Cas protein guide RNA, and with or without the donor polynucleotide). As another example, a pre-formed complex (RNP) of the Cas protein and Cas protein guide RNA of this disclosure can be introduced into cells (e.g., eukaryotic cells) (e.g., by injection, by nuclear transfection; by conjugation to one or more protein transduction domains (PTDs), such as conjugation to the Cas protein, conjugation to the guide RNA, conjugation to the Cas protein of this disclosure, and guide RNA, etc.).

[0355] In some embodiments, fusion proteins comprising Cas proteins (e.g., dCas fused to a functional domain) are provided as nucleic acids encoding the fusion protein (e.g., mRNA, DNA, plasmids, expression vectors, viral vectors, etc.). In some embodiments, the fusion proteins of this disclosure are provided directly as proteins (e.g., not with an associated guide RNA or with an associated guide RNA, i.e., as a ribonucleoprotein complex (RNP)). Fusion proteins of the Cas proteins of this disclosure can be introduced into cells by any convenient method known to those skilled in the art. As an illustrative example, the fusion proteins of this disclosure can be directly injected into cells (e.g., with or without a nucleic acid encoding a guide RNA, and with or without a donor polynucleotide). As another example, a pre-formed complex of the fusion protein of this disclosure and the Cas protein guide RNA (RNP) can be introduced into cells (e.g., by injection; by nuclear transfection; by a protein transduction domain (PTD) conjugated to one or more components, such as conjugated to the fusion protein, conjugated to the guide RNA, conjugated to the fusion protein and guide RNA of this disclosure, etc.).

[0356] In some embodiments, polynucleotides (e.g., Cas protein guide RNA; polynucleotides containing a nucleotide sequence encoding the Cas protein of this disclosure, etc.) are delivered to cells (e.g., target host cells) and / or to proteins (e.g., Cas proteins; fusion proteins containing Cas proteins) in or bound to particles. In some embodiments, the CRISPR-Cas system of this disclosure is delivered to cells in or associated with particles. The terms "particle" and "nanoparticle" are used interchangeably as appropriate. Recombinant expression vectors containing a nucleotide sequence encoding the Cas protein of this disclosure and / or Cas protein guide RNA, mRNA containing a nucleotide sequence encoding the Cas protein of this disclosure, and guide RNA can be delivered simultaneously using particles or lipid envelopes; for example, the Cas protein and guide RNA, as a complex (e.g., a ribonucleoprotein (RNP) complex), can be delivered via particles, for example, via delivery particles containing lipids or lipid-like substances and hydrophilic polymers (e.g., cationic lipids and hydrophilic polymers).

[0357] The Cas protein of this disclosure (or mRNA containing a nucleotide sequence encoding the Cas protein of this disclosure; or a recombinant expression vector containing a nucleotide sequence encoding the Cas protein of this disclosure) and / or Cas protein guide RNA (or nucleic acid, such as one or more expression vectors encoding Cas protein guide RNA) can be delivered simultaneously using particles or lipid envelopes. For example, biodegradable core-shell nanoparticles having a poly(β-amino ester) (PBAE) core encapsulated by a phospholipid bilayer can be used. In some embodiments, poly(β-amino alcohol) (PBAA) can be used to deliver the Cas protein of this disclosure, the fusion protein of this disclosure, the ribonucleoprotein (RNP) complex of this disclosure, the polynucleotide of this disclosure, or the CRISPR-Cas system of this disclosure to target cells. U.S. Patent Publication No. 20130302401 relates to a class of poly(β-amino alcohol) (PBAA) prepared using combinatorial polymerization.

[0358] In some embodiments, lipid nanoparticles (LNPs) are used to deliver the Cas protein, fusion peptide, RNP, polynucleotide, or CRISPR-Cas system of this disclosure to target cells. Negatively charged polymers (such as RNA) can be loaded into the LNPs at low pH values ​​(e.g., pH 4), where the ionizable lipids exhibit a positive charge. However, at physiological pH values, the LNPs exhibit a low surface charge compatible with longer cycle times.

[0359] In some embodiments, LNPs can be used to deliver DNA molecules (e.g., sequences encoding CasY7 protein and / or guide RNA) and / or RNA molecules (e.g., CasY7 protein, guide RNA). In some embodiments, LNPs can be used to deliver CasY7 / guide RNA RNP complexes. Cationic lipids form a complex with mRNA, forming a lipid complex, which is then endocytosed by the cell. In exemplary embodiments, LNPs comprise cationic lipids, helper lipids, cholesterol, and polyethylene glycol (PEG). In some embodiments, nanoparticles can be developed according to selective organ targeting (SORT), wherein multiple classes of lipid nanoparticles are systematically engineered to specifically edit extrahepatic tissues via the addition of supplemental SORT molecules. In one embodiment, lipid nanoparticles comprise cationic lipids, neutral lipids, cholesterol, PEG lipids, or combinations thereof, wherein a plurality of lipid nanoparticles optionally have an average particle size between 80 nm and 160 nm; and wherein the lipid nanoparticles comprise one or more polynucleotides encoding at least one polypeptide of the present disclosure (e.g., a CasY7 polypeptide).

[0360] As used herein, "nanoparticle" refers to any particle having a diameter of less than 1000 nm. In some embodiments, nanoparticles suitable for delivering the Cas protein, fusion protein, RNP, polynucleotide, or CRISPR-Cas system of this disclosure to target cells have a diameter of 500 nm or less, for example, 25 nm to 35 nm, 35 nm to 50 nm, 50 nm to 75 nm, 75 nm to 100 nm, 100 nm to 150 nm, 150 nm to 200 nm, 200 nm to 300 nm, 300 nm to 400 nm, or 400 nm to 500 nm. In some embodiments, nanoparticles suitable for delivering the Cas protein, fusion protein, RNP, polynucleotide, or CRISPR-Cas system of this disclosure to target cells have a diameter of 25 nm to 200 nm. In some embodiments, nanoparticles suitable for delivering the Cas protein, fusion protein, RNP, polynucleotide, or CRISPR-Cas system of this disclosure to target cells have a diameter of 100 nm or less. In some cases, nanoparticles suitable for delivering the Cas protein, fusion protein, RNP, polynucleotide, or CRISPR-Cas system of this disclosure to target cells have a diameter of 35 nm to 60 nm.

[0361] Nanoparticles suitable for delivering the Cas protein, fusion protein, RNP, polynucleotide, or CRISPR-Cas system of this disclosure to target cells can be provided in various forms, such as as solid nanoparticles (e.g., metallic (such as silver, gold, iron, titanium), nonmetallic, lipid-based solids, polymers), suspensions of nanoparticles, or combinations thereof. Metallic, dielectric, and semiconductor nanoparticles, as well as hybrid structures (e.g., core-shell nanoparticles), can be prepared. If nanoparticles made of semiconductor materials are small enough (typically less than 10 nm) to allow for the quantization of electronic energy levels, they can also be labeled with quantum dots. Such nanoparticles are used as drug carriers or imaging agents in biomedical applications and are suitable for similar purposes as described in this disclosure.

[0362] In some embodiments, semi-solid and soft nanoparticles are also suitable for delivering the Cas proteins, fusion proteins, RNPs, polynucleotides, or CRISPR-Cas systems of this disclosure to target cells. Prototype nanoparticles with semi-solid properties are liposomes.

[0363] In some embodiments, exosomes can be used to deliver the Cas protein, fusion protein, RNP, nucleic acid, or CRISPR-Cas system of this disclosure to target cells. Exosomes are endogenous nanovesicles that transport RNA and proteins and can deliver RNA to the brain and other target organs.

[0364] In some embodiments, liposomes can be used to deliver the Cas proteins, fusion proteins, RNPs, polynucleotides, or CRISPR-Cas systems of this disclosure to target cells. Liposomes are spherical vesicle structures consisting of a single or multiple lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes can be made from several different types of lipids; however, phospholipids are most commonly used for liposome formation. Although liposome formation is spontaneous when the lipid membrane is mixed with an aqueous solution, it can be accelerated by applying force in the form of shaking using a homogenizer, ultrasonic disruptor, or extrusion device. Several other additives can be added to liposomes to modify their structure and properties. For example, cholesterol or sphingomyelin can be added to liposome mixtures to help stabilize the liposome structure and prevent leakage of the liposome contents (inner cargo). Liposome formulations may consist primarily of natural phospholipids and lipids, such as 1,2-distearate-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, lecithin, and monosialotetrazolium ganglioside.

[0365] In some embodiments, lipids can be formulated using the CRISPR-Cas system of this disclosure or one or more components thereof or nucleic acids encoding them to form lipid nanoparticles (LNPs).

[0366] In some embodiments, the CRISPR-Cas system or its components of this disclosure may be delivered encapsulated in PLGA microspheres, such as those further described in U.S. Publications 20130252281, 20130245107, and 20130244279.

[0367] In some embodiments, pressurizing proteins can be used to deliver the Cas proteins, fusion proteins, RNPs, polynucleotides, or CRISPR-Cas systems of this disclosure to target cells. Supercharged proteins are a class of engineered or naturally occurring proteins with exceptionally high net positive or negative theoretical charges. Both supernegative and superpositive proteins exhibit resistance to heat- or chemically induced aggregation. Superpositive proteins are also capable of penetrating mammalian cells. Associating cargo with these proteins (such as plasmid DNA, RNA, or other proteins) enables the functional delivery of these macromolecules to mammalian cells both in vitro and in vivo.

[0368] In some embodiments, when the target of delivery is plant cells, cell-penetrating peptides (CPPs) can be used to deliver the Cas protein, fusion protein, RNP, polynucleotide, or CRISPR-Cas system of this disclosure to the target cells. CPPs typically have an amino acid composition containing a high relative abundance of positively charged amino acids (such as lysine or arginine), or a sequence containing alternating patterns of polar / charged amino acids and nonpolar hydrophobic amino acids.

[0369] In some embodiments, an implantable device may be used to deliver the Cas protein, fusion protein, RNP, nucleic acid (e.g., Cas protein guide RNA, nucleic acid encoding Cas protein guide RNA, polynucleotide encoding Cas protein, donor template, etc.) of this disclosure to target cells (e.g., in vivo target cells, wherein the target cells are circulating target cells, tissue target cells, organ target cells, etc.). An implantable device suitable for delivering the Cas protein, fusion protein, RNP, polynucleotide, or CRISPR-Cas system of this disclosure to target cells (e.g., in vivo target cells, wherein the target cells are circulating target cells, tissue target cells, organ target cells, etc.) may include a container (e.g., a reservoir, matrix, etc.) containing the Cas protein, fusion protein, RNP, or CRISPR-Cas system (or components thereof, such as the nucleic acid of this disclosure).

[0370] Suitable implantable devices may include, for example, a polymeric substrate (such as a matrix) serving as the device body, and in some cases, additional scaffold materials (such as metals or other polymers), as well as materials that enhance visibility and imaging. Implantable delivery devices can advantageously provide local and long-term release, wherein the peptides and / or nucleic acids to be delivered are released directly to target sites, such as the extracellular matrix (ECM), the vascular system surrounding a tumor, diseased tissue, etc. Suitable implantable delivery devices include devices suitable for delivery to cavities such as the peritoneal cavity and / or any other type of administration in which the drug delivery system is not anchored or attached, comprising biostable and / or degradable and / or bioabsorbable polymeric matrices, which may optionally be, for example, a matrix. In some cases, suitable implantable drug delivery devices contain degradable polymers, wherein the primary release mechanism is bulk erosion. In some cases, suitable implantable drug delivery devices contain non-degradable or slowly degradable polymers, wherein the primary release mechanism is diffusion rather than bulk erosion, such that the outer portion acts as a membrane and its inner portion acts as a drug reservoir, which, in effect, is unaffected by the surrounding environment for a long period (e.g., from about a week to about several months). Alternatively, combinations of different polymers with different release mechanisms can be used. Throughout the effective period of total release, the concentration gradient can remain effectively constant, and therefore the diffusion rate is effectively constant (referred to as “zero-mode” diffusion). As used herein, the term “constant” means that the diffusion rate is maintained above a lower threshold of therapeutic effectiveness, but it still optionally features an initial burst and / or may fluctuate, e.g., increase and decrease to a certain extent. The diffusion rate can be maintained in this way for a long period, and the diffusion rate can be considered constant at a certain level to optimize the therapeutic duration, e.g., an effective period of silence.

[0371] In some embodiments, the implantable delivery system is designed to protect the nucleotide-based therapeutic agent from degradation, whether chemical or due to attack by enzymes and other factors in the subject's body. The implantation site or target site of the device can be selected to obtain maximum therapeutic efficacy. For example, the delivery device can be implanted within or near the tumor environment, or within or near the tumor-associated blood supply. Insertion methods (such as implantation) may optionally have been used for other types of tissue implantation and / or for insertion and / or for tissue sampling, optionally without modification, or alternatively with only optional minor modifications in such methods. Such methods optionally include, but are not limited to, brachytherapy, biopsy, endoscopy with and / or without ultrasound (such as stereotactic approaches to access brain tissue), and laparoscopy (including laparoscopic implantation into joints, abdominal organs, bladder walls, and body cavities). In some embodiments, the delivery device may be selected from gene guns, pre-filled syringes, autoinjectors, aerosol sprays, spray cans, or patch syringes.

[0372] In some implementations, the delivery medium is nanoparticles, liposomes, exosomes, microvesicles, or a gene gun.

[0373] carrier

[0374] As used herein, a vector is a nucleic acid molecule capable of delivering another nucleic acid molecule linked to it. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. A vector can be introduced into a host cell through transformation, transduction, or transfection, thereby enabling the expression of its carried genetic material elements in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including protein variants, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. The vector may also contain a replication initiation site.

[0375] Vectors include plasmids and viral vectors. A plasmid is a circular double-stranded DNA loop in which another DNA fragment can be inserted, for example, using standard molecular cloning techniques. A viral vector contains a virus-derived DNA or RNA sequence within a vector used to package the virus; viruses include, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. Viral vectors also contain polynucleotides carried by a virus intended for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and augmented mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.

[0376] Other vectors (e.g., non-attachment mammalian vectors) integrate into the host cell's genome after introduction and thereby replicate along with the host genome. Furthermore, some vectors can direct the expression of genes they are operatively linked to. Such vectors are called "expression vectors."

[0377] In some embodiments, the vector (e.g., a viral vector or a non-viral vector, such as a lentiviral vector or plasmid) can be delivered to the target tissue via, for example, intramuscular injection, intravenous administration, percutaneous administration, intranasal administration, oral administration, or mucosal administration. The delivery can be performed via a single dose or multiple doses. Those skilled in the art will understand that the actual dose to be delivered herein can vary considerably depending on a variety of factors, including but not limited to the choice of vector, target cells, organism, tissue, general condition of the subject to be treated, the degree of transformation / modification sought, the route of administration, the manner of administration, and the type of transformation / modification sought.

[0378] In some aspects, this disclosure provides a vector comprising the polynucleotide of this disclosure, or a nucleotide encoding the guide RNA of this disclosure.

[0379] In some embodiments, the vector includes plasmids and viral vectors;

[0380] In some embodiments, the viral vector is selected from the group consisting of adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpesvirus, SV40, poxvirus, or combinations thereof;

[0381] In some implementations, the vector includes cloning vectors, transformation vectors, expression vectors, shuttle vectors, integration vectors, and multifunctional vectors;

[0382] In some embodiments, the carrier comprises:

[0383] (1) A first regulatory element, operatively linked to a nucleotide sequence encoding a Cas protein disclosed herein or a nucleotide sequence encoding a fusion protein disclosed herein; and

[0384] (2) A second regulatory element, operatively linked to a nucleotide sequence encoding a guide RNA, the guide RNA comprising:

[0385] (a) A direct repeat (DR) sequence capable of forming a complex with the Cas protein or fusion protein of this disclosure, and

[0386] (b) A spacer sequence that can hybridize with the target sequence of the target DNA, thereby guiding the complex to the target DNA.

[0387] In some embodiments, the first control element and the second control element are located on the same or different carriers.

[0388] In some implementations, the first regulating element and / or the second regulating element is a promoter, such as an inductive promoter.

[0389] In some embodiments, the vector includes one or more promoters operatively linked to the nucleic acid sequence, enhancer, transcription termination signal, polyadenylation sequence, origin of replication, selectivity marker, nucleic acid restriction site, and / or homologous recombination site.

[0390] In one aspect, this disclosure provides a CRISPR-Cas system comprising one or more vectors, said one or more vectors comprising:

[0391] (1) A first regulatory element, operatively linked to a nucleotide sequence encoding the Cas protein or a variant thereof, or a nucleotide sequence encoding the fusion protein; and

[0392] (2) A second regulatory element, operatively linked to a nucleotide sequence encoding the guide RNA, the guide RNA comprising:

[0393] (a) Spacer sequences capable of hybridizing with the target sequence of the target nucleic acid, and

[0394] (b) A direct repeat (DR) sequence attached to the spacer sequence that guides the Cas protein or a variant thereof to bind to the guide RNA to form a CRISPR-Cas complex targeting the target sequence.

[0395] The first control element and the second control element are located on the same or different carriers of the CRISPR-Cas carrier system.

[0396] In one embodiment, the first or second regulatory element includes a promoter, which includes one or more of an inductive promoter, a constitutive promoter, or a tissue-specific promoter.

[0397] In one embodiment, the promoter includes one or more of T7, SP6, T3, CMV, EF1a, SV40, PGK1, humanβ-actin, CAG, U6, H1, T7, T7lac, araBAD, trp, lac, or Ptac.

[0398] In one embodiment, the first control element and the second control element are located on the same or different carriers.

[0399] In one embodiment, the vector includes a retroviral vector, a lentiviral vector, an adenovirus vector, an adeno-associated virus vector, a herpes simplex vector, or a phage vector. In one embodiment, the vector includes a plasmid vector.

[0400] In some aspects, this disclosure provides a delivery composition comprising a delivery medium and one or more of the following: the Cas protein of this disclosure, the fusion protein of this disclosure, the polynucleotide of this disclosure, the vector of this disclosure, the complex of this disclosure, the CRISPR-Cas system of this disclosure, the CRISPR-Cas composition of this disclosure, or the system of this disclosure.

[0401] host cells

[0402] The system disclosed herein can be introduced into a host cell and cause the host cell to alter the production of one or more cellular products (such as antibodies, starch, ethanol, or any other desired product). Such host cells and their progeny are within the scope of this disclosure. As used herein, “host cell” means eukaryotic cells (e.g., animal cells, plant cells, fungal cells, etc.), prokaryotic cells (e.g., some microbial cells, Escherichia coli, Substance Bacillus, etc.), or cells derived from multicellular organisms (e.g., cell lines) cultured as single-celled entities, said cells serving as recipients of nucleic acids (e.g., expression vectors), and includes progeny of the original cells that have been genetically modified with nucleic acids.

[0403] It should be understood that the offspring of a single cell may be attributed to natural, accidental, or intentional mutations and may not necessarily have the exact same morphology or genome as the original parent cell. A “recombinant host cell” (also called a “genetically modified host cell”) is a host cell in which a heterologous nucleic acid, such as an expression vector, has been introduced. Those skilled in the art will understand that the design of the expression vector can depend on factors such as the selection of the host cell to be transformed and the desired expression level.

[0404] In one aspect, this disclosure provides a host cell or its progeny comprising the Cas protein of this disclosure, or the fusion protein of this disclosure, or the polynucleotide of this disclosure, or the vector system of this disclosure, or the CRISPR-Cas system of this disclosure, or the composition of this disclosure.

[0405] In some embodiments, the cell is a stem cell. In some embodiments, the cell is not a human embryonic stem cell. In some embodiments, the cell is not a human germ cell.

[0406] In some implementations, the cells are prokaryotic cells.

[0407] In some embodiments, the cell is a eukaryotic cell (e.g., animal cell, vertebrate cell, mammalian cell, non-human mammalian cell, non-human primate cell, rodent (e.g., mouse or rat) cell, human cell, plant cell, or yeast cell) or a prokaryotic cell (e.g., bacterial cell).

[0408] In some implementations, the cells are derived from plants or animals.

[0409] In some embodiments, the plant is a dicotyledonous plant. In some embodiments, the dicotyledonous plant is selected from soybean, cabbage (e.g., Chinese cabbage), rapeseed, Brassica, watermelon, cantaloupe, potato, tomato, tobacco, eggplant, pepper, cucumber, cotton, alfalfa, and grape. In some embodiments, the plant is a monocotyledonous plant. In some embodiments, the monocotyledonous plants are selected from rice, corn, wheat, barley, oats, sorghum, millet, grasses, Poaceae, Zizania, Avena, Coix, Hordeum, Oryza, Panicum (e.g., millet), Secale, Setaria (e.g., foxtail millet), Sorghum, Triticum, Zea, Cymbopogon, Saccharum (e.g., sugarcane), Phyllostachys, Dendrocalamus, Bambusa, and Yushania.

[0410] In some implementations, the animal is selected from pigs, cattle, sheep, goats, mice, rats, alpacas, monkeys, rabbits, chickens, ducks, geese, and fish (e.g., zebrafish).

[0411] In some embodiments, the cells are eukaryotic cells, injected with mammalian cells, including human cells (primary human cells or established human cell lines). In some embodiments, the cells are non-human mammalian cells, such as cells from non-human primates (e.g., monkeys), cows / bulls / cattle, sheep, goats, pigs, horses, dogs, cats, rodents (e.g., rabbits, mice, rats, hamsters, etc.). In some embodiments, the cells are derived from fish (e.g., salmon), birds (e.g., poultry, including chickens, ducks, geese), reptiles, shellfish (e.g., oysters, clams, lobsters, shrimp), insects, worms, yeast, etc. In some embodiments, the cells are derived from plants, such as monocots or dicots. In some embodiments, the plants are food crops, such as barley, cassava, cotton, peanuts or peanuts, corn, millet, oil palm fruit, potatoes, dried beans, rapeseed or canola, rice, rye, sorghum, soybeans, sugarcane, sugar beets, sunflowers, and wheat. In some embodiments, the plant is a cereal (barley, corn, millet, rice, rye, sorghum, and wheat). In some embodiments, the plant is a tuber (cassava and potato). In some embodiments, the plant is a sugar crop (beet and sugarcane). In some embodiments, the plant is an oilseed crop (soybean, peanut, rapeseed or low-erucic acid rapeseed, sunflower, and oil palm fruit). In some embodiments, the plant is a fiber crop (cotton). In some embodiments, the plant is a tree (such as a peach or nectarine tree, apple or pear tree, nut tree (such as an almond or walnut or pistachio tree), or citrus tree (e.g., orange, grapefruit, or lemon tree)), grass, vegetable, fruit, or algae. In some implementations, the plants are plants of the genera *Solanum*, *Brassica*, *Lactuca*, *Spinacia*, *Capsicum*, cotton, tobacco, asparagus, carrots, cabbage, broccoli, cauliflower, tomatoes, eggplants, peppers, lettuce, spinach, strawberries, blueberries, raspberries, blackberries, grapes, coffee, cocoa, etc.

[0412] In one embodiment, the host cell includes non-human mammals, humans, insects, birds, reptiles, amphibians, rodents, fish, worms, nematodes, or yeast cells.

[0413] In another aspect, the present invention also provides a multicellular organism comprising the aforementioned cells or their descendants.

[0414] In one embodiment, the multicellular organism is an animal or plant model used for the relevant disease.

[0415] In another aspect, this disclosure provides a cell modified by the system or method of this disclosure. In some embodiments, the cell is a eukaryotic organism. In some embodiments, the cell is a human cell. In some embodiments, the cell is modified in vitro, in vivo, or ex vivo.

[0416] enzyme preparations

[0417] In one aspect, this disclosure provides an enzyme preparation comprising the Cas protein of this disclosure, the fusion protein of this disclosure, the complex of this disclosure, the CRISPR-Cas system of this disclosure, the CRISPR-Cas composition of this disclosure, or the system or delivery composition of this disclosure.

[0418] In some embodiments, the enzyme preparation includes injectable and / or lyophilized formulations.

[0419] Modification methods

[0420] The CRISPR-Cas system disclosed herein comprises the Cas protein or fusion protein of this disclosure and has multiple uses, including modifying (e.g., cleaving, deletion, insertion, translocation, inactivation, or activation) target DNA in various cell types. The methods and / or systems of this disclosure can be used to modify target DNA, for example, to modify the translation and / or transcription of one or more genes in a cell. For example, modification can lead to an increase in gene transcription / translation / expression. In other embodiments, modification can lead to a decrease in gene transcription / translation / expression.

[0421] As used herein, cleavage refers to a break in the DNA of the target nucleic acid produced by the Cas protein of this disclosure. In some embodiments, cleavage is a double-stranded DNA break. In some embodiments, cleavage is a single-stranded DNA break. In some embodiments, cleavage is the non-spacer complementary strand of the double-stranded target nucleic acid creating a nick. In some embodiments, cleavage is the cutting of the two strands of the double-stranded nucleic acid at different sites, resulting in staggered cleavage.

[0422] In one aspect, this disclosure provides a method for targeting and editing or cutting a target gene, comprising: contacting the Cas protein of this disclosure, or the fusion protein of this disclosure, or the complex of this disclosure, or the CRISPR-Cas system of this disclosure, the composition of this disclosure, or the system of this disclosure, or the delivery composition of this disclosure, or the enzyme preparation of this disclosure, or the kit of this disclosure with target DNA, or delivering it to a cell containing target DNA, wherein the target sequence is present in the target DNA.

[0423] In another aspect, this disclosure also provides a method for inducing changes in cell state, comprising contacting a Cas protein of this disclosure, or a fusion protein of this disclosure, or a complex of this disclosure, or a CRISPR-Cas system of this disclosure, a composition of this disclosure, or a system of this disclosure, or a delivery composition of this disclosure, or an enzyme preparation of this disclosure, or a kit of this disclosure with a target gene in a cell. In some embodiments, the cell state includes apoptosis or dormancy.

[0424] In another aspect, this disclosure also provides a method for altering the expression of a gene product, comprising: contacting a Cas protein of this disclosure, or a fusion protein of this disclosure, or a complex of this disclosure, or a CRISPR-Cas system of this disclosure, a composition of this disclosure, or a system of this disclosure, or a delivery composition of this disclosure, or an enzyme preparation of this disclosure, or a kit of this disclosure with a nucleic acid molecule encoding a gene product, or delivering it to a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

[0425] In another aspect, this disclosure provides a method for modifying target DNA, the method comprising contacting the target DNA with a Cas protein of this disclosure, or a fusion protein of this disclosure, or a complex of this disclosure, or a CRISPR-Cas system of this disclosure, a composition of this disclosure, or a system of this disclosure, or a delivery composition of this disclosure, or an enzyme preparation of this disclosure, or a kit of this disclosure. In some embodiments, the target DNA is intracellular. In some embodiments, the modification includes one or more of the following: cleavage of the target DNA, base editing, repair, and insertion or integration of a foreign sequence.

[0426] In some embodiments, the method of modifying target DNA according to this disclosure includes contacting the target DNA with: the Cas protein of this disclosure, the guide RNA of this disclosure, and a donor nucleic acid, such as a donor template. In some embodiments, the donor template is a double-stranded or single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear or circular (e.g., a plasmid). In some instances, the donor template is a foreign nucleic acid molecule. In some embodiments, the donor template is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, gene recombination, specifically homologous recombination, can be achieved using the donor template.

[0427] Pharmaceutical Composition

[0428] In one aspect, this disclosure provides a pharmaceutical composition comprising a complex of this disclosure, a CRISPR-Cas system of this disclosure, a composition of this disclosure, a system of this disclosure, a delivery composition of this disclosure, or a host cell of this disclosure, and a pharmaceutically acceptable excipient.

[0429] Suitable pharmaceutically acceptable excipients typically include inert substances that facilitate administration of the pharmaceutical composition to a subject, facilitate formulation of the pharmaceutical composition into a deliverable formulation, or facilitate storage of the pharmaceutical composition prior to administration. Pharmaceutically acceptable carriers can include agents that can stabilize, optimize, or otherwise modify the form, consistency, viscosity, pH, pharmacokinetics, or solubility of a formulation. Such agents include buffers, wetting agents, emulsifiers, diluents, encapsulating agents, and skin penetration enhancers. For example, carriers can include, but are not limited to, saline, buffered saline, dextran, arginine, sucrose, water, glycerol, ethanol, sorbitol, dextran, sodium carboxymethyl cellulose, and combinations thereof.

[0430] Some non-limiting examples of substances that can be used as pharmaceutically acceptable carriers include: (1) sugars, such as lactose, glucose and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose and its derivatives, such as sodium carboxymethyl cellulose, methyl cellulose, ethyl cellulose, microcrystalline cellulose and cellulose acetate; (4) powdered tragacanth gum; (5) malt; (6) gelatin; (7) lubricants, such as magnesium stearate, sodium lauryl sulfonate and talc; (8) excipients, such as cocoa butter and suppository wax; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil and soybean oil; (10) glycols, such as propylene glycol. (11) Polyols, such as glycerol, sorbitol, mannitol and polyethylene glycol (PEG); (12) Esters, such as ethyl oleate and ethyl laurate; (13) Agar; (14) Buffers, such as magnesium hydroxide and aluminum hydroxide; (15) Alginate; (16) Pyrogen-free water; (17) Isotonic saline; (18) Ringer's solution; (19) Ethanol; (20) pH buffer solution; (21) Polyesters, polycarbonates and / or polyanhydrides; (22) Fillers, such as peptides and amino acids; (23) Serum alcohols, such as ethanol; and (24) Other non-toxic and compatible substances used in pharmaceutical formulations. The formulation may also contain wetting agents, colorants, release agents, coating agents, sweeteners, flavoring agents, fragrances, preservatives and antioxidants.

[0431] Pharmaceutical compositions may contain one or more pH buffering compounds to maintain the pH of the formulation at a predetermined level reflecting physiological pH, such as in the range of about 5.0 to about 8.0. pH buffering compounds for aqueous liquid formulations may be amino acids or mixtures of amino acids, such as histidine or mixtures of amino acids (such as histidine and glycine). Alternatively, the pH buffering compound is preferably an agent that maintains the pH of the formulation at a predetermined level (such as in the range of about 5.0 to about 8.0) and does not chelate calcium ions. Illustrative examples of such pH buffering compounds include, but are not limited to, imidazole and acetate ions. The pH buffering compound may be present in any amount suitable for maintaining the pH of the formulation at the predetermined level.

[0432] The pharmaceutical composition may also contain one or more osmotic modifiers, which adjust the osmotic properties of the formulation (e.g., tone, osmotic pressure, and / or osmolarity) to a level acceptable to the recipient individual's blood flow and blood cells. The osmotic modifier may be an agent that does not chelate calcium ions. The osmotic modifier may be any compound known or available to those skilled in the art for adjusting the osmotic properties of the formulation. Those skilled in the art can empirically determine the suitability of a given osmotic modifier in the formulation of the present invention. Illustrative examples of suitable types of osmotic modifiers include, but are not limited to: salts, such as sodium chloride and sodium acetate; sugars, such as sucrose, dextran, and mannitol; amino acids, such as glycine; and mixtures of one or more of these agents and / or dosage forms. One or more osmotic modifiers may be present at any concentration sufficient to adjust the osmotic properties of the formulation.

[0433] Treatment

[0434] In one aspect, this disclosure provides a method for diagnosing, preventing, or treating a disease in a subject in need, the method comprising administering (e.g., a therapeutically effective dose) to the subject a Cas protein of this disclosure, a fusion protein of this disclosure, a polynucleotide of this disclosure, a vector of this disclosure, a complex of this disclosure, a CRISPR-Cas system of this disclosure, a CRISPR-Cas composition of this disclosure, or a system of this disclosure, or a kit of this disclosure, or a delivery composition of this disclosure, or an enzyme preparation of this disclosure, or a cassette of this disclosure, or a pharmaceutical composition of this disclosure, wherein the disease is associated with a target DNA, wherein the spacer sequence is capable of hybridizing with a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnoses, prevents, or treats the disease.

[0435] As used herein, the term “treatment” means treating or curing a subject’s condition, delaying the onset of symptoms of the condition, and / or delaying the severity of the condition.

[0436] In some embodiments, suitable routes of administration include, but are not limited to: local, subcutaneous, transdermal, intradermal, intralesional, intra-articular, intraperitoneal, intrabladder, transmucosal, gingival, intradental, intracochlear, transtympanic membrane, intra-organ, epidural, intrasheath, intramuscular, intravenous, intravascular, intraosseous, periorbital, intratumoral, intracerebral, and intraventricular administration. In some embodiments, administration to the subject is by injection, via catheter, via suppository, or via implantation, said implantation being a porous, non-porous, or gel-like material, including membranes such as salivary membranes or fibers.

[0437] In some implementations, the target DNA encodes mRNA, tRNA, ribosomal RNA, microRNA (miRNA), non-coding RNA, long non-coding (lnc) RNA, nuclear RNA, interfering RNA, small interfering RNA (siRNA), ribozyme, riboswitch, satellite RNA, microswitch, yeast microzyme, or viral RNA.

[0438] In some embodiments, the target DNA is eukaryotic DNA. In some embodiments, the eukaryotic DNA is mammalian DNA (such as non-human mammalian DNA), non-human primate DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, rodent (e.g., mouse, rat) DNA, fish DNA, nematode DNA, or yeast DNA.

[0439] In some implementations, the target DNA is located within eukaryotic cells (e.g., human cells, non-human primate cells, or mouse cells).

[0440] In some embodiments, administration includes local or systemic application. In some embodiments, administration is by injection or infusion. In some embodiments, the subject is a human, a non-human primate, or a mouse.

[0441] In some implementations, the level of the target DNA transcript (e.g., mRNA) in the subject is decreased or increased by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the target DNA transcript (e.g., mRNA) in the subject prior to administration.

[0442] In some embodiments, the level of the target DNA expression product (e.g., protein) is decreased or increased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the target DNA expression product (e.g., protein) in the subject prior to administration. In some embodiments, the expression product is a functional mutant of the target DNA expression product.

[0443] In some embodiments, the disease is selected from treating non-dividing cell diseases, and in one embodiment, the gene or transcript to be corrected is located in a non-dividing cell. Exemplary non-dividing cells are muscle cells or neurons.

[0444] In some embodiments, the disease is selected from diseases that affect the eye, such as glaucoma or retinal degenerative diseases. In one embodiment, a lentiviral vector is selected for administration to the eye.

[0445] In some embodiments, the disease is selected from muscle diseases and cardiovascular diseases. In some embodiments, the disease includes Duchenne muscular dystrophy (DMD).

[0446] In some embodiments, the disease is selected from liver and kidney diseases. In some embodiments, the disease may also be selected from epithelial diseases, lung diseases, skin diseases, and cancer.

[0447] In some embodiments, the disease is selected from familial hypercholesterolemia (FH), atherosclerosis (ASCVD), transthyretin amyloidosis (ATTR), α-1-antitrypsin deficiency (AATD), primary hyperoxaluria (PH1), hereditary angioedema (HAE), Angelman syndrome (AS), Alzheimer's disease (AD), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), diabetes, progressive pseudohypertrophic muscular dystrophy, Duchenne muscular dystrophy (DMD), Becker muscle dystrophy. Benign diseases (BMD), spinal muscular atrophy (SMA), Pompe disease, myotonic dystrophy, Huntington's disease (HTT), Fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, thalassemia (e.g., β-thalassemia), Parkinson's disease (PD), myelodysplastic syndromes (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, non-alcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), hypercholesterolemia, heart disease (e.g., hypertrophic cardiomyopathy (HCM)), and cancer.

[0448] In some implementations, the median survival of subjects who have the disease but receive the treatment is 5 days, 10 days, 20 days, 30 days, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 1.5 years, 2 years, 2.5 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, or longer than the median survival of subjects who have the disease but do not receive the treatment or the subject population.

[0449] The effective therapeutic dose can be administered via a single dose or multiple doses. Those skilled in the art will understand that the actual dose can vary considerably depending on a variety of factors, such as the choice of carrier, target cells, organism, tissue, general condition of the subject to be treated, the degree of transformation / modification sought, route of administration, mode of administration, and type of transformation / modification sought.

[0450] In one embodiment, administration further includes administering the CRISPR-Cas composition to the subject or to ex vivo cells of the subject.

[0451] Detection methods

[0452] In another aspect, this disclosure provides a method for detecting the presence of a target nucleic acid molecule in a sample, characterized in that the method includes contacting the sample with a Cas protein, a fusion protein, or a complex of the present disclosure, a CRISPR-Cas system, a CRISPR-Cas composition, or a system, a kit, a delivery composition, or an enzyme preparation of the present disclosure, and contacting a non-target sequence, detecting a detectable signal generated by the cleavage of the non-target sequence, thereby detecting the target nucleic acid molecule, wherein the non-target sequence does not hybridize with guide RNA. In some embodiments, modifications are used to generate a detectable signal, such as a fluorescent signal. In some embodiments, a reporter nucleic acid is used to detect target DNA; the reporter nucleic acid is a molecule that can be cleaved or otherwise deactivated by an activated CRISPR system protein as described herein. The reporter nucleic acid comprises a nucleic acid element that can be cleaved by a CRISPR protein (e.g., a single-stranded non-target nucleic acid molecule with different reporter groups or label molecules at both ends). Cleavage of the nucleic acid element generates a detectable signal. Before cleavage, or while the reporter nucleic acid is in an "active" state, the reporter nucleic acid prevents the generation or detection of a positive detectable signal. It will be understood that, in some exemplary embodiments, minimal background signal may be generated in the presence of an active reporter nucleic acid. The positive detectable signal can be any signal detectable using optical, fluorescent, chemiluminescent, electrochemical, or other detection methods known in the art. For example, in some embodiments, a first signal (i.e., a negative detectable signal) may be detected in the presence of a reporter nucleic acid, and then converted into a second signal (e.g., a positive detectable signal) upon detection of the target molecule and upon cleavage or attenuation by an activated CRISPR protein. The reporter nucleic acid can be a single-stranded DNA molecule, a single-stranded RNA molecule, or a single-stranded DNA-RNA hybrid.

[0453] The detection method described in this invention can be used for the quantitative detection of target nucleic acids. The quantitative detection index can be determined based on the signal strength of the reporter group, such as the luminescence intensity of the fluorescent group or the width of the colored band.

[0454] use

[0455] In one aspect, the use of the Cas protein, fusion protein, polynucleotide, vector, complex, CRISPR-Cas system, CRISPR-Cas composition, system, kit, delivery composition, enzyme preparation, or cassette disclosed herein for the preparation of a drug or formulation for nucleic acid editing (e.g., gene or genome editing).

[0456] In another aspect, the use of the Cas protein, fusion protein, polynucleotide, vector, complex, CRISPR-Cas system, CRISPR-Cas composition, system, kit, delivery composition, enzyme, or cassette disclosed herein for the preparation of a drug or formulation for use in one or more of the following groups:

[0457] (i) In vitro gene or genome editing;

[0458] (ii) Detection of isolated single-stranded DNA;

[0459] (iii) Editing target sequences in target loci to modify biological or non-human organisms;

[0460] (iv) Treating conditions caused by defects in target sequences at target loci;

[0461] (v) Treat the symptoms or diseases of the subject in need.

[0462] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.

[0463] Unless otherwise specified, the reagents and materials used in the embodiments of this invention are all commercially available products.

[0464] Example 1. Obtaining the Cas protein

[0465] Metagenomic analysis of uncultured organisms, including redundancy removal and protein clustering, identified a novel Cas protein. Blast analysis showed low sequence identity with previously reported Cas proteins, and this Cas protein was named CasY7.

[0466] The amino acid sequence of the CasY7 protein is shown in SEQ ID NO.1, and its nucleotide coding sequence after human codon optimization is shown in SEQ ID NO.2.

[0467] Analysis of the direct repeat (DR) sequence of the guide RNA corresponding to the CasY7 protein using CRISPR locus annotation revealed the following results:

[0468] The DNA sequence encoding the direct repeat (DR) sequence of the guide RNA corresponding to the CasY7 protein is as follows:

[0469] CAAGTTGAATCCGTCTATAACTGACGG (SEQ ID NO: 5).

[0470] RNA secondary structure of the DR sequence in pre-crRNA was further analyzed using RNAfold, and the results are as follows: Figure 2 As shown, any "T" in the sequence represents "U" when referring to the DR sequence.

[0471] PAM library depletion experiments (conducted according to any of the embodiments of Karvelis et al. Methods. 2017 May 15; 121-122:3-8 (the entire contents of which are incorporated herein by reference) were used to identify the PAM sequence specificity of CasY7. Analysis revealed that the PAM corresponding to CasY7 is 5'-TTN, with N being A / T / C / G. The sgRNA (also known as crRNA) sequence of the CasY7 protein consists of a spacer sequence and a direct repeat (DR) sequence. Based on these findings, the CasY7 of this invention belongs to the Cas12 protein family.

[0472] Example 2. Validation of Cas protein cleavage activity

[0473] 1. Plasmid construction

[0474] (1) The TTR gene was used as the target, and a spacer sequence (Target-TTR-spacer1) was designed based on the target sequence of the TTR gene: GCATCTCCCCATTCCATGAG (SEQ ID NO.6).

[0475] Based on the DR sequences of CasY7 and LbCpf1 proteins, sgRNA sequences targeting TTR target genes were designed, as detailed in the table below:

[0476]

[0477] According to the requirements of vector expression, a T7 promoter and an rrnB T2 terminator were added to the 5' and 3' ends of the aforementioned sgRNA sequences of CasY7 and LbCpf1, respectively, to obtain the CasY7-sgRNA1 expression fragment sequence:

[0478] The single-underlined sequence is the CasY7DR sequence, the double-underlined sequence is the spacer sequence, the italicized sequence is the T7 promoter, the wavy-underlined sequence is the rrnB T2 terminator sequence, the sequence between the spacer and rrnB T2 terminator sequences is the linker sequence, the dashed sequence is the MfeI restriction site, the bold sequence is the MluI restriction site, and CACCG is the linker.

[0479] In the same manner, the LbCpf1-sgRNA1 expression fragment sequence was synthesized:

[0480] The sequence is divided into several parts: the single underlined part is the LbCpf1DR sequence, the double underlined part is the spacer sequence, the italicized part is the T7 promoter, the wavy underlined part is the rrnB T2 terminator sequence, the spacer sequence and the rrnB T2 terminator sequence are the linker sequence, the dashed part is the MfeI restriction site, the bold part is the MluI restriction site, and CACCG is the linker.

[0481] To protect sequence integrity, AGC was introduced at the 5' end and ATA was introduced at the 3' end as protective bases when synthesizing the CasY7-sgRNA1 expression fragment sequence and the LbCpfl-sgRNA1 expression fragment sequence.

[0482] (2) The nucleotide sequence encoding CasY7 protein (as shown in SEQ ID NO. 2) was synthesized by Suzhou Hongxun Biotechnology Co., Ltd., and the synthesized nucleotide sequence encoding CasY7 protein was constructed into positions 466-5160 of the ABE8e plasmid (Addgene, Plasmid #138489) to obtain the CasY7 recombinant expression plasmid (see Figure 1 for the plasmid map). Figure 1A ).

[0483] The encoding nucleotide sequence (SEQ ID NO. 20) of LbCpf1 (amino acid sequence as shown in SEQ ID NO. 19), synthesized and optimized by Suzhou Hongxun Biotechnology Co., Ltd., was used to construct the LbCpf1 recombinant expression plasmid in the same manner (see Figure 1 for plasmid map). Figure 1B ).

[0484] (3) The sgRNA expression sequence fragments (CasY7-sgRNA1 expression fragment sequence and LbCpf1-sgRNA1 expression fragment sequence) described in step (1) were synthesized by Suzhou Hongxun Biotechnology Co., Ltd., and the sgRNA expression fragment sequence was double-digested (MfeI / MluI) and inserted into the CasY7 recombinant expression plasmid vector that was also double-digested (MfeI / MluI) to obtain the recombinant expression plasmid CasY7+sgRNA1 expressing CasY7 and sgRNA; the LbCpf1+sgRNA1 expression plasmid was constructed in the same way.

[0485] (4) Construct the Target plasmid with the target sequence. The construction process is as follows:

[0486] The araC-pBAD-CCDB fragment (SEQ ID NO.18) carrying the TTR target sequence (SEQ ID NO.6) was synthesized by Suzhou Hongxun Biotechnology Co., Ltd., and inserted into positions 1284-1300 of the pKESK22 (Addgene, Plasmid#64857) plasmid to obtain the Target plasmid. The sequence of the Target plasmid is shown in SEQ ID NO.4, and the plasmid map is shown in [link to plasmid map]. Figure 3 .

[0487] 2. Preparation and transformation of competent Escherichia coli cells

[0488] The Target plasmid was transferred into DH5α competent cells and isolated by streaking with an inoculating loop onto LB agar containing 50 μg / ml kanamycin sulfate. The cells were incubated overnight at 37°C. The next day, a single colony was picked from the plate and inoculated into a test tube containing 4 ml of 50 μg / ml kanamycin sulfate (Sangon Biotech, A100408-0100) in LB liquid medium. The colony was incubated overnight at 37°C with shaking at 200 rpm. The following day, 4 ml of the bacterial culture was inoculated into a 2 L shake flask containing 400 ml of 50 μg / ml kanamycin sulfate in LB liquid medium and incubated at 37°C with shaking at 200 rpm for 2-3 hours.

[0489] When the OD600nm value of the bacterial suspension reaches 0.3-0.5, remove the shake flask and place it on ice for 10-15 minutes. Under aseptic conditions, pour the bacterial suspension into a pre-chilled 500ml centrifuge bottle, centrifuge at 4℃ and 3000rpm for 8 minutes, discard the supernatant, add approximately 200ml of pre-chilled CaCl2 solution, mix by pipetting to resuspend the bacterial cells, and incubate on ice for 30 minutes. Then, centrifuge the bacterial suspension at 4℃ and 3000rpm for 8 minutes, discard the supernatant, add approximately 8ml of pre-chilled CaCl2 solution to resuspend the bacterial cells, and aliquot the resuspended bacterial cells into 1.5ml EP tubes (110μl per tube) and store at -80℃ for later use.

[0490] 3. Determination of in vivo editing efficiency of Escherichia coli

[0491] The CasY7+sgRNA1 expression plasmid and the LbCpf1+sgRNA1 expression plasmid were respectively transformed into the competent cells prepared in step 2. The specific procedure is as follows:

[0492] (1) Remove competent cells from -80℃ and quickly place them on ice. After about 5 minutes, the bacterial block will thaw. Add the CasY7+sgRNA expression plasmid, then gently mix by tapping the bottom of the centrifuge tube. Let it stand on ice for 25 minutes. Heat shock in a 42℃ water bath for 45 seconds, then quickly return to ice and let it stand for 2 minutes. Add 900 μl of antibiotic-free sterile LB medium to the centrifuge tube, mix well, and incubate at 37℃ and 220 rpm for 60 minutes. Take 100 μl of bacterial suspension and spread it on LB agar plates containing 30 μg / ml carbenicillin resistance (Sangon Biotech, A100358-0001) (referred to as C-LB medium) and LB agar plates containing 30 μg / ml carbenicillin resistance and 10 mM L-arabinose (Sangon Biotech, A610071-0100) (referred to as CL-LB medium). Invert two LB agar plates and place them in an incubator overnight at 37°C.

[0493] (2) Detection of in vivo editing efficiency of Escherichia coli

[0494] like Figure 3 As shown, the Target plasmid contains a PBAD promoter that can be induced by L-arabinose and a CCDB gene that is regulated by the PBAD promoter. The CCDB gene can express CCDB toxic protein. As a DNA gyrase inhibitor, the CCDB toxic protein can lock the DNA gyrase and the broken double-stranded DNA complex, preventing the DNA gyrase from functioning and ultimately leading to cell death.

[0495] Based on this, the applicant designed a method for detecting the in vivo editing efficiency of E. coli:

[0496] In the presence of L-arabinose in the culture medium, if the CasY7 or LbCpf1 proteins, guided by sgRNA, can specifically target the target sequence (SEQ ID NO. 6) of the TTR gene on the Target plasmid and exert a cleavage effect, the PBAD promoter's regulatory pathway for CCDB toxic protein expression is interrupted, and the host cell survives because it does not produce ccdB toxic protein. Conversely, if the CasY7 or LbCpf1 proteins cannot specifically target the TTR target sequence on the Target plasmid, the host E. coli cell dies because the L-arabinose-induced PBAD promoter regulates the expression of CCDB toxic protein by the CCDB gene.

[0497] Therefore, the editing efficiency of CasY7 protein in targeting and cleaving TTR target genes in Escherichia coli can be calculated based on the ratio of the number of bacterial clones on CL-LB medium to the number of bacterial clones on C-LB medium in step (2).

[0498] The results are as follows Figure 4 As shown, after counting and calculating the ratio of E. coli clones, the editing efficiency of CasY7 protein was 45.5%, while that of LbCpf1 protein was 11.1%. The editing efficiency of CasY7 protein was significantly higher than that of LbCpf1 protein.

[0499] Example 3. Detection of editing efficiency in HEK293T cells

[0500] 1. Construction of TTR-sgRNA expression plasmid

[0501] (1) Design TTR-sgRNA sequences based on the target sequence of the TTR gene (Target-TTR-spacer2) tagaagggatatacaaagtg (SEQ ID NO.11), and synthesize oligonucleotides:

[0502] CasY7-TTR-sgRNA2 sequence:

[0503] CAAGTTGAATCCGTCTATAACTGACGG tagaagggatatacaaagtg (SEQ ID NO.12), the underlined part of the sequence is the DR sequence, and the rest of the sequence is the spacer sequence.

[0504] LbCpf1-TTR-sgRNA2 sequence:

[0505] TAATTTCTACTAAGTGTAGATtagaagggatatacaaagtg (SEQ ID NO.13), the underlined part of the sequence is the DR sequence, and the rest of the sequence is the spacer sequence.

[0506] (2) Add the CACC sequence to the 5' end of the upstream sequence of the TTR-sgRNA and the AAAA sequence to the 5' end of the downstream sequence, and synthesize oligos. The specific sequences are as follows:

[0507]

[0508]

[0509] After the synthesis of the aforementioned TTR-sgRNA upstream and downstream primers, annealing was performed using a pre-defined program (95℃, 5 min; 95℃-85℃ at -2℃ / s; 85℃-25℃ at -0.1℃ / s; held at 4℃). The annealed product was then ligated into the PHK09T vector linearized using BsmBI (NEB, #R0580L). The sequence of the PHK09T vector is shown in SEQ ID NO.3. The plasmid map can be found in [link to plasmid map]. Figure 5 .

[0510] The linearization of the PHK09T vector and its ligation with the TTR-sgRNA annealing product are as follows:

[0511] First, the PHK09T vector was linearized. The linearization system is as follows:

[0512] 3 μg PHK09T vector; 6 μL buffer (NEB: R0539L); 2 μL BsmBI; ddH2O to bring the total to 60 μL; digest overnight at 50°C.

[0513] TTR-gRNA annealing product and linearized vector ligation system:

[0514] 1 μL of T4 ligase buffer (NEB, #M0202L), 20 ng of linearized vector, 5 μL of annealed oligo fragment, 0.5 μL of T4 ligase (NEB, #M0202L), and ddH2O to a final volume of 10 μL were added. The mixture was ligated overnight at 16°C to obtain the CasY7-TTR-sgRNA2 expression plasmid and the LbCpf1-TTR-sgRNA2 expression plasmid.

[0515] (3) The CasY7-TTR-sgRNA2 expression plasmid and LbCpf1-TTR-sgRNA2 expression plasmid obtained in step (2) were transformed into Escherichia coli DH5α competent cells (Weidi Bio, DL1001), respectively. The specific steps are as follows:

[0516] After removing DH5α competent cells from the -80°C freezer, quickly place them on ice. After 5 minutes, once the bacterial block has thawed, add the ligation product and gently mix by tapping the bottom of the centrifuge tube. Incubate on ice for 25 minutes. Perform a heat shock at 42°C for 45 seconds, then quickly return to ice and incubate for 2 minutes. Add 700 μl of antibiotic-free sterile LB medium to the centrifuge tube, mix well, and incubate at 37°C, 200 rpm for 60 minutes. Centrifuge at 3000 rpm for one minute to collect the bacteria. Retain approximately 100 μl of the supernatant, gently resuspend the bacterial block by pipetting, and spread it onto LB medium containing Amp antibiotics. Incubate the plates inverted at 37°C overnight. Pick single colonies, confirm them by sequencing, and then shake positive clones to extract plasmids (using an endotoxin-free plasmid large-scale extraction kit, TIANGEN: DP120-01). Determine the plasmid concentration and store at -20°C for later use.

[0517] 2. Cell-level editing efficiency detection

[0518] (1) HEK293T cell culture

[0519] HEK293T cells (purchased from ATCC) were seeded in DMEM medium (Gibco, 11965092) supplemented with 10% FBS (v / v) containing 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) and cultured in a cell culture incubator at 37°C with 5% CO2. Cells for transfection were seeded in 24-well cell culture plates the day before transfection, and observed the cells the next day. Transfection was performed when the cell density reached approximately 80%.

[0520] (2) CasY7 recombinant expression plasmid (see Figure 1 for plasmid map) Figure 1A ), CasY7-TTR-sgRNA2 expression plasmid, LbCpf1 recombinant expression plasmid (see Figure 1 for plasmid map). Figure 1B The LbCpf1-TTR-sgRNA2 expression plasmid and the EGFP-C1 (Addgene, Plasmid, #54759) plasmid were transfected into HEK293T cells.

[0521] The amount of plasmids used for transfection per well in a 24-well plate was as follows: 0.3 μg of nuclease expression plasmid (CasY7 recombinant expression plasmid or LbCpf1 recombinant expression plasmid), 0.3 μg of sgRNA expression plasmid (CasY7-TTR-sgRNA2 expression plasmid or LbCpf1-TTR-sgRNA2 expression plasmid), and 0.3 μg of EGFP-C1 plasmid. The specific transfection procedure is as follows:

[0522] The CasY7 expression plasmid, CasY7-TTR-sgRNA2 expression plasmid, and EGFP-C1 plasmid were mixed separately and then used in 25 μl of water. Dilute serum-depleted transfection medium (Yuanpei Biotechnology, L530KJ) with 2 μl of Lipofectamine 3000 (Invitrogen, L3000015) reagent, mix well by pipetting, and let stand for 5 minutes. Simultaneously, dilute 2 μl of Lipofectamine 3000 transfection reagent (Invitrogen, L3000015) with 25 μl of... Dilute and mix the serum-depleted transfection medium (Yuanpei Biotechnology, L530KJ) as reagent B, and let stand for 5 minutes.

[0523] Mix reagent A and reagent B thoroughly and let stand for 20 minutes. After standing, add the mixed reagent dropwise to the cells in the 24-well plate to be transfected, and return to a 37°C, 5% CO2 incubator. Six hours after transfection, change the culture medium to DMEM medium containing 10% FBS.

[0524] Using the same method, the LbCpf1 recombinant expression plasmid, LbCpf1-TTR-sgRNA2 expression plasmid, and EGFP-C1 plasmid were transfected into HEK293T cells.

[0525] (3) Editing efficiency test

[0526] Forty-eight hours after transfection, the expression of EGFP fluorescent protein indicated successful cell transfection. Cells expressing EGFP were sorted for editing efficiency testing. Genomic DNA extraction was performed on the cells (using a TIANGEN DP304-03 genomic DNA extraction kit). Identification primers were designed according to experimental requirements, and the sequences of the primers used are shown in the table below.

[0527]

[0528]

[0529] Using the genome as a template, PCR amplification was performed on sequences near the target site using the primers listed in the table above. The PCR amplification system is as follows:

[0530] 25 μL of 2×Taq Master Mix (Vazyme, P112-03); 1 μL of Primer-F (TTR-F) (10 pmol / μL); 1 μL of Primer-R (TTR-R) (10 pmol / μL); 1 μL of template; ddH2O to bring the total to 50 μL.

[0531] The amplified PCR products were used for editing efficiency identification using high-throughput deep sequencing (Qingke Biotechnology Co., Ltd.) or Sanger sequencing (Platinum Biotechnology (Shanghai) Co., Ltd.).

[0532] The editing efficiency of CasY7 and LbCpf1 was tested and evaluated, and the results are as follows: Figure 6 As shown, in 293T cells, the editing efficiency of CasY7 at the TTR gene was 33%, while that of LbCpf1 at the TTR gene was only 22%, indicating that the editing efficiency of CasY7 was much higher than that of LbCpf1.

[0533] Example 4. Application of CasY7 in base editing

[0534] 1. Obtaining catalytically inactive CaSY7

[0535] To obtain dCasY7 with no catalytic activity (i.e., loss of cleavage activity), the inventors constructed CasY7 mutants with single-point mutations of D592A, D643A, E820A, and D992A, respectively: D592A-dCasY7, D643-dCasY7, E820A-dCasY7, and D992A-dCasY7. The amino acid sequences of the above mutants are shown in SEQ ID NO.47-50, respectively. The specific construction method is as follows:

[0536] The CasY7+sgRNA1 expression plasmid obtained in step 1(3) of Example 2 was subjected to point mutation, and the amino acid of CasY7 was modified at four sites, namely aspartic acid at position 592 (Asp, D), aspartic acid at position 643 (Asp, D), glutamic acid at position 820 (Glu, E), and aspartic acid at position 992 (Asp, D) in SEQ ID No. 1. The amino acids at the above sites were mutated to alanine (Ala, A). The codons before and after the mutation of each amino acid are shown in the table below:

[0537] Pre-mutation amino acid Codon Post-mutation amino acid Codon Aspartic acid (Asp, D) GAC Alanine (Ala, A) GCA Aspartic acid (Asp, D) GAC Alanine (Ala, A) GCA Glutamic acid (Glu, E) GAG Alanine (Ala, A) GCA Aspartic acid (Asp, D) GAC Alanine (Ala, A) GCA

[0538] Forward and reverse primers were designed and synthesized for the amino acids and their codons listed in the table above. PCR amplification was then performed using the CasY7+sgRNA1 expression plasmid as a template. After amplification, the amplified products were purified using a universal DNA purification kit (TIANGEN Biotechnology (Beijing) Co., Ltd., DP214). The purified products were transformed into *E. coli* Dh5a competent cells (Weidi Biotechnology, DL1001) and cultured overnight at 37°C. The next day, single clones were picked and sequenced. After sequencing confirmation, positive clones were cultured and plasmids were extracted (TIANGEN, DP120-01), and their concentrations were determined. The plasmids were then stored at -20°C for later use.

[0539] The recombinant plasmids with point mutations were named as follows: D592A-dCasY7+sgRNA expression plasmid, D643A-dCasY7+sgRNA expression plasmid, E820A-dCasY7+sgRNA expression plasmid, and D992A-dCasY7+sgRNA expression plasmid.

[0540] Subsequently, the constructed D592A-dCasY7+sgRNA expression plasmids, D643A-dCasY7+sgRNA expression plasmids, E820A-dCasY7+sgRNA expression plasmids, and D992A-dCasY7+sgRNA expression plasmids were subjected to in vivo editing efficiency assays in *E. coli*. The assay methods and calculations were the same as in steps 2 and 3 of Example 2. The experimental results are as follows: Figure 7 As shown, by counting and calculating the ratio of E. coli clones, it is believed that D592A-dCasY7, D643A-dCasY7, E820A-dCasY7, and D992A-dCasY7 have lost their catalytic activity (cleavage activity). In other words, point mutations in D592A, D643A, E820A, and D992A have caused the CasY7 protein to lose its cleavage activity.

[0541] 2. Cell-level base editing efficiency detection

[0542] (1) Construction of plasmids CasY7-TTR-sgRNA3 and CasY7-TTR-sgRNA4

[0543] Based on the target sequence of the TTR gene, sgRNAs were designed: CasY7-TTR-sgRNA3 and CasY7-TTR-sgRNA4 sequences, and oligonucleotides were synthesized.

[0544] CasY7-TTR-sgRNA3: CAAGTTGAATCCGTCTATAACTGACGG tatatcccttctacaaattc(SEQID NO.23);

[0545] CasY7-TTR-sgRNA4: CAAGTTGAATCCGTCTATAACTGACGG gtgtctatttccactttgta(SEQID NO.24), where the underlined part of the sequence is the DR sequence and the rest of the sequence is the spacer sequence.

[0546] (2) Add a CACC sequence to the 5' end of the upstream sequence of each sgRNA and an AAAA sequence to the 5' end of the downstream sequence, as follows:

[0547]

[0548]

[0549] Following the method in step 1 of Example 3, the upstream and downstream sequences of CasY7-TTR-sgRNA3 and CasY7-TTR-sgRNA4 were annealed and ligated into the PHK09T vector to obtain the CasY7-TTR-sgRNA3 expression plasmid (named CasY7-TTR-sgRNA' expression plasmid) and the CasY7-TTR-sgRNA4 expression plasmid (named CasY7-TTR-sgRNA" expression plasmid). These plasmids were then transformed into E. coli DH5α competent cells for plasmid amplification culture. After confirming the sequencing results and determining the concentration, the plasmids were stored for later use.

[0550] (3) Construction of plasmids for base editor (using deaminase 005V1-10-3 as an example)

[0551] The selected adenosine deaminase catalytic domain was chosen from a mutant of the amino acid sequence shown in SEQ ID NO.31 (named 005V1-10-3): Q148G+Q149M+P150R. The amino acid sequence of this mutant is shown in SEQ ID NO.32, and the nucleotide sequence encoding deaminase 005V1-10-3 is shown in SEQ ID NO.33. A base editor fusion protein composed of deaminase 005V10-3 and CasY7 protein was constructed by homologous recombination. The specific operation is as follows:

[0552] First, Suzhou Hongxun Biotechnology Co., Ltd. synthesized the 005V1-10-3 nucleotide fragment containing a homologous arm sequence and a linker:

[0553] The bold part is the 005V1-10-3 nucleotide sequence, the italic part is the left and right homologous arm regions, and the wavy part is the linker sequence.

[0554] The D992A-dCasY7+sgRNA expression plasmid obtained in step 1 was linearized by PCR amplification to obtain a linearized expression vector. The primers used are shown in the table below:

[0555]

[0556] Homologous recombination was performed using the 005V1-10-3 nucleotide fragment (SEQ ID NO.34) containing the homologous arm sequence and linker, and the linearized expression vector D992A-dCasY7+sgRNA. The reaction was carried out using Gibson Assembly Master Mix (NEB, E2611S). After the reaction, the ligation product was transformed into *E. coli* DH5α competent cells (Weidi Bio, DL1001). The specific procedure is as follows:

[0557] After removing DH5α competent cells from the -80℃ freezer, immediately place them on ice. After 5 minutes, once the bacterial block has thawed, add the ligation product and gently mix by tapping the bottom of the centrifuge tube. Incubate on ice for 25 minutes. Perform a heat shock at 42℃ for 45 seconds, then immediately return to ice and incubate for 2 minutes. Add 700 μl of sterile LB medium to the centrifuge tube, mix well, and incubate at 37℃, 200 rpm for 60 minutes. Centrifuge at 5000 rpm for one minute to collect the bacteria. Retain approximately 100 μl of supernatant, gently resuspend the bacterial block by pipetting, and spread it onto LB medium containing Amp antibiotic. Invert the plate and incubate overnight at 37℃. Single colonies were picked, and after sequencing confirmation, positive clones were cultured and the base editor plasmid was extracted using an endotoxin-free plasmid large-scale extraction kit (TIANGEN: DP120-01). The concentration was then determined, and the plasmid was stored at -20℃ for later use. The nucleotide sequence encoding the base editor fusion protein 005V1-10-3-D992A-dCasY7 is shown in SEQ ID NO.37, and the amino acid sequence is shown in SEQ ID NO.46.

[0558] (4) Following the method in step 2 of Example 3, the base editor plasmid and EGFP-C1 (Addgene, Plasmid#54759) plasmid were co-transfected with CasY7-TTR-sgRNA' expression plasmid and CasY7-TTR-sgRNA" expression plasmid into 293T cells.

[0559] Forty-eight hours after transfection, the expression of EGFP fluorescent protein indicated successful cell transfection. Cells expressing EGFP were sorted for editing efficiency testing. The genome of the 293T cells was extracted using a kit (TIANGEN, DP304-03).

[0560] (5) Base editing efficiency was tested according to the method in step (3) of Example 3.

[0561] Primers were designed according to experimental requirements, and the sequences of the identification primers used are shown in the table below:

[0562] Primer name Specific sequence TTR-F’ gggtgtattactttgccatg (SEQ ID NO. 29) TTR-R’ aacctttggtcattcatcaccttc (SEQ ID NO. 30)

[0563] The result is shown in Figure 8. Figure 8A ,8B As shown, the base editor composed of D992A-dCasY7 can achieve effective editing at multiple sites, from Figure 8A It can be seen that under the guidance of CasY7-TTR-sgRNA', effective editing exists at +4, +13, +15, +16, and +17 sites, with editing efficiency at +13, +15, +16, and +17 sites reaching close to 5% to close to 30%.

[0564] from Figure 8B It can be seen that under the guidance of "CasY7-TTR-sgRNA", effective editing exists at +7, +13, and +20 sites. The editing efficiency at +7 site is close to 10%, the base editing efficiency at +13 site is close to 20%, and the editing efficiency at +20 site is over 20%.

[0565] Example 5. Screening of CasY7 mutants

[0566] To screen for CasY7 mutants with high cleavage activity, the 7-684 region of the CasY7 protein (SEQ ID NO.1) was engineered to generate 31 mutants.

[0567] The mutant was obtained by recombining the CasY7 expression plasmid from Example 2 (see Figure 1 for the plasmid map). Figure 1A Using a template, PCR primers are designed with the mutation site as the center. The mutated nucleotide sequence is introduced onto the PCR primers, and then each CasY7 mutant plasmid is obtained by site-directed PCR mutagenesis.

[0568] The mutation modes of each mutant are shown in the table below:

[0569]

[0570]

[0571] Based on the target sequences of the TTR, PD-1, and Trac genes, gRNA targeting sequences were designed. Expression plasmids for each targeting sgRNA were constructed using the same method as in Example 3. The CasY7 recombinant expression plasmid or its mutant recombinant expression plasmid, along with the TTR-sgRNA2 expression plasmid, were transfected into HEK293 cells using PEI transfection. A negative control group (using other targeting sgRNAs, the NT-sgRNA targeting sequence is shown in SEQ ID NO.52 and marked as "NT") and a blank control group were also double-transfected. Primers were designed and identified based on the targeting of each gRNA.

[0572]

[0573] Genomic DNA was extracted from cells after 48 hours of cell culture and amplified by PCR using identification primers. The PCR products were then used for high-throughput deep sequencing (Qingke Biotechnology Co., Ltd.) to identify the editing efficiency.

[0574] The cleavage efficiencies of CasY7 wild-type (WT) and its various mutants against TTR-1, PD-1, Trac-1, Trac-2, Trac-3, and Trac-4 targets are as follows: Figures 9-14 As shown in the figure. Analysis revealed that, compared to wild-type CasY7, the editing efficiency of each mutant of CasY7 at each target gene was close to or significantly higher than that of wild-type CasY7. For example, at the TTR gene site, except for the C06354 variant, other variants showed higher cleavage activity compared to wild-type (WT), especially C04346, whose cleavage activity was more than three times that of the parent protein. Similarly, at the PD-1 and Trac gene sites, several variants showed higher cleavage activity compared to wild-type (WT). The negative control (NT) group and blank control group set up for each target showed no or only background-level cleavage activity.

[0575] Example 6: Effect of Direct Repeat (DR) Sequences on CSY7 Cleavage Activity

[0576] To test the effects of different DR sequences on CasY7 activity, deletions and mismatches were designed at different positions in the DR sequences, resulting in a new DR variant: DR-1 (SEQ ID NO.71). The secondary structure of DR-1 (…) Figure 15 (Any "T" in the sequence in the figure represents "U" when referring to the DR sequence) and the secondary structure of the wild-type DR sequence (SEQ ID NO.5) Figure 2 The basic sequence is the same for both DR-1 and wild-type DR. The stem sequence is the same, consisting of 5'-CCGTC-3' and its complementary strand. However, there are some differences in the sequence in other regions.

[0577] Using the hTTR gene as a target, a spacer sequence (Target-TTR-spacer2, SEQ ID NO.11) targeting the TTR gene sequence was designed, along with a DR-1-TTR-sgRNA (SEQ ID NO.72) containing the DR-1 sequence. Following the method described in Example 3, the synthesized DR-1-TTR-sgRNA (SEQ ID NO.72) sequence fragment was constructed into the PHK09T vector to obtain the DR-1-CasY7-TTR-sgRNA expression plasmid.

[0578] Following steps 2(2) and (3) of Example 3, the DR-1-CasY7-TTR-sgRNA expression plasmid and the CasY7 expression plasmid were transfected into cells. Editing efficiency was assessed using upstream and downstream primers TTR-F (SEQ ID NO. 21) and TTR-R (SEQ ID NO. 22). Analysis revealed that the DR-1 sequence can mediate higher editing activity of CasY7 (…). Figure 16 This also demonstrates that CasY7 is adaptable to different DR sequences.

[0579] Example 7: Screening of highly active CaSY7 mutants

[0580] 1. To screen for CasY7 mutants with high cleavage activity, 17 mutants of CasY7 protein (SEQ ID NO.1) were constructed through directed mutagenesis and screening.

[0581] The mutant was obtained by recombining the CasY7 expression plasmid from Example 2 (see Figure 1 for the plasmid map). Figure 1A Using a template, PCR primers are designed with the mutation site as the center. The mutated nucleotide sequence is introduced onto the PCR primers, and then each CasY7 mutant plasmid is obtained by site-directed PCR mutagenesis.

[0582]

[0583]

[0584] To efficiently and sensitively detect the cleavage activity of various mutants, a dual-fluorescence reporter system vector was constructed. Figure 17B It contains the coding sequence for Fluc luciferase and the coding sequence for LUxUC. Figure 17A In LUxUC, the LU and UC sequences encode the 119 bp 5' end and 469 bp 3' end of NanoLuc luciferase, respectively, with a total overlap of 60 bp. An insert sequence containing a target sequence for cleavage activity detection is designed within the LUxUC sequence. Target sequences are designed based on TTR and PD-1' targets, and insert sequences containing these target sequences are designed accordingly. In this embodiment, the insert sequences, target sequences, and corresponding LUxUC sequences used are shown in the table below:

[0585] Target Insert sequence Target sequence Corresponding LUxUC sequence TTR SEQ ID NO. 76 SEQ ID NO. 11 SEQ ID NO. 73 PD-1’ SEQ ID NO. 77 SEQ ID NO. 75 SEQ ID NO. 74

[0586] In this embodiment, a dual-fluorescent reporter system vector for detecting TTR target cleavage by various variants was constructed, the nucleotide sequence of which is shown in SEQ ID NO.97, wherein positions 6759-7380 are TTR-LUxUC sequences, the TTR-LUxUC sequence is shown in SEQ ID NO.73, the TTR-LUxUC sequence contains a TTR-insertion sequence (SEQ ID NO.76), and the target sequence (SEQ ID NO.11) is designed in the middle of the TTR-insertion sequence; similarly, a dual-fluorescent reporter system vector for detecting PD-1' target cleavage by various variants was also constructed (replacing positions 6759-7380 of the reporter system plasmid shown in SEQ ID NO.97 with the PD-1'-LUxUC sequence), the PD-1'-LUxUC sequence, the PD-1' target sequence, and the PD-1' insert sequence are shown in the table above.

[0587] The target sequence contains a TAG early stop codon. When this codon is cleaved, LUxUC utilizes a recombination mechanism to generate the correct NanoLuc coding frame, expressing NanoLuc luciferase, which then catalyzes the substrate to emit fluorescence. The intensity of this fluorescence indicates cleavage activity. Fluoroluciferase serves as an internal control, and its fluorescence indicates whether the reporter vector has been successfully transfected into the host cell.

[0588] Using the same method as in Example 1, recombinant expression plasmids for each mutant were constructed. sgRNA expression plasmids targeting TTR and PD-1' were constructed using the method described in Example 3. Following the transfection method in Example 3, the recombinant expression plasmids of CasY7 and each mutant, along with their corresponding dual-luciferase reporter system vectors and sgRNA expression plasmids, were co-transfected into the HEK293 cell line. After 24 hours of culture, the expression levels of Fluc and NanoLuc were detected using a dual-luciferase reporter gene assay kit (Beyotime, RG028) on a microplate reader.

[0589] In addition, a negative control group was set up for CasY7 and its mutants. The negative control group used an sgRNA expression plasmid with a non-target sequence (SEQ ID NO.78). The analysis showed that the cleavage activity of CasY7 and its mutants is shown in the table below, while no cleavage activity was detected in the negative control group. The cleavage activity of each mutant was much higher than that of CasY7.

[0590] Mutant Targeted TTR cleavage activity (%) Targeted PD-1’ cleavage activity (%) C05440 4.05 8.39 C03897 33.92 31.64 C06865 32.49 32.51 C04558 30.10 27.15 C08675 23.80 35.92 C01144 29.49 31.38 C06623 29.33 34.29 C03023 29.75 39.67 C07879 30.38 35.04 C03803 25.31 38.54 C06467 27.84 18.66 C08724 26.15 36.45 C03563 26.11 35.15 C05664 28.69 40.12 C02742 21.18 35.02 C02009 27.09 32.39 C09752 20.31 35.04 CasY7 2.05 4.45

[0591] 2. Using the same method as step 1 of this embodiment, expression plasmids expressing 39 mutants (as shown in the table below) were constructed.

[0592]

[0593]

[0594]

[0595]

[0596]

[0597] Following the method described in Example 3, expression plasmids for sgRNAs targeting the TTR gene and the Hao1 gene were synthesized, respectively. The sgRNA sequences and target sequences are shown in the table below:

[0598]

[0599] The cleavage activity of the above variants was detected by selecting the target TTR and hHao1. A dual-fluorescence reporter system vector (as shown in SEQ ID NO. 97, the reporter system vector for detecting the cleavage activity of the target TTR is obtained by replacing positions 6759-7380 with the hHao1-LUxUC sequence) was used to detect its cleavage activity. The target sequence, insert sequence, and LUxUC sequence are shown in the table below:

[0600] Target Insert sequence Target sequence Corresponding LUxUC sequence TTR SEQ ID NO. 76 SEQ ID NO. 11 SEQ ID NO. 73 hHao1 SEQ ID NO. 98 SEQ ID NO. 101 SEQ ID NO. 99

[0601] Analysis revealed the cleavage activities of CasY7 and its various mutants as shown in the table below, while no cleavage activity was detected in the negative control (NT) group.

[0602]

[0603]

[0604] Analysis revealed that, compared to the wild-type CasY7 protein, the variants exhibited higher cleavage activity at the TTR and hHao1 gene targets. In the TTR gene, the variants showed at least 16-fold cleavage activity compared to the wild-type CasY7 protein, with a maximum of approximately 23-fold. In the hHao1 gene, the variants showed at least 8.5-fold cleavage activity compared to the wild-type CasY7 protein, with a maximum of approximately 22-fold.

[0605] Example 8. Fusion with T5 exonuclease to enhance cleavage activity

[0606] To improve the cleavage activity of the CRISPR-Cas system, the CasY7 variant C05440 (mutation method as shown in Example 6, its amino acid sequence as shown in SEQ ID NO. 82, and its nucleotide coding sequence as shown in SEQ ID NO. 83) is fused with T5 exonuclease (its amino acid sequence as shown in SEQ ID NO. 81, and its nucleotide coding sequence as shown in SEQ ID NO. 100). This allows the T5 exonuclease to be linked to the C-terminus or N-terminus of C05440.

[0607] In this embodiment, the amino acid sequences and nucleotide coding sequences of the fusion proteins C05440-T5 and T5-C05440 are shown below:

[0608] Fusion protein Amino acid sequence Nucleotide coding sequence C05440-T5 SEQ ID NO. 84 SEQ ID NO. 85 T5-C05440 SEQ ID NO. 86 SEQ ID NO. 87

[0609] C05440-T5 mRNA transcription template (SEQ ID NO.88) and T5-C05440 mRNA transcription template (SEQ ID NO.89) were synthesized by Nanjing Genscript Biotech. The T7 High Yield RNA Synthesis Kit (NEB, E2040S) was used to perform in vitro transcription to obtain C05440-T5 mRNA and T5-C05440 mRNA.

[0610] The hHAO1 gene was selected as the target gene for editing. A spacer sequence was designed based on the target sequence, and hHAO1-crRNA was synthesized by Nanjing Genscript Biotech. The crRNA sequence is as follows:

[0611] AGAAATCCGTCCAAAGCTGACGG ggacagagggtcagcatgcc (SEQ ID NO.90), where the underlined sequence is the DR-1 sequence and the other sequences are spacer sequences.

[0612] A four-component LNP lipid delivery system (purchased from Avitol (Shanghai) Pharmaceutical Technology Co., Ltd.) was used. Specifically, Yoltech Lipid1 (compound 10), DSPC, cholesterol, and PEG-DMG were dissolved in anhydrous ethanol at a molar ratio of 50:10:38.5:1.5. C05440-T5 mRNA, T5-C05440 mRNA, and hHAO-crRNA targeting the hHAO1 gene (mass ratio 1:1) were dissolved in 100 mM enzyme-free citrate buffer at pH 4 (RNA concentration 0.2 mg / mL). The ethanol solution of the lipid carrier and the buffer of mRNA were mixed at a 1:3 (volume / volume) ratio (where the mass ratio of total lipid to mRNA was 40:1), and nucleic acid lipid nanoparticles were obtained by passing the mixture through a microfluidic nanomedicine manufacturing system (NanoAssemblr Ignite, Canada) at a flow rate of 12 ml / min. The obtained nucleic acid lipid nanoparticles were immediately diluted 40-fold in 1×DPBS buffer.

[0613] HepG2 cells (purchased from ATCC) were seeded in DMEM medium (Gibco, 11965092) supplemented with 10% FBS (v / v) containing 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) and cultured in a 37°C cell culture incubator containing 5% CO2. Cells for transfection were seeded in 96-well cell culture plates the day before transfection and observed the following day. LNP transfection was performed when the cell density reached approximately 80%. LNP@mRNA (transfection doses of 5 ng / well, 10 ng / well, 20 ng / well, and 40 ng / well) was added to HepG2 cells. Cells were collected 48 hours after transfection, and genomic DNA was extracted from the collected cells (TIANGEN, DP304-03) for cleavage activity testing. Figure 18 The identification primer sequences used were hHAO1-F (SEQ ID NO. 91) and hHAO1-R (SEQ ID NO. 92).

[0614] Analysis shows that C05440-T5 has significantly higher cleavage activity than T5-C0544 and C05440 at all dosages, while T5-C0544 has the lowest cleavage activity at all dosages. This indicates that the ligation site of the T5 exonuclease is related to its cleavage activity.

[0615] The synthesis method of Yoltech Lipid1 (compound 10) is as follows:

[0616] 5-[(2-Butyl-1-oxylidene octyl)oxy]valerate-7-butyl-21-(10-butyl-3,9-dioxylidene-2,8-dioxahexadecane-1-yl)-19-[3-(diethylamino)propyl]-8-oxylidene-19-aza-9-oxadodecane-22-yl ester

[0617]

[0618] Step 1: Synthesis of Compounds 1-2

[0619] In a 500 mL round-bottom flask, cyclohexyl ester (25.00 g, 249.70 mmol, 1.0 eq), distilled water (20 mL), ethanol (200 mL), and sodium hydroxide (10.99 g, 274.67 mmol, 1.1 eq) were added. After reacting at 70 °C for 3 hours, the solvent was removed by concentration under reduced pressure. Then, 200 mL of acetone, tetrabutylammonium iodide (4.61 g, 12.48 mmol, 0.05 eq), and benzyl bromide (51.25 g, 299.64 mmol, 1.2 eq) were slowly added to the flask, and the reaction was continued overnight at 70 °C. The reaction was quenched with 500 mL of water, and the mixture was extracted twice with 500 mL of ethyl acetate. The organic phases were combined, washed with saturated brine, dried over anhydrous sodium sulfate, concentrated under reduced pressure, and purified by column chromatography to give benzyl 5-hydroxyvalerate (37.00 g, 71.2% yield).

[0620] Step 2: Synthesis of compounds 1-4

[0621] In a 500 mL round-bottom flask, benzyl 5-hydroxypentanoate (37.00 g, 177.67 mmol, 1.0 eq), 2-butyloctanoic acid (35.59 g, 177.67 mmol, 1.0 eq), 250 mL of dichloromethane, and 4-dimethylaminopyridine (21.70 g, 177.67 mmol, 1.0 eq) were added, followed by 1-(3-dimethylaminopropyl)-3-ethylcarbodiimide hydrochloride (51.09 g, 266.50 mmol, 1.5 eq). The reaction was carried out at room temperature for 4 hours. The mixture was diluted with 500 mL of water and extracted twice with 500 mL of dichloromethane. The organic phases were combined, washed with saturated brine, dried over anhydrous sodium sulfate, concentrated under reduced pressure, and purified by column chromatography to give 2-butyloctanoate-5-(benzyloxy)-5-oxylidenepentyl ester (64.00 g, 92.2% yield).

[0622] Step 3: Synthesis of compounds 1-5

[0623] In a 250 mL round-bottom flask, 2-butyloctanoic acid-5-(benzyloxy)-5-oxylidene pentyl ester (64.00 g, 163.87 mmol, 1.0 eq), methanol (75 mL), tetrahydrofuran (75 mL), and finally Pd / C (3.49 g, 32.78 mmol, 0.2 eq, 10% purity) were added. The reaction was carried out at room temperature under a hydrogen atmosphere at one atmosphere for 16 hours. The mixture was then filtered and concentrated to give compound 5-[(2-butyl-1-oxylidene octyl)oxy]pentanoic acid (45.00 g, yield 91.4%).

[0624] Step 4: Synthesis of compounds 1-7

[0625] At room temperature, 5-[(2-butyl-1-oxomylidene octyl)oxy]valerate (10.00 g, 33.29 mmol, 1.0 eq), 2-hydroxymethylpropane-1,3-diol (3.53 g, 33.29 mmol, 1.0 eq), 4-dimethylaminopyridine (0.81 g, 6.66 mmol, 0.2 eq), N-(3-dimethylaminopropyl)-N'-ethylcarbodiimide hydrochloride (9.57 g, 49.94 mmol, 1.5 eq) and N,N-diisopropylethylamine (8.60 g, 66.58 mmol, 2.0 eq) were added to a round-bottom flask containing 100 mL of dichloromethane, and the mixture was stirred at room temperature for 4 hours. The reaction solution was quenched with 200 mL of water, extracted twice with 200 mL of dichloromethane, the organic phases were combined, washed with brine, dried over anhydrous sodium sulfate, filtered, concentrated, and purified by column chromatography to obtain 2-butyloctanoic acid-18-butyl-8-(hydroxymethyl)-5,11,17-trioxane-6,10,16-trioxane-1-yl ester (7.80 g, yield 69.9%).

[0626] Step 5: Synthesis of compounds 1-8

[0627] At room temperature, 3.90 g (5.81 mmol, 1.0 eq) of compound 2-butyloctanoic acid-18-butyl-8-(hydroxymethyl)-5,11,17-trioxylidene-6,10,16-trioxacotetraco-1-yl ester and triethylamine (1.76 g, 17.43 mmol, 3.0 eq) were added to 30 mL of dichloromethane. Methanesulfonic anhydride (2.02 g, 11.62 mmol, 2.0 eq) was slowly added at 0°C, and the mixture was slowly brought back to room temperature for 4 hours. The reaction mixture was quenched with 30 mL of water and extracted twice with 50 mL of dichloromethane. The organic phases were combined, washed with brine, dried over anhydrous sodium sulfate, filtered, concentrated, and purified by column chromatography to give methanesulfonic acid-12-butyl-2-(10-butyl-3,9-dioxylidene-2,8-dioxahexadecane-1-yl ester.

[0628] 5,11-dioxane-4,10-dioxaoctadecane-1-yl ester (3.85 g, yield 88.4%).

[0629] Step 6: Synthesis of compounds 1-10

[0630] Compounds 1-8 (600.0 mg, 0.80 mmol, 1.0 eq), 3-amino-1-propanol (300.0 mg, 3.99 mmol, 5.0 eq), potassium carbonate (280.0 mg, 2.00 mmol, 2.5 eq), and potassium iodide (130.0 mg, 0.80 mmol, 1.0 eq) were added to 10 mL of acetonitrile under nitrogen protection and heated to 90 °C for 16 hours. The reaction solution was concentrated, diluted with water, and extracted three times with ethyl acetate. The organic phases were combined, washed with saturated brine, dried over anhydrous sodium sulfate, concentrated, and purified by column chromatography to give compound 5-[(2-butyl-1-oxoylideneoctyl)oxy]valerate-12-butyl-2-{[(3-hydroxypropyl)amino]methyl}-5,11-dioxoylidene-4,10-dioxaoctadecane-1-yl ester (210.0 mg, 36.11%). MS: m / z [M+H] + =728.6.

[0631] Step 7: Synthesis of Compound 10

[0632] At room temperature, 2-butyloctanoic acid-8-(10-butyl-3,9-dioxayne-2,8-dioxahexadecane-1-yl)-14-ethyl-5-oxayne-10,14-diaza-6-oxahexadecane-1-yl ester (500.0 mg, 0.64 mmol, 1.0 eq), 2-butyloctanoic acid-9-bromononyl ester (390.0 mg, 0.96 mmol, 1.5 eq), potassium carbonate (270.0 mg, 1.92 mmol, 3.0 eq), and potassium iodide (110.0 mg, 0.64 mmol, 1.0 eq) were added to 20 mL of acetonitrile, heated to 90 °C under nitrogen protection, and reacted overnight. The reaction solution was concentrated, diluted with water, and extracted three times with dichloromethane. The organic phases were combined, washed with saturated brine, dried over anhydrous sodium sulfate, concentrated, and purified by column chromatography to give 5-[(2-butyl-1-oxoylidene octyl)oxy]valerate-7-butyl-21-(10-butyl-3,9-dioxoylidene-2,8-dioxahexadecane-1-yl)-19-[3-(diethylamino)propyl]-8-oxoylidene-19-aza-9-oxadodecane-22-yl ester (132.8 mg, yield 18.8%). MS: m / z [M+H] + =1107.9. 1H NMR(300MHz, CDCl3)δ4.15-4.01(m,10H),3.44-3.20(m,4H),2.71-2.50(m,6H), 2.39-2.23(m,12H),2.02-1.40(m,28H),1.38-1.22(m,48H),0.92-0.75(m,18H).

[0633] Example 9. Verification of CasY7 off-target activity at the Trac gene

[0634] To evaluate the application of CasY7 in human cells, a spacer sequence Target-Trac-spacer3 (SEQ ID NO. 59) targeting the Trac gene was designed based on the PAM sequence of CasY7. Twenty-nine gene sites that may experience off-target cleavage, along with their off-target prototype spacers, were predicted, as shown in the table below.

[0635]

[0636]

[0637] In this embodiment, the DR-1 sequence is used as a direct repeat sequence. The DR-1-Trac-crRNA sequence is shown below:

[0638] AGAAATCCGTCCAAAGCTGACGGttgctccaggccacagcact (SEQ ID NO. 93). The gRNA expression plasmid was constructed using the method described in Example 3. Both the CasY7 recombinant expression plasmid and the aforementioned gRNA expression plasmid were co-transfected into HEK293 cells using PEI transfection. Deep sequencing was used to detect on-target and off-target activities. Analysis showed that CasY7 exhibited virtually no off-target cleavage activity at the predicted sites. Figure 19 On Target represents on-target cleavage activity, and OT1-29 represents predicted off-target gene sites. The CRISPR-Cas system disclosed herein has a wide range of applications, exhibiting high cleavage activity and low collateral activity, and can be used for drug screening, disease diagnosis and prognosis, and treatment of various genetic diseases.

[0639] Example 10. Cleavage activity of CasY7 in plant cells

[0640] To evaluate the targeted cleavage of CasY7 in plant cells, the applicant selected the rice endogenous gene OsBEL as the target gene. Based on the PAM (5'-TTN-3') of CasY7, an OsBEL-crRNA was designed, using the DR-1 sequence. The OsBEL-crRNA sequence is shown below:

[0641] AGAAATCCGTCCAAAGCTGACGGATCTCCTTCTAGAAGCACAA(SEQ ID NO.94)

[0642] The purified CasY7 protein (60 μg) was combined with OsBEL-crRNA (120 μg) at 37°C to form an RNP. The RNP was then transferred into rice protoplasts using PEG4000. After 24 hours of culture, editing activity was assessed, and analysis revealed that CasY7 cleaved approximately 75% of the target site. This indicates that the CRIPSR-CasY7 system has potential applications in the plant field.

[0643] Sequence information:

[0644]

[0645]

[0646]

[0647]

[0648]

[0649]

[0650]

[0651]

[0652]

[0653]

[0654]

[0655]

[0656]

[0657]

[0658]

[0659]

[0660]

[0661]

[0662]

[0663]

[0664]

[0665]

[0666]

[0667]

[0668]

[0669]

[0670]

[0671]

[0672]

[0673]

[0674]

[0675]

[0676]

[0677]

[0678]

[0679]

[0680]

[0681] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A Cas protein, characterized in that, The protein is selected from the following group: (1) A polypeptide with the amino acid sequence shown in SEQ ID NO.1; (2) The polypeptide contains mutations at one or more sites relative to the amino acid sequence shown in SEQ ID NO.

1. The mutation mode of the Cas protein relative to SEQ ID NO.1 is as follows: (a) M231K+D234V; (b) S240A+S241R+Q242P+E243L; (c) L227I+L229R, S225P+L227A+L229R; (d) S240L+S241A+Q242S+E243H+I244L+L307R+Y308F+S309A, D283S+A285P+L307R+Y30 8F+S309A, L307R+Y308F+S309A+E648R+L650I+A651P+Y652F, D283S+A285P+L307R+ Y308F+S309A+I373V+E748G+S751G+K752R+S787F, S303T+I304M+L307R+Y308A+S30 9I, S196N+K197F+S198N+A199T+I304A+L307R+Y308A+S309A, L307R+Y308F+S309A; (e) S250W+F251A+E252A+K253R+V254R; (f) K260R+T261P+E263H; (g) R315L+E316R+T317R+I318V+I319A, Y282F+D283Q+A285T+R315A+E316Q+T317R+I31 8L+I319Q, Y282F+D283Q+A285T+R315S+E316R+T317K+I318K+I319W+E648S+A651L; (h) L354F+S355G+N356D+L357I+K358R; (i) E648S+A651L, E648S+A651T+Y652W, Y282F+D283Q+A285T+E648T+A651S+ Y652F, Y282F+D283Q+A285T+E648L+A651G, N276S+D283N+E648S+A651L; (j) T681V+N682G+E683R+S684G、T681P+N682T+E683G+S684R、T681S+N682P+E683R+S684G、T681A+N682H+E683Y+S684R、Y282F+D283Q+A285T+T681P+N682P+E683Q+S684C、Y282F+D283Q+A285T+G416S+I417H+E418Q+F419T+D420V+L501P+T681N+N682P+E683G+S684A、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+T681D+N682A+E683R+S684D、Y165W+S166R+S198N+G416R+I417V+E418Q+F419V+D420T+T681S+N682R+E683R、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509N+P510G+V511G+L512S+T681S+N682R+E683R、G416R+I417V+E418Q+F419V+D420T+T681S+N682R+E683R+G727V+K728C+N729G+E730G、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509V+P510R+V511L+L512A+T681S+N682R+E683R、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509Q+P510A+V511A+L512R+T681S+N682R+E683R+A661V、S110G+Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509A+P510A+V511R+L512Q+T681S+N682R+E683R、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509L+P510V+V511S+L512F+T681S+N682R+E683R、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509G+P510K+L512C+T681S+N682R+E683R、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+T634V+T635Y+K636S+N637R+T681S+N682R+E683R;、 (k) Y165W+S166R、Y165W+S166V+S167I+F168I、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T、Y165W+S166R+G416R+I417E+E418Q+F419V+D420A、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+N682G+E683A+D896N、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+G513L+N514A+R515V+V516M+N682G+E683A+S684R+D896N、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+G513C+R515Y+V516M+N682G+E683A+S684R+D896N、Y165W+S166R+G416A+I417R+E418V+F419C+D420Q+H486S+S488W+N682G+E683A+S684R+D896N、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+G513C+R515Y+V516M+T635S+K636G+N637L+N682G+E683A+S684R+D896N、 Y165W+S166R+G416A+I417R+E418V+F419C+D420Q+H486S+S488W+K781T+G783S+E784F+C785L+D896N、Y165W+S166R+G416A+I417R+E418V+F419C+D420Q+H486S+S488W+D638F+R639P+G640N+E641Y+F642L、 Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+N682G+E683A+S684R+P1018L+S1019R+R1020P+N1021R+S1022V; (l) Q160V+E161I+Y162T+N163S+C164V、Q160S+E161S+Y162L+N163C+C164V、Q160P+E161V+Y162H+N163S+C164T、Q160L+E161S+Y162F+N163S+C164A; (m) V7H+T11R+S12M, V7F+R8L+T11R+S12V, V7I+T11V+S12M+Y282F+D283Q+A285T; (n) I149M+N150S+H151C+N152T+L153Y; (o) K195R+K197A+S198A+A199G; (p) T216L+A217T+K219R, C215V+T216S+A217S+K219R; (q) Y282F+D283Q+A285T, D283I+A285R+N292L, Y282F+D283Q+A285T+K865S+P866S+Y867H+N868F+I872L、Y282F+D283Q+A285T+M957V+F958L+Q960V+W961V、Y282F+D283Q+A285T+V348I+I349V+E350L+P351T、Y282F+D283Q+A285T+E748A+G749S+S751G、Y282F+D283Q+A285T+G416S+I417H+E418Q+F419T+D420V、Y282F+D283Q+A285T+G416L+I417Q+E418M+F419R+D420A、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q、Y282F+D283Q+A285T+G416S+I417H+E418Q+F419T+D420V+L501P+N936S、Y282F+D283Q+A285T+G416L+I417Q+E418M+F419R+D420A+I455G+R456E+V458E、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+L525R+I526V+N527D+K528P+K529G+N682G+E683A+S684R+D896N、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+N682G+E683A+S684R+G727E+K728H+N729Q+E730R、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+K781H+L782I+G783F+E784R+C785S、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+D638P+E641M+F642V、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+P510D+V511H+L512V+P1018S+S1019R+R1020V+N1021M+S1022P、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633N+T634E+T635E+K636G+N637L+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633F+T634P+T635F+K636P+N637C+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633S+T634C+T635G+K636H+N637F+N682G+E683A+S684R+S743T、Y282F+D283Q+A285T+G416A+I417R+E418V+F419C+D420Q+R509C+P510S+L512F+C633R+T634Y+T635L+K636V+N637D+N682G+E683A+S684R+S743T;、 (r) G416R+I417W+E418T+F419R+D420V, R509C+P510S+L512F+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+N682G+E683A+S684R+D896N; (3) The following mutations exist relative to the amino acid sequence shown in SEQ ID NO.1: D592A, D643A, E820A or D992A.

2. The Cas protein as described in claim 1, characterized in that, The amino acid sequence of the Cas protein is shown in any one of SEQ ID NO.47-50.

3. The Cas protein as described in claim 1, characterized in that, The mutation mode of the Cas protein relative to SEQ ID NO.1 is as follows: Y282F+D283Q+A285T+T681P+N682P+E683Q+S684C, Y282F+D283Q+A285T+E748A+G749S+S751G, N276S +D283N+E648S+A651L, Y282F+D283Q+A285T+R315S+E316R+T317K+I318K+I319W+E648S+A651L, V7I+ T11V+S12M+Y282F+D283Q+A285T, S240L+S241A+Q242S+E243H+I244L+L307R+Y308F+S309A, S196N+K197F+S198N+A199T+I304A+L307R+Y308A+S309A, Y282F+D283Q+A285T+M957V+F958L+Q960V+W961V, or Y165W+S166R+S198N+G416R+I417V+E418Q+F419V+D420T+T681S+N682R+E683R, Y165W+S166R+G416R+I417V+E418Q+F419V+ D420T+N682G+E683A+D896N, R509C+P510S+L512F+G416R+I417V+E418Q+F419V+D420T+H486A+S488W+N682G+E683A+S684R+ D896N、Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509N+P510G+V511G+L512S+T681S+N682R+E683R、Y165W+S166R+ G416R+I417V+E418Q+F419V+D420T+G513L+N514A+R515V+V516M+N682G+E683A+S684R+D896N, Y165W+S166R+G416R+I417V+ E418Q+F419V+D420T+G513C+R515Y+V516M+N682G+E683A+S684R+D896N, G416R+I417V+E418Q+F419V+D420T+T681S+N682R+ E683R+G727V+K728C+N729G+E730G, S110G+Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+R509A+P510A+V511R+L512Q+ T681S+N682R+E683R、 Y165W+S166R+G416R+I417V+E418Q+F419V+D420T+G513C+R515Y+V516M+T635S+K636G+N637L+N682G+E683A+S684R+D896N.

4. A fusion protein, characterized in that, It comprises the Cas protein according to any one of claims 1-3; and one or more functional domains; The functional domains are selected from nuclear localization signal (NLS), nuclear output signal (NES), reporter protein, epitope tag, deaminase, and T5 exonuclease.

5. The fusion protein as described in claim 4, characterized in that, The functional domain is selected from the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain.

6. The fusion protein as described in claim 5, characterized in that, The adenosine deaminase catalytic domain or cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

7. The fusion protein as described in claim 5, characterized in that, The amino acid sequence of the adenosine deaminase catalytic domain is shown in SEQ ID NO.31 or SEQ ID NO.

70.

8. The fusion protein as described in claim 4, characterized in that, The functional structural domain is TadA8e.

9. The fusion protein as described in claim 4, characterized in that, The functional structural domains include the nuclear localization signal (NLS) and / or the nuclear output signal (NES).

10. The fusion protein as described in claim 9, characterized in that, The sequence of the nuclear positioning signal is shown in any of SEQ ID NO. 38-45.

11. The fusion protein as described in claim 9 or 10, characterized in that, The sequence of the nuclear localization signal is located at, near, or close to the N-terminus or C-terminus of the Cas protein of claim 1.

12. The fusion protein as described in claim 9, characterized in that, The nuclear output signal includes protein tyrosine kinase 2.

13. The fusion protein as described in claim 4, characterized in that, The reporter proteins include glutathione S-transferase, horseradish peroxidase, chloramphenicol acetyltransferase, β-galactosidase, β-glucuronidase, and autofluorescent protein.

14. The fusion protein as described in claim 13, characterized in that, The autofluorescent proteins include green fluorescent protein, HcRed, DsRed, cyan fluorescent protein, yellow fluorescent protein, and blue fluorescent protein.

15. The fusion protein as described in claim 4, characterized in that, The epitope tags include histidine tags, V5 tags, FLAG tags, influenza virus hemagglutinin tags, Myc tags, VSV-G tags, thioredoxin tags, and streptavidin tags.

16. The fusion protein according to claim 4, characterized in that, The functional domain is a T5 exonuclease.

17. The fusion protein of claim 16, characterized in that, The amino acid sequence of the T5 exonuclease is shown in SEQ ID NO.

81.

18. The fusion protein as described in claim 4, characterized in that, The functional domain is attached to the N-terminus and / or C-terminus of the Cas protein.

19. The fusion protein as described in claim 4, characterized in that, One or more functional domains are connected to the N-terminus and / or C-terminus of the Cas protein via adapters.

20. The fusion protein according to claim 4, characterized in that, The amino acid sequence of the fusion protein is shown in any one of SEQ ID NO. 46, 84, and 86.

21. An isolated polynucleotide, characterized in that, The polynucleotide encodes the Cas protein as described in any one of claims 1-3 or the fusion protein as described in any one of claims 4-20.

22. The polynucleotide of claim 21, characterized in that, The polynucleotide has been codon-optimized for expression in eukaryotic cells.

23. The polynucleotide of claim 21, characterized in that, The sequence of the polynucleotide is shown in any one of SEQ ID NO. 2, 37, 83, 85, or 87.

24. A guide RNA (gRNA), characterized in that, It contains (1) A direct repeat (DR) sequence that is capable of forming a complex with the Cas protein of any one of claims 1-3 or the fusion protein of any one of claims 4-20; The direct repeat (DR) sequence is as shown in either SEQ ID NO. 5 or 71; and (2) A spacer sequence that can hybridize with the target sequence of the target DNA, thereby guiding the complex to the target DNA.

25. The guide RNA as described in claim 24, characterized in that, The spacer sequence is at least 15 nt in length.

26. The guide RNA as described in claim 25, characterized in that, The length of the interval sequence is 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nt.

27. The guide RNA as claimed in claim 25, characterized in that, The length of the interval sequence is 15-50 nt.

28. The guide RNA as described in claim 25, characterized in that, The length of the interval sequence is 18-41 nt.

29. The guide RNA as described in claim 25, characterized in that, The length of the interval sequence is 18-27 nt.

30. The guide RNA as described in claim 25, characterized in that, The length of the interval sequence is 18-24 nt.

31. The guide RNA as described in claim 25, characterized in that, The length of the interval sequence is 18-22 nt.

32. The guide RNA according to any one of claims 24-31, characterized in that, The 5' end of the same-direction repeating sequence is connected to the interval sequence.

33. The guide RNA as described in claim 24, characterized in that, The spacer sequence is any one of the nucleotide sequences in SEQ ID NO. 6, 11, 53, 55, 57, 59, 61, 101.

34. An isolated nucleic acid molecule, characterized in that, It contains, or consists of, sequences selected from, the following: (i) The sequence shown in SEQ ID NO: 5 or 71; (ii) The complementary sequence of the sequence described in (i).

35. The isolated nucleic acid molecule as described in claim 34, characterized in that, The isolated nucleic acid molecule is RNA.

36. A carrier, characterized in that, It comprises the polynucleotide of any one of claims 21-23, and / or the nucleotide encoding the guide RNA of any one of claims 24-33.

37. The carrier as described in claim 36, characterized in that, The vector is a plasmid or a viral vector.

38. The carrier as described in claim 37, characterized in that, The viral vector is selected from the group consisting of adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpesvirus, SV40, poxvirus, or combinations thereof.

39. The carrier as described in claim 36, characterized in that, The vectors are selected from the following group: cloning vectors, transformation vectors, expression vectors, shuttle vectors, integration vectors, and multifunctional vectors.

40. The carrier according to any one of claims 36-39, characterized in that, The carrier comprises: (1) A first regulatory element, operatively linked to a nucleotide sequence encoding a Cas protein of any one of claims 1-3 or a nucleotide sequence encoding a fusion protein of claim 4; and (2) A second regulatory element, the second regulatory element being operatively linked to a nucleotide sequence encoding the guide RNA of any one of claims 24-33.

41. The carrier as described in claim 40, characterized in that, The first control element and the second control element are located on the same or different carriers.

42. The carrier as described in claim 40, characterized in that, The first control element and / or the second control element are promoters.

43. The carrier as described in claim 36, characterized in that, The vector contains one or more promoters operatively linked to the polynucleotide and / or nucleotide, enhancer, transcription termination signal, polyadenylation sequence, origin of replication, selectivity marker, nucleic acid restriction site, and / or homologous recombination site.

44. A complex, characterized in that, Include: (i) A protein component selected from the group consisting of: the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, or a combination thereof; and (ii) The guide RNA according to any one of claims 24-33.

45. A CRISPR-Cas system, characterized in that, Include: (i) the Cas protein of any one of claims 1-3 or the fusion protein of any one of claims 4-20, or the nucleotide encoding said Cas protein or fusion protein; and (ii) The guide RNA of any one of claims 24-33, or the nucleotide encoding the guide RNA.

46. ​​A CRISPR-Cas composition, characterized in that, Include: (i) A first component selected from the group consisting of: the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, the nucleotide sequence encoding the Cas protein of any one of claims 1-3 or the fusion protein of any one of claims 4-20, and any combination thereof; and (ii) A second component comprising one or more guide RNAs according to any one of claims 24-33, or a nucleotide sequence encoding the guide RNA comprising one or more of claims 24-33.

47. A CRISPR-Cas system, characterized in that, It comprises one or more carriers, wherein the one or more carriers comprise: (i) a first nucleic acid, which is a nucleotide sequence encoding the Cas protein of any one of claims 1-3 or the fusion protein of any one of claims 4-20; and (ii) A second nucleic acid, which is a nucleotide sequence encoding the guide RNA comprising any one of claims 24-33; in: The first nucleic acid and the second nucleic acid may exist on the same or different vectors.

48. A reagent kit, characterized in that, It includes one or more components selected from the following: the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, the polynucleotide of any one of claims 21-23, the vector of any one of claims 36-43, the complex of claim 44, the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, or the CRISPR-Cas system of claim 47.

49. A delivery composition, characterized in that, It comprises a delivery medium and one or more of the following: the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, the polynucleotide of any one of claims 21-23, the vector of any one of claims 36-43, the complex of claim 44, the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, or the CRISPR-Cas system of claim 47.

50. A host cell, characterized in that, The invention comprises the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, the polynucleotide of any one of claims 21-23, the vector of any one of claims 36-43, the complex of claim 44, the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, the CRISPR-Cas system of claim 47, or the delivery composition of claim 49.

51. An enzyme preparation, characterized in that, The enzyme preparation comprises the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, or the complex of claim 44.

52. A medicine box, characterized in that, include: A first container, and a compound of claim 44 or a CRISPR-Cas system of claim 45, a CRISPR-Cas composition of claim 46 or a CRISPR-Cas system of claim 47, or a drug containing the compound of claim 44 or the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46 or the CRISPR-Cas system of claim 47, located in the first container.

53. A medicine box, characterized in that, include: (a1) A first container, and the Cas protein of any one of claims 1-3 or the fusion protein of any one of claims 4-20, or its encoding gene or expression vector, located in the first container, or a drug containing the Cas protein of any one of claims 1-3 or the fusion protein of any one of claims 4-20, or its encoding gene or expression vector; (b1) A second container, and the guide RNA or its expression vector of any one of claims 24-33 located in the second container, or a drug containing the guide RNA or its expression vector of any one of claims 24-33.

54. A method for non-diagnostic and non-therapeutic targeting and editing or cleaving of target genes in vitro, characterized in that, include: The Cas protein of any one of claims 1-3, or the fusion protein of any one of claims 4-20, or the complex of claim 44, or the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, or the CRISPR-Cas system of claim 47, or the delivery composition of claim 49, or the enzyme preparation of claim 51, or the cassette of claim 52 or 53, is contacted with the target gene or delivered to a cell containing the target gene, wherein the target sequence is present in the target gene.

55. A method for non-diagnostic and non-therapeutic alteration of gene product expression in vitro, characterized in that, include: The Cas protein of any one of claims 1-3, or the fusion protein of any one of claims 4-20, or the complex of claim 44, or the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, or the CRISPR-Cas system of claim 47, or the delivery composition of claim 49, or the enzyme preparation of claim 51, or the cassette of claim 52 or 53, is contacted with a nucleic acid molecule encoding the gene product, or delivered to a cell containing the nucleic acid molecule, wherein the target sequence is present in the nucleic acid molecule.

56. Use of the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, the polynucleotide of any one of claims 21-23, the vector of any one of claims 36-43, the complex of claim 44, the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, or the CRISPR-Cas system of claim 47, or the kit of claim 48, or the delivery composition of claim 49, or the enzyme preparation of claim 51, or the cassette of claim 52 or 53, characterized in that, Used to prepare formulations for nucleic acid editing.

57. Use of the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, the polynucleotide of any one of claims 21-23, the vector of any one of claims 36-43, the complex of claim 44, the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, or the CRISPR-Cas system of claim 47, or the kit of claim 48, or the delivery composition of claim 49, or the enzyme preparation of claim 51, or the cassette of claim 52 or 53, characterized in that, For the preparation of formulations for one or more uses selected from the group consisting of: (i) In vitro gene or genome editing; (ii) Detection of isolated single-stranded DNA.

58. A method for non-diagnostic and non-therapeutic detection of the presence of target nucleic acid molecules in a sample, characterized in that, The method includes contacting a sample with the Cas protein of any one of claims 1-3, the fusion protein of any one of claims 4-20, or the complex of claim 44, the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, or the CRISPR-Cas system of claim 47, the kit of claim 48, or the delivery composition of claim 49, or the enzyme preparation of claim 51, and a non-target sequence, wherein the non-target sequence is a reporter nucleic acid; when a target nucleic acid molecule is present, detecting a detectable signal generated by the cleavage of the non-target sequence to detect the target nucleic acid molecule, wherein the non-target sequence does not hybridize with the guide RNA.

59. A pharmaceutical composition, characterized in that, The invention comprises the vector of any one of claims 36-43, the complex of claim 44, the CRISPR-Cas system of claim 45, the CRISPR-Cas composition of claim 46, or the CRISPR-Cas system of claim 47, or the kit of claim 48, or the delivery composition of claim 49, or the host cell of claim 48, or the enzyme of claim 51, and a pharmaceutically acceptable carrier or excipient.

Citation Information

Patent Citations

  • Adenosine deaminase, base editor fusion protein, base editor system and application

    CN114634923A

  • Formulation and delivery of PLGA microspheres

    US20130244279A1

  • Dlin-MC3-DMA lipid nanoparticle delivery of modified polynucleotides

    US20130245107A1

  • Formulation and delivery of PLGA microspheres

    US20130252281A1

  • Poly(beta-amino alcohols), their preparation, and uses thereof

    US20130302401A1