Programmable nuclease and application thereof

By developing a novel nuclease containing HNH, RuvC, and REC leaves, the performance limitations of the CRISPR-Cas system in DNA targeting and editing have been addressed, achieving highly efficient DNA targeting and editing effects suitable for a variety of applications.

CN121737086APending Publication Date: 2026-03-27YOLTECH THERAPEUTICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411353459.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems are inadequate in DNA targeting and editing, failing to meet diverse needs.

Method used

A novel programmable nuclease is provided, comprising an HNH domain and a RuvC domain, which binds to REC leaves, has an amino acid sequence length between 600 and 800 amino acids, and forms a complex with a guide nucleic acid, enabling it to specifically edit target DNA.

Benefits of technology

It achieves efficient targeting and editing of DNA, and is characterized by its small molecular size and easy delivery, making it suitable for a variety of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005062935420000221
    Figure BDA0005062935420000221
  • Figure BDA0005062935420000231
    Figure BDA0005062935420000231
  • Figure BDA0005062935420000251
    Figure BDA0005062935420000251
Patent Text Reader

Abstract

The invention provides a novel nuclease and a method for gene editing and nucleic acid detection by using the nuclease. In some embodiments, the disclosure provides uses of the programmable nuclease. The disclosure also provides compositions, systems, and methods comprising the programmable nuclease.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure provides a targeting DNA-guided nucleic acid comprising a targeting sequence, and programmable nucleases and applications thereof. The present disclosure further provides methods of targeting, binding, cleaving, labeling or modifying a target DNA. BACKGROUND

[0002] The CRISPR-Cas system recognizes specific DNA targets through a guide nucleic acid and can specifically edit the DNA target through a nuclease, based on which, the CRISPR-Cas system (including Cas9, Cas12, Cas13 system) has been developed as a genome editing tool. However, the performance of each nuclease system still has its own advantages and disadvantages, therefore, there is still a need to develop new nucleases and corresponding CRISPR-Cas systems to meet the diversified needs.

[0003] The citation or identification of any document in this disclosure is not an admission that such document is available as prior art to the present disclosure. Each reference recited herein is incorporated by reference in its entirety. SUMMARY

[0004] The present disclosure meets the above needs by providing new programmable nucleases and CRISPR-Cas systems.

[0005] In some aspects, the present disclosure provides a programmable nuclease comprising an HNH domain and a RuvC domain, further comprising a REC lobe having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 12, the REC lobe comprising a REC domain having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 21, the nuclease amino acid sequence length is between 600-800 amino acids, preferably, between 600-700 amino acids, more preferably, between 650-700 amino acids.

[0006] In some aspects, the present disclosure provides a nuclease system comprising:

[0007] (1) a programmable nuclease of the present disclosure or a polynucleotide encoding the programmable nuclease; and

[0008] (2) a guide nucleic acid or a polynucleotide encoding the guide nucleic acid, the guide nucleic acid comprising:

[0009] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0010] (ii) a DNA targeting sequence capable of hybridizing to a target sequence of a target DNA, thereby directing the complex to the target DNA;

[0011] Optionally, wherein the protein binding sequence is at the 3’ end of the DNA targeting sequence; and / or

[0012] Optionally, wherein the guide nucleic acid is a guide RNA (gRNA).

[0013] In some aspects, the present disclosure provides a nucleic acid molecule encoding the programmable nuclease of the present disclosure.

[0014] In some aspects, the present disclosure provides a nucleic acid molecule comprising one or more polynucleotide sequences encoding a guide nucleic acid, wherein the guide nucleic acid comprises:

[0015] (i) a DNA targeting sequence comprising a nucleotide sequence that is complementary to a target sequence in a target DNA; and

[0016] (ii) a protein binding sequence that can bind to the programmable nuclease of the present disclosure to form a complex;

[0017] Optionally, the one or more polynucleotide sequences encoding the guide nucleic acid are operably linked to a promoter that functions in a eukaryotic cell.

[0018] In some aspects, the present disclosure provides a vector comprising the polynucleotide of the present disclosure; optionally, wherein the vector encodes the guide nucleic acid of the present disclosure; optionally, wherein the vector is a plasmid vector, a recombinant AAV (rAAV) vector, or a recombinant lentivirus vector.

[0019] In some aspects, the present disclosure provides a ribonucleoprotein comprising the programmable nuclease of the present disclosure and the guide nucleic acid of the present disclosure.

[0020] In some aspects, the present disclosure provides a programmable nuclease composition comprising:

[0021] (1) the programmable nuclease of the present disclosure or a polynucleotide encoding the programmable nuclease; and

[0022] (2) a guide nucleic acid or a polynucleotide encoding the guide nucleic acid, the guide nucleic acid comprising:

[0023] (i) a protein binding sequence capable of forming a complex with the programmable nuclease; and

[0024] (ii) a DNA targeting sequence capable of hybridizing to a target sequence of a target DNA, thereby directing the complex to the target DNA;

[0025] Optionally, wherein the protein-binding sequence is at the 3’ end of the DNA-targeting sequence; and / or

[0026] Optionally, wherein the guide nucleic acid is a guide RNA (gRNA).

[0027] In some aspects, the disclosure also provides a delivery system comprising the programmable nuclease of the disclosure, or the nucleic acid molecule of the disclosure, or the programmable nuclease composition of the disclosure.

[0028] In one embodiment, the delivery system further comprises a delivery vehicle, the delivery vehicle comprising a nanoparticle, a liposome, an exosome, a microvesicle, a gene gun, or an electroporation device.

[0029] In some aspects, the disclosure provides a lipid nanoparticle comprising the programmable nuclease of the disclosure or the system of the disclosure.

[0030] In some aspects, the disclosure provides a cell comprising

[0031] (1) the system of the disclosure;

[0032] (2) the nucleic acid molecule of the disclosure;

[0033] (3) the vector of the disclosure;

[0034] (4) the ribonucleoprotein of the disclosure; or

[0035] (5) the lipid nanoparticle of the disclosure.

[0036] In some aspects, the disclosure provides a pharmaceutical composition comprising (1) the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, or the cell of the disclosure; and (2) a pharmaceutically acceptable excipient.

[0037] In some aspects, the disclosure provides a method of targeting, binding, cleaving, labeling, or modifying a target DNA, the method comprising contacting the target DNA with a complex, the complex comprising:

[0038] (1) the programmable nuclease of the disclosure or a polynucleotide encoding the programmable nuclease; and

[0039] a guide nucleic acid or a polynucleotide encoding the guide nucleic acid, the guide nucleic acid comprising:

[0040] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0041] (ii) a DNA targeting sequence capable of hybridizing to a target sequence of a target DNA, thereby directing the complex to the target DNA;

[0042] In some aspects, the disclosure provides a method for diagnosing, preventing, or treating a disease in a subject in need thereof, the method comprising administering to the subject the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, the cell of the disclosure, or the pharmaceutical composition of the disclosure, wherein the disease is associated with a target DNA, wherein the spacer sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnoses, prevents, or treats the disease.

[0043] In some aspects, the disclosure provides a method of detecting a target DNA, the method comprising contacting the target DNA with the system of the disclosure, wherein the target DNA is modified by the complex, and wherein the modification detects the target DNA; optionally, wherein the modification generates a detectable signal, e.g., a fluorescent signal.

[0044] In some aspects, the disclosure provides a kit comprising the programmable nuclease of the foregoing, the nucleic acid molecule of the foregoing, the composition of the foregoing, the system of the foregoing, the vector of the foregoing, the ribonucleoprotein of the foregoing, the lipid nanoparticle of the foregoing, the cell of the foregoing, or the pharmaceutical composition of the foregoing.

[0045] In a preferred embodiment, the components of the kit are in the same or different containers.

[0046] In some aspects, the disclosure also provides a container comprising the kit of the foregoing.

[0047] In one embodiment, the container comprises a sterile container;

[0048] In one embodiment, the container comprises a syringe.

[0049] According to WIPO Standard ST.26, the symbol "t" is used to represent both T in DNA and U in RNA. Thus in the present sequence listing prepared according to ST.26, T in a sequence should be read as U when the sequence is RNA. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 A phylogenetic tree of YTGE-1001 is depicted, showing that YTGE-1001 diverged clade away from Cas9 systems as references.

[0051] Figure 2An exemplary domain distribution of YTGE-1001 is described, from which it can be seen that YTGE-1001 is smaller than Cas9 (SaCas9 or SpCas9) and has a different domain distribution, for example, there is a Linker structure between the BH domain and the REC domain.

[0052] Figure 3 A predicted 3D structure of YTGE-1001 is described. From the figure, it can be seen that there is a clear Linker structure in YTGE-1001.

[0053] Figure 4 A secondary structure prediction of the protein binding sequence (scaffold sequence) corresponding to YTGE-1001 is described.

[0054] Figure 5 YTGE-1001 and PCSK9-sgRNA expression plasmid maps are described.

[0055] Figure 6 A map of the target plasmid as a target object is described.

[0056] Figure 7 Endonuclease activity of YTGE-1001 is described, indicated by the points falling outside the right dashed line.

[0057] Figure 8 PAM analysis results corresponding to YTGE-1001 are described. DETAILED DESCRIPTION

[0058] The details of one or more embodiments of the disclosure are set forth in the accompanying description below. Other features or advantages of the disclosure will be apparent from the following drawings and from the detailed description of several embodiments, which should be considered in conjunction with the claims. It should be understood that any of the features, aspects, or embodiments of the disclosure can be combined with any of the other features, aspects, or embodiments of the disclosure, to form another embodiment of the disclosure, either explicitly or implicitly disclosed herein.

[0059] SUMMARY

[0060] The present disclosure provides a new programmable nuclease, which has low similarity with other nucleases compared with known nucleases, has DNA (e.g., ssDNA or dsDNA) cleavage activity, has a small molecule, and is easy to deliver.

[0061] The programmable nuclease of the present disclosure has certain functional similarity with existing nucleases (e.g., Cas9), but it does not belong to the Cas9 protein, which comprises multiple functional domains such as HNH, RuvC, BH, REC lobe, WED, PI, etc., wherein the REC lobe comprises a REC domain, and there is a Linker structure between the BH domain and the REC domain, which is different from existing nucleases, showing that it is a new type of nuclease.

[0062] The guide nucleic acid that guides the nuclease of the present disclosure can be RNA, which is referred to as guide RNA (gRNA) or crRNA. The guide nucleic acid comprises a scaffold sequence capable of interacting with the nuclease of the present disclosure to form a complex (protein-RNA complex) and a spacer sequence capable of hybridizing with a target sequence in a target DNA, thereby guiding the complex to the target DNA.

[0063] The term

[0064] The present disclosure will be described with respect to specific embodiments, but the disclosure is not limited thereto and only limited by the claims. Unless otherwise indicated, the terms set forth below are to be understood in their common meaning.

[0065] As used herein, the singular forms "a", "an" and "the" include both singular and plural referents unless the context clearly dictates otherwise.

[0066] As used herein, the term "about" is used herein to mean approximately, around, or in the region of, and can refer to a value or composition that is within an acceptable error range for the specific value or composition determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined. For example, as used herein, the expression "about 100" includes all values between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).

[0067] As used herein, the term "comprising" or "including" can be open, semi-closed and closed. In other words, the term also includes "consisting essentially of" or "consisting of".

[0068] As used herein, the term "optionally" means that the event, condition or substituent subsequently described can or can not occur, and the description includes instances where the event or condition occurs and instances where it does not.

[0069] The terms "substantially" and "essentially," as used herein, mean to a degree or extent that is about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more of a degree, amount, level, value, number, frequency, percentage, dimension, size, quantity, weight, or length, as compared to a reference amount, level, value, number, frequency, percentage, dimension, size, quantity, weight, or length. For example, as used herein, a "substantially identical sequence" can refer to a sequence, polynucleotide, or polypeptide that has a certain degree of identity to a reference sequence.

[0070] The term "and / or," as used herein, means either or both of the alternatives exist.

[0071] As used herein, the terms "nucleic acid," "nucleic acid molecule," "polynucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides or their analogs, in either single- or double-stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can exist in a cell-free environment. A polynucleotide can be a gene or a fragment of a gene. A polynucleotide can be DNA or RNA. A polynucleotide can include one or more analogs of a natural nucleotide (e.g., altered backbone, sugar or nucleobase). If present, modifications to the nucleotide structure can be imparted before or after assembly into a polymer. Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acid, heteronomous nucleic acid, morpholino, locked nucleic acid, glycerol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, a number of loci defined as a contiguous stretch of DNA (one locus), exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. A nucleotide sequence can be interspersed with non-nucleotide components. A polynucleotide can include a mixture of nucleotides found in nature and nucleotide analogs (e.g., synthetic nucleotide analogs).

[0072] As used herein, the terms "peptide," "polypeptide," and "protein" are used interchangeably herein and generally refer to a polymer of at least two amino acid residues linked by peptide bonds. The terms do not denote a particular length of the polymer, nor are they intended to be limited by the particular amino acid content of the polymer. The terms also encompass amino acid polymers in which one or more amino acid residues are anabolic or catabolic modifications of a naturally occurring amino acid. In some embodiments, the polymer can be interspersed with non-amino acids. The terms include amino acid chains of any length including full-length proteins as well as proteins with or without secondary and / or tertiary structure (e.g., domains). The terms also encompass amino acid polymers that have been modified; for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation and any other manipulation, such as conjugation with a labeling component.

[0073] As used herein, the term "identity" refers to the overall relatedness between polymer molecules, e.g., between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. Sequence identity (or homology) is determined by comparing two aligned sequences over a predetermined comparison window, which can be 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the length of a reference nucleotide sequence or protein, and determining the number of positions at which the same residue occurs in both sequences. Typically, this is expressed as a percentage. In the case of polypeptide sequences, if the type of mutation is one or more of the following: substitution / replacement of one or more amino acids / nucleotides, insertion within a sequence, and deletion within a sequence, the total number of residues is calculated as the larger of the two molecules being compared. If the type of mutation also includes an insertion (extension) at either or both ends of a sequence or a deletion (truncation) at either or both ends of a sequence, the number of amino acids inserted or deleted at either or both ends (e.g., less than 20 inserted or deleted at both ends) is not counted in the total number of residues. In calculating the percent identity, the sequences being compared are aligned to produce the maximum match of sequences, gaps, if any, in the alignment are addressed by particular algorithms. The same is true for nucleotide identity calculations.

[0074] As used herein, the term "ortholog" refers to a gene (and the protein encoded by the gene) that is inferred to be the product of a speciation event that formed a new species from an ancestral sequence: when one species splits into two separate species, the copies of a single gene in these two resulting species are said to be orthologs. Orthologs or orthologous genes are genes in different species that are descended from a single gene in their last common ancestor.

[0075] As used herein, the terms "amino acid" and "amino acids" generally refer to natural and unnatural amino acids, including but not limited to modified amino acids and amino acid analogs. Modified amino acids can include natural amino acids and unnatural amino acids that have been chemically modified to include a group or chemical moiety that does not naturally occur on an amino acid. An amino acid analog can refer to an amino acid derivative. The term "amino acid" includes D-amino acids and L-amino acids.

[0076] As used herein, the term "nuclease" refers to a polypeptide that is capable of cleaving the phosphodiester bond between nucleotide subunits of a nucleic acid; the term "endonuclease" refers to a polypeptide that is capable of catalyzing (e.g., cleaving) a phosphodiester bond within a polynucleotide (e.g., DNA or RNA) strand.

[0077] As used herein, the term "exonuclease" refers to a protein or polypeptide that is capable of digesting a nucleic acid (e.g., RNA or DNA) from a free end.

[0078] As used herein, the term“non-native” can generally refer to a nucleic acid or polypeptide sequence that is not found in a naturally occurring nucleic acid or protein. Non-native can refer to an affinity tag. Non-native can refer to a fusion. Non-native can refer to a naturally occurring nucleic acid or polypeptide sequence that includes a mutation, insertion, and / or deletion. A non-native sequence can exhibit and / or encode an activity (e.g., an enzymatic activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) that can also be exhibited by a nucleic acid and / or polypeptide sequence that is fused to the non-native sequence. A non-native nucleic acid or polypeptide sequence can be linked by genetic engineering to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to produce a chimeric nucleic acid and / or polypeptide sequence encoding a chimeric nucleic acid and / or polypeptide.

[0079] As used herein, the term“fusion protein” refers to a hybrid polypeptide that includes protein domains from at least two different proteins. A protein can be positioned at an amino-terminal (N-terminal) portion or at a carboxy-terminal (C-terminal) protein of a fusion protein, thus forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively. A protein can include different domains, for example, a nucleic acid binding domain (e.g., a gRNA binding domain of a nuclease that directs binding of the protein to a target site) and a nucleic acid cleavage domain, or a catalytic domain of a nucleic acid editing protein. In some embodiments, there can be a linker between the proteins. In some embodiments, a protein includes a protein portion (e.g., an amino acid sequence that constructs a nucleic acid binding domain) and an organic compound (e.g., a compound that can act as a nucleic acid cleavage agent). In some embodiments, a protein is complexed or associated with a nucleic acid (e.g., an RNA or DNA).

[0080] As used herein, the term "linker" refers to any means, entity, or moiety for joining two or more entities. In some embodiments, the linker is a covalent linker. In some embodiments, the linker is a non-covalent linker. Examples of covalent linkers include covalent bonds or linker moieties covalently attached to one or more proteins or domains to be connected. In some embodiments, the linker is a non-covalent bond, such as an organometallic bond through a metal center, such as a platinum atom. The joining can be permanent or reversible. For covalent attachment, various functional groups can be used, such as amide groups, including carbonic acid derivatives, ethers, esters (including organic and inorganic esters), amines, carbamates, ureas, and the like. To provide attachment, the domains can be modified to provide coupling sites by oxidation, hydroxylation, substitution, reduction, and the like. Conjugation methods are well known to those of skill in the art and are encompassed for use in the present application. Linker moieties include, but are not limited to, chemical linker moieties, or, for example, peptide linker moieties (linker sequences). The length and type of linker can be designed as desired. In some embodiments, the linker can be selected from an artificially synthesized amino acid sequence or a naturally occurring polypeptide sequence. It will be appreciated that modifications that do not significantly reduce the functionality of the RNA-binding domain and the effector domain are preferred.

[0081] As used herein, the term "complementary" refers to the ability of a nucleobase of a first polynucleotide sequence (e.g., a guide sequence) to base pair with a nucleobase of a second polynucleotide sequence (such as a target sequence) through traditional Watson-Crick base pairing. In some embodiments, a first polynucleotide can be substantially complementary to a second polynucleotide, i.e., have at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementarity to the second polynucleotide. In some embodiments, a first polynucleotide is fully complementary to a second polynucleotide, i.e., has 100% complementarity to the second polynucleotide. "Substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% complementary over an 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions. The term "stringent conditions" in relation to hybridization refers to conditions under which one nucleic acid having complementarity to a target sequence will hybridize primarily to that target sequence and not to non-target sequences. Stringent conditions are typically sequence dependent, and depend on many factors. Generally, the longer the sequence, the higher the temperature at which the sequence will specifically hybridize to its target sequence. "Hybridize" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues. The complex can comprise two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction can constitute a step in a more extensive process (such as the initiation of PCR, or cleavage of a polynucleotide by an enzyme). A sequence capable of hybridizing to a given sequence is referred to as the "complement" of the given sequence.

[0082] As used herein, the term "domain" or "protein domain" refers to a portion of a protein sequence that can exist and function independently of the remainder of the protein chain.

[0083] As used herein, the term "operably linked" refers to functional linkage between a regulatory sequence and a target nucleic acid sequence, thereby inducing expression of the latter (e.g., in an in vitro transcription / translation system or in a host cell when the vector has been introduced into the host cell). For example, a first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence.

[0084] As used herein, the term "conservative amino acid substitution" refers to the replacement of an amino acid with a chemically or functionally similar amino acid without affecting the normal function of the protein, e.g., the interchangeability of amino acid residues with similar side chains: a group of amino acids having aliphatic side chains is comprised of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains is comprised of serine and threonine; a group of amino acids having amide-containing side chains is comprised of asparagine and glutamine; a group of amino acids having aromatic side chains is comprised of phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains is comprised of lysine, arginine, and histidine; a group of amino acids having sulfur-containing side chains is comprised of cysteine and methionine. In some embodiments, considered conservative amino acid substitutions include: aspartate-glutamate, lysine-arginine-histidine, serine-threonine-asparagine-glutamine, glycine-alanine-valine-leucine-isoleucine, cysteine-methionine-proline, phenylalanine-tyrosine-tryptophan. In some embodiments, considered conservative amino acid substitutions include: valine-leucine-methionine-isoleucine, alanine-serine-threonine.

[0085] As used herein, the term "Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meaning generally understood by those of skill in the art, and generally includes a transcription product or other element associated with the expression of a CRISPR-associated ("Cas") gene, or a transcription product or other element capable of directing the activity of the Cas gene.

[0086] As used herein, the term "complex" refers to a combination of two or more molecules. In some embodiments, a complex includes a polypeptide and a nucleic acid molecule that interact with each other (e.g., bind, contact, adhere). For example, the term "complex" can refer to a combination of a guide RNA and a polypeptide (e.g., a nuclease). Alternatively, the term "complex" can refer to a combination of a guide RNA, a nuclease, and a complementary region of a target nucleic acid.

[0087] As used herein, the term "cleave" refers to the breakage of the covalent backbone of a DNA molecule. Cleavage can be carried out by various methods, including but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Single-strand cleavage and double-strand cleavage are both possible, and a double-strand cleavage can occur as a result of two distinct single-strand cleavage events. DNA cleavage can result in the production of a blunt end or a sticky end.

[0088] The term "cleave" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond within the backbone of a target nucleotide sequence, which can result in a single- or double-stranded break within the target sequence. For example, a nuclease of the disclosure, or a variant polypeptide thereof, can cleave nucleotides within a polynucleotide as an endonuclease, or can be an exonuclease (removing contiguous nucleotides from the end (5' and / or 3' end) of a polynucleotide). In some embodiments, cleavage of a target polynucleotide by a nuclease of the disclosure, or a variant polypeptide thereof, can result in staggered breaks or blunt ends.

[0089] As used herein, the term "protospacer adjacent motif" or "PAM" refers to a short sequence (or motif) adjacent to a protospacer sequence on a non-target strand of a target nucleic acid recognized by a CRISPR complex. The target nucleic acid is double-stranded DNA (dsDNA), one strand comprising a target sequence adjacent to the PAM and is referred to as the "PAM strand" (e.g., non-target strand or non-spacer complement strand), while the other complementary strand is referred to as the "non-PAM strand" (e.g., target strand or spacer complement strand). As used herein, the term "adjacent" includes instances in which the RNA guide of the complex specifically binds to, interacts with, or associates with the target sequence immediately next to the PAM. In such instances, there are no nucleotides between the target sequence and the PAM. The term "adjacent" also includes instances in which there are a small number (e.g., 1, 2, 3, 4, or 5) of nucleotides between the target sequence bound by the targeting moiety and the PAM.

[0090] As used herein, the terms “guide nucleic acid,” “RNA guide,” “RNA guide sequence,” “guide RNA (gRNA),” “single guide RNA (sgRNA)” are used interchangeably to refer to a nucleic acid-based molecule capable of forming a complex with a CRISPR-nuclease (e.g., a programmable nuclease of the present disclosure) and comprising a sequence (e.g., a guide sequence) sufficiently complementary to a target nucleic acid to hybridize to the target nucleic acid and direct the complex to the target nucleic acid, including but not limited to RNA-based molecules, e.g., guide RNAs. A guide nucleic acid can comprise a segment that can be referred to as a “nucleic acid targeting segment” or “nucleic acid targeting sequence,” which can comprise a sub-segment that can be referred to as a “protein binding sequence” or “protein binding sequence” or “nuclease binding segment.” A guide nucleic acid can be a DNA molecule, an RNA molecule, or a DNA / RNA hybrid molecule. A “DNA / RNA hybrid molecule” refers to a nucleic acid comprising one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, whether contiguous or not. However, a “DNA molecule” or “RNA molecule” can also refer to a DNA molecule containing one or more modified or unmodified ribonucleotides, whether contiguous or not, or an RNA molecule containing one or more modified or unmodified deoxyribonucleotides, whether contiguous or not.

[0091] As used herein, the term “activity” refers to biological activity. In some embodiments, nuclease activity includes enzymatic activity, e.g., catalytic ability, of a nuclease. For example, nuclease activity can include nuclease activity. In some embodiments, nuclease activity includes binding activity, e.g., binding activity of a nuclease to an RNA guide and / or a target nucleic acid.

[0092] As used herein, the term “protein binding sequence” is used interchangeably with “scaffold sequence,” “tracrRNA” to refer to an RNA comprising a sequence that forms a structure required for a CRISPR-associated protein to bind a particular target nucleic acid.

[0093] As used herein, the terms “upstream” and “downstream” refer to relative positions within individual nucleic acid (e.g., DNA) sequences in a nucleic acid molecule. “Upstream” and “downstream” relate to the 5’ to 3’ direction in which RNA transcription occurs. A first sequence is upstream of a second sequence when the 3’ end of the first sequence precedes the 5’ end of the second sequence. A first sequence is downstream of a second sequence when the 5’ end of the first sequence follows the 3’ end of the second sequence.

[0094] As used herein, "regulatory elements" include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals, poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some embodiments, regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells, as well as those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired tissue of interest, such as muscle, neuronal, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). In other embodiments, regulatory elements can also direct expression in a temporal manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which can or can not be tissue- or cell type-specific. As used herein, the term "promoter" refers to a non-coding nucleotide sequence that is located upstream from a gene and initiates the expression of the downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or defining a gene product, will result in the production of the gene product in a cell under most or all physiological conditions. An inducible promoter refers to a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of an endogenous or exogenous stimulus, such as by a chemical compound (chemical inducer), or in response to an environmental, hormonal, chemical, and / or developmental signal. Inducible or regulated promoters include, for example, promoters that are induced or regulated by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounding, or chemicals such as ethanol, abscisic acid (ABA), jasmonate, salicylic acid, or safeners.

[0095] As used herein, the term "on-target" refers to a predetermined or intended region of DNA that is bound, cleaved, and / or edited by a nuclease of the disclosure or a variant polypeptide thereof.

[0096] As used herein, the term "off-target" refers to a non-predetermined or non-intended region of DNA that is bound, cleaved, and / or edited, for example, by a nuclease of the disclosure or a variant polypeptide thereof. In some embodiments, a region of DNA is an off-target region when it differs from a predetermined or intended region of DNA that is bound, cleaved, and / or edited by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. In some embodiments, such sites are detected using targeted sequencing of sites predicted via in silico modeling or by other methods known in the art.

[0097] As used herein, the term "in vivo" refers to events that occur within a multicellular organism, such as a human or non-human animal. In the case of cell-based systems, it can be used to refer to events that occur within a living cell (as opposed to, for example, in vitro systems).

[0098] As used herein, the term "ex vivo / in vitro" refers to events that occur in a cell or tissue that is grown outside of a multicellular organism, rather than within a multicellular organism.

[0099] As used herein, the term "cell" is to be understood as referring not only to a particular individual cell, but also to progeny or potential progeny of the cell. Since certain modifications can occur in successive generations due to mutations or environmental influences, such progeny can not, in principle, be identical to the parent cell, but are still included in the scope of the term.

[0100] As used herein, the terms "subject," "individual," and "patient" are used interchangeably herein and refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Also encompassed are tissues, cells, and progeny of biological entities obtained in vivo or cultured in vitro.

[0101] As used herein, the term "disease" refers to any condition or disorder that impairs or interferes with the normal functioning of a cell, tissue, or organ, including, but not limited to, those specific diseases that have been medically or clinically defined.

[0102] As used herein, the term "treatment" refers to the administration of a therapeutic molecule (e.g., a CRISPR-Cas system described herein) that partially or completely alleviates, ameliorates, relieves, inhibits, delays onset of, reduces severity of, and / or reduces incidence of one or more symptoms or features of a disease, disorder, and / or condition. Such treatment can be of a subject who does not exhibit signs of the relevant disease, disorder, and / or condition and / or of a subject who exhibits only early signs of the disease, disorder, and / or condition. Alternatively or additionally, such treatment can be of a subject who exhibits one or more established signs of the relevant disease, disorder, and / or condition.

[0103] As used herein, a "biological sample" can contain whole cells and / or viable cells and / or cell fragments. A biological sample can comprise (or be derived from) a "body fluid." In some embodiments, a body fluid can be selected from amniotic fluid, aqueous humor, vitreous fluid, bile, blood serum, breast milk, cerebrospinal fluid, cerumen, chyle, chyme, endolymph, perilymph, exudate, feces, female ejaculate, gastric acid, gastric juice, lymph, mucus (including nasal drainage and sticky sputum), pericardial fluid, peritoneal fluid, pleural fluid, pus, rheum, saliva, sebum (skin oil), semen, sputum, synovial fluid, sweat, tears, urine, vaginal discharge, vomit, and mixtures of one or more thereof. Biological samples include cell cultures, body fluids, cell cultures from body fluids. Body fluids can be obtained from a mammal, for example, by puncture or other collection or sampling procedures.

[0104] As used herein, "expression of a genomic locus" or "gene expression" is the process by which information from a gene is used in the synthesis of a functional gene product. The product of gene expression is usually a protein, but in non-protein coding genes such as rRNA genes or tRNA genes, the product is functional RNA. All known forms of life—eukaryotes (including multicellular organisms), prokaryotes (bacteria and archaea), and viruses—use processes of gene expression to produce functional products, thereby enabling life. As used herein, "expression" of a gene or nucleic acid encompasses not only the transcription and translation of cellular gene expression but also transcription and translation of nucleic acids in a cloning context and any other context. As used herein, "expression" also refers to the process by which a polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene product." If the polynucleotide is derived from genomic DNA, expression can include splicing of mRNA in eukaryotic cells.

[0105] As used herein, the term "RuvC domain" refers to a conserved domain or motif of amino acids having nuclease (e.g., endonuclease) activity. As used herein, a protein having split RuvC domains refers to a protein having two or more RuvC motifs at different sequential positions within the sequence that interact in tertiary structure to form a RuvC domain. A RuvC domain or segment thereof (e.g., RuvC I, RuvC II, or RuvC III) can generally be identified by alignment to known domain sequences, alignment to structures of proteins having annotated domains, or by comparison to known domain sequences.

[0106] As used herein, the term "HNH domain" generally refers to an endonuclease domain with characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to known domain sequences, structural alignment to proteins with annotated domains, or by comparison to known domain sequences.

[0107] As used herein, the term "BH domain," "bridge helix domain," or "bridge helix" generally refers to a rich arginine helix domain found in Cas enzymes and plays an important role in initiating cleavage activity after target DNA binding.

[0108] As used herein, the term "recognition domain" or "REC domain" generally refers to a domain believed to interact with the repeat:anti-repeat duplex of the crRNA and mediate formation of the Cas endonuclease / crRNA complex.

[0109] As used herein, the term "WED domain" or "wedge domain" generally refers to a fold comprising a twisted five-stranded beta sheet flanked by four alpha helices, which is generally responsible for recognition of the distorted repeat:anti-repeat duplex of the nuclease. The WED domain can be responsible for recognition of a single guide RNA scaffold.

[0110] As used herein, "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that is capable of catalyzing a hydrolytic deamination reaction that converts an adenine (or an adenine portion of a molecule) to a hypoxanthine (or a hypoxanthine portion of a molecule). In some embodiments, the functional domain is selected from an adenosine deaminase. In some embodiments, the adenosine deaminase comprises is capable of deaminating an adenine (A) to a hypoxanthine (I). In some embodiments, deamination of adenine to hypoxanthine converts an adenosine (A) or deoxyadenosine (dA) containing an adenine to a guanosine (G) or deoxyguanosine (dG).

[0111] As used herein, the term "cytosine deaminase" or "cytosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that is capable of catalyzing a hydrolytic deamination reaction that converts a cytosine (or a cytosine portion of a molecule) to a uracil (or a uracil portion of a molecule). In some embodiments, the cytosine-containing molecule is cytosine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule can be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0112] In some embodiments, the cytosine deaminase is selected from an APOBEC (e.g., APOBEC3, e.g., APOBEC3A, APOBEC3B, APOBEC3C).

[0113] Various embodiments are described below. It should be noted that specific embodiments are not intended as exhaustive descriptions or as limitations on the broader aspects discussed herein. An aspect described in connection with a particular embodiment is not necessarily limited to that embodiment and may be practiced in conjunction with any other embodiment. Throughout the specification, references to “one embodiment,” “implementation,” “some embodiments,” “implementation,” “example embodiment,” or “example embodiment” refer to specific features, structures, or characteristics described in connection with an embodiment that are included in at least one embodiment of the invention. Therefore, the phrases “in one embodiment,” “in one embodiment,” or “example embodiment” appearing throughout the specification do not necessarily all refer to the same embodiment, but may. Furthermore, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner, as will be apparent to those skilled in the art based on this disclosure. Moreover, although some embodiments described herein include some but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention. For example, in the appended claims, any claimed embodiment may be used in any combination.

[0114] Programmable nuclease

[0115] In some respects, this disclosure provides an engineered programmable nuclease.

[0116] In some technical solutions, the programmable nuclease disclosed herein includes an HNH domain and a RuvC domain.

[0117] In some embodiments, the programmable nuclease of this disclosure further includes a REC leaf having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 12, the REC leaf including a REC domain having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 21, and the nuclease having an amino acid sequence length between 600 and 800 amino acids, preferably between 600 and 700 amino acids, more preferably between 650 and 700 amino acids.

[0118] In some embodiments, the programmable nuclease of the present disclosure further comprises a WED domain having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 13; and a PI domain having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 14.

[0119] In some embodiments, the REC lobe further comprises a Linker structure located between the BH domain and the REC domain; preferably, the Linker structure has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 20.

[0120] In some embodiments, the RuvC domain of the programmable nuclease of the present disclosure comprises a RuvC-I, a RuvC-II, and a RuvC-III domain, wherein the RuvC-I domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 15; the RuvC-II domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 18; and the RuvC-III domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 19; and the HNH domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 16.

[0121] In some embodiments, the programmable nuclease of the present disclosure further comprises a BH domain, preferably, the BH domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 17.

[0122] In some embodiments, the programmable nuclease of the present disclosure has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 1.

[0123] In some aspects, the present disclosure provides a nucleic acid molecule encoding the programmable nuclease of the present disclosure.

[0124] In some embodiments, the nucleic acid molecule of the present disclosure comprises a nucleotide sequence that is at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity to SEQ ID NO. 2.

[0125] Fusion protein

[0126] In some aspects, the present disclosure provides a fusion protein comprising a programmable nuclease of the present disclosure and a functional domain fused thereto.

[0127] In some embodiments, the functional domain is selected from a nuclear localization signal (NLS), a nuclear export signal (NES), a reporter protein (e.g., a fluorescent protein), a nuclease targeting moiety, a DNA binding domain (e.g., Lex A DBD, Gal4 DBD, Sp1 DBD), an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, etc.), a transcriptional activation domain (e.g., VP64, VPR, p65, Rta), a transcriptional repression domain (e.g., KRAB domain, SID domain, NuE domain, NcoR domain, or SID4X domain), a nuclease, a deaminase (e.g., an adenosine deaminase or a cytidine deaminase), a methylase (e.g., a DNA methylase DNMT), a demethylase, a transcription release factor, an HDAC, a lytic activity polypeptide, a ligase, an integrase, a transposase, a recombinase, a polymerase, an exonuclease (e.g., T5E), and a base excision repair inhibitor (e.g., a uracil-DNA glycosylase inhibitor (UGI)).

[0128] In some embodiments, the functional domain comprises one or more of the following enzymatic activities on a target sequence: methylase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, desumoylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), and deglycosylation activity;

[0129] In some embodiments, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain;

[0130] In some embodiments, the adenosine deaminase catalytic domain or cytidine deaminase catalytic domain comprises one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

[0131] In some embodiments, the localization signal comprises a nuclear localization signal (NLS) and / or a nuclear export signal (NES);

[0132] In some embodiments, the sequence of the nuclear localization signal is selected from the group consisting of PKKKRKV, KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, PAAKRVKLD, RQRRNELKRSP, NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY, RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV, VSRKRPRP, PPKKARED, PQPKKKPL, SALIKKKKKMAP, DRLRR, PKQKKRK, RKLKKKIKKL, REKKKFLKRR, KRKGDEVDGVDEVAKKKSKK, RKCLQAGMNLEARKTKK.

[0133] In some embodiments, the sequence of the nuclear localization signal is located at, near, or proximal to the end (e.g., N- or C-terminus) of the programmable nuclease;

[0134] In some embodiments, the nuclear export signal comprises protein tyrosine kinase 2 (e.g., human protein tyrosine kinase 2);

[0135] In some embodiments, the reporter protein comprises glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, a self-fluorescent protein;

[0136] In some embodiments, the self-fluorescent protein comprises a green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, a cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyanl, etc.), a yellow fluorescent protein (e.g., (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), a blue fluorescent protein (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire);

[0137] In some embodiments, the DNA binding domain comprises a methylation binding protein, a LexA DBD, a Gal4 DBD;

[0138] In some embodiments, the epitope tag comprises a histidine tag, a V5 tag, a FLAG tag, an influenza virus hemagglutinin tag, a Myc tag, a VSV-G tag, a thioredoxin tag, a streptavidin tag;

[0139] In some embodiments, the transcription activation domain comprises VP64 and / or VPR;

[0140] In some embodiments, the transcription repression domain comprises KRAB and / or SID;

[0141] In some embodiments, the nuclease comprises Fokl;

[0142] In some embodiments, the cleavage active polypeptide comprises a polypeptide having single-stranded RNA cleavage activity, a polypeptide having double-stranded RNA cleavage activity, a polypeptide having single-stranded DNA cleavage activity, or a polypeptide having double-stranded DNA cleavage activity;

[0143] In some embodiments, the ligase comprises a DNA ligase and / or an RNA ligase;

[0144] In some embodiments, the exonuclease is selected from a TREX2 protein, a TREX1 protein, an APE1 protein, an Artemis protein, a CtIP protein, an Exol protein, an Mrel 1 protein, a RAD1 protein, a RAD9 protein, a Tp53 protein, a WRN protein, an Exonuclease V, a T5 exonuclease, or a T7 exonuclease, or a variant thereof;

[0145] In some embodiments, the functional domain is linked to the N-terminus and / or the C-terminus of the programmable nuclease;

[0146] In some embodiments, the functional domain is inserted between the N-terminus and the C-terminus of the programmable nuclease;

[0147] In some embodiments, the one or more functional domains are optionally linked to the N-terminus and / or the C-terminus of the programmable nuclease by a linker.

[0148] System

[0149] In some aspects, the present disclosure provides a nuclease system, comprising:

[0150] (1) a programmable nuclease of the present disclosure or a polynucleotide encoding the programmable nuclease; and

[0151] (2) a guide nucleic acid or a polynucleotide encoding the guide nucleic acid, the guide nucleic acid comprising:

[0152] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0153] (ii) a DNA-targeting sequence capable of hybridizing to a target sequence of a target DNA, thereby directing the complex to the target DNA;

[0154] optionally, wherein the protein-binding sequence is at the 3’ end of the DNA-targeting sequence; and / or

[0155] optionally, wherein the guide nucleic acid is a guide RNA (gRNA).

[0156] In some embodiments, the protein-binding sequence has a secondary structure that is substantially identical to the secondary structure of SEQ ID NO. 3;

[0157] In some embodiments, wherein the protein-binding sequence:

[0158] (1) comprises the polynucleotide sequence of SEQ ID NO. 3; or

[0159] (2) comprises a polynucleotide sequence having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity to SEQ ID NO. 3;

[0160] In some embodiments, the target sequence comprises about or at least about 16 contiguous nucleotides of the target DNA, for example, comprises about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more contiguous nucleotides of the target DNA, optionally, the target sequence comprises about 15-70 contiguous nucleotides of the target DNA; optionally, the target sequence comprises about 16 to about 22 contiguous nucleotides, or about 17 to about 22 contiguous nucleotides of the target DNA; optionally, wherein the target sequence comprises about 20 contiguous nucleotides of the target DNA.

[0161] In some embodiments, wherein the programmable nuclease comprises one or more nuclear localization sequences (NLS) adjacent to the N-terminus or C-terminus of the endonuclease, preferably, the nuclear localization sequence (NLS) is selected from the group consisting of PKKKRKV, KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, PAAKRVKLD, RQRRNELKRSP, NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY, RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV, VSRKRPRP, PPKKARED, PQPKKKPL, SALIKKKKKMAP, DRLRR, PKQKKRK, RKLKKKIKKL, REKKKFLKRR, KRKGDEVDGVDEVAKKKSKK, RKCLQAGMNLEARKTKK.

[0162] In some embodiments, the target DNA is a dsDNA or a ssDNA. In some embodiments, the target DNA is present in a bacterial cell, an archaeal cell, a unicellular eukaryotic organism, a plant cell, an invertebrate animal cell, or a vertebrate animal cell.

[0163] Guide nucleic acid

[0164] In some aspects, the disclosure provides a guide nucleic acid, comprising:

[0165] a) a DNA-targeting sequence comprising a nucleotide sequence that is complementary to a target sequence in a target DNA; and

[0166] b) a protein-binding sequence that binds to a programmable nuclease of the present disclosure to form a complex.

[0167] In some embodiments, the protein-binding sequence comprises a polynucleotide sequence having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity to SEQ ID NO. 3.

[0168] In some embodiments, the DNA-targeting sequence is about or at least about 16 contiguous nucleotides in length, e.g., about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more contiguous nucleotides, optionally, about 15-70 contiguous nucleotides; optionally, about 16 to about 22 contiguous nucleotides, or about 17 to about 22 contiguous nucleotides; optionally, about 20 contiguous nucleotides.

[0169] In some embodiments, the DNA-targeting sequence is about at least 70%, about 80%, about 90%, about 99%, or 100% complementary to the target DNA.

[0170] In some embodiments, the one or more polynucleotide sequences encoding the guide nucleic acid are operably linked to a promoter that functions in a eukaryotic cell. In some embodiments, the promoter is selected from a constitutive promoter, an inducible promoter, a ubiquitous promoter, a cell type-specific promoter, or a tissue-specific promoter.

[0171] Base editing

[0172] In some aspects, the disclosure provides a fusion protein comprising a deaminase or catalytic domain thereof, the fusion protein comprising a programmable nuclease of the disclosure, the programmable nuclease having an amino acid sequence with at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity to SEQ ID NO. 1, or comprising one or more amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO. 1, and the programmable nuclease has reduced or substantially lacking dsDNA cleavage activity relative to a parent nuclease (e.g., the programmable nuclease set forth in SEQ ID NO. 1). In some embodiments, the programmable nuclease is not completely endonuclease-deficient, but the endonuclease activity is not directed to the double strand of the dsDNA, but to one strand (sense or anti-sense; or target or non-target) of the dsDNA or ssDNA, which means that the programmable nuclease is essentially incapable of acting as a dsDNA endonuclease that cleaves the double strand of the target DNA or the non-target DNA, but is essentially capable of acting as a ssDNA endonuclease (or called nickase) that cleaves ssDNA or “creates a nick” in one strand of the dsDNA.

[0173] In some embodiments, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain;

[0174] In some embodiments, the adenosine deaminase catalytic domain or cytidine deaminase catalytic domain comprises one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

[0175] In some embodiments, the fusion protein further comprises a reverse transcriptase (RT) or catalytic domain thereof. In some embodiments, the guide nucleic acid further comprises or is used in combination with a reverse transcription donor RNA (RT donor RNA) comprising a primer binding site (PBS) and a template sequence.

[0176] Delivery

[0177] Various modes of delivery can be applied to the programmable nuclease of the disclosure, the guide nucleic acid, or the system of the disclosure.

[0178] In some aspects, the disclosure provides a delivery system comprising (1) a nuclease of the disclosure, a polynucleotide of the disclosure, or a system of the disclosure; and (2) a delivery vehicle.

[0179] In some embodiments, the delivery vehicle comprises a nanoparticle, a liposome, an exosome, a microvesicle, a gene gun, or an electroporation device.

[0180] Further, when the delivery subject is a plant cell, delivery can also be performed using a cell-penetrating peptide (CPP), for example. In one particular embodiment, the nuclease and / or at least one guide RNA is coupled to one or more CPPs, thereby effectively transporting the CPPs coupled to the nuclease and / or guide RNA into a plant cell (e.g., into a protoplast). CPPs are short peptides of less than 35 amino acids that are derived from a protein or from a chimeric sequence, and are capable of transporting biological molecules across a cell membrane in a non-receptor dependent manner. CPPs can be cationic peptides, peptides with hydrophobic sequences, amphipathic peptides, peptides with proline-rich and antimicrobial sequences, and chimeric or bipartite peptides. CPPs are capable of penetrating biological membranes and thus trigger the movement of different biological molecules across the cell membrane into the cytoplasm, improve their intracellular passage, and thus facilitate the interaction of the biological molecules with the target.

[0181] Exemplary CPPs include Tat (a transactivation protein required for viral replication by HIV type 1), penetratin, Kaposi fibroblast growth factor (FGF) signal peptide sequence, integrin beta 3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, a sweet arrow peptide, and the like.

[0182] In yet another aspect, the disclosure provides a vector comprising a nucleic acid molecule of the disclosure. In some embodiments, the vector encodes a guide nucleic acid as defined in the disclosure. In some embodiments, the vector is a plasmid vector, a recombinant AAV (rAAV) vector (vector genome), or a recombinant lentivirus vector.

[0183] In some aspects, the disclosure provides a recombinant AAV viral particle comprising a rAAV vector genome of the disclosure. A brief introduction to AAV for delivery can be found at “Guide to Adeno-Associated Virus (AAV)” (addgene.org / guides / aav / ).

[0184] In some aspects, the disclosure provides a ribonucleoprotein (RNP) comprising a programmable nuclease of the disclosure and optionally a guide nucleic acid as defined in the disclosure.

[0185] In some aspects, the present disclosure provides a lipid nanoparticle (LNP) comprising a programmable nuclease of the present disclosure or a system of the present disclosure.

[0186] In some embodiments, the lipid nanoparticle comprises a programmable nuclease of the present disclosure and a RNA (e.g., mRNA) guide nucleic acid of the present disclosure.

[0187] Methods of modification

[0188] In some aspects, the present disclosure provides a method of targeting, binding, cleaving, labeling, or modifying a target DNA, the method comprising contacting the target DNA with a complex comprising:

[0189] (1) a programmable nuclease of the present disclosure or a polynucleotide encoding the programmable nuclease; and

[0190] a guide nucleic acid or a polynucleotide encoding the guide nucleic acid, the guide nucleic acid comprising:

[0191] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0192] (ii) a DNA targeting sequence capable of hybridizing to a target sequence of a target DNA, thereby directing the complex to the target DNA;

[0193] Cells

[0194] In some aspects, the present disclosure provides a cell comprising

[0195] (1) a system as defined by the present disclosure;

[0196] (2) any nucleic acid molecule as defined by the present disclosure;

[0197] (3) a vector of the present disclosure;

[0198] (4) a ribonucleoprotein of the present disclosure; or

[0199] (5) a lipid nanoparticle of the present disclosure;

[0200] Pharmaceutical compositions

[0201] In some aspects, the present disclosure provides a pharmaceutical composition comprising:

[0202] (1) a system of the present disclosure, a vector of the present disclosure, a ribonucleoprotein of the present disclosure, or a lipid nanoparticle of the present disclosure, or a cell of the present disclosure; and (2) a pharmaceutically acceptable excipient.

[0203] In some embodiments, the dosage form of the composition is selected from the group consisting of a lyophilized formulation, a liquid formulation, or a combination thereof.

[0204] In some embodiments, the dosage form of the composition is a liquid formulation.

[0205] In some embodiments, the dosage form of the composition is an injectable dosage form.

[0206] In some embodiments, the composition is a cell preparation.

[0207] Suitable pharmaceutically acceptable excipients generally include inert substances that aid in the administration of a pharmaceutical composition to a subject, aid in the processing of a pharmaceutical composition into a deliverable formulation, or aid in the storage of a pharmaceutical composition prior to administration. Pharmaceutically acceptable excipients can include agents that can stabilize, optimize, or otherwise alter the form, consistency, viscosity, pH, pharmacokinetics, solubility of a formulation. Such agents include buffers, wetting agents, emulsifiers, diluents, encapsulants, and skin penetration enhancers. For example, the carrier can include, but is not limited to, saline, buffered saline, dextrose, arginine, sucrose, water, glycerol, ethanol, sorbitol, dextran, sodium carboxymethylcellulose, and combinations thereof.

[0208] Some non-limiting examples of substances that can be used as pharmaceutically acceptable carriers include: (1) sugars, such as lactose, dextrose, and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose and its derivatives, such as sodium carboxymethyl cellulose, methyl cellulose, ethyl cellulose, microcrystalline cellulose, and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants, such as magnesium stearate, sodium lauryl sulfate, and talc; (8) excipients, such as cocoa butter and suppository waxes; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; (10) glycols, such as propylene glycol; (11) polyols, such as glycerin, sorbitol, mannitol, and polyethylene glycol (PEG); (12) esters, such as ethyl oleate and ethyl laureate; (13) agar; (14) buffering agents, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer's solution; (19) ethyl alcohol; (20) pH buffered solutions; (21) polyesters, polycarbonates, and / or poly anhydrides; (22) fillers such as polypeptides and amino acids; (23) serum albumin such as ethanol; and (24) other non-toxic compatible substances employed in pharmaceutical formulations. Wetting agents, coloring agents, releasing agents, coating agents, sweetening, flavoring, perfuming agents, preservatives, and antioxidants can also be present in the formulations.

[0209] The pharmaceutical composition can include one or more pH buffering compounds to maintain the pH of the formulation at a predetermined level that reflects physiological pH, such as in the range of about 5.0 to about 8.0. The pH buffering compounds for aqueous liquid formulations can be an amino acid or mixture of amino acids, such as histidine or a mixture of amino acids such as histidine and glycine. Alternatively, the pH buffering compounds are preferably agents that maintain the pH of the formulation at a predetermined level, such as in the range of about 5.0 to about 8.0, and do not sequester calcium ions. Illustrative examples of such pH buffering compounds include, but are not limited to, imidazole and acetate ions. The pH buffering compounds can be present in any amount suitable to maintain the pH of the formulation at a predetermined level.

[0210] The pharmaceutical composition can also contain one or more tonicity adjusting agents, i.e., agents that adjust the tonicity (e.g., the tension, osmolarity, and / or osmotic pressure) of the formulation to a level that is acceptable to the blood stream and blood cells of the recipient individual. The tonicity adjusting agents can be agents that do not sequester calcium ions. The tonicity adjusting agents can be any compound known or available to those skilled in the art that adjusts the tonicity of the formulation. The suitability of a given tonicity adjusting agent in the formulations of the present application can be determined empirically by those skilled in the art. Illustrative examples of suitable tonicity adjusting agent types include, but are not limited to: salts, such as sodium chloride and sodium acetate; sugars, such as sucrose, dextrose, and mannitol; amino acids, such as glycine; and mixtures of one or more of these agents and / or dosage forms. The one or more tonicity adjusting agents can be present in any concentration sufficient to adjust the tonicity of the formulation.

[0211] Methods of treatment

[0212] In some aspects, the present disclosure provides a method for diagnosing, preventing, or treating a disease in a subject in need thereof, the method comprising administering to the subject a system of the present disclosure, a vector of the present disclosure, a ribonucleoprotein of the present disclosure, a lipid nanoparticle of the present disclosure, a cell of the present disclosure, or a pharmaceutical composition of the present disclosure, wherein the disease is associated with a target DNA, wherein a spacer sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnoses, prevents, or treats the disease.

[0213] In some embodiments, the disease is selected from hypercholesterolemia, familial hypercholesterolemia (FH), atherosclerosis, Angelman syndrome (AS), Alzheimer’s disease (AD), transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), alpha-1-antitrypsin deficiency, primary hyperoxaluria, Pompe disease, myotonic muscular dystrophy, Huntington’s disease (HTT), fragile X syndrome, Friedreich’s ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, thalassemia (e.g., beta thalassemia), Parkinson’s disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, nonalcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), heart disease (e.g., hypertrophic cardiomyopathy (HCM)), hepatitis B, and cancer.

[0214] Detection methods

[0215] In some aspects, the disclosure provides a method of detecting a target DNA, comprising contacting the target DNA with a system of the disclosure, wherein the target DNA is modified by the complex, and wherein the modification detects the target DNA. In some embodiments, the modification generates a detectable signal, e.g., a fluorescent signal.

[0216] Kits

[0217] In some aspects, the disclosure provides a kit comprising a programmable nuclease of the disclosure, a system of the disclosure, a polynucleotide of the disclosure, a vector of the disclosure, an RNP of the disclosure, an LNP of the disclosure, a delivery system of the disclosure, a cell of the disclosure, or a pharmaceutical composition of the disclosure, or any one, two, or all components thereof.

[0218] In some embodiments, the kit further comprises instructions for using one or more components contained therein, and / or instructions for combination with additional component(s) that can be obtained or necessary elsewhere.

[0219] In some embodiments, the kit further comprises one or more buffers that can be used to solubilize any of the one or more components contained therein, and / or to provide suitable reaction conditions for one or more of the one or more components. Such buffers can include one or more of the following: PBS, HEPES, Tris, MOPS, Na2CO3, NaHCO3, NaB, or combinations thereof. In some embodiments, the reaction conditions include an appropriate pH, such as a basic pH. In some embodiments, the pH is between 7-10.

[0220] In some embodiments, any one or more of the kit components can be stored in a suitable container or at a suitable temperature, for example, 4 degrees Celsius.

[0221] Further implementations are illustrated in the following examples, which are for illustrative purposes only and are not intended to limit the scope of the disclosure. Where RNA is indicated in the sequence listing, a “t” or “T” therein shall be taken as “U”.

[0222] The experimental methods used in the following examples are routine methods unless otherwise specified.

[0223] The materials, reagents, and the like used in the following examples are commercially available unless otherwise specified.

[0224] Examples

[0225] Example 1. Discovery of a new nuclease by metagenomics

[0226] A new nuclease (amino acid sequence as shown in SEQ ID NO. 1, nucleotide coding sequence as shown in SEQ ID NO. 2) was identified by analyzing the uncultured metagenome, which was named YTGE-1001. Using ClustalW to align with reference proteins (e.g., SaCas9, SpCas9), a phylogenetic tree was inferred using RAxML ( Figure 1 ), and its domain distribution was analyzed as shown in Figure 2 , and its protein structure was predicted by alphafold2 ( Figure 3 ). By comparing its representative domains with those of reference proteins, it was identified and analyzed, and it was found to have the following characteristics compared with reference proteins:

[0227] Table 1. Properties of YTGE-1001 described herein and its comparison with known nucleases

[0228]

[0229]

[0230] As the above analysis shows, although the nuclease YTGE-1001 contains several domains / structural parts similar to Cas9 (SaCas9, SpCas9) (RuvC domain (RuvC-Ⅰ, RuvC-Ⅱ, RuvC-Ⅲ domain sequences are shown in SEQ ID NO. 15, 18, 19, respectively), HNH domain (sequence shown in SEQ ID NO. 16), BH domain (sequence shown in SEQ ID NO. 17), REC leaf (sequence shown in SEQ ID NO. 12), WED domain (sequence shown in SEQ ID NO. 13), PI domain (sequence shown in SEQ ID NO. 14)), the composition, structure, and size of its multiple domains / structural parts are different. Among them, the REC leaf consists of consecutive tandem Linker structures (sequence shown in SEQ ID NO. 20) and the REC domain (sequence shown in SEQ ID NO. 14)... Composed of NO.21, with a total length of 135aa, the WED domain contains only 3 α-helices, and the PI domain contains only β-sheets and lacks α-helices. YTGE-1001 is considered to be a CRISPR nuclease that is different from known nucleases.

[0231] Example 2. Structural prediction of protein-binding sequence folding

[0232] Locus annotation of samples containing YTGE-1001 was performed using CRISPRCasFinder and CMsearch, obtaining its scaffold sequence (SEQ ID NO.3), and its secondary structure was predicted. Figure 4 ).

[0233] Example 3. Detection of the mediating activity of guide RNAs with different protein binding sequences

[0234] 1. To detect the cleavage activity of nucleases, expression plasmids capable of expressing nuclease YTGE-1001 and guide RNA, as well as target plasmids for detecting cleavage activity, were constructed.

[0235] DNA targeting sequences were designed based on the human PCSK9 target sequence (SEQ ID NO.4). Expression plasmids (SEQ ID NO.8, plasmid map shown) were constructed, containing the coding sequence of the nuclease YTGE-1001 (SEQ ID NO.1) regulated by the T7 promoter (SEQ ID NO.2) and the coding sequence of PCSK9-sgRNA (SEQ ID NO.5) regulated by the T7 promoter (SEQ ID NO.6). Figure 5 ).

[0236] The PCSK9-sgRNA consists of a PCSK9 targeting sequence (SEQ ID NO. 7) and a scaffold sequence in the 5'-3' direction.

[0237] The target plasmid for nuclease cleavage (SEQ ID NO. 9, see plasmid map in Figure 6 ) contains a Kana resistance gene, and a target sequence (SEQ ID NO. 4) for guide RNA targeting. Considering that there might be a target sequence recognition bias for programmable nucleases, which is affected by the PAM sequence next to the 3' end of the target sequence, a 6N random PAM sequence next to the 3' end of the target sequence was also set in the target cleavage plasmid, where N represents nucleotides A, T, C or G, i.e. giving 4^6 = 4096 PAMs, thus constituting a target cleavage plasmid library containing all possible PAMs.

[0238] 2. The expression plasmid (100 ng) and the target plasmid library (100 ng) were transformed into 100 μL of competent E. coli cells (purchased from Vazyme) by electroporation. After transformation, the E. coli was cultured at 37°C for 1 hour, then plated on medium coated with ampicillin (Amp) and kanamycin (Kan), ensuring that the number of colonies was greater than 100 times the number of libraries. After 14-16 hours of growth, all colonies were recovered and plasmids were extracted.

[0239] As a control, the target plasmid library 100 ng was electroporated into 100 ul of competent E. coli cells, and after transformation, the E. coli was cultured at 37°C for 1 hour, then plated on medium coated with kanamycin (Kan), and after 14-16 hours of growth, all colonies were recovered and plasmids were extracted.

[0240] When the expression plasmid and the target plasmid library are co-transformed into host E. coli, the expression plasmid expresses programmable nuclease YTGE-1001 and the corresponding guide RNA, which form a complex (RNP). The guide RNA targets the target sequence contained in the target plasmid, and guides the programmable nuclease YTGE-1001 to target and cleave the target plasmid. If YTGE-1001 has endonuclease activity and can recognize the PAM at the 3' end of the target sequence, it will cleave the target sequence, resulting in degradation of the target cleavage plasmid and failure to express the kanamycin resistance gene, thus preventing the host E. coli from growing on medium containing kanamycin.

[0241] 3. The fragment with the PAM sequence was amplified by PCR, and PE150 sequencing was performed using the primers PCSK9-PAM-SF (SEQ ID NO. 10) and PCSK9-PAM-SR (SEQ ID NO. 11).

[0242] The number of PAM sequences in the experimental group and the control group was counted respectively, and the PAM sequence number was standardized. For a PAM sequence, when log2 (control group standardized value / experimental group standardized value) > 3, it was considered that the PAM sequence was significantly consumed, and the PAM domain was predicted by the significantly consumed PAM sequence. The sequencing results are summarized as Figure 7 As shown in the figure, each point on the graph represents one of the 4096 PAMs. If the point representing a certain PAM falls on the right lower side of the curve, it means that for the PAM and its adjacent target sequence represented by the point, the endonuclease YTGE-1001 recognizes and causes statistically significant cleavage, resulting in degradation of the target cleavage plasmid, non-expression of the resistance gene, and further death of the experimental group of E. coli.

[0243] Figure 7 The statistical results show that the difference in PAM abundance is statistically significant, indicating that YTGE-1001 exhibits endonuclease activity; through analysis, the PAM sequence of the nuclease YTGE-1001 is found (see Figure 8 ), which can be used to cut the DNA target sequence with the PAM at the 3' end.

[0244] Sequence information:

[0245]

[0246]

[0247]

[0248]

[0249] Although preferred embodiments of the application have been shown and described herein, it should be apparent that these embodiments are provided by way of example only. The application is not intended to be limited by the particular examples disclosed in the specification. While the application has been described with reference to the above explanation in connection with the illustrated embodiments, the embodiments of the application described herein are not intended to be exhaustive or to be limited to the precise forms disclosed. Many modifications and variations of the present application can be possible in light of the above teachings, the preferred embodiments being chosen for illustration only.

[0250] All documents referred to in this disclosure are incorporated by reference herein in their entirety, as if each document were individually incorporated by reference. In addition, it should be understood that various modifications and changes can be made to the present application by those skilled in the art in light of the foregoing teachings without departing from the spirit or scope of the application.

Claims

1. A programmable nuclease comprising an HNH domain and a RuvC domain, characterized in that, the programmable nuclease further comprises a REC lobe having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 12, the REC lobe comprising a REC domain having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 21, the nuclease having an amino acid length between 600-800 amino acids, preferably between 600-750 amino acids, more preferably between 600-700 amino acids.

2. The programmable nuclease of claim 1, further comprising a WED domain having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO. 13; and a PI domain having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO.

14.

3. The programmable nuclease of claim 1 or 2, the REC lobe further comprising a Linker structure located between the BH domain and the REC domain; preferably the Linker structure has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO.

20.

4. The programmable nuclease of claim 1, having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity to SEQ ID NO.

1.

5. The programmable nuclease of claim 4, wherein the programmable nuclease further comprises a functional domain fused thereto. Preferably, the functional domain is selected from a nuclear localization signal (NLS), a nuclear export signal (NES), a reporter protein (e.g., a fluorescent protein), a nuclease targeting moiety, a DNA binding domain (e.g., Lex A DBD, Gal4 DBD, Sp1 DBD), an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, etc.), a transcriptional activation domain (e.g., VP64, VPR, p65, Rta), a transcriptional repression domain (e.g., KRAB domain, SID domain, NuE domain, NcoR domain, or SID4X domain), a nuclease, a deaminase (e.g., adenosine deaminase or cytidine deaminase), a methylase (e.g., DNA methylase DNMT), a demethylase, a transcriptional release factor, an HDAC, a lytic activity polypeptide, a ligase, an integrase, a transposase, a recombinase, a polymerase, an exonuclease (e.g., T5E), and a base excision repair inhibitor (e.g., uracil-DNA glycosylase inhibitor (UGI)); Preferably, the functional domain comprises one or more enzymatic activities on a target sequence: methylase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, de-ribosylation activity, myristoylation activity, de-myristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), and deglycosylation activity; Preferably, the functional domain is selected from an adenosine deaminase catalytic domain or a cytidine deaminase catalytic domain; Preferably, the adenosine deaminase catalytic domain or cytidine deaminase catalytic domain comprises one or more of ADAR1, ADAR2, APOBEC, AID, or TAD; Preferably, the localization signal comprises a nuclear localization signal (NLS) and / or a nuclear export signal (NES); Preferably, the sequence of the nuclear localization signal is located at, near, or proximal to the end (e.g., N- or C-terminus) of the programmable nuclease; Preferably, the nuclear export signal comprises protein tyrosine kinase 2 (e.g., human protein tyrosine kinase 2); Preferably, the reporter protein comprises glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, autofluorescent protein; Preferably, the autofluorescent protein includes green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyanl, etc.), yellow fluorescent protein (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), and blue fluorescent protein (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire); Preferably, the DNA binding domain includes methylation-binding proteins, LexADBD, and Gal4DBD; Preferably, the epitope tag includes histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, thioredoxin tag, and streptavidin tag; Preferably, the transcriptional activation domain includes VP64 and / or VPR; Preferably, the transcriptional repression domain includes KRAB and / or SID; Preferably, the nuclease comprises FokI; Preferably, the cleavage-active polypeptide includes a polypeptide with single-stranded RNA cleavage activity, a polypeptide with double-stranded RNA cleavage activity, a polypeptide with single-stranded DNA cleavage activity, or a polypeptide with double-stranded DNA cleavage activity. Preferably, the ligase comprises DNA ligase and / or RNA ligase; Preferably, the exonuclease is selected from TREX2 protein, TREX1 protein, APE1 protein, Artemis protein, CtIP protein, Exo1 protein, Mre11 protein, RAD1 protein, RAD9 protein, Tp53 protein, WRN protein, exonuclease V, T5 exonuclease or T7 exonuclease or variants thereof. Preferably, the functional structural domain is connected to the N-terminus and / or C-terminus of the programmable nuclease; Preferably, the functional structural domain is inserted between the N-terminus and C-terminus of the programmable nuclease; Preferably, the one or more functional domains are optionally connected to the N-terminus and / or C-terminus of the programmable nuclease via a connector.

6. A nuclease system comprising: (1) The programmable nuclease or the polynucleotide encoding the programmable nuclease according to any one of claims 1-5; and (2) A guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, wherein the guiding nucleic acid comprises: (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and (ii) capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the DNA target sequence of the target DNA; Optionally, the protein-binding sequence is located at the 3' end of the DNA targeting sequence; and / or Optionally, the guiding nucleic acid is a guiding RNA (gRNA).

7. The nuclease system of claim 6, the protein-binding sequence has a secondary structure that is substantially identical to the secondary structure of SEQ ID NO. 3; optionally, wherein the protein-binding sequence: (1) comprises the polynucleotide sequence of SEQ ID NO. 3; or (2) comprises a polynucleotide sequence having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity to SEQ ID NO.

3.

8. The system of claim 6 or 7, wherein the target sequence comprises about or at least about 16 contiguous nucleotides of the target DNA, e.g., about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more contiguous nucleotides, optionally, about 15-70 contiguous nucleotides; optionally, about 16 to about 22 contiguous nucleotides, or about 17 to about 22 contiguous nucleotides; optionally, about 20 contiguous nucleotides.

9. The system of claim 8, wherein the programmable nuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the programmable nuclease.

10. The system of any one of claims 6-9, wherein the target DNA is a dsDNA, optionally, the target DNA is present in a bacterial cell, an archaeal cell, a unicellular eukaryote, a plant cell, an invertebrate animal cell, or a vertebrate animal cell.

11. A nucleic acid molecule encoding the programmable nuclease of any one of claims 1-5.

12. A nucleic acid molecule comprising one or more polynucleotide sequences encoding a guide nucleic acid, wherein the guide nucleic acid comprises: (i) a DNA-targeting sequence comprising a nucleotide sequence that is complementary to a target sequence in a target DNA; and (ii) a protein-binding sequence that can bind to the programmable nuclease of any one of claims 1-5 to form a complex; Optionally, the one or more polynucleotide sequences encoding the guide nucleic acid are operably linked to a promoter that functions in a eukaryotic cell.

13. The nucleic acid molecule of claim 12, the promoter is selected from a constitutive promoter, an inducible promoter, a ubiquitous promoter, a cell type-specific promoter, or a tissue-specific promoter.

14. The nucleic acid molecule of claim 12 or 13, wherein the target DNA is a dsDNA, optionally, the target DNA is present in a bacterial cell, an archaeal cell, a unicellular eukaryote, a plant cell, an invertebrate animal cell, or a vertebrate animal cell.

15. The nucleic acid molecule of claim 14, wherein the DNA-targeting sequence is about or at least about 16 contiguous nucleotides in length, for example, comprises about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more contiguous nucleotides of the target DNA, optionally, the target sequence comprises about 15-70 contiguous nucleotides of the target DNA; optionally, the target sequence comprises about 16 to about 22 contiguous nucleotides, or about 17 to about 22 contiguous nucleotides of the target DNA; optionally, wherein the target sequence comprises about 20 contiguous nucleotides of the target DNA.

16. The nucleic acid molecule of claim 15, the DNA-targeting sequence is about at least 70%, about 80%, about 90%, about 99%, or 100% complementary to the target DNA.

17. A vector comprising the nucleic acid molecule of claim 11; optionally, wherein the vector encodes a guide nucleic acid as defined in any one of claims 12-15; optionally, wherein the vector is a plasmid vector, a recombinant AAV (rAAV) vector, or a recombinant lentivirus vector.

18. A ribonucleoprotein comprising the programmable nuclease of any one of claims 1-5 and a guide nucleic acid as defined in any one of claims 12-15.

19. A lipid nanoparticle comprising the programmable nuclease of any one of claims 1-5 or the system of any one of claims 6-10.

20. A cell comprising (1) the system of any one of claims 6-10; (2) the nucleic acid molecule of claim 11, and the nucleic acid molecules of claims 12-16; (3) the vector of claim 17; (4) the ribonucleoprotein of claim 18; or (5) the lipid nanoparticle of claim 19.

21. A pharmaceutical composition comprising (1) the system of any one of claims 6-10, the vector of claim 17, the ribonucleoprotein of claim 18, or the lipid nanoparticle of claim 19, or the cell of claim 20; and (2) a pharmaceutically acceptable excipient.

22. A method of targeting, binding, cleaving, labeling, or modifying a target DNA, the method comprising contacting the target DNA with a complex, the complex comprising: (1) the programmable nuclease of any one of claims 1-5 or a polynucleotide encoding the programmable nuclease; and and a guide nucleic acid or a polynucleotide encoding the guide nucleic acid, the guide nucleic acid comprising: (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and (ii) a DNA-targeting sequence capable of hybridizing to a target sequence of a target DNA, thereby directing the complex to the target DNA.

23. A method for diagnosing, preventing, or treating a disease in a subject in need thereof, the method comprising administering to the subject the system of any one of claims 6-10, the vector of claim 17, the ribonucleoprotein of claim 18, the lipid nanoparticle of claim 19, the cell of claim 20, or the pharmaceutical composition of claim 21, wherein the disease is associated with a target DNA, wherein a spacer sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnoses, prevents, or treats the disease.

24. The method of claim 23, wherein the disease is selected from hypercholesterolemia, familial hypercholesterolemia (FH), atherosclerosis, Angelman syndrome (AS), Alzheimer’s disease (AD), transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), alpha-1-antitrypsin deficiency, primary hyperoxaluria, Pompe disease, myotonic muscular dystrophy, Huntington’s disease (HTT), fragile X syndrome, Friedreich’s ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, thalassemia (e.g., beta thalassemia), Parkinson’s disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, nonalcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), heart disease (e.g., hypertrophic cardiomyopathy (HCM)), hepatitis B, and cancer.

25. A method of detecting a target DNA, the method comprising contacting the target DNA with the system of any one of claims 6-10, wherein the target DNA is modified by the complex, and wherein the modification detects the target DNA; optionally, wherein the modification generates a detectable signal, e.g., a fluorescent signal.