Novel programmable nuclease and application thereof

By developing a novel programmable nuclease containing HNH, RuvC, REC leaf and BH domains, the diverse needs of the CRISPR-Cas system in DNA targeting and editing were addressed, achieving highly efficient DNA targeting and editing effects.

CN121737085APending Publication Date: 2026-03-27SHANGHAI JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411353456.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems have their own advantages and disadvantages in DNA targeting and editing, making it difficult to meet diverse needs.

Method used

A novel programmable nuclease is provided, comprising an HNH domain and a RuvC domain, combined with a REC leaflet and a BH domain, having an amino acid sequence length between 600 and 800 amino acids, and performing targeting and editing by guiding the hybridization of nucleic acid with target DNA.

Benefits of technology

It achieves efficient DNA targeting and editing, and features small molecular size and easy delivery, making it suitable for a variety of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005062935190000221
    Figure BDA0005062935190000221
  • Figure BDA0005062935190000231
    Figure BDA0005062935190000231
  • Figure BDA0005062935190000251
    Figure BDA0005062935190000251
Patent Text Reader

Abstract

The invention provides a novel nuclease and a method for gene editing and nucleic acid detection by using the nuclease. In some embodiments, the disclosure provides uses of the programmable nuclease. The disclosure also provides compositions, systems, and methods comprising the programmable nuclease.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure provides guide nucleic acids comprising target DNA containing a target sequence, as well as programmable nucleases and their applications. This disclosure further provides methods for targeting, binding, cleaving, labeling, or modifying target DNA. Background Technology

[0002] The CRISPR-Cas system guides nucleic acids to recognize specific DNA targets, and through nucleases, these targets can be specifically edited. Based on this, CRISPR-Cas systems (including Cas9, Cas12, and Cas13 systems) have been developed as genome editing tools. However, each nuclease system still has its own advantages and disadvantages. Therefore, it is still necessary to develop new nucleases and corresponding CRISPR-Cas systems to meet diverse needs.

[0003] Any reference or designation of any document in this disclosure is not an admission that such document is available as prior art. Every reference mentioned or cited in this disclosure is incorporated herein by way of its entirety. Summary of the Invention

[0004] This invention addresses the above-mentioned needs by providing novel programmable nucleases and CRISPR-Cas systems.

[0005] In some aspects, this disclosure provides a programmable nuclease comprising an HNH domain and a RuvC domain, the programmable nuclease further comprising a BH domain and a REC leaflet, wherein the BH domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 18; the REC leaflet comprises a REC domain, the REC domain comprising two parts, REC1 and REC2, wherein the REC1 part has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 13, and the REC2 part has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 13, and the REC2 part has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 18. NO.14 has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity; the amino acid sequence length of the nuclease is between 600 and 800 amino acids, preferably between 600 and 750 amino acids.

[0006] In some aspects, this disclosure provides a nuclease system comprising:

[0007] (1) The programmable nuclease of this disclosure or the polynucleotide encoding said programmable nuclease; and

[0008] (2) A guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, wherein the guiding nucleic acid comprises:

[0009] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0010] (ii) capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the DNA target sequence of the target DNA;

[0011] Optionally, the protein-binding sequence is located at the 3' end of the DNA targeting sequence; and / or

[0012] Optionally, the guiding nucleic acid is a guiding RNA (gRNA).

[0013] In some aspects, this disclosure provides a nucleic acid molecule that encodes a programmable nuclease of this disclosure.

[0014] In some aspects, this disclosure provides a nucleic acid molecule comprising one or more polynucleotide sequences encoding a guiding nucleic acid, wherein the guiding nucleic acid comprises:

[0015] (i) a DNA targeting sequence comprising a nucleotide sequence complementary to a target sequence in the target DNA; and

[0016] (ii) A protein-binding sequence that can bind to a programmable nuclease disclosed herein to form a complex;

[0017] Optionally, the one or more polynucleotide sequences encoding the guiding nucleic acid are operatively linked to a promoter that functions in eukaryotic cells.

[0018] In some aspects, this disclosure provides a vector comprising the nucleic acid molecule of this disclosure; optionally, wherein the vector encodes the guide nucleic acid of this disclosure; optionally, wherein the vector is a plasmid vector, a recombinant AAV (rAAV) vector, or a recombinant lentiviral vector.

[0019] In some aspects, this disclosure provides a ribonucleoprotein comprising a programmable nuclease of this disclosure and a guide nucleic acid of this disclosure.

[0020] In some aspects, this disclosure provides a programmable nuclease composition comprising:

[0021] (1) The programmable nuclease or polynucleotide encoding the programmable nuclease described in this disclosure; and

[0022] (2) A guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, wherein the guiding nucleic acid comprises:

[0023] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0024] (ii) capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the DNA target sequence of the target DNA;

[0025] Optionally, the protein-binding sequence is located at the 3' end of the DNA targeting sequence; and / or

[0026] Optionally, the guiding nucleic acid is a guiding RNA (gRNA).

[0027] In some aspects, this disclosure also provides a delivery system comprising the programmable nuclease, the nucleic acid molecule, or the programmable nuclease composition described herein.

[0028] In one embodiment, the delivery system further includes a delivery medium, which includes nanoparticles, liposomes, exosomes, microbubbles, gene guns, or electroporation devices.

[0029] In some aspects, this disclosure provides a lipid nanoparticle comprising a programmable nuclease or a system of this disclosure.

[0030] In some aspects, this disclosure provides a cell comprising

[0031] (1) The system disclosed herein;

[0032] (2) The nucleic acid molecules disclosed herein;

[0033] (3) The medium of this disclosure;

[0034] (4) The ribonucleoprotein of this disclosure; or

[0035] (5) The lipid nanoparticles disclosed herein.

[0036] In some aspects, this disclosure provides a pharmaceutical composition comprising (1) the system of this disclosure, the carrier of this disclosure, the ribonucleoprotein of this disclosure, the lipid nanoparticle of this disclosure, or the cell of this disclosure; and (2) a pharmaceutically acceptable excipient.

[0037] In some aspects, this disclosure provides a method for targeting, binding, cleaving, labeling, or modifying target DNA, the method comprising contacting the target DNA with a complex, the complex comprising:

[0038] (1) The programmable nuclease of this disclosure or the polynucleotide encoding said programmable nuclease; and

[0039] The guiding nucleic acid or the polynucleotide encoding the guiding nucleic acid, the guiding nucleic acid comprising:

[0040] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0041] (ii) capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the DNA target sequence of the target DNA;

[0042] In some aspects, this disclosure provides a method for diagnosing, preventing, or treating a disease in a subject of need, the method comprising administering to the subject a system of this disclosure, a carrier of this disclosure, a ribonucleoprotein of this disclosure, a lipid nanoparticle of this disclosure, a cell of this disclosure, or a pharmaceutical composition of this disclosure, wherein the disease is associated with a target DNA, wherein a spacer sequence is capable of hybridizing with a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnoses, prevents, or treats the disease.

[0043] In some aspects, this disclosure provides a method for detecting target DNA, the method comprising contacting the target DNA with a system of the present disclosure, wherein the target DNA is modified by the complex, and wherein the modification detects the target DNA; optionally, wherein the modification generates a detectable signal, such as a fluorescence signal.

[0044] In some aspects, this disclosure provides a kit comprising the aforementioned programmable nuclease, the aforementioned nucleic acid molecule, the aforementioned composition, the aforementioned system, the aforementioned carrier, the aforementioned ribonucleoprotein, the aforementioned lipid nanoparticles, the aforementioned cells, or the aforementioned pharmaceutical composition.

[0045] In a preferred embodiment, the components of the kit are in the same or different containers.

[0046] In some respects, this disclosure also provides a container that contains the aforementioned reagent kit.

[0047] In one embodiment, the container includes a sterile container;

[0048] In one embodiment, the container includes a syringe.

[0049] According to WIPO standard ST.26, the symbol "t" is used to represent T in DNA and U in RNA. Therefore, in this sequence listing prepared according to ST.26, when the sequence is RNA, T in the sequence should be considered as U. Attached Figure Description

[0050] Figure 1The phylogenetic tree of YTGE-1006 is described, showing that the YTGE-1006 divergent clade is far from the Cas9 system used as a reference.

[0051] Figure 2 An exemplary domain distribution of YTGE-1006 is described. As can be seen from the figure, YTGE-1006 is smaller and has a different domain distribution compared to Cas9 (SaCas9 or SpCas9).

[0052] Figure 3 The predicted 3D structure of YTGE-1006 is described.

[0053] Figure 4 The secondary structure prediction of the scaffold sequence corresponding to YTGE-1006 is described.

[0054] Figure 5 The expression plasmid maps of YTGE-1006 and PCSK9-sgRNA are described.

[0055] Figure 6 A map depicting the target plasmid as the target object is described.

[0056] Figure 7 The endonuclease activity of YTGE-1006 is described, indicated by the dots falling outside the dashed line on the right.

[0057] Figure 8 The PAM analysis results corresponding to YTGE-1006 are described. The analysis shows that its corresponding PAM is 5'-NGG. Detailed Implementation

[0058] Details of one or more embodiments of this disclosure are set forth in the following description. Other features or advantages of this disclosure should be apparent from the following drawings and detailed description of several embodiments, and from the claims. It should be understood that, unless otherwise indicated, any aspect or embodiment of this disclosure may be combined with any other aspect or embodiment of this disclosure to constitute another embodiment explicitly or implicitly disclosed herein.

[0059] Overview

[0060] This disclosure provides a novel programmable nuclease that, compared to known nucleases, has low similarity to other nucleases, possesses DNA (e.g., ssDNA or dsDNA) cleavage activity, and is small in size and easy to deliver.

[0061] The programmable nuclease disclosed herein has certain functional similarities to existing nucleases (such as Cas9), but it does not belong to the Cas9 protein. It contains multiple functional domains such as HNH, RuvC, REC, BH, WED, and PI. The distribution and size of its domains are different from those of existing nucleases, indicating that it is a novel nuclease.

[0062] The guide nucleic acid for the nucleases of this disclosure can be RNA, referred to as guide RNA (gRNA) or crRNA. The guide nucleic acid includes a scaffold sequence capable of interacting with the nucleases of this disclosure to form a complex (protein-RNA complex) and a spacer sequence capable of hybridizing with a target sequence in the target DNA, thereby guiding the complex to the target DNA.

[0063] the term

[0064] This disclosure will be described with respect to specific embodiments, but is not limited thereto by the claims. Unless otherwise stated, the terms set forth herein should generally be understood in their common meaning.

[0065] As used herein, unless the context clearly indicates otherwise, the singular forms “a / an” and “the” include both a single referent and plural referents.

[0066] As used herein, the term “about” is used to mean approximately, roughly, about, or around, and can refer to a value or composition within an acceptable margin of error for a particular value or composition as determined by one of ordinary skill in the art, depending in part on how the value or composition is measured or determined. For example, as used herein, the expression “about 100” includes all values ​​between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).

[0067] As used herein, the terms “containing” or “including (comprise)” can be open-ended, semi-closed, or closed. In other words, the terms also include “consistently made of” or “composed of”.

[0068] As used herein, the term “optionally” means that the event, situation or substituent described below may or may not occur, and the description includes examples of the event or situation occurring as well as examples of the event or situation not occurring.

[0069] As used in this text, the term "substantially / essentially" means approximately 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or higher in degree, quantity, level, value, quantity, frequency, percentage, scale, size, quantity, weight, or length compared to a reference sequence. For example, as used herein, "substantially identical sequence" can refer to a sequence, polynucleotide, or polypeptide that has a certain degree of identity with a reference sequence.

[0070] As used herein, the term “and / or” refers to any one or both of the alternatives.

[0071] As used herein, the terms “nucleic acid,” “nucleic acid molecule,” and “polynucleotide” are used interchangeably to refer to a polymeric form of nucleotides of any length, deoxyribonucleotides or ribonucleotides or their analogues, in single-stranded, double-stranded, or multi-stranded form. Polynucleotides can be exogenous or endogenous to cells. Polynucleotides can exist in cell-free environments. Polynucleotides can be genes or segments thereof. Polynucleotides can be DNA or RNA. Polynucleotides can include one or more analogues (e.g., altered backbone, sugars, or nucleobases). If present, modifications to the nucleotide structure can be conferred before or after polymer assembly. Some non-limiting examples of analogues include: 5-bromouracil, peptide nucleic acids, heteronucleotides, morpholino, locked nucleic acids, glycerol nucleic acids, threonine nucleic acids, dideoxynucleotides, cordycepin, 7-denitro-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogues, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, brassinoside, and woyoside. Non-restricted examples of polynucleotides include coding or non-coding regions of genes or gene fragments, multiple loci (one locus) as defined by ligation analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated DNA of any sequence, cell-free polynucleotides containing cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. Nucleotide sequences may contain interspersed non-nucleotide components. Polynucleotides may include mixtures of naturally occurring nucleotides and nucleotide analogs (e.g., synthetic nucleotide analogs).

[0072] As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably and generally refer to a polymer of at least two amino acid residues linked by peptide bonds. This term does not indicate a specific length of the polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or is naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. In some embodiments, the polymer may contain intercalated non-amino acid components. The term encompasses amino acid chains of any length, including full-length proteins and proteins with or without secondary and / or tertiary structures (e.g., domains). The term also covers modified amino acid polymers; for example, through disulfide bond formation, glycosylation, esterification, acetylation, phosphorylation, oxidation, and any other operations, such as conjugation with a labeled component.

[0073] As used herein, the term "identity" refers to the overall correlation between polymer molecules, such as nucleic acid molecules (e.g., DNA and / or RNA molecules) and / or polypeptide molecules. Sequence identity (or homology) is determined by comparing two aligned sequences along a predetermined comparison window (which may be 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the length of a reference nucleotide sequence or protein) and determining the number of positions where the same residues occur. Typically, this is expressed as a percentage. For example, in the case of polypeptide sequences, if the mutation type is one or more of the following: substitution / replacement of one or more amino acids / nucleotides, insertion within the sequence, and deletion within the sequence, the total number of residues is calculated based on the larger of the molecules being compared. If the mutation type also includes insertions (extensions) at either end or both ends of the sequence or deletions (truncations) at either end or both ends of the sequence, the number of amino acids inserted or deleted at either end or both ends (e.g., fewer than 20 insertions or deletions at both ends) is not included in the total number of residues. When calculating the percentage of identity, the sequences being compared are aligned in a way that produces the maximum match between the sequences, and gaps in the alignment are resolved (if any) using a specific algorithm. The calculation of nucleotide identity is similar.

[0074] As used herein, the term "ortholog" refers to genes (and the proteins encoded by those genes) that are presumed to be descendants of the same ancestral sequence from which speciation occurred: when a species diverges into two separate species, copies of a single gene in the resulting two species are said to be orthologous. Orthologs, or orthologous genes, are genes in different species that have originated vertically from a single gene of a last common ancestor.

[0075] As used herein, the terms "amino acid" and "amino acids" generally refer to natural and non-natural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids can include natural and non-natural amino acids that have been chemically modified to include groups or chemical moieties not naturally present on the amino acid. Amino acid analogs can refer to amino acid derivatives. The term "amino acid" includes D-amino acids and L-amino acids.

[0076] As used herein, the term "nuclease" refers to a polypeptide that can cleave phosphodiester bonds between nucleotide subunits of nucleic acids; the term "endonuclease" refers to a polypeptide that can catalyze (e.g., cleave) phosphodiester bonds within a polynucleotide chain (e.g., DNA or RNA).

[0077] As used herein, the term "exonuclease" refers to a protein or polypeptide that can digest nucleic acids (e.g., RNA or DNA) from their free ends.

[0078] As used herein, the term "non-natural" can generally refer to a nucleic acid or polypeptide sequence not found in native nucleic acids or proteins. Non-natural can refer to an affinity tag. Non-natural can refer to a fusion. Non-natural can refer to a naturally occurring nucleic acid or polypeptide sequence, including mutations, insertions, and / or deletions. Non-natural sequences can exhibit and / or encode activities that can also be exhibited by nucleic acid and / or polypeptide sequences fused with non-natural sequences (e.g., enzyme activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.). Non-natural nucleic acid or polypeptide sequences can be genetically engineered to link with naturally occurring nucleic acid or polypeptide sequences (or variants thereof) to produce chimeric nucleic acids and / or polypeptide sequences encoding chimeric nucleic acids and / or polypeptides.

[0079] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains derived from at least two different proteins. A protein may be located at the N-terminal (N-terminal) portion of the fusion protein or at the C-terminal (C-terminal) portion of the fusion protein, thus forming an N-terminal fusion protein or a C-terminal fusion protein, respectively. Proteins may comprise different domains, such as nucleic acid-binding domains (e.g., the gRNA-binding domain of a nuclease, which guides the protein to a target site for binding) and nucleic acid-cleaving domains, or catalytic domains of nucleic acid-editing proteins. In some embodiments, a linker may be present between the proteins. In some embodiments, the protein comprises a protein moiety (e.g., an amino acid sequence that forms the nucleic acid-binding domain) and an organic compound (e.g., a compound that can act as a nucleic acid cleavage agent). In some embodiments, the protein complexes or associates with a nucleic acid (e.g., RNA or DNA).

[0080] As used herein, the term "connector" refers to any means, entity, or portion used to join two or more entities. In some embodiments, a connector is a covalent connector. In some embodiments, a connector is a non-covalent connector. Examples of covalent connectors include connector portions covalently bonded or covalently attached to one or more proteins or domains to be joined. In some embodiments, connectors are non-covalent, such as organometallic bonds through a metal center (such as a platinum atom). Joining can be permanent or reversible. For covalent connections, various functional groups can be used, such as amide groups, including carbonate derivatives, ethers, esters (including organic and inorganic esters), amino groups, carbamates, ureas, etc. To provide a connection, domains can be modified by oxidation, hydroxylation, substitution, reduction, etc., to provide coupling sites. Conjugation methods are well known to those skilled in the art and are covered herein for use. Connector portions include, but are not limited to, chemical connector portions, or, for example, peptide connector portions (connector sequences). The length and type of connector can be designed as needed. In some embodiments, the connector can be selected from synthetically produced amino acid sequences or naturally occurring polypeptide sequences. It should be understood that modifications that do not significantly reduce the function of RNA-binding domains and effector domains are preferred.

[0081] As used herein, the term "complementarity" refers to the ability of a first polynucleotide sequence (e.g., a guide sequence) to pair with a second polynucleotide sequence (such as a target sequence) via conventional Watson-Crick base pairing. In some embodiments, the first polynucleotide may be substantially complementary to the second polynucleotide, i.e., having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementarity with the second polynucleotide. In some embodiments, the first polynucleotide is completely complementary to the second polynucleotide, i.e., having 100% complementarity with the second polynucleotide. "Substantially complementary" refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% on a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or 100% nucleotides, or it refers to two nucleic acids hybridizing under stringent conditions. The term "stringent conditions" associated with hybridization refers to conditions under which a nucleic acid complementary to a target sequence hybridizes primarily with that target sequence and substantially does not hybridize to non-target sequences. Stringent conditions are typically sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence. "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonds between the bases of these nucleotide residues. The complex can contain two strands forming a double helix, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can constitute a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.

[0082] As used herein, the term "domain" or "protein domain" refers to a portion of a protein sequence that can exist and function independently of the rest of the protein chain.

[0083] As used herein, the term "operably linked" refers to a functional connection between a regulatory sequence and a target nucleic acid sequence that results in the expression of the latter (e.g., in an in vitro transcription / translation system or in a host cell when a vector is introduced into the host cell). For example, the first nucleic acid sequence is operably linked to the second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For example, the promoter is operably linked to the coding sequence if it affects the transcription or expression of the coding sequence.

[0084] As used herein, the term "conservative amino acid substitution" refers to the substitution of a chemically or functionally similar amino acid without affecting the normal function of the protein. Examples include interchangeability in proteins with amino acid residues having similar side chains: a group of amino acids with aliphatic side chains consisting of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids with aliphatic-hydroxy side chains consisting of serine and threonine; a group of amino acids with amide-containing side chains consisting of asparagine and glutamine; a group of amino acids with aromatic side chains consisting of phenylalanine, tyrosine, and tryptophan; a group of amino acids with basic side chains consisting of lysine, arginine, and histidine; and a group of amino acids with sulfur-containing side chains consisting of cysteine ​​and methionine. In some embodiments, considered conserved amino acid substitutions include: aspartic acid-glutamic acid, lysine-arginine-histidine, serine-threonine-asparagine-glutamine, glycine-alanine-valine-leucine-isoleucine, cysteine-methionine-proline, and phenylalanine-tyrosine-tryptophan. In some embodiments, the amino acid substitutions considered to be conserved include: valine-leucine-methionine-isoleucine and alanine-serine-threonine.

[0085] As used herein, the terms “regularly clustered short palindromic repeats (CRISPR)-CRISPR-related (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which typically includes transcripts or other elements relating to the expression of CRISPR-related (“Cas”) genes, or transcripts or other elements capable of directing the activity of said Cas genes.

[0086] As used herein, the term "complex" refers to a combination of two or more molecules. In some embodiments, a complex includes polypeptide and nucleic acid molecules that interact with each other (e.g., bind, contact, adhere). For example, the term "complex" may refer to a combination of guide RNA and a polypeptide (e.g., a nuclease). Alternatively, the term "complex" may refer to a combination of complementary regions of guide RNA, a nuclease, and a target nucleic acid.

[0087] As used herein, the term "cut" refers to the breakage of the covalent backbone of a DNA molecule. Cutting can be performed by a variety of methods, including but not limited to enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand and double-strand cutting are possible, and double-strand cutting can occur as a result of two distinct single-strand cutting events. DNA cutting can result in blunt or sticky ends.

[0088] The term "cleave" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond within the backbone of a target nucleotide sequence, resulting in a single-stranded or double-stranded break within that target sequence. For example, the nucleases or variant peptides of this disclosure can act as endonucleases or exonucleases (removing consecutive nucleotides from the ends (5' and / or 3') of a polynucleotide) to cleave nucleotides within a polynucleotide. In some embodiments, the cleavage of the target polynucleotide by the nucleases or variant peptides of this disclosure can result in staggered breaks or blunt ends.

[0089] As used herein, the term “protospacer adjacent motif” or “PAM” refers to a short sequence (or motif) adjacent to the protospacer sequence on the non-target strand of a target nucleic acid recognized by the CRISPR complex. The target nucleic acid is double-stranded DNA (dsDNA), one strand containing the target sequence adjacent to the PAM and referred to as the “PAM strand” (e.g., the non-target strand or non-spacer complementary strand), while the other complementary strand is referred to as the “non-PAM strand” (e.g., the target strand or spacer complementary strand). As used herein, the term “adjacent” includes cases where the RNA guide of the complex specifically binds to, interacts with, or associates with the target sequence immediately adjacent to the PAM. In such cases, there are no nucleotides between the target sequence and the PAM. The term “adjacent” also includes cases where there are a few (e.g., 1, 2, 3, 4, or 5) nucleotides between the target sequence that binds to the target portion and the PAM.

[0090] As used herein, the terms “guide nucleic acid,” “RNA guide,” “RNA guide sequence,” “guide RNA (gRNA),” and “single guide RNA (sgRNA)” are used interchangeably to refer to a nucleic acid-based molecule capable of forming a complex with a CRISPR-nuclease (e.g., the programmable nuclease of this disclosure) and comprising a sequence (e.g., a guide sequence) that is sufficiently complementary to a target nucleic acid to hybridize with the target nucleic acid and guide the complex to the target nucleic acid. The nucleic acid includes, but is not limited to, RNA-based molecules, such as guide RNA. The guide nucleic acid may comprise a segment that may be referred to as a “nucleic acid targeting segment” or “nucleic acid targeting sequence,” and the nucleic acid targeting segment may comprise a sub-segment that may be referred to as a “protein-binding sequence” or “protein-binding sequence” or “nuclease-binding segment.” The guide nucleic acid may be a DNA molecule, an RNA molecule, or a DNA / RNA mixture molecule. A “DNA / RNA mixture molecule” refers to a nucleic acid comprising one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, whether contiguous or not. However, "DNA molecule" or "RNA molecule" can also refer to a DNA molecule containing one or more modified or unmodified ribonucleotides, whether continuous or not, or an RNA molecule containing one or more modified or unmodified deoxyribonucleotides, whether continuous or not.

[0091] As used herein, the term "activity" refers to biological activity. In some embodiments, nuclease activity includes the enzymatic activity of the nuclease, such as catalytic ability. For example, nuclease activity may include nuclease activity itself. In some embodiments, nuclease activity includes binding activity, such as the binding activity of the nuclease to RNA guides and / or target nucleic acids.

[0092] In this article, the term "protein-binding sequence" is used interchangeably with "scaffold sequence" and "tracrRNA," referring to RNA sequences that include the structures required for CRISPR-related proteins to bind to specific target nucleic acids.

[0093] As used herein, the terms “upstream” and “downstream” refer to relative positions within a single nucleic acid (e.g., DNA) sequence in a nucleic acid molecule. “Upstream” and “downstream” respectively relate to the 5’ to 3’ direction in which RNA transcription occurs. When the 3’ end of the first sequence precedes the 5’ end of the second sequence, the first sequence is upstream of the second sequence. When the 5’ end of the first sequence follows the 3’ end of the second sequence, the first sequence is downstream of the second sequence.

[0094] As used herein, “regulatory elements” include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals, poly-U sequences), for which detailed description can be found in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some embodiments, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters may primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In other embodiments, regulatory elements may also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell type-specific. As used herein, the term "promoter" refers to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operatively linked to a polynucleotide encoding or defining a gene product, will result in the production of that gene product in the cell under most or all physiological conditions. An inducible promoter is a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of endogenous or exogenous stimuli, such as by a chemical compound (chemical inducer), or in response to environmental, hormone, chemical, and / or developmental signals. Inducible or regulatory promoters include promoters that are induced or regulated, for example, by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounds, or chemicals (such as ethanol, abscisic acid (ABA), jasmonic acid esters, salicylic acid, or safeners).

[0095] As used herein, the term “target” refers to a predetermined or intended region of DNA that is bound, cleaved, and / or edited by the nuclease or a variant polypeptide of the present disclosure.

[0096] As used herein, the term "off-target" refers to a non-predetermined or unintended region of DNA that is bound, cleaved, and / or edited by, for example, a nuclease or a variant peptide of the present disclosure. In some embodiments, a DNA region is an off-target region when it differs from a DNA region intended or expected to be bound, cleaved, and / or edited by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. In some embodiments, such sites are detected using targeted sequencing of off-target sites predicted via computer simulation or by other methods known in the art.

[0097] As used herein, the term "in vivo" refers to events that occur within a multicellular organism such as a human or non-human animal. In the context of cell-based systems, it can be used to refer to events that occur within living cells (as opposed to, for example, in vitro systems).

[0098] As used herein, the term “in vitro” refers to an event that occurs in a cell or tissue that grows outside a multicellular organism rather than inside it.

[0099] As used herein, the term "cell" should be understood not only to a specific single cell, but also to the cell's offspring or potential offspring. Because certain modifications may occur in offspring due to mutations or environmental influences, such offspring may indeed differ from the parent cell, but are still included within the scope of this term.

[0100] As used herein, the terms “subject,” “individual,” and “patient” are used interchangeably and refer to a vertebrate, preferably a mammal, and more preferably a human. Mammals include, but are not limited to, rodents, apes, humans, farm animals, sporting animals, and pets. It also encompasses tissues, cells, and their progeny from biological entities obtained in vivo or cultured in vitro.

[0101] As used herein, the term “disease” refers to any disorder or condition that impairs or interferes with the normal functioning of cells, tissues, or organs, including but not limited to those specific diseases that have been medically or clinically defined.

[0102] As used herein, the term "treatment" refers to the administration of a therapeutic molecule (e.g., the CRISPR-Cas system described herein) that partially or completely reduces, improves, alleviates, inhibits, or inhibits one or more symptoms or features of a disease, condition, and / or symptom, delays the onset of one or more symptoms or features of a disease, condition, and / or symptom, reduces the severity of one or more symptoms or features of a disease, condition, and / or symptom, and / or reduces the incidence of one or more symptoms or features of a disease, condition, and / or symptom. Such treatment may be for subjects who do not have symptoms of the relevant disease, condition, and / or symptom and / or for subjects who only have early symptoms of the disease, condition, and / or symptom. Alternatively, such treatment may be for subjects who have one or more identified symptoms of the relevant disease, condition, and / or symptom.

[0103] As used herein, a “biological sample” may contain whole cells and / or live cells and / or cell debris. A biological sample may contain (or be derived from) “body fluids.” In some embodiments, body fluids may be selected from amniotic fluid, aqueous humor, vitreous humor, bile, serum, breast milk, cerebrospinal fluid, cerumen (earwax), chyle, chyme, endolymph, perilymph, exudate, feces, female ejaculation, gastric acid, gastric juice, lymph, mucus (including nasal drainage and sputum), pericardial fluid, peritoneal fluid, pleural fluid, pus, thin mucus, saliva, sebum (skin oil), semen, sputum, synovial fluid, sweat, tears, urine, vaginal secretions, vomit, and one or more mixtures thereof. Biological samples include cell cultures, body fluids, and cell cultures derived from body fluids. Body fluids may be obtained from mammals, for example, by puncture or other collection or sampling procedures.

[0104] As used herein, “expression of a genomic locus” or “gene expression” is the process by which information from a gene is used to synthesize a functional gene product. The product of gene expression is typically a protein, but in non-protein-coding genes such as rRNA or tRNA genes, the product is functional RNA. All known life forms—eukaryotes (including multicellular organisms), prokaryotes (bacteria and archaea), and viruses—use gene expression processes to produce functional products for survival. As used herein, “expression” of a gene or nucleic acid encompasses not only cellular gene expression but also the transcription and translation of nucleic acids in cloning systems and any other context. As used herein, “expression” also refers to the process by which polynucleotides are transcribed from a DNA template (such as into mRNA or other RNA transcripts) and / or the subsequent translation of transcribed mRNA into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as “gene products.” If the polynucleotide is derived from genomic DNA, expression may include the splicing of mRNA in eukaryotic cells.

[0105] As used herein, the term "RuvC domain" refers to a conserved domain or motif of an amino acid having nuclease (e.g., endonuclease) activity. As used herein, a protein with a split-type RuvC domain is a protein having two or more RuvC motifs at sequentially different sites within its sequence, these RuvC motifs interacting at a tertiary level to form the RuvC domain. RuvC domains or segments thereof (e.g., RuvC_I, RuvC_II, or RuvC_III) can generally be identified by alignment with known domain sequences, by structural alignment with proteins having annotated domains, or by comparison with sequences based on known domain sequences.

[0106] As used herein, the term "HNH domain" generally refers to a nuclease domain containing characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment with known domain sequences, alignment with the structure of proteins with annotated domains, or by comparison with sequences based on known domain sequences.

[0107] As used herein, the terms “BH domain,” “bridged helical domain,” or “bridged helix” generally refer to an arginine-rich helical domain present in Cas enzymes and playing an important role in the initiation of cleavage activity after target DNA binding.

[0108] As used herein, the term "recognition domain" or "REC domain" generally refers to the domain that is believed to interact with the repeat:anti-repeat double helix of crRNA and mediate the formation of the Cas endonuclease / crRNA complex.

[0109] As used herein, the term "WED domain" or "wedge domain" generally refers to a fold containing a twisted five-stranded β-sheet and four α-helices on either side, which is typically responsible for the recognition of the modified repeating of the nuclease: the anti-repeating double helix. The WED domain can also be responsible for the recognition of unidirectional guide RNA scaffolds.

[0110] As used herein, "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that catalyze a hydrolytic deamination reaction that converts adenine (or the adenine portion of a molecule) to hypoxanthine (or the hypoxanthine portion of a molecule). In some embodiments, the functional domain is selected from adenosine deaminases. In some embodiments, the adenosine deaminase comprises the ability to deaminate adenine (A) to hypoxanthine (I). In some embodiments, the deamination of adenine to hypoxanthine converts adenosine (A) or deoxyadenosine (dA) containing adenine to guanosine (G) or deoxyguanosine (dG).

[0111] As used herein, the term "cytidine deaminase" or "cytidine deaminase protein" refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that catalyzes a hydrolytic deamination reaction that converts cytosine (or the cytosine portion of a molecule) to uracil (or the uracil portion of a molecule). In some embodiments, the cytosine-containing molecule is cytidine (C), and the uracil-containing molecule is uridine (U). The cytosine-containing molecule may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0112] In some embodiments, the cytidine deaminase is selected from APOBEC (e.g., APOBEC3, such as APOBEC3A, APOBEC3B, APOBEC3C).

[0113] Various embodiments are described below. It should be noted that specific embodiments are not intended as exhaustive descriptions or as limitations on the broader aspects discussed herein. An aspect described in connection with a particular embodiment is not necessarily limited to that embodiment and may be practiced in conjunction with any other embodiment. Throughout the specification, references to “one embodiment,” “implementation,” “some embodiments,” “implementation,” “example embodiment,” or “example embodiment” refer to specific features, structures, or characteristics described in connection with an embodiment that are included in at least one embodiment of the invention. Therefore, the phrases “in one embodiment,” “in one embodiment,” or “example embodiment” appearing throughout the specification do not necessarily all refer to the same embodiment, but may. Furthermore, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner, as will be apparent to those skilled in the art based on this disclosure. Moreover, although some embodiments described herein include some but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention. For example, in the appended claims, any claimed embodiment may be used in any combination.

[0114] Programmable nuclease

[0115] In some respects, this disclosure provides an engineered programmable nuclease.

[0116] In some technical solutions, the programmable nuclease disclosed herein includes an HNH domain and a RuvC domain.

[0117] In some embodiments, the programmable nuclease of this disclosure further includes a BH domain and a REC leaf, wherein the BH domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 18; the REC leaf contains a REC domain, wherein the REC domain comprises two parts, REC1 and REC2, wherein the REC1 part has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 13, and the REC2 part has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 14; the amino acid sequence length of the nuclease is between 600 and 800 amino acids, preferably between 600 and 750 amino acids.

[0118] In some embodiments, the HNH domain of the programmable nuclease disclosed herein has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 17.

[0119] In some embodiments, the programmable nuclease of this disclosure includes a WED domain that has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 11.

[0120] In some embodiments, the RuvC domain of the programmable nuclease of this disclosure includes RuvC-Ⅰ, RuvC-Ⅱ, and RuvC-Ⅲ domains, wherein the RuvC-Ⅰ domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.16; the RuvC-Ⅱ domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.19; the RuvC-Ⅲ domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.15; and the PI domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.15; and the PI domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.16. NO.12 has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity.

[0121] In some embodiments, the programmable nuclease of this disclosure has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.1.

[0122] In some respects, this disclosure provides a nucleic acid molecule that encodes a programmable nuclease of this disclosure.

[0123] In some embodiments, the nucleic acid molecule of this disclosure comprises a nucleotide sequence that is at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) identical to the sequence of SEQ ID NO.2.

[0124] Fusion protein

[0125] In some aspects, this disclosure provides a fusion protein comprising a programmable nuclease of this disclosure and a functional domain fused thereto.

[0126] In some embodiments, the functional domain is selected from nuclear localization signals (NLS), nuclear output signals (NES), reporter proteins (e.g., fluorescent proteins), nuclease targeting regions, DNA binding domains (e.g., Lex A DBD, Gal4 DBD, Sp1 DBD), epitope tags (e.g., His, myc, V5, FLAG, HA, VSV-G, etc.), transcriptional activation domains (e.g., VP64, VPR, p65, Rta), transcriptional repression domains (e.g., KRAB domain, SID domain, NuE domain, NcoR domain, or SID4X domain), nucleases, deaminases (e.g., adenosine deaminase or cytidine deaminase), methyltransferases (e.g., DNA methyltransferase DNMT), demethylases, transcription release factors, HDAC, cleavage-active peptides, ligases, integrases, transposases, recombinases, polymerases, exonucleases (e.g., T5E), and base excision repair inhibitors (e.g., uracil-DNA glycosyltransferase inhibitors (UGI)).

[0127] In some embodiments, the functional domain includes one or more of the following enzymatic activities against the target sequence: methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, demyristylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), and deglycosylation activity;

[0128] In some embodiments, the functional domain is selected from the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain;

[0129] In some embodiments, the adenosine deaminase catalytic domain or cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

[0130] In some implementations, the positioning signal includes a nuclear positioning signal (NLS) and / or a nuclear output signal (NES);

[0131] In some implementations, the sequence of the nuclear localization signal is selected from PKKKRKV, KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, PAAKRVKLD, RQRRNELKRSP, NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY, RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV, VSRKRPRP, PPKKARED, PQPKKKPL, SALIKKKKKMAP, DRLRR, PKQKKRK, RKLKKKIKKL, REKKKFLKRR, KRKGDEVDGVDEVAKKKSKK, RKCLQAGMNLEARKTKK.

[0132] In some embodiments, the sequence of the nuclear localization signal is located at, near, or close to the end (e.g., the N-terminus or C-terminus) of the programmable nuclease;

[0133] In some embodiments, the nuclear output signal includes protein tyrosine kinase 2 (such as human protein tyrosine kinase 2);

[0134] In some embodiments, the reporter protein includes glutathione S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, and autofluorescent protein;

[0135] In some embodiments, the autofluorescent protein includes green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, etc.), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), and blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire).

[0136] In some embodiments, the DNA binding domain includes methylation-binding proteins, LexADBD, and Gal4DBD;

[0137] In some implementations, the epitope tags include histidine tags, V5 tags, FLAG tags, influenza virus hemagglutinin tags, Myc tags, VSV-G tags, thioredoxin tags, and streptavidin tags;

[0138] In some implementations, the transcriptional activation domain includes VP64 and / or VPR;

[0139] In some implementations, the transcriptional repression domain includes KRAB and / or SID;

[0140] In some embodiments, the nuclease includes FokI;

[0141] In some embodiments, the cleavage-active polypeptide includes a polypeptide with single-stranded RNA cleavage activity, a polypeptide with double-stranded RNA cleavage activity, a polypeptide with single-stranded DNA cleavage activity, or a polypeptide with double-stranded DNA cleavage activity.

[0142] In some embodiments, the ligase includes DNA ligase and / or RNA ligase;

[0143] In some embodiments, the exonuclease is selected from TREX2 protein, TREX1 protein, APE1 protein, Artemis protein, CtIP protein, Exo1 protein, Mre11 protein, RAD1 protein, RAD9 protein, Tp53 protein, WRN protein, exonuclease V, T5 exonuclease or T7 exonuclease or variants thereof.

[0144] In some implementations, the functional domain is connected to the N-terminus and / or C-terminus of the programmable nuclease;

[0145] In some implementations, the functional domain is inserted between the N-terminus and C-terminus of the programmable nuclease;

[0146] In some implementations, the one or more functional domains are optionally connected to the N-terminus and / or C-terminus of the programmable nuclease via a connector.

[0147] system

[0148] In some aspects, this disclosure provides a nuclease system comprising:

[0149] (1) The programmable nuclease of this disclosure or the polynucleotide encoding said programmable nuclease; and

[0150] (2) A guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, wherein the guiding nucleic acid comprises:

[0151] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0152] (ii) capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the DNA target sequence of the target DNA;

[0153] Optionally, the protein-binding sequence is located at the 3' end of the DNA targeting sequence; and / or

[0154] Optionally, the guiding nucleic acid is a guiding RNA (gRNA).

[0155] In some embodiments, the protein-binding sequence has a secondary structure substantially identical to that of SEQ ID NO.3;

[0156] In some embodiments, the protein binding sequence is:

[0157] (1) Contains the polynucleotide sequence shown in SEQ ID NO.3; or

[0158] (2) Contains a polynucleotide sequence having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO. 3;

[0159] In some embodiments, the protein-binding sequence comprises the polynucleotide sequence shown in SEQ ID NO.3.

[0160] In some embodiments, the target sequence comprises about or at least about 16 consecutive nucleotides of the target DNA, for example, about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 6...) of the target DNA. 8, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more consecutive nucleotides; optionally, the target sequence comprises about 15-70 consecutive nucleotides of the target DNA; optionally, the target sequence comprises about 16 to about 22 consecutive nucleotides, or about 17 to about 22 consecutive nucleotides of the target DNA; optionally, the target sequence comprises about 20 consecutive nucleotides of the target DNA.

[0161] In some embodiments, the inverse complementary sequence of the target sequence is adjacent to the 3' end of the prototype spacer adjacent motif (PAM); optionally, the PAM is 5'-NGG, where N is A, T, G, or C.

[0162] In some embodiments, the DNA target sequence is about or at least about 16 consecutive nucleotides long, for example, about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, ...). 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more consecutive nucleotides, optionally about 15-70 consecutive nucleotides; optionally about 16 to about 22 consecutive nucleotides, or about 17 to about 22 consecutive nucleotides; optionally about 20 consecutive nucleotides.

[0163] In some embodiments, the target DNA is dsDNA or ssDNA. In some embodiments, the target DNA is present in bacterial cells, archaea cells, single-celled eukaryotes, plant cells, invertebrate cells, or vertebrate cells.

[0164] Guidance on nucleic acid testing

[0165] In some respects, this disclosure provides a guide for nucleic acids, which include:

[0166] a) A DNA targeting sequence, said DNA targeting sequence comprising a nucleotide sequence complementary to a target sequence in the target DNA; and

[0167] b) A protein-binding sequence that binds to a programmable nuclease disclosed herein to form a complex.

[0168] In some embodiments, the protein-binding sequence comprises a polynucleotide sequence having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO. 3.

[0169] In some embodiments, the DNA target sequence is about or at least about 16 consecutive nucleotides long, for example, about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, ...). 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more consecutive nucleotides, optionally about 15-70 consecutive nucleotides; optionally about 16 to about 22 consecutive nucleotides, or about 17 to about 22 consecutive nucleotides; optionally about 20 consecutive nucleotides.

[0170] In some embodiments, the DNA targeting sequence is at least 70%, about 80%, about 90%, about 99%, or 100% complementary to the target DNA.

[0171] In some embodiments, the one or more polynucleotide sequences encoding the guiding nucleic acid are operatively linked to a promoter that functions in eukaryotic cells. In some embodiments, the promoter is selected from constitutive promoters, inducible promoters, ubiquitous promoters, cell type-specific promoters, or tissue-specific promoters.

[0172] Base editing

[0173] In some aspects, this disclosure provides a fusion protein comprising a deaminase or its catalytic domain, the fusion protein comprising a programmable nuclease of this disclosure, the programmable nuclease having an amino acid sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) with respect to the amino acid sequence shown in SEQ ID NO. 1, or containing one or more amino acid substitutions relative to the amino acid sequence shown in SEQ ID NO. 1, and the programmable nuclease having reduced or substantially absent dsDNA cleavage activity relative to the parental nuclease (e.g., the programmable nuclease shown in SEQ ID NO. 1). In some embodiments, the programmable nuclease is not entirely endonuclease-deficient, but its endonuclease activity is not against the double strand of dsDNA, but against one strand of dsDNA or ssDNA (meaning or senseless; or target or non-target strand). This means that the programmable nuclease is essentially unable to act as a dsDNA endonuclease that cuts the double strand of dsDNA against target DNA or against non-target DNA, but is essentially able to act as an ssDNA endonuclease (or nicking enzyme) that cuts ssDNA or "creates a nick" in one strand of dsDNA.

[0174] In some embodiments, the functional domain is selected from the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain;

[0175] In some embodiments, the adenosine deaminase catalytic domain or cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD.

[0176] In some embodiments, the fusion protein further comprises a reverse transcriptase (RT) or its catalytic domain. In some embodiments, the guiding nucleic acid further comprises or is used in combination with a reverse transcription donor RNA (RT donor RNA), which contains a primer binding site (PBS) and a template sequence.

[0177] deliver

[0178] Various delivery methods can be applied to the programmable nuclease, the directed nucleic acid, or the system disclosed herein.

[0179] In some aspects, this disclosure provides a delivery system comprising (1) a nuclease, a polynucleotide, or a system of this disclosure; and (2) a delivery vehicle.

[0180] In some implementations, the delivery medium includes nanoparticles, liposomes, exosomes, microvesicles, gene guns, or electroporation devices.

[0181] Furthermore, when the target of delivery is plant cells, delivery methods such as cell-penetrating peptides (CPPs) are employed. For example, in one specific embodiment, a nuclease and / or at least one guide RNA is coupled to one or more CPPs, thereby efficiently transporting the CPP coupled with the nuclease and / or guide RNA into plant cells (e.g., protoplasts). CPPs are short peptides of fewer than 35 amino acids, derived from proteins or chimeric sequences, capable of transporting biomolecules across cell membranes in a receptor-independent manner. CPPs can be cationic peptides, peptides with hydrophobic sequences, amphiphilic peptides, peptides rich in proline and antimicrobial sequences, and chimeric or dipeptides. CPPs can penetrate biological membranes and thus trigger the transmembrane movement of various biomolecules into the cytoplasm, improving their intracellular pathways and thus promoting biomolecule-target interactions.

[0182] For example, CPP includes Tat (a nuclear transcription activation protein required for viral replication by HIV type 1), penetrin, Kaposi's fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, sweet arrow peptide, etc.

[0183] In another aspect, this disclosure provides a vector comprising the nucleic acid molecule of this disclosure. In some embodiments, the vector encodes a guide nucleic acid as defined in this disclosure. In some embodiments, the vector is a plasmid vector, a recombinant AAV (rAAV) vector (vector genome), or a recombinant lentiviral vector.

[0184] In some aspects, this disclosure provides a recombinant AAV viral particle, said rAAV viral particle containing the rAAV vector genome of this disclosure. A brief introduction to AAV for delivery can be found in the "Adeno-Associated Virus (AAV) Guide" (addgene.org / guides / aav / ).

[0185] In some aspects, this disclosure provides a ribonucleoprotein (RNP) comprising a programmable nuclease of this disclosure and optionally a guiding nucleic acid as defined in this disclosure.

[0186] In some aspects, this disclosure provides a lipid nanoparticle (LNP) comprising a programmable nuclease or system encoded by this disclosure.

[0187] In some embodiments, the lipid nanoparticles comprise the programmable nuclease of this disclosure and the RNA (e.g., mRNA) that guides the nucleic acid of this disclosure.

[0188] Modification methods

[0189] In some aspects, this disclosure provides a method for targeting, binding, cleaving, labeling, or modifying target DNA, the method comprising contacting the target DNA with a complex, the complex comprising:

[0190] (1) The programmable nuclease of this disclosure or the polynucleotide encoding said programmable nuclease; and

[0191] The guiding nucleic acid or the polynucleotide encoding the guiding nucleic acid, the guiding nucleic acid comprising:

[0192] (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and

[0193] (ii) capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the DNA target sequence of the target DNA;

[0194] cell

[0195] In some respects, this disclosure provides a cell comprising

[0196] (1) The system defined in this disclosure;

[0197] (2) Any nucleic acid molecule as defined in this disclosure;

[0198] (3) The medium of this disclosure;

[0199] (4) The ribonucleoprotein of this disclosure; or

[0200] (5) The lipid nanoparticles disclosed herein;

[0201] Pharmaceutical Composition

[0202] In some aspects, this disclosure provides a pharmaceutical composition comprising:

[0203] (1) the system of this disclosure, the carrier of this disclosure, the ribonucleoprotein of this disclosure, or the lipid nanoparticle of this disclosure or the cell of this disclosure; and (2) pharmaceutically acceptable excipients.

[0204] In some embodiments, the dosage form of the composition is selected from the group consisting of lyophilized formulations, liquid formulations, or combinations thereof.

[0205] In some embodiments, the composition is in the form of a liquid formulation.

[0206] In some embodiments, the composition is in the form of an injection.

[0207] In some embodiments, the composition is a cell preparation.

[0208] Suitable pharmaceutically acceptable excipients typically include inert substances that facilitate administration of the pharmaceutical composition to a subject, facilitate formulation of the pharmaceutical composition into a deliverable formulation, or facilitate storage of the pharmaceutical composition prior to administration. Pharmaceutically acceptable excipients can include agents that can stabilize, optimize, or otherwise alter the form, consistency, viscosity, pH, pharmacokinetics, or solubility of a formulation. Such agents include buffers, wetting agents, emulsifiers, diluents, encapsulating agents, and skin penetration enhancers. For example, carriers can include, but are not limited to, saline, buffered saline, dextran, arginine, sucrose, water, glycerol, ethanol, sorbitol, dextran, sodium carboxymethyl cellulose, and combinations thereof.

[0209] Some non-limiting examples of substances that can be used as pharmaceutically acceptable carriers include: (1) sugars, such as lactose, glucose and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose and its derivatives, such as sodium carboxymethyl cellulose, methyl cellulose, ethyl cellulose, microcrystalline cellulose and cellulose acetate; (4) powdered tragacanth gum; (5) malt; (6) gelatin; (7) lubricants, such as magnesium stearate, sodium lauryl sulfonate and talc; (8) excipients, such as cocoa butter and suppository wax; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil and soybean oil; (10) glycols, such as propylene glycol. (11) Polyols, such as glycerol, sorbitol, mannitol and polyethylene glycol (PEG); (12) Esters, such as ethyl oleate and ethyl laurate; (13) Agar; (14) Buffers, such as magnesium hydroxide and aluminum hydroxide; (15) Alginate; (16) Pyrogen-free water; (17) Isotonic saline; (18) Ringer's solution; (19) Ethanol; (20) pH buffer solution; (21) Polyesters, polycarbonates and / or polyanhydrides; (22) Fillers, such as peptides and amino acids; (23) Serum alcohols, such as ethanol; and (24) Other non-toxic and compatible substances used in pharmaceutical formulations. The formulation may also contain wetting agents, colorants, release agents, coating agents, sweeteners, flavoring agents, fragrances, preservatives and antioxidants.

[0210] Pharmaceutical compositions may contain one or more pH buffering compounds to maintain the pH of the formulation at a predetermined level reflecting physiological pH, such as in the range of about 5.0 to about 8.0. pH buffering compounds for aqueous liquid formulations may be amino acids or mixtures of amino acids, such as histidine or mixtures of amino acids (such as histidine and glycine). Alternatively, the pH buffering compound is preferably an agent that maintains the pH of the formulation at a predetermined level (such as in the range of about 5.0 to about 8.0) and does not chelate calcium ions. Illustrative examples of such pH buffering compounds include, but are not limited to, imidazole and acetate ions. The pH buffering compound may be present in any amount suitable for maintaining the pH of the formulation at the predetermined level.

[0211] The pharmaceutical composition may also contain one or more osmotic modifiers, which adjust the osmotic properties of the formulation (e.g., tone, osmotic pressure, and / or osmolarity) to a level acceptable to the recipient individual's blood flow and blood cells. The osmotic modifier may be an agent that does not chelate calcium ions. The osmotic modifier may be any compound known or available to those skilled in the art for adjusting the osmotic properties of the formulation. Those skilled in the art can empirically determine the suitability of a given osmotic modifier in the formulation of the present invention. Illustrative examples of suitable types of osmotic modifiers include, but are not limited to: salts, such as sodium chloride and sodium acetate; sugars, such as sucrose, dextran, and mannitol; amino acids, such as glycine; and mixtures of one or more of these agents and / or dosage forms. One or more osmotic modifiers may be present at any concentration sufficient to adjust the osmotic properties of the formulation.

[0212] Treatment

[0213] In some aspects, this disclosure provides a method for diagnosing, preventing, or treating a disease in a subject of need, the method comprising administering to the subject a system of this disclosure, a carrier of this disclosure, a ribonucleoprotein of this disclosure, a lipid nanoparticle of this disclosure, a cell of this disclosure, or a pharmaceutical composition of this disclosure, wherein the disease is associated with a target DNA, wherein a spacer sequence is capable of hybridizing with a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnoses, prevents, or treats the disease.

[0214] In some embodiments, the disease is selected from hypercholesterolemia, familial hypercholesterolemia (FH), atherosclerosis, Angelman's syndrome (AS), Alzheimer's disease (AD), transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, progressive pseudohypertrophic muscular dystrophy, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), α-1-antitrypsin deficiency, primary hyperoxaluria, Pompe disease, myotonic dystrophy, and Hennepin County. Tinton's disease (HTT), Fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, thalassemia (e.g., β-thalassemia), Parkinson's disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, non-alcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), heart disease (e.g., hypertrophic cardiomyopathy (HCM)), hepatitis B, and cancer.

[0215] Detection methods

[0216] In some aspects, this disclosure provides a method for detecting target DNA, comprising contacting the target DNA with a system of the present disclosure, wherein the target DNA is modified by the complex, and wherein the modification detects the target DNA. In some embodiments, the modification generates a detectable signal, such as a fluorescence signal.

[0217] Reagent test kit

[0218] In some aspects, this disclosure provides a kit comprising the programmable nuclease of this disclosure, the system of this disclosure, the polynucleotide of this disclosure, the vector of this disclosure, the RNP of this disclosure, the LNP of this disclosure, the delivery system of this disclosure, the cells of this disclosure, or the pharmaceutical composition of this disclosure, or any one, two or all of the components thereof.

[0219] In some embodiments, the kit further includes instructions for using one or more components contained therein, and / or instructions for combining with one or more additional components that may be available elsewhere or are required.

[0220] In some embodiments, the kit further includes one or more buffer solutions that can be used to dissolve any of the one or more components contained therein, and / or to provide suitable reaction conditions for one or more of the components. Such buffer solutions may include one or more of the following: PBS, HEPES, Tris, MOPS, Na₂CO₃, NaHCO₃, NaB, or combinations thereof. In some embodiments, reaction conditions include an appropriate pH, such as an alkaline pH. In some embodiments, the pH is between 7 and 10.

[0221] In some implementations, any one or more of the kit components may be stored in a suitable container or at a suitable temperature, such as 4 degrees Celsius.

[0222] Further embodiments are described in the following examples, which are for illustrative purposes only and are not intended to limit the scope of this disclosure. When representing RNA in sequence, "t" or "T" should be treated as "U".

[0223] Unless otherwise specified, the experimental methods used in the following examples are conventional methods.

[0224] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.

[0225] Example

[0226] Example 1. Discovering novel nucleases through metagenomics

[0227] Analysis of the metagenomics of uncultured organisms identified a novel nuclease (amino acid sequence as shown in SEQ ID NO. 1, nucleotide coding sequence as shown in SEQ ID NO. 2), named YTGE-1006. ClustalW alignment was used to predict and compare it with reference proteins (e.g., SaCas9, SpCas9), and RAxML was used to infer the phylogenetic tree. Figure 1 ), analyze its structural domain distribution as follows Figure 2 As shown, its protein structure was predicted using alphafold2. Figure 3 By comparing its representative domains with those of a reference protein, it was identified and analyzed, revealing the following characteristics compared to the reference protein:

[0228] Table 1. Properties of YTGE-1006 described in this paper and its comparison with known nucleases.

[0229]

[0230]

[0231] The above analysis shows that although the nuclease YTGE-1006 contains several domains / structural parts similar to Cas9 (SaCas9, SpCas9) (RuvC domain (RuvC-Ⅰ, RuvC-Ⅱ, RuvC-Ⅲ domain sequences are shown in SEQ ID NO. 16, 19, 15, respectively), HNH domain (sequence shown in SEQ ID NO. 17), BH domain (sequence shown in SEQ ID NO. 18), REC leaflet (sequence shown in SEQ ID NO. 10), WED domain (sequence shown in SEQ ID NO. 11), PI domain (sequence shown in SEQ ID NO. 12)), the composition, structure, and distribution of its multiple domains / structural parts are different. Therefore, YTGE-1006 is considered a CRISPR nuclease that is different from known nucleases.

[0232] Example 2. Structural prediction of protein-binding sequence folding

[0233] Locus annotation of samples containing YTGE-1006 was performed using CRISPRCasFinder and CMsearch, and the YTGE-1006-protein binding sequence (SEQ ID NO.3) was analyzed, with its secondary structure predicted. Figure 4 ).

[0234] Example 3. Detection of the mediating activity of guide RNAs with different protein binding sequences

[0235] 1. To detect the cleavage activity of nucleases, an expression plasmid capable of expressing nuclease YTGE-1006 and guide RNA, as well as a target plasmid for detecting cleavage activity, were constructed.

[0236] A DNA targeting sequence was designed based on the human PCSK9 target sequence (SEQ ID NO.4), and a PCSK9-sgRNA was designed based on the YTGE-1006 backbone sequence: PCSK9-sgRNA sequence (SEQ ID NO.5).

[0237] The PCSK9-sgRNA consists of a DNA targeting sequence (targeting PCSK9, the sequence is the same as SEQ ID NO.4) in the 5'-3' direction and a scaffold sequence (YTGE-1006-protein binding sequence, SEQ ID NO.3). When representing RNA, "t" or "T" should be regarded as "U".

[0238] Expression plasmids containing the coding sequence of nuclease YTGE-1006 (SEQ ID NO:1) regulated by the T7 promoter (SEQ ID NO:2) and the coding sequence of PCSK9-sgRNA regulated by the T7 promoter (SEQ ID NO:5) (SEQ ID NO:6, plasmid map see [link]) were constructed. Figure 5 ).

[0239] Construct a target plasmid for nuclease cleavage (SEQ ID NO.7, plasmid map see [link]). Figure 6 The plasmid contains the Kana resistance gene and a target sequence (SEQ ID NO.4) for guiding RNA targeting. Considering the potential target sequence recognition bias of programmable nucleases, and influenced by the PAM sequence immediately adjacent to the 3' end of the target sequence, a 6N random PAM sequence immediately adjacent to the 3' end of the target sequence is also included in the target cleavage plasmid. Here, N represents nucleotides A, T, C, or G, resulting in 4^6 = 4096 possible PAM sequences, thus forming a target cleavage plasmid library containing all possible PAM sequences.

[0240] 2. The expression plasmid (100 ng) and the target plasmid library (100 ng) were electroporated into 100 μL of competent Escherichia coli cells (purchased from Weidi Biotechnology). After transformation, the E. coli cells were cultured at 37°C with shaking for 1 hour, and then plated on a medium coated with ampicillin (Amp) and kanamycin (Kan) to ensure that the colony count was greater than 100 times the library count. After growing for 14-16 hours, all colonies were recovered and plasmids were extracted.

[0241] As a control, 100 ng of the target plasmid library was electrotransformed into 100 μL of competent E. coli cells. After transformation, the E. coli cells were cultured at 37°C with shaking for 1 hour, and then plated on a medium coated with kanamycin (Kan). After growing for 14-16 hours, all colonies were recovered and plasmids were extracted.

[0242] When the nuclease and guide RNA expression plasmid are co-transformed into host *E. coli*, the nuclease and guide RNA expression plasmid express the programmable nuclease YTGE-1006 and the corresponding sgRNA, forming a complex (RNP). The guide RNA targets the target sequence contained in the target plasmid, and the programmable nuclease YTGE-1006 targets and cleaves the target plasmid. If YTGE-1006 has endonuclease activity and can recognize the PAM at the 3' end of the target sequence, it will cleave the target sequence, causing the target cleavage plasmid to degrade and unable to express the kanamycin resistance gene. Consequently, the host *E. coli* cannot grow on media containing kanamycin.

[0243] 3. The fragment containing the PAM sequence was amplified by PCR and sequenced by PE150. The primers used were PCSK9-PAM-SF (SEQ ID NO.8) and PCSK9-PAM-SR (SEQ ID NO.9).

[0244] The occurrence frequency of PAM sequences in 4096 combinations was then counted in both the experimental and control groups. The sequences were then standardized using the PAM sequence count. For a given PAM sequence, a value greater than 3 (log2(standardized value in control group / standardized value in experimental group)) was considered significantly consumed. The PAM domains were predicted from these significantly consumed PAM sequences. The sequencing results are summarized as follows: Figure 7 As shown, each point on the graph represents one of 4096 PAMs. If the point representing a certain PAM falls to the lower right of the curve, it means that for the PAM represented by that point and its adjacent target sequence, the endonuclease YTGE-1006 recognizes and causes statistically significant cleavage, leading to the degradation of the target cleavage plasmid, the absence of expression of the resistance gene, and consequently, the death of E. coli in the experimental group.

[0245] PAM analysis results showed that, mediated by PCSK9-sgRNA, nuclease YTGE-1006 exhibited a significant PAM: 5'-NGG-3' ( Figure 8 ).

[0246] The following is a partial sequence from this article:

[0247]

[0248]

[0249]

[0250]

[0251]

[0252] While preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The invention is not intended to be limited to the specific embodiments provided in the specification. Although the invention has been described with reference to the foregoing description, the description and illustration of the embodiments herein are not intended to be restrictive. Many variations, modifications, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific descriptions, configurations, or relative proportions set forth herein, and are subject to various conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. Therefore, the invention should also be considered to cover any such alternatives, modifications, variations, or equivalents. The appended claims are intended to define the scope of the invention and thereby cover the methods and structures within the scope of these claims and their equivalents.

[0253] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A programmable nuclease comprising an HNH domain and a RuvC domain, characterized in that, The programmable nuclease further includes a BH domain and a REC leaflet. The BH domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.

18. The REC leaflet includes a REC domain, which comprises two parts, REC1 and REC2. The REC1 part has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO. 13, and the REC2 part has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.

14. The amino acid sequence length of the nuclease is between 600 and 800 amino acids, preferably between 650 and 800 amino acids.

2. The programmable nuclease of claim 1, wherein the HNH domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.

17.

3. The programmable nuclease of claim 1 or 2 further comprises a WED domain having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.

11.

4. The programmable nuclease as described in claim 3, wherein the RuvC domain comprises RuvC-Ⅰ, RuvC-Ⅱ, and RuvC-Ⅲ domains, wherein, The RuvC-Ⅰ domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.16; the RuvC-Ⅱ domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.19; the RuvC-Ⅲ domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.15; and the PI domain has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.

12.

5. The programmable nuclease of claim 1, wherein the programmable nuclease has at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% sequence identity with SEQ ID NO.

1.

6. A programmable nuclease according to claim 5, wherein the programmable nuclease further comprises a functional structural domain fused thereto; Preferably, the functional domains are selected from nuclear localization signals (NLS), nuclear output signals (NES), reporter proteins (e.g., fluorescent proteins), nuclease targeting regions, DNA binding domains (e.g., Lex A DBD, Gal4 DBD, Sp1 DBD), epitope tags (e.g., His, myc, V5, FLAG, HA, VSV-G, etc.), transcriptional activation domains (e.g., VP64, VPR, p65, Rta), transcriptional repression domains (e.g., KRAB domain, SID domain, NuE domain, NcoR domain, or SID4X domain), nucleases, deaminases (e.g., adenosine deaminase or cytidine deaminase), methyltransferases (e.g., DNA methyltransferase DNMT), demethylases, transcription release factors, HDAC, cleavage active peptides, ligases, integrases, transposases, recombinases, polymerases, exonucleases (e.g., T5E), and base excision repair inhibitors (e.g., uracil-DNA glycosylation inhibitors (UGI)). Preferably, the functional domain includes one or more of the following enzymatic activities against the target sequence: methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, demyristylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), and deglycosylation activity; Preferably, the functional domain is selected from the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain; Preferably, the adenosine deaminase catalytic domain or the cytidine deaminase catalytic domain includes one or more of ADAR1, ADAR2, APOBEC, AID, or TAD; Preferably, the positioning signal includes a nuclear positioning signal (NLS) and / or a nuclear output signal (NES); Preferably, the sequence of the nuclear localization signal is located at, near, or close to the end (e.g., the N-terminus or C-terminus) of the programmable nuclease; Preferably, the nuclear output signal includes protein tyrosine kinase 2 (such as human protein tyrosine kinase 2); Preferably, the reporter protein includes glutathione S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, and autofluorescent protein; Preferably, the autofluorescent protein includes green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, CopGFP, AceGFP, etc.), HcRed, DsRed, cyan fluorescent protein (e.g., eCFP, Cerulean, CyPet, AmCyanl, etc.), yellow fluorescent protein (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, etc.), and blue fluorescent protein (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire); Preferably, the DNA binding domain includes methylation-binding proteins, LexADBD, and Gal4DBD; Preferably, the epitope tag includes histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, thioredoxin tag, and streptavidin tag; Preferably, the transcriptional activation domain includes VP64 and / or VPR; Preferably, the transcriptional repression domain includes KRAB and / or SID; Preferably, the nuclease comprises FokI; Preferably, the cleavage-active polypeptide includes a polypeptide with single-stranded RNA cleavage activity, a polypeptide with double-stranded RNA cleavage activity, a polypeptide with single-stranded DNA cleavage activity, or a polypeptide with double-stranded DNA cleavage activity. Preferably, the ligase comprises DNA ligase and / or RNA ligase; Preferably, the exonuclease is selected from TREX2 protein, TREX1 protein, APE1 protein, Artemis protein, CtIP protein, Exo1 protein, Mre11 protein, RAD1 protein, RAD9 protein, Tp53 protein, WRN protein, exonuclease V, T5 exonuclease or T7 exonuclease or variants thereof. Preferably, the functional structural domain is connected to the N-terminus and / or C-terminus of the programmable nuclease; Preferably, the functional structural domain is inserted between the N-terminus and C-terminus of the programmable nuclease; Preferably, the one or more functional domains are optionally connected to the N-terminus and / or C-terminus of the programmable nuclease via a connector.

7. A nuclease system comprising: (1) The programmable nuclease or the polynucleotide encoding the programmable nuclease according to any one of claims 1-6; and (2) A guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, wherein the guiding nucleic acid comprises: (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and (ii) capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the DNA target sequence of the target DNA; Optionally, the protein-binding sequence is located at the 3' end of the DNA targeting sequence; and / or Optionally, the guiding nucleic acid is a guiding RNA (gRNA).

8. The nuclease system of claim 7, wherein the protein binding sequence has a secondary structure substantially identical to that of SEQ ID NO. 3; Optionally, the protein binding sequence is: (1) Contains the polynucleotide sequence shown in SEQ ID NO.3; or (2) A polynucleotide sequence comprising having at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO.

3.

9. The system of claim 7 or 8, wherein the target sequence comprises about or at least about 16 consecutive nucleotides of the target DNA, for example, about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66) of the target DNA. The target sequence comprises about 15-70 consecutive nucleotides of the target DNA; optionally, the target sequence comprises about 16 to about 22 consecutive nucleotides, or about 17 to about 22 consecutive nucleotides of the target DNA; optionally, the target sequence comprises about 20 consecutive nucleotides of the target DNA.

10. The system of any one of claims 7-9, wherein the inverse complementary sequence of the target sequence is immediately adjacent to the 3' end of the prototype spacer neighboring motif (PAM); optionally, the PAM is 5'-NGG, wherein N is A, T, G or C.

11. The system of any one of claims 7-10, wherein the DNA targeting sequence is about or at least about 16 consecutive nucleotides in length, for example, about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, ...). 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more consecutive nucleotides, optionally about 15-70 consecutive nucleotides; optionally about 16 to about 22 consecutive nucleotides, or about 17 to about 22 consecutive nucleotides; optionally about 20 consecutive nucleotides.

12. The system of any one of claims 7-11, wherein the target DNA is dsDNA, optionally, the target DNA is present in bacterial cells, archaea cells, single-celled eukaryotes, plant cells, invertebrate cells or vertebrate cells.

13. A nucleic acid molecule, said nucleic acid molecule encoding a programmable nuclease as described in any one of claims 1-6.

14. A nucleic acid molecule comprising one or more polynucleotide sequences encoding a guiding nucleic acid, wherein the guiding nucleic acid comprises: (i) a DNA targeting sequence comprising a nucleotide sequence complementary to a target sequence in the target DNA; and (ii) A protein-binding sequence that can bind to a programmable nuclease as described in any one of claims 1-6 to form a complex; Optionally, the one or more polynucleotide sequences encoding the guiding nucleic acid are operatively linked to a promoter that functions in eukaryotic cells.

15. The nucleic acid molecule of claim 14, wherein the promoter is selected from constitutive promoters, inducible promoters, ubiquitous promoters, cell type-specific promoters, or tissue-specific promoters.

16. The nucleic acid molecule of claim 14 or 15, wherein the target DNA is dsDNA, optionally, the target DNA is present in bacterial cells, archaea cells, single-celled eukaryotes, plant cells, invertebrate cells, or vertebrate cells.

17. The nucleic acid molecule of claim 16, wherein the DNA targeting sequence is about or at least about 16 consecutive nucleotides in length, for example, about or at least about 15-100 (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 6...). 0, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100) or more consecutive nucleotides, optionally about 15-70 consecutive nucleotides; optionally about 16 to about 22 consecutive nucleotides, or about 17 to about 22 consecutive nucleotides; optionally about 20 consecutive nucleotides.

18. The nucleic acid molecule of claim 17, wherein the DNA targeting sequence is at least 70%, about 80%, about 90%, about 99%, or 100% complementary to the target DNA.

19. A vector comprising the polynucleotide of claim 13; optionally, wherein the vector encodes a guide nucleic acid as defined in any one of claims 14-18; optionally, wherein the vector is a plasmid vector, a recombinant AAV (rAAV) vector, or a recombinant lentiviral vector.

20. A ribonucleoprotein comprising a programmable nuclease as described in any one of claims 1-6 and a guiding nucleic acid as defined in any one of claims 14-18.

21. A lipid nanoparticle comprising a programmable nuclease according to any one of claims 1-6 or a system according to any one of claims 7-12.

22. A cell, said cell comprising (1) The system according to any one of claims 7-12; (2) The nucleic acid molecule of claim 13 and the nucleic acid molecules of claims 14-18; (3) The carrier according to claim 19; (4) The ribonucleoprotein of claim 20; or (5) The lipid nanoparticles according to claim 21.

23. A pharmaceutical composition comprising (1) the system of any one of claims 7-12, the carrier of claim 19, the ribonucleoprotein of claim 20, or the lipid nanoparticles of claim 21, or the cell of claim 22; and (2) a pharmaceutically acceptable excipient.

24. A method for targeting, binding, cleaving, labeling, or modifying target DNA, the method comprising contacting the target DNA with a complex, the complex comprising: (1) The programmable nuclease or the polynucleotide encoding the programmable nuclease according to any one of claims 1-6; as well as The guiding nucleic acid or the polynucleotide encoding the guiding nucleic acid, the guiding nucleic acid comprising: (i) a protein-binding sequence capable of forming a complex with the programmable nuclease; and (ii) It is capable of hybridizing with the target sequence of the target DNA, thereby guiding the complex to the DNA target sequence of the target DNA.

25. A method for diagnosing, preventing, or treating a disease in a subject in need, the method comprising administering to the subject the system of any one of claims 7-12, the carrier of claim 19, the ribonucleoprotein of claim 20, the lipid nanoparticle of claim 21, the cell of claim 22, or the pharmaceutical composition of claim 23, wherein the disease is associated with a target DNA, wherein a spacer sequence is capable of hybridizing with a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnoses, prevents, or treats the disease.

26. The method of claim 25, wherein the disease is selected from hypercholesterolemia, familial hypercholesterolemia (FH), atherosclerosis, Angelman's syndrome (AS), Alzheimer's disease (AD), transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, progressive pseudohypertrophic muscular dystrophy, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), α-1-antitrypsin deficiency, primary hyperoxaluria, Pompe disease, and myotonic dystrophy. Adverse diseases, Huntington's disease (HTT), Fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, thalassemia (e.g., β-thalassemia), Parkinson's disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, non-alcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), heart disease (e.g., hypertrophic cardiomyopathy (HCM)), hepatitis B, and cancer.

27. A method for detecting target DNA, the method comprising contacting the target DNA with a system according to any one of claims 7-12, wherein the target DNA is modified by the complex, and wherein the modification is detected in the target DNA; optionally, wherein the modification generates a detectable signal, such as a fluorescence signal.