Class 2 Type II CRISPR systems
By developing a small molecular weight SMART nuclease system, the delivery difficulties caused by the large size of Class 2 Cas effectors have been solved, improving their feasibility and efficiency in therapeutic applications.
Patent Information
- Application Number
- JP2024144487
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-19
- Filing Date
- 2024-08-26
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2041-03-30
Smart Images

Figure 0007785139000004 
Figure 0007785139000005 
Figure 0007785139000006
Abstract
Description
[Technical Field]
[0001] <Cross reference> This application claims the benefit of U.S. Provisional Patent Application No. 63 / 116,149, filed November 19, 2020, and entitled "CLASS II, TYPE II CRISPR SYSTEMS," and U.S. Provisional Patent Application No. 63 / 003,159, filed March 31, 2020, and entitled "CLASS II, TYPE II CRISPR SYSTEMS," both of which are incorporated herein in their entireties.
[0002] <Sequence Listing> This application contains a Sequence Listing, which has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy was created on March 27, 2021, has the file name 55921-711_601_SL.txt, and is 2,235,526 bytes in size. [Background technology]
[0003] Cas enzymes, along with their associated clustered regularly interspaced short palindromic repeats (CRISPR) guide ribonucleic acid (RNA), are widespread components of prokaryotic immune systems (~45% of bacteria, ~84% of archaea) and appear to play a role in protecting such microorganisms from non-self nucleic acids, such as infectious viruses and plasmids, through CRISPR-RNA-guided nucleic acid cleavage. While deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid-interacting domains. While CRISPR DNA elements were observed as early as 1987, the programmable endonucleolytic cleavage capabilities of CRISPR / Cas complexes were recognized more recently, leading to the use of recombinant CRISPR / Cas systems in diverse DNA manipulation and gene editing applications. The utility of these enzymes has led to their repurposing in a wide variety of bioengineering, gene editing, and therapeutic applications. Due to their single-effector architecture, the majority of systems currently being repurposed for genome engineering belong to the CRISPR class 2 type II and class 2 type V categories. Summary of the Invention
[0004] The large size (greater than approximately 1200 amino acids) of many Class 2 Cas effectors makes their delivery for therapeutic applications challenging. Thus, described herein are methods, compositions, and systems related to novel putative guided dsDNA nucleases, termed SMART (Small ARchaeal-associated) nuclease systems. These endonuclease effectors are defined by their small size (400 aa-1050 aa), the presence of RuvC and HNH catalytic domains, and other predicted protein features that collectively suggest a novel biochemical mechanism.
[0005] In some aspects, the present disclosure provides an engineered nuclease system, the engineered nuclease system comprising: (a) an uncultivated microorganism comprising a RuvC domain and an HNH domain; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease, wherein the endonuclease has a molecular weight of approximately 96 kDa or less. In some embodiments, the endonuclease is an archaeal endonuclease. In some embodiments, the endonuclease is a Class 2 Type II Cas endonuclease. In some embodiments, the endonuclease has a sequence having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain having PF14239 homology. In some embodiments, the arginine-rich region or the domain having PF14239 homology has at least 85%, at least 90%, or at least 95% identity to an arginine-rich region or a domain having PF14239 homology in any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease further comprises an REC (recognition) domain. In some embodiments, the REC domain has at least 85%, at least 90%, or at least 95% identity to the REC domain in any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668.In some embodiments, the endonuclease further comprises a BH (bridge helix) domain, a WED (wedge) domain, and a PI (PAM interaction) domain. In some embodiments, the BH domain, the WED domain, or the PI domain has at least 85%, at least 90%, or at least 95% identity to the BH domain, the WED domain, and / or the PI domain of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668.
[0006] In some aspects, the present disclosure provides an engineered nuclease system, the engineered nuclease system comprising: (a) an endonuclease comprising a RuvC-I domain and an HNH domain; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to the endonuclease, wherein the endonuclease comprises a sequence having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease is an archaeal endonuclease. In some embodiments, the endonuclease is a Class 2 Type II Cas endonuclease. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain having PF14239 homology. In some embodiments, the arginine-rich region or the domain having PF14239 homology has at least 85%, at least 90%, or at least 95% identity to the arginine-rich region of any one of SEQ ID NOs: 1-198, 221-459, 463-612, and 617-668. In some embodiments, the endonuclease further comprises an REC (recognition) domain. In some embodiments, the REC domain has at least 85%, at least 90%, or at least 95% identity to the REC domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, and 617-668. In some embodiments, the endonuclease further comprises a BH domain, a WED domain, and a PI domain.In some embodiments, the BH domain, the WED domain, or the PI domain has at least 85%, at least 90%, or at least 95% identity to the BH domain, the WED domain, and / or the PI domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease is derived from an unculturable microorganism. In some embodiments, the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 199-200, 460-461, or 669-673, or comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 201-203 or 613-616. In some embodiments, the guide nucleic acid structure comprises a sequence having at least 80% identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 201-203, 613-616.
[0007] In some aspects, the present disclosure provides an engineered nuclease system, the engineered nuclease system comprising: (a) an engineered guide ribonucleic acid structure, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to an endonuclease, wherein the ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 199-200, 460-461, or 669-673, or a sequence having at least 80% sequence identity to a non-variable nucleotide of any one of SEQ ID NOs: 201-203 or 613-616; and (b) an RNA-guided endonuclease configured to bind to the engineered guide ribonucleic acid. In some embodiments, the RNA-guided endonuclease is an archaeal endonuclease. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less. In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some embodiments, the engineered guide ribonucleic acid structure comprises a single ribonucleic acid polynucleotide comprising the guide ribonucleic acid sequence and the tracr ribonucleic acid sequence. In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 205-220.In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template, the single-stranded or double-stranded DNA repair template comprising, from 5' to 3', a first homology arm comprising a sequence of at least 20 nucleotides 5' to the target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homology arm comprising a sequence of at least 20 nucleotides 3' to the target sequence. In some embodiments, the first homology arm or the second homology arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides. In some embodiments, the engineered nuclease system comprises Mg. 2+In some embodiments, the endonuclease and the tracr ribonucleic acid sequence are derived from distinct bacterial species within the same phylum. In some embodiments, the endonuclease comprises a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 2-24, and the guide RNA structure comprises an RNA sequence predicted to comprise a hairpin comprising a stem and a loop, wherein the stem comprises at least 12 pairs of ribonucleotides. In some embodiments, the guide RNA structure further comprises a second stem and a second loop, wherein the second stem comprises at least 5 pairs of ribonucleotides. In some embodiments, the guide RNA structure further comprises an RNA structure comprising at least two hairpins. In some embodiments, the endonuclease comprises a sequence having at least 70% sequence identity to SEQ ID NO: 1, and the guide RNA structure comprises an RNA sequence predicted to comprise at least four hairpins comprising stems and loops. In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 1, 2, 10, 17, or 613-616; and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 199-200 or 669-673, or to the non-variable nucleotides of any one of SEQ ID NOs: 201-203 or 613-616. In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 1-24, 462-488, or 501-612, and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 199-200 or 669-673, or to the non-variable nucleotides of any one of SEQ ID NOs: 201-203 or 613-616.In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 2, 10, or 17, and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of the non-variable nucleotides of SEQ ID NOs: 202-203 or 613-614. In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 25-198, 221-459, or 489-580, and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to a Class 2 Type II sgRNA or tracr sequence. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. In some embodiments, the sequence identity is determined by the BLASTP homology search algorithm using the parameters wordlength (W) of 3, expectation (E) of 10, and scoring matrix BLOSUM62 with gap costs set to existence of 11, extension of 1, and a conditional compositional score matrix adjustment. In some embodiments, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease.In some embodiments, the endonuclease has less than 80% identity to the Cas9 endonuclease.
[0008] In some aspects, the present disclosure provides a single engineered guide ribonucleic acid polynucleotide, the single engineered guide ribonucleic acid polynucleotide comprising: a) a DNA-targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other by an intervening nucleotide, wherein the engineered guide ribonucleic acid polynucleotide is configured to form a complex with an endonuclease comprising a variant having at least 75% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the DNA-targeting segment is located 5' to both of the two complementary stretches of nucleotides. In some embodiments, a) the protein-binding segment comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 199-200 or 669-673, and b) the protein-binding segment comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to a non-variable nucleotide of any one of SEQ ID NOs: 201-203 or 613-616. In some embodiments, a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 2, 10, or 17, and b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one non-variable nucleotide of SEQ ID NO: 200, or SEQ ID NOs: 202-203 or 613-614.In some embodiments, a) the endonuclease comprises a sequence at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 25-198, 221-459, or 489-580, and b) the guide RNA structure comprises a sequence at least 70%, at least 80%, or at least 90% identical to a Class 2 Type II sgRNA. In some embodiments, the endonuclease further comprises a base editor or a histone editor linked to the endonuclease. In some embodiments, the base editor is an adenosine deaminase. In some embodiments, the adenosine deaminase comprises ADAR1 or ADAR2. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4.
[0009] In some aspects, the present disclosure provides a deoxyribonucleic acid polynucleotide that encodes any of the engineered guide ribonucleic acid polynucleotides described herein.
[0010] In some aspects, the present disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes a Class 2 Type II Cas endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from a fastidious microorganism, and wherein the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, 60 kDa or less, or 30 kDa or less. In some embodiments, the endonuclease comprises SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668, or a variant having at least 70% sequence identity thereto. In some embodiments, the endonuclease further comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N- or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 205-220. In some embodiments, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human. In some embodiments, the organism is a prokaryote or bacterium, and is a different organism from the organism from which the endonuclease is derived. In some embodiments, the organism is not a fastidious microorganism.
[0011] In some aspects, the present disclosure provides a vector comprising a nucleic acid sequence encoding an RNA-guided endonuclease comprising a RuvC-I domain and an HNH domain, wherein the endonuclease is derived from a fastidious microorganism, and wherein the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less, and wherein the RNA-guided endonuclease is optionally archaeal. In some embodiments, the endonuclease further comprises an arginine-rich region comprising an RRxRR motif or a domain with PF14239 homology. In some embodiments, the endonuclease further comprises an REC (recognition) domain. In some embodiments, the endonuclease further comprises a BH domain, a WED domain, and a PI domain.
[0012] In some aspects, the present disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: a) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and b) a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.
[0013] In some aspects, the present disclosure provides a cell comprising any of the vectors described herein. In some embodiments, the cell is a bacterial, archaeal, fungal, eukaryotic, mammalian, or plant cell. In some embodiments, the cell is a bacterial cell.
[0014] In some aspects, the present disclosure provides a method of producing an endonuclease, the method comprising culturing any of the cells described herein.
[0015] In some aspects, the disclosure provides methods for binding, cleaving, labeling, or modifying a double-stranded deoxyribonucleic acid polynucleotide, the methods comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with a Class 2 Type II Cas endonuclease complexed with an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide; and (b) the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), wherein the endonuclease has a molecular weight of about 120 kDa or less, 100 kDa or less, 90 kDa or less, or 60 kDa or less. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide, wherein the PAM comprises NGG. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide 6 to 8 nucleotides or 7 nucleotides from the PAM. In some embodiments, the endonuclease comprises a variant having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668.
[0016] In some aspects, the present disclosure provides methods for binding, cleaving, labeling, or modifying double-stranded deoxyribonucleic acid polynucleotides, the methods comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with an RNA-guided archaeal endonuclease complexed with an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and wherein the endonuclease comprises a variant having at least 70%, at least 75%, at least 80%, or at least 90% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide, wherein the PAM comprises NGG. In some embodiments, the endonuclease cleaves the double-stranded deoxyribonucleic acid polynucleotide 6 to 8 nucleotides or 7 nucleotides from the PAM. In some embodiments, the Class 2 Type II Cas endonuclease is not Cas9 endonuclease, Cas14 endonuclease, Cas12a endonuclease, Cas12b endonuclease, Cas12c endonuclease, Cas12d endonuclease, Cas12e endonuclease, Cas13a endonuclease, Cas13b endonuclease, Cas13c endonuclease, or Cas13d endonuclease. In some embodiments, the Class 2 Type II Cas endonuclease is derived from a fastidious microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, bacterial, eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, or bacterial double-stranded deoxyribonucleic acid polynucleotide from a species other than the species from which the endonuclease is derived.
[0017] In some aspects, the present disclosure provides a method for modifying a target nucleic acid locus, the method comprising delivering any of the engineered nuclease systems described herein to the target nucleic acid locus, wherein the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure, wherein the complex is configured to modify the target nucleic acid locus upon binding to the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or labeling the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic eukaryotic DNA, archaeal DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid comprises bacterial DNA, wherein the bacterial DNA is from a bacterial or archaeal species different from the species from which the endonuclease is derived. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is within a cell. In some embodiments, the endonuclease and the engineered guide nucleic acid structure are encoded by separate nucleic acid molecules. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, an archaeal cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, the cell is from a species different from the species from which the endonuclease is derived. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering any of the nucleic acids described herein or any of the vectors described herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease.In some embodiments, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing the open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid structure operably linked to a ribonucleic acid (RNA) pol III promoter. In some embodiments, the endonuclease creates a single-strand or double-strand break at or proximal to the target locus. In some embodiments, the endonuclease creates a double-strand break 5' from a protospacer adjacent motif (PAM) and proximal to the target locus. In some embodiments, the endonuclease causes a double-stranded break 6 to 8 nucleotides, or 7 nucleotides 5', from the PAM. In some embodiments, the engineered nuclease system causes a chemical modification of a nucleotide base within or proximal to the target locus, or causes a chemical modification of a histone within or proximal to the target locus. In some embodiments, the chemical modification is deamination of an adenosine or cytosine nucleotide. In some embodiments, the endonuclease further comprises a base editor linked to the endonuclease. In some embodiments, the base editor is an adenosine deaminase. In some embodiments, the adenosine deaminase comprises ADAR1 or ADAR2. In some embodiments, the base editor is a cytosine deaminase.In some embodiments, the cytosine deaminase comprises APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4.
[0018]
[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure have been shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its various details can be modified in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0019] <Incorporated by reference> All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]
[0020] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figures" and "FIGs.")
[0021] [Figure 1]Dendrograms showing the homology relationships of various classes and types of CRISPR / Cas loci are presented. Here, SMART I and II Cas enzyme classes are described in comparison with class 2 type II-A, II-B, and II-C Cas systems, demonstrating that these systems group into separate classes rather than type II-A, II-B, and II-C. (A) shows the SMART phylogenetic tree in the context of the Cas9 reference sequence, where SMART effectors cluster far from the Cas9 reference sequence (type II-A, II-B, and II-C). (B) shows the SMART phylogenetic tree illustrating subgroups of SMART enzymes. [Figure 2] Figure 1 shows the length distribution of the SMART effectors described herein, demonstrating that SMART I and II enzymes cluster at lower molecular weights than Cas9-like enzymes. SMART nucleases show a bimodal distribution, with one peak at about 400 aa (SMART II) and a second peak at about 750 aa (SMART I). Cas9 nucleases also show a bimodal distribution, with peaks at about 1,100 aa (e.g., SaCas9) and 1,300 aa (e.g., SpCas9). [Figure 3]Genomic context of the "small" type II nucleases MG33-1 and MG35-236. SMART nucleases and CRISPR accessory proteins are shown as dark gray arrows, while other genes are shown as light gray arrows. Predicted domains for all genes in the genomic fragment are shown as gray boxes below the arrows. In the figure, (A) is the genomic context of the CRISPR locus encoded upstream from the SMART I nuclease MG33-1 and SMART II nuclease MG35-236, showing the predicted insertion sequence with transposases TnpA and TnpB downstream from SMART II; (B) is the genomic context of the SMART I nuclease MG34-1, where environmentally expressed sequencing reads are shown aligned below the CRISPR array and predicted tracrRNA, and transcriptome coverage for the region is illustrated above the contig sequence; (C) is the genomic context of the SMART I nuclease MG34-16, where environmentally expressed sequencing reads are shown aligned below the CRISPR array and predicted tracrRNA, and transcriptome coverage for the region is illustrated above the contig sequence; and (D) is the genomic context of the MG34-16 in the figure, where environmentally expressed sequencing reads are shown aligned below the CRISPR array and predicted tracrRNA, and transcriptome coverage for the region is illustrated above the contig sequence. Genomic fragment targeted by CRISPR array-derived spacer 7, where the genomic fragment was identified as phage-derived based on the terminase and portal of virus-specific gene annotation. The inset shows the location of MG34-16 spacer 7, which targets the C-terminus of a viral gene of unknown function; the putative NGG PAM for MG34-16 is highlighted by a gray box downstream from the spacer match. [Figure 4]Multiple sequence alignment of exemplary SMART endonucleases (MG33-1 (SEQ ID NO: 1), MG33-2 (SEQ ID NO: 463), MG33-3 (SEQ ID NO: 464), MG34-1 (SEQ ID NO: 2), MG34-9 (SEQ ID NO: 10), MG34-16 (SEQ ID NO: 17), MG102-1 (SEQ ID NO: 581), MG102-2 (SEQ ID NO: 582), MG35-1 (SEQ ID NO: 25), MG35-2 (SEQ ID NO: 26) , MG35-3 (SEQ ID NO: 27), MG35-102 (SEQ ID NO: 126), MG35-236 (SEQ ID NO: 284), MG35-419 (SEQ ID NO: 222), MG35-420 (SEQ ID NO: 223), and MG35-421 (SEQ ID NO: 224), where the sequence of SaCas9 was used as the reference domain and is shown as a rectangle below the reference sequence, and catalytic residues are shown as squares above each sequence. In the figure, (A) is an alignment of the endonuclease region encompassing the RuvC-I and bridge helix domains, (B) is an alignment of the region encompassing the RuvC-III domain, and (C) is an alignment of the region encompassing the RuvCII and HNH domains. [Figure 5] An example of the domain organization for SMART I endonucleases is presented using MG34-1 as an example. In the figure, (A) is a diagram showing the predicted domain architecture of SMART I nucleases, which consist of three RuvC domains, showing the bridge helix ("BH"), a domain with homology to Pfam PF14239, interrupted by a recognition domain ("REC"), an HNH endonuclease domain ("HNH"), a wedge domain ("WED"), and a PAM-interacting domain (PI), and (B) is an overview of the multiple sequence alignment of two SMART I nucleases against the reference Cas9 nuclease sequence, where the catalytic residues of RuvC and HNH are shown as black bars above each sequence, regions that align in 3D space with the SaCas crystal structure are represented by rounded boxes, and dashed lines represent regions of poor or no alignment in 3D space between the SMART and SaCas9 3D structure predictions. [Figure 6] An example of the domain organization for SMART II endonucleases is shown using MG35 family enzymes (MG35-3, MG35-4) as an example. (A) Diagram showing the predicted domain architecture of SMART II nucleases, consisting of three RuvC domains, a domain with homology to Pfam PF14239, an HNH endonuclease domain, an unknown domain, and a recognition domain (REC). (B) Overview of multiple sequence alignments of two SMART II nucleases against the reference Cas9 nuclease sequence, where RuvC and HNH catalytic residues are shown as black bars above each sequence, regions that align with the SaCas crystal structure in 3D space are represented by rounded boxes, and residues identified from the 3D structure prediction that may be involved in recognizing guide / target / PAM sequences are represented by dark gray boxes (within the RRXRR and REC domains) above the MG35-419 sequence. [Figure 7]
[0023] Figure 1 illustrates various features of SMART enzymes, in which (A) is a dot plot showing the identity of the SMART I domains of various enzymes described herein to that of spCas9, showing that they share up to about 35% sequence identity, and (B) is a dot plot of the lengths of the individual SMART I domains of the enzymes described herein. [Figure 8]Figure 1 illustrates the distribution of counts of various SMART-specific motifs relative to motifs predicted in Cas9 nuclease sequences, showing that these motifs are more frequently found in SMART enzymes. Motifs were predicted in 803 reference Cas9 sequences (types II-A, II-B, and II-C), 84 SMART I sequences, and 471 SMART II sequences. (A) Boxplot of the count frequencies of Zn-binding ribbon motifs (CX[2-4]C and CX[2-4]H) in various types of class 2 Cas enzymes, and (B) histogram of the count frequencies of RRXRR motifs in various types of class 2 Cas enzymes. In (A) and (B), lines track the average count values, while outliers are represented by dots. [Figure 9] Illustrated are predicted guide RNA structures of single guide RNAs (sgRNAs) designed for cleavage activity by SMART I endonucleases, where (A) is MG34-1 sgRNA 1, (B) is MG34-1 sgRNA 2, (C) is MG34-9 sgRNA 1, and (D) is MG34-16 sgRNA 1. [Figure 10]Figure 1 depicts the characterization of SMART I nuclease cleavage as described in Example 1. (A) shows an Agilent TapeStation gel of the ligation products of a cleavage assay for MG34-1 with two sgRNA designs compared to a negative control. Lane L3 is a ladder. Lane A4 is Apo, no sgRNA. Lanes B4 and C4 are the MG34-1 sgRNAs tested (sg1: SEQ ID NO: 612, sg2: 613). Cleavage product bands are labeled with arrows. Lanes G3 and H3 are grayed out and are not relevant to this experiment. (B) shows a PCR gel of the ligation products, demonstrating the activity of MG34-1, 34-9, and 34-16. Lane 1 is a ladder. Lanes 2-7 are sgRNA designs with six spacer lengths for MG34-1. Lanes 8 and 9 are sgRNA designs for 34-9 and 34-16, respectively. Arrows indicate cleavage confirmation bands. [Figure 11] Sequence cleavage preferences are illustrated for MG34 nuclease. (A) shows a SeqLogo representation of the consensus PAM sequence (NGGN) for MG34-1 with sgRNA 1 (top, SEQ ID NO: 612) and sgRNA 2 (bottom, SEQ ID NO: 613). (B) shows a histogram depicting the location of the cleavage site for MG34-1, demonstrating that MG34-1 prefers cleavage around position 7 from the PAM. (C) shows a chromatogram from Sanger sequencing, showing the NGG PAM (highlighted in a box) preferred by MG34-9. The arrow indicates the cleavage site at position 7 from the PAM. [Figure 12]Figure 1 illustrates the results of plasmid targeting experiments in E. coli for MG34-1. (A) shows replica plating of E. coli strains demonstrating plasmid cleavage. E. coli expressing MG34-1 and sgRNA were transformed with a kanamycin resistance plasmid containing a target for the sgRNA (+sp). The quadrants showing growth defect (+sp) versus negative controls (no target and PAM (-sp)) represent successful targeting and cleavage by the enzyme. The experiment was replicated twice and performed in triplicate. (B) shows a graph of colony-forming unit (cfu) measurements from the replica plating experiment in (A) showing growth inhibition in the targeting condition (+sp) versus the non-targeting control (-sp), demonstrating that the plasmid was cleaved. [Figure 13] An example of the genomic context of the SMART system is shown for MG35-419. SMART nucleases are shown as dark gray arrows, while other genes are represented as lighter gray arrows. Predicted domains for all genes in the genomic fragment are shown as gray boxes below the arrows. Sequencing reads of environmental expression are shown aligned below the CRISPR array in (A) and upstream from the effector in (B). Transcriptome coverage for the region showing expression is illustrated above the contig sequence. (A) shows the genomic context of the SMART II MG35-419 effector and nearby encoded CRISPR loci. (B) shows the genomic context of the SMART II effector MG35-3, showing the transcribed 5'UTR. [Figure 14]Figure 1 shows a predicted 3D structure for SMART II MG35-419. This 3D model aligns well with regions of the SaCas9 crystal structure, despite being less than half the size. Regions aligned with the SaCas9 template include the catalytic lobe (RuvC-I, HNH, and RuvC-III domains) and a short region of the recognition (REC) lobe. SMART II-specific domains include a domain encompassing the RRXRR motif and homology to Pfam PF14239, as well as a domain of unknown function. [Figure 15] Figure 1 shows the results of a preliminary cleavage assay for the SMART II effector. MG35-420 (SEQ ID NO: 223) protein preparation was tested for cleavage activity in TXTL extracts in which the entire locus was expressed. The experiment involved incubating a protein preparation with the PAM library (dsDNA target), the predicted repeat region in both forward and reverse orientations (fw and rv), and an intergenic region encoding a potentially required cofactor. Lanes 2-9 (non-cr array) are control experiments lacking the repeat region. Apo is the protein preparation with the target PAM library alone. Labels 1-2.5 represent seven different intergenic regions. -IG is a control without the intergenic region. A PCR gel of the ligation products shows the putative cleavage bands (arrows) suggesting dsDNA cleavage.
[0022] <Brief explanation of the sequence listing> The Sequence Listing filed herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems of the present disclosure. Below are exemplary descriptions of the sequences in the Sequence Listing.
[0023] MG33 nuclease
[0024] SEQ ID NOs: 1 and 463-486 show the full-length peptide sequence of MG33 nuclease.
[0025] SEQ ID NOs: 199 and 669-670 show the nucleotide sequences of tracrRNA predicted to function with MG33 nuclease.
[0026] SEQ ID NO: 201 shows the nucleotide sequence of a predicted single guide RNA (sgRNA) sequence predicted to function with MG33 nuclease. "N" denotes a variable residue, and non-N residues represent scaffold sequences.
[0027] MG34 nuclease
[0028] SEQ ID NOs: 2-24 and 487-488 show the full-length peptide sequences of MG1 nuclease.
[0029] SEQ ID NO: 200 shows the nucleotide sequence of an sgRNA predicted to function with MG4 nuclease.
[0030] SEQ ID NOs: 202, 203, and 613-616 show the nucleotide sequences of predicted single guide RNA (sgRNA) sequences predicted to function with MG34 nuclease. "N" denotes a variable residue, and non-N residues represent scaffold sequences.
[0031] MG35 nuclease
[0032] SEQ ID NOs: 25-198, 221-459, 489-580, and 617-668 show the full-length peptide sequences of MG35 nuclease.
[0033] SEQ ID NOs: 460-461 show the nucleotide sequences of MG35tracrRNAs, which are derived from the same locus as the MG35 nuclease.
[0034] SEQ ID NO: 462 shows the repeat of the MG35 nuclease described herein.
[0035] MG102 nuclease
[0036] SEQ ID NOs: 581-612 show the full-length peptide sequences of MG102 nuclease.
[0037] SEQ ID NOs: 672-673 show the nucleotide sequence of MG102 tracrRNA, which is derived from the same locus as MG102 nuclease.
[0038] SEQ ID NOs: 205-220 show sequences of examples of nuclear localization sequences (NLS) that can be added to nucleases according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0039] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It will be understood that various alternatives to the embodiments of the invention described herein may be utilized.
[0040] The implementation of some methods disclosed herein utilizes immunological, biochemical, chemical, molecular biology, microbiology, cell biology, genomics, and recombinant DNA techniques, unless otherwise specified. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FM Ausubel, et al. eds.); the series Methods in Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RI Freshney, ed. (2010)) (incorporated herein by reference in their entirety).
[0041] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or variations thereof are used in either the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0042] The term "about" or "approximately" means within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which error range depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean 1 or more than 1 standard deviation per practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of any value.
[0043] As used herein, "cell" generally refers to a biological cell. A cell can be the basic structural, functional, and / or biological unit of an organism. A cell may originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-celled eukaryotic cells, protozoan cells, plant cells (e.g., cells of crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, liverworts, and mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardti, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens C. Agardh, etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, invertebrate (e.g., Drosophila, cnidaria, echinoderms, nematodes, etc.) cells, vertebrate (e.g., fish, amphibians, reptiles, birds, mammals) cells, mammalian (e.g., pig, cow, goat, sheep, rodent, rat, mouse, non-human primate, human, etc.) cells, etc. Cells may not originate from a natural organism (e.g., cells may be synthetically created and sometimes referred to as artificial cells).
[0044] The term "nucleotide," as used herein, generally refers to a base-sugar-phosphate combination. Nucleotides may include synthetic nucleotides. Nucleotides may include synthetic nucleotide analogs. Nucleotides may be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term "nucleotide" may refer to dideoxyribonucleoside triphosphate (ddNTP) and its derivatives. Illustrative examples of dideoxyribonucleoside triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or may be detectably labeled, for example, by using a moiety that contains an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels.Fluorescent labels for nucleotides may include, but are not limited to, fluorescein, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer (Foster City, Calif.); FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham (Arlington Heights, Ill.); and Boehringer Mannheim (Indianapolis, Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Molecular Probes (Ind.); and Chromosome Labeled Nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, and Oregon Green available from Molecular Probes (Eugene, Oregon). Examples include 488-5-dUTP, rhodamine Green-5-UTP, rhodamine Green-5-dUTP, tetramethylrhodamine 6-UTP, tetramethylrhodamine 6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP.Nucleotides can also be labeled or marked by chemical modification. The chemically modified single nucleotide can be a biotin-dNTP. Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP). Nucleotides can include nucleotide analogs. In some embodiments, nucleotide analogs can include the structure of natural nucleotides that are modified at any position to change certain chemical properties of the nucleotide, but still retain the ability of the nucleotide analog to perform its intended function (e.g., hybridization to other nucleotides in RNA or DNA). Examples of positions of nucleotides that can be derivatized include the 5-position (e.g., 5-(2-amino)propyl uridine, 5-bromo uridine, 5-propyne uridine, 5-propenyl uridine, etc.), the 6-position (e.g., 6-(2-amino)propyl uridine), the 8-position of adenosine and / or guanosine, e.g., 8-bromo guanosine, 8-chloro guanosine, 8-fluoroguanosine, etc.Nucleotide analogs also include deazanucleotides, e.g., 7-deaza-adenosine, O- and N-modified (e.g., alkylated, e.g., N-methyl adenosine, otherwise known in the art) nucleotides, and other heterocyclically modified nucleotide analogs, such as those described in Herdewijn, Antisense Nucleic Acid Drug Dev., 2000 Aug. 10(4):297-310. Nucleotide analogs may also include modifications to the sugar portion of the nucleotide. For example, the 2'OH group may be replaced with a group selected from H, OR, R, F, Cl, Br, I, SH, SR, NH, NHR, NR, COOR, or OR, where R is a substituted or unsubstituted C-C alkyl, alkenyl, alkynyl, aryl, etc. Other possible modifications include those described in U.S. Pat. Nos. 5,858,988 and 6,291,438. Examples of nucleotide positions that can be derivatized include the 5-position, such as 5-(2-amino)propyluridine, 5-bromouridine, 5-propyneuridine, 5-propenyluridine, etc., the 6-position, such as 6-(2-amino)propyluridine, and the 8-position of adenosine and / or guanosine, such as 8-bromoguanosine, 8-chloroguanosine, 8-fluoroguanosine, etc. Nucleotide analogs also include deazanucleotides, such as 7-deaza-adenosine, O- and N-modified (e.g., alkylated, e.g., N-6-methyl adenosine, otherwise known in the art) nucleotides, and other heterocyclically modified nucleotide analogs, such as those described in Herdewijn, Antisense Nucleic Acid Drug Dev., 2000 Aug. 10(4):297-310. Nucleotide analogs may also include modifications to the sugar moiety of the nucleotide.For example, the 2'OH group may be replaced with a group selected from H, OR, R, F, Cl, Br, I, SH, SR, NH, NHR, NR, COOR, or OR, where R is a substituted or unsubstituted C-C alkyl, alkenyl, alkynyl, aryl, etc. Other possible modifications include those described in U.S. Patent Nos. 5,858,988 and 6,291,438.
[0045] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are generally used interchangeably to refer to a polymeric form of nucleotides of any length (either deoxyribonucleotides or ribonucleotides), or analogs thereof, in either single-stranded, double-stranded, or multi-stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can be present in a cell-free environment. A polynucleotide can be a gene or a fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three-dimensional structure. A polynucleotide may contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, xenonucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queusine, and wyosine.Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of nucleotides may be interrupted by non-nucleotide components.
[0046] The terms "transfection" or "transfected" generally refer to the introduction of nucleic acid into a cell, either by non-viral or viral-based methods. The nucleic acid molecule can be a genetic sequence encoding a complete protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88 (incorporated herein by reference in its entirety).
[0047] The terms "peptide," "polypeptide," and "protein" are generally used interchangeably herein to refer to a polymer of at least two amino acid residues linked by a peptide bond. The term does not imply a particular length of the polymer, and is not intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical synthesis, enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins, and proteins with or without secondary and / or tertiary structure (e.g., domains). The term also encompasses amino acid polymers modified, for example, by disulfide bond formation, glycosylation, lipid modification, acetylation, phosphorylation, oxidation, and other manipulations, such as conjugation with a labeling moiety. The term "amino acid," as used herein, generally refers to natural amino acids and non-natural amino acids, including modified amino acids and amino acid analogs. Modified amino acids can include natural and unnatural amino acids, which are chemically modified to include groups or chemical moieties not naturally found on amino acids. Amino acid analogs can also refer to amino acid derivatives. The term "amino acid" includes both D- and L-amino acids.
[0048] As used herein, the term "non-naturally occurring" refers to a nucleic acid or polypeptide sequence that is not typically found in naturally occurring nucleic acids or proteins. Non-naturally occurring may refer to an affinity tag. Non-naturally occurring may refer to a fusion. Non-naturally occurring may refer to a naturally occurring nucleic acid or polypeptide sequence that contains mutations, insertions, and / or deletions. A non-naturally occurring sequence may exhibit and / or encode an activity (e.g., enzymatic activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.) that may be exhibited by the nucleic acid and / or polypeptide sequence to which it is fused. A non-naturally occurring nucleic acid or polypeptide sequence may be joined by genetic engineering to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to generate a chimeric nucleic acid and / or polypeptide sequence encoding the chimeric nucleic acid and / or polypeptide.
[0049] The term "promoter," as used herein, typically refers to a regulatory DNA region that controls the transcription or expression of a gene and may be located adjacent to or overlapping the nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often referred to as transcription factors, which promote the binding of RNA polymerase to DNA and cause gene transcription. A "basal promoter," also referred to as a "core promoter," typically refers to a promoter that contains all the basic elements necessary to promote the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters typically, but not necessarily, contain a TATA box and / or a CAAT box.
[0050] The term "expression," as used herein, generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcript) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide are sometimes collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in eukaryotic cells.
[0051] As used herein, "operably linked," "operably linked," or "operably linked," or grammatical equivalents thereof, generally refer to the juxtaposition of genetic elements, e.g., promoters, enhancers, polyadenylation sequences, etc., in a relationship permitting these elements to function in their expected manner. For example, a regulatory element, which may include a promoter and / or enhancer sequence, is operably linked to a coding region if the regulatory element helps initiate transcription of the coding sequence. There may be intervening residues between the regulatory element and the coding region so long as this functional relationship is maintained.
[0052] As used herein, a "vector" generally refers to a polymer or an association of polymers that contains or associates with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally contains genetic elements, such as regulatory elements, operably linked to a gene to promote expression of the gene in a target.
[0053] As used herein, "expression cassette" and "nucleic acid cassette" are generally used interchangeably to refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to a combination of regulatory elements and genes to which they are operably linked for expression.
[0054] A "functional fragment" of a DNA or protein sequence generally refers to a fragment that retains a biological activity (functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be its ability to affect expression in a manner known to be attributed to the full-length sequence.
[0055] As used herein, "engineered" generally refers to an object that has been modified by human intervention. By way of non-limiting example, a nucleic acid may be modified by changing its sequence to one that does not occur in nature, a nucleic acid may be modified by ligating it to a nucleic acid with which it is not naturally associated so that the ligated product possesses a function not present in the original nucleic acid, an engineered nucleic acid may be synthesized in vitro with a sequence that does not occur in nature, a protein may be modified by changing its amino acid sequence to a sequence that does not occur in nature, and an engineered protein may acquire a new function or property. An "engineered" system includes at least one engineered component.
[0056] As used herein, the term "optimally aligned" generally refers to an alignment of two amino acid sequences that exhibits the highest percent identity score or maximizes the number of matching residues.
[0057] As used herein, "synthetic" and "artificial" are used interchangeably to refer to proteins or domains thereof that have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to naturally occurring human proteins. For example, the VPR and VP64 domains are synthetic transactivation domains.
[0058] The term "tracrRNA" or "tracr sequence," as used herein, may generally refer to a nucleic acid having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, Staphylococcus aureus, etc., or SEQ ID NOs: 5476-5511) and / or a sequence similar to that wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, S. aureus, etc., or SEQ ID NOs: 199-203). A tracrRNA may refer to a nucleic acid having up to about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity to and / or similar to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from Streptococcus pyogenes, Staphylococcus aureus, etc.). A tracrRNA may also refer to modified forms of a tracrRNA that contain nucleotide changes, such as deletions, insertions, or substitutions, variants, mutations, or chimeras. A tracrRNA may also refer to a nucleic acid that is at least about 60% identical to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from Streptococcus pyogenes, Staphylococcus aureus, etc.) over a stretch of at least six contiguous nucleotides. For example, the tracrRNA sequence is at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, or 100% identical to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from Streptococcus pyogenes, Staphylococcus aureus, etc.) over a stretch of at least six contiguous nucleotides. Type II tracrRNA sequences can be predicted on a genomic sequence by identifying regions that have complementarity to portions of the repeat sequences in adjacent CRISPR arrays.
[0059] As used herein, "guide nucleic acid" may generally refer to a nucleic acid that can hybridize to another nucleic acid. A guide nucleic acid may be RNA. A guide nucleic acid may be DNA. A guide nucleic acid may be programmed to site-specifically bind to a nucleic acid sequence. A targeted or target nucleic acid may comprise nucleotides. A guide nucleic acid may comprise nucleotides. A portion of a target nucleic acid may be complementary to a portion of a guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to and hybridizes with a guide nucleic acid may be referred to as a complementary strand. A strand of a double-stranded target polynucleotide that is complementary to a complementary strand and therefore may not be complementary to the guide nucleic acid may be referred to as a noncomplementary strand. A guide nucleic acid may comprise one polynucleotide strand and may be referred to as a single guide nucleic acid. A guide nucleic acid may comprise two polynucleotide strands and may be referred to as a double guide nucleic acid. Unless otherwise specified, the term "guide nucleic acid" is inclusive and can refer to both single-guide and double-guide nucleic acids. A guide nucleic acid can include a segment, sometimes referred to as a "nucleic acid targeting segment" or "nucleic acid targeting sequence." A nucleic acid targeting segment can include a subsegment, sometimes referred to as a "protein binding segment," "protein binding sequence," or "Cas protein binding segment."
[0060] The terms "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences generally refer to two (e.g., pairwise alignment) or more (e.g., multiple sequence alignment) sequences that are the same or have a specified percentage of the same amino acid residues or nucleotides, when compared or aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP, using parameters of the BLOSUM62 scoring matrix, which sets a word length (W) of 3, an expectation (E) of 10, and an existence of 11, a gap cost of 1, and a conditional compositional score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP, using parameters of the PAM30 scoring matrix, which sets a word length (W) of 2, an expectation (E) of 1,000,000, and a gap cost of 9 for opening a gap and 1 for extending a gap for sequences shorter than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); or a match of 2, a mismatch of -1, and a gap of -1. These include CLUSTALW using Smith-Waterman homology search algorithm parameters; MUSCLE using default parameters; MAFFT using parameters of 2 retree and 1000 maxiterations; Novafold using default parameters; and HMMER hmmalign using default parameters.
[0061] As used herein, the term "RuvC_III domain" generally refers to the third, non-contiguous segment of the RuvC endonuclease domain (the RuvC nuclease domain, which is composed of three non-contiguous segments: RuvC_I, RuvC_II, and RuvC_III). RuvC domains or segments thereof can generally be identified by alignment to known domain sequences, structural alignment to proteins with annotated domains, or by comparison with hidden Markov models (HMMs) constructed based on known domain sequences (e.g., Pfam HMM PF18541 for RuvC_III).
[0062] As used herein, the term "HNH domain" generally refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to known domain sequences, structural alignment to proteins with annotated domains, or by comparison with a hidden Markov model (HMM) constructed based on known domain sequences (e.g., Pfam HMM PF01844 for the HNH domain).
[0063] As used herein, the term "bridge helix domain" or "BH domain" generally refers to an arginine-rich helical domain present in Cas enzymes that plays a key role in binding target DNA and simultaneously generating cleavage activity.
[0064] As used herein, the term "recognition domain" or "REC domain" generally refers to the domain that is thought to interact with the repeat:anti-repeat duplex of a gRNA to mediate the formation of a Cas endonuclease / gRNA complex.
[0065] As used herein, the term "wedge domain" or "WED domain" generally refers to a fold generally comprising a twisted five-stranded beta sheet flanked by four alpha helices, and generally responsible for the recognition of distorted repeat:anti-repeat duplexes for Cas enzymes. The WED domain may be responsible for the recognition of the scaffold of a single guide RNA.
[0066] As used herein, the term "PAM-interacting domain" or "PI domain" generally refers to a domain found in Cas enzymes that is positioned in an endonuclease-DNA complex to recognize a PAM sequence in the non-complementary DNA strand of a guide RNA.
[0067] <Summary>
[0068] The discovery of new Cas enzymes with unique functions and structures offers the potential to further disrupt deoxyribonucleic acid (DNA) editing technologies, improving their speed, specificity, functionality, and ease of use. Given the predicted prevalence of clustered regularly interspaced short palindromic repeats (CRISPR) systems in microorganisms and the vast diversity of microbial species, relatively few functionally characterized CRISPR / Cas enzymes exist in the literature. This is in part due to the fact that the vast number of microbial species may not be easily cultivated under laboratory conditions. Metagenomic sequencing from natural environmental niches representing many microbial species will exponentially increase the number of known new CRISPR / Cas systems and potentially facilitate the discovery of novel oligonucleotide editing functions. A recent example of the utility of such an approach is illustrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.
[0069] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repeat sequences (30-40 bp) separated by an equally short spacer sequence that encodes an RNA-based targeting element, and (ii) an ORF encoding a Cas, which encodes a nuclease polypeptide directed by the RNA-based targeting element along with accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (the target seed) and the crRNA guide and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (the PAM is generally a sequence not commonly represented in the host genome). Depending on the exact function and organization of the system, CRISPR-Cas systems are generally organized into two classes, five types, and 16 subtypes based on shared functional properties and evolutionary similarities.
[0070] Class I CRISPR-Cas systems have large multi-subunit effector complexes and include types I, III, and IV.
[0071] Type I CRISPR-Cas systems are considered to be of intermediate complexity in terms of components. In type I CRISPR-Cas systems, an array of RNA-targeting elements is transcribed as a long precursor crRNA (pre-crRNA) that is processed at the repeat element, liberating a short, mature crRNA. This short, mature crRNA, when followed by an appropriate short consensus sequence called a protospacer adjacent motif (PAM), directs a nuclease complex to the nucleic acid target. This processing occurs via the endoribonuclease subunit (Cas6) of a larger endonuclease complex called Cascade, which further contains the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Cas I nuclease primarily functions as a DNA nuclease.
[0072] Type III CRISPR systems may be characterized by the presence of a central nuclease known as Cas10, along with a repeat-associated mysterious protein (RAMP) containing Csm or Cmr protein subunits. As in type I systems, mature crRNA is processed from the pre-crRNA using enzymes such as Cas6. Unlike type I and type II systems, type III systems appear to target and cleave DNA-RNA duplexes (such as the DNA strand used as a template for RNA polymerase).
[0073] Type IV CRISPR-Cas systems have an effector complex consisting of a highly reduced large subunit nuclease (csf1) and two genes for RAMP proteins of the Cas5 (csf3) and Cas7 (csf2) families, and optionally a gene for a predicted small subunit; such systems are typically found on endogenous plasmids.
[0074] Class II CRISPR-Cas systems generally have a single polypeptide multidomain nuclease effector and include types II, V, and VI.
[0075] Type II CRISPR-Cas systems are considered the simplest in terms of components. In type II CRISPR-Cas systems, processing of the CRISPR array into mature crRNA does not require the presence of a specialized endonuclease subunit, but rather a small trans-encoded crRNA (tracrRNA) with a region complementary to the array repeat sequence. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure. This precursor dsRNA structure is cleaved by endogenous RNAse III to generate the mature effector enzyme loaded with both the tracrRNA and crRNA. Cas II nucleases are known as DNA nucleases. Type II effectors generally exhibit an architecture consisting of a RuvC-like endonuclease domain adopting an RNase H fold with an unrelated HNH nuclease domain inserted within the RuvC-like nuclease domain fold. The RuvC-like domain is responsible for cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is responsible for cleavage of the displaced DNA strand.
[0076] Type V CRISPR-Cas systems feature a nuclease effector (e.g., Cas12) structure similar to that of type II effectors, including a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use a tracrRNA to process pre-crRNA into mature crRNA. However, unlike type II systems, which require RNAse III to cleave pre-crRNA into multiple crRNAs, type V systems can use the effector nuclease itself to cleave pre-crRNA. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems are also known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to have robust single-stranded, nonspecific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.
[0077] Type VI CRIPSR-Cas systems possess an RNA-guided RNA endonuclease. Instead of a RuvC-like domain, the single polypeptide effector of type VI systems (e.g., Cas13) contains two HEPN ribonuclease domains. Unlike both type II and type V systems, type VI systems do not appear to require a tracrRNA to process pre-crRNA into crRNA. However, like type V systems, some type VI systems (e.g., C2C2) appear to possess robust single-stranded nonspecific nuclease (ribonuclease) activity that is activated by the initial crRNA-directed cleavage of the target RNA.
[0078] Due to their simpler structure, class II CRISPR-Cas have been most widely adopted for engineering and development as designer nuclease / genome editing applications.
[0079] One of the earliest adaptations of such a system for in vitro use can be found in Jinek et al. (Science. 2012 Aug 17;337(6096):816-21, incorporated herein by reference in its entirety). Jinek's study involved (i) recombinantly expressed and purified full-length Cas9 (e.g., a Class II, Type II Cas enzyme) isolated from S. pyogenes SF370, (ii) purified mature ∼42 nt crRNA (the entire crRNA transcribed in vitro from a synthetic DNA template with a T7 promoter sequence) with a ∼20 nt 5' sequence complementary to the target DNA sequence desired to be cleaved, followed by a 3' tracr binding sequence, (iii) purified tracrRNA transcribed in vitro from a synthetic DNA template with a T7 promoter sequence, and (iv) Mg 2+ A system comprising (ii) a crRNA and (iii) a nucleotide sequence was first described. Jinek subsequently described an improved, engineered system in which (ii) a crRNA is joined to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fused synthetic guide RNA (sgRNA) that can itself direct Cas9 to a target (compare the top and bottom panels of Figure 2).
[0080] Mali et al. (Science. 2013 Feb 15; 339(6121):823-826.) (which is hereby incorporated by reference in its entirety) subsequently adapted this system for use in mammalian cells by providing a DNA vector encoding (i) an ORF encoding codon-optimized Cas9 (e.g., a class II type II Cas enzyme) under a suitable mammalian promoter having a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal), and (ii) an ORF encoding an sgRNA (having a 5' sequence starting with G, followed by 20 nt of a complementary targeting nucleic acid sequence, a 3' tracr-binding sequence bound thereto, a linker, and a tracrRNA sequence) under a suitable polymerase III promoter (e.g., U6 promoter).
[0081] <MG enzyme>
[0082] In certain aspects, the present disclosure provides an engineered nuclease system. The engineered nuclease system may include (a) an endonuclease. In some cases, the endonuclease includes a RuvC domain and an HNH domain. The endonuclease may be derived from a fastidious microorganism. The endonuclease may be a Cas endonuclease. The endonuclease may be a Class 2 endonuclease. The endonuclease may be a Class 2 Type II Cas endonuclease. The engineered nuclease system may include (b) an engineered guide ribonucleic acid structure. The engineered guide ribonucleic acid structure may be configured to form a complex with the endonuclease. In some cases, the engineered guide ribonucleic acid structure configured to form a complex with the endonuclease includes a guide ribonucleic acid sequence. The guide ribonucleic acid sequence may be configured to hybridize to a target deoxyribonucleic acid sequence. In some cases, the engineered guide ribonucleic acid structure configured to form a complex with an endonuclease comprises a tracr ribonucleic acid sequence. The tracr ribonucleic acid sequence may be configured to bind to the endonuclease. In some cases, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 100 kDa or less, about 90 kDa or less, about 80 kDa or less, about 70 kDa or less, about 60 kDa or less, about 50 kDa or less, about 40 kDa or less, about 30 kDa or less, about 20 kDa or less, or about 10 kDa or less.
[0083] In some cases, the endonuclease comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668.
[0084] In certain aspects, the present disclosure provides an engineered nuclease system. The engineered nuclease system may include (a) an endonuclease. The endonuclease may include a RuvC-1 domain or a RuvC domain. The endonuclease may include an HNH domain. The endonuclease may include a RuvC-1 domain and an HNH domain. The endonuclease may be a Cas endonuclease. The endonuclease may be a Class 2 endonuclease. The endonuclease may be a Class 2 Type II Cas endonuclease. The engineered nuclease system may include (b) an engineered guide ribonucleic acid. The engineered guide ribonucleic acid structure may be configured to form a complex with the endonuclease. The engineered guide ribonucleic acid structure configured to form a complex with the endonuclease may include a guide ribonucleic acid sequence. The guide ribonucleic acid sequence may be configured to hybridize to a target deoxyribonucleic acid sequence. The engineered guide ribonucleic acid structure configured to form a complex with an endonuclease can include a tracr ribonucleic acid sequence. The tracr ribonucleic acid sequence can be configured to bind to an endonuclease. The endonuclease can include a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least about 99% sequence identity to any one of 1-198, 221-459, 463-612, or 617-668. The endonuclease can be an archaeal endonuclease. The endonuclease can be a Class 2, Type II Cas endonuclease.The endonuclease may contain an arginine-rich region containing an RR motif or a domain with PF14239 homology. The arginine-rich region or domain having PF14239 homology may comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to an arginine-rich region or domain having PF14239 homology of any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. The domain boundaries of the arginine-rich domain or the domain with PF14239 homology can be identified by optimal alignment to MG34-1 or MG34-9. The endonuclease may contain a REC domain. The REC domain may comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least about 99% sequence identity to the REC domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. The domain boundaries of the REC domain can be identified by optimal alignment with MG34-1 or MG34-9. The endonuclease may comprise a BH (bridge helix) domain.The BH domain can comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least about 99% sequence identity to the BH domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. Domain boundaries of the BH domain can be identified by optimal alignment to MG34-1 or MG34-9.
[0085] The endonuclease may comprise a WED (wedge) domain. The WED domain may comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least about 99% sequence identity to the WED domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. The domain boundaries of the WED domain can be identified by optimal alignment to MG34-1 or MG34-9. The endonuclease may comprise a PI (PAM-interacting) domain. The PI domain may comprise a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least about 99% sequence identity to the PI domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668. The domain boundaries of the PI domain can be identified by optimal alignment to MG34-1 or MG34-9.
[0086] In some cases, the endonuclease is derived from a fastidious microorganism. In some cases, the tracr ribonucleic acid sequence has at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to at least 50, at least 60, at least 70, or at least 80 contiguous nucleotides from any one of SEQ ID NOs: 199-200, 460-461, or 669-673. or a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to at least 50, at least 60, at least 70, at least 80 contiguous nucleotides of any one of SEQ ID NOs: 201-203 or 613-616.
[0087] In some cases, the guide nucleic acid structure comprises SEQ ID NO: 201. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 202. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 203. In some cases, the guide nucleic acid structure comprises SEQ ID NOs: 201-203. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 613. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 614. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 615. In some cases, the guide nucleic acid structure comprises SEQ ID NO: 616.
[0088] In certain aspects, the present disclosure provides an engineered nuclease system. The engineered nuclease system may include (a) an engineered guide ribonucleic acid structure. The engineered guide ribonucleic acid structure may include a guide ribonucleic acid sequence. The guide ribonucleic acid sequence may be configured to hybridize to a target deoxyribonucleic acid sequence. The engineered guide ribonucleic acid structure may include a tracr ribonucleic acid sequence. The tracr ribonucleic acid sequence may be configured to bind to an endonuclease. Optionally, the tracr ribonucleic acid sequence comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to at least 50, at least 60, at least 70, at least 80 contiguous nucleotides from any one of SEQ ID NOS: 199-200, 460-461, or 669-673, or to SEQ ID NOS: 201-203 or comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity over at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, or at least 80 consecutive nucleotides of any one of 613-616 non-variable nucleotides.
[0089] In some cases, the engineered nuclease system includes an endonuclease. The endonuclease can be a Class 2 endonuclease. The endonuclease can be a Cas endonuclease. The endonuclease can be a Class 2 Type II Cas endonuclease.
[0090] In some cases, the endonuclease has a specific molecular weight range. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some cases, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some cases, the endonuclease comprises a specific number of residues. The endonuclease may contain about 1,100 or fewer residues, about 1,000 or fewer residues, about 950 or fewer residues, about 900 or fewer residues, about 850 or fewer residues, about 800 or fewer residues, about 750 or fewer residues, about 700 or fewer residues, about 650 or fewer residues, about 600 or fewer residues, about 550 or fewer residues, about 500 or fewer residues, about 450 or fewer residues, about 400 or fewer residues, or about 350 or fewer residues. The endonuclease may contain about 700 to about 1,100 residues. The endonuclease may contain about 400 to about 600 residues. In some cases, the engineered guide ribonucleic acid structure comprises a single ribonucleic acid polynucleotide. The single ribonucleic acid polynucleotide may comprise a guide ribonucleic acid sequence and a tracr ribonucleic acid sequence.
[0091] In some cases, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some cases, the guide ribonucleic acid sequence is complementary to a sequence of a prokaryotic genome. In some cases, the guide ribonucleic acid sequence is complementary to a sequence of a bacterial genome. In some cases, the guide ribonucleic acid sequence is complementary to a sequence of an archaeal genome. In some cases, the guide ribonucleic acid sequence is complementary to a sequence of a eukaryotic genome. In some cases, the guide ribonucleic acid sequence is complementary to a sequence of a fungal genome. In some cases, the guide ribonucleic acid sequence is complementary to a sequence of a plant genome. In some cases, the guide ribonucleic acid sequence is complementary to a sequence of a mammalian genome. In some cases, the guide ribonucleic acid sequence is complementary to a sequence of a human genome.
[0092] In some cases, the guide ribonucleic acid targeting sequence or spacer is 10-30 nucleotides in length, 12-28 nucleotides in length, or 15-24 nucleotides in length. In some cases, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some cases, the NLS comprises a sequence selected from SEQ ID NOs: 205-220.
[0093] [Table 1]
[0094] The present disclosure also includes variants of any of the enzymes described herein that contain one or more conservative amino acid substitutions. Conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by positioning mutated amino acid residues (e.g., non-conserved residues) between species without altering the basic function of the encoded protein. Such conservatively substituted variants can include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, including at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of the endonuclease protein sequences described herein. In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with substitutions such that the activity of one or more critical active site residues or guide RNA binding residues of the endonuclease is not disrupted. In some embodiments, a functional variant of any of the proteins described herein lacks at least one substitution of a conserved or functional residue listed in Figure 4. In some embodiments, a functional variant of any of the proteins described herein lacks substitutions of all conserved or functional residues listed in Figure 4. The present disclosure also provides modified activity variants of any of the nucleases described herein.Such altered activity mutants may contain inactivating mutations in one or more catalytic residues identified herein (e.g., in Figure 4) or generally described for the RuvC domain. Such altered activity mutants may contain change switch mutations in catalytic residues of the RuvCI, RuvCII, or RuvCIII domains.
[0095] Conservative substitution tables providing functionally similar amino acids are available from various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (W.H. Freeman & Co.; 2nd edition (December 1993))). The following eight groups each include amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G), 2) Aspartic acid (D), glutamic acid (E), 3) Asparagine (N), Glutamine (Q), 4) Arginine (R), Lysine (K), 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V), 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W), 7) serine (S), threonine (T), and 8) Cysteine (C), Methionine (M)
[0096] The present disclosure also encompasses variants of any of the endonucleases described herein that share a specific domain identity. The domain may be an arginine-rich domain (e.g., a domain with PF14239 homology), a REC (recognition) domain, a BH (bridge-helix) domain, a WED (wedge) domain, a PI (PAM-interacting) domain, a PF14239 homology domain, or any other domain described herein. In some embodiments, one or more of the residues encompassing these domains are identified in a protein by alignment with one of the following proteins (e.g., when the protein of interest is optimally aligned with one of the following proteins), where the residue boundaries of example domains are described:
[0097] [Table 2]
[0098] In some cases, the engineered nuclease system further comprises a single-stranded DNA repair template. In some cases, the engineered nuclease system further comprises a double-stranded DNA repair template. In some cases, the single-stranded or double-stranded DNA repair template comprises a first homology arm, 5' to 3', 5' to the target deoxyribonucleic acid sequence, comprising a sequence of at least 20 nucleotides. In some cases, the single-stranded or double-stranded DNA repair template comprises a synthetic DNA sequence, 5' to 3', of at least 10 nucleotides. In some cases, the single-stranded or double-stranded DNA repair template comprises a second homology arm, 5' to 3', 3' to the target sequence, comprising a sequence of at least 20 nucleotides. In some cases, the single-stranded or double-stranded DNA repair template comprises, from 5' to 3', a first homology arm comprising a sequence of at least 20 nucleotides 5' of a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, or a second homology arm comprising a sequence of at least 20 nucleotides 3' of said target sequence.
[0099] In some cases, the first homology arm comprises a sequence of at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides. 2+ In some cases, the endonuclease and the tract ribonucleic acid sequence are derived from different bacterial species. In some cases, the endonuclease and the tract ribonucleic acid sequence are derived from distinct bacterial species within the same phylum.
[0100] In some cases, the endonuclease comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOS: 1-24 or 462-488. In some cases, the guide RNA structure comprises an RNA sequence predicted to include a hairpin. In some cases, the hairpin comprises a stem and a loop. In some cases, the stem comprises at least 12 pairs, at least 14 pairs, at least 16 pairs, or at least 18 pairs of ribonucleotides.
[0101] In some cases, the guide RNA structure may further comprise a second stem and a second loop. In some cases, the second stem comprises at least 5 pairs, at least 6 pairs, at least 7 pairs, at least 8 pairs, at least 9 pairs, or at least 10 pairs of ribonucleotides. In some cases, the guide RNA structure comprises an RNA structure, and the RNA structure comprises at least two hairpins. In some cases, the endonuclease comprises a sequence having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1, and the guide RNA structure comprises an RNA sequence that is determined to include at least four hairpins. In some cases, each of these four hairpins comprises a stem and a loop.
[0102] In some cases, the engineered nuclease system comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:1. In some cases, the engineered nuclease system comprises a guide RNA structural sequence that comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least one non-variable nucleotide of SEQ ID NO:199 or SEQ ID NO:201.
[0103] In some cases, the engineered nuclease system comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 1-24 or 462-488. In some cases, the engineered nuclease system comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the non-variable nucleotides of any one of SEQ ID NOs: 199-200 or 669-673, or any one of SEQ ID NOs: 201-203 or 613-616.
[0104] In some cases, sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or Smith-Waterman homology search algorithm parameters using CLUSTALW. In some cases, sequence identity is determined by the aforementioned BLASTP homology search algorithm using the parameters wordlength (W) of 3, expectation (E) of 10, and scoring matrix BLOSUM62 with gap costs set to existence of 11 and extension of 1, as well as a conditional composition score matrix adjustment.
[0105] In some cases, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some cases, the endonuclease has less than 80% identity, less than 75% identity, less than 70% identity, less than 65% identity, less than 60% identity, less than 55% identity, or less than 50% identity to a Cas9 endonuclease.
[0106] In one aspect, the present disclosure provides an engineered guide RNA comprising (a) a DNA targeting segment. In some cases, the DNA targeting segment comprises a nucleotide sequence complementary to a target sequence in a target DNA molecule. In some cases, the engineered single guide ribonucleic acid polynucleotide comprises a protein-binding segment. The protein-binding segment comprises two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex. In some cases, the two complementary stretches of nucleotides are covalently linked to each other by an intervening nucleotide. In some cases, the engineered guide ribonucleic acid polynucleotide is configured to form a complex with an endonuclease that comprises a variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668.
[0107] In some cases, the DNA targeting segment is located 5' of both of the two complementary stretches of nucleotides. In some cases, the protein-binding segment comprises a sequence that is at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the non-variable nucleotides of any one of SEQ ID NOs: 199-200 or 669-673, or any one of SEQ ID NOs: 201-203 or 613-616. In some cases, the deoxyribonucleic acid polynucleotide encodes an engineered guide ribonucleic acid polynucleotide described herein.
[0108] In one aspect, the present disclosure provides a nucleic acid comprising an engineered nucleic acid sequence. In some cases, the engineered nucleic acid sequence is optimized for expression in an organism. In some cases, the nucleic acid encodes an endonuclease. The endonuclease can be a Cas endonuclease. The endonuclease can be a class 2 endonuclease. The endonuclease can be a class 2 type II Cas endonuclease. In some cases, the endonuclease includes a RuvC domain and an HNH domain. In some cases, the endonuclease is derived from a fastidious microorganism. In some cases, the endonuclease has a specific molecular weight range. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some cases, the endonuclease comprises a specific number of residues. The endonuclease may contain about 1,100 or fewer residues, about 1,000 or fewer residues, about 950 or fewer residues, about 900 or fewer residues, about 850 or fewer residues, about 800 or fewer residues, about 750 or fewer residues, about 700 or fewer residues, about 650 or fewer residues, about 600 or fewer residues, about 550 or fewer residues, about 500 or fewer residues, about 450 or fewer residues, about 400 or fewer residues, or about 350 or fewer residues. The endonuclease may contain about 700 to about 1,100 residues. The endonuclease may contain about 400 to about 600 residues.In some cases, the endonuclease comprises SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668, or a variant thereof having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. In some cases, the endonuclease further comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some cases, the NLS comprises a sequence selected from SEQ ID NOs: 205-220.
[0109] In some cases, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human. In some cases, the organism is a prokaryote. In some cases, the organism is a bacterium. In some cases, the organism is an archaea. In some cases, the organism is a fungus. In some cases, the organism is a plant. In some cases, the organism is a mammal. In some cases, the organism is a fungus. In some cases, the organism is a human. If the organism is a prokaryote or bacterium, the organism can be a different organism from the organism from which the endonuclease is derived. In some cases, the organism is not a fastidious microorganism.
[0110] In one aspect, the present disclosure provides a vector comprising a nucleic acid sequence. In some cases, the nucleic acid sequence encodes an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 endonuclease. In some cases, the endonuclease is a class 2 type II Cas endonuclease. In some cases, the endonuclease comprises a RuvC-I domain and an HNH domain. In some cases, the endonuclease is derived from a fastidious microorganism. In some cases, the endonuclease has a specific molecular weight range. In some embodiments, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 105 kDa or less, about 100 kDa or less, 95 kDa or less, about 90 kDa or less, about 95 kDa or less, about 80 kDa or less, about 75 kDa or less, about 70 kDa or less, about 65 kDa or less, about 60 kDa or less, about 55 kDa or less, about 50 kDa or less, about 45 kDa or less, about 40 kDa or less, about 35 kDa or less, about 30 kDa or less, about 25 kDa or less, about 20 kDa or less, about 15 kDa or less, or about 10 kDa or less. In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some cases, the endonuclease comprises a specific number of residues. The endonuclease may contain about 1,100 or fewer residues, about 1,000 or fewer residues, about 950 or fewer residues, about 900 or fewer residues, about 850 or fewer residues, about 800 or fewer residues, about 750 or fewer residues, about 700 or fewer residues, about 650 or fewer residues, about 600 or fewer residues, about 550 or fewer residues, about 500 or fewer residues, about 450 or fewer residues, about 400 or fewer residues, or about 350 or fewer residues. The endonuclease may contain about 700 to about 1,100 residues. The endonuclease may contain about 400 to about 600 residues.
[0111] In some embodiments, the present disclosure provides an endonuclease described herein configured to create a double-stranded break 5' from a protospacer adjacent motif (PAM) and proximal to the target locus. The endonuclease may create a double-stranded break 6-8 nucleotides from the PAM or 7 nucleotides from the PAM. In some embodiments, the present disclosure provides an endonuclease described herein configured to create a single-stranded break 5' from a protospacer adjacent motif (PAM) and proximal to the target locus. The endonuclease may create a double-stranded break 6-8 nucleotides from the PAM or 7 nucleotides from the PAM. In some cases, the endonuclease configured to create a single-stranded break comprises an inactivating mutation in one or more catalytic residues of the endonuclease described herein.
[0112] In some embodiments, the present disclosure provides an endonuclease described herein configured to cause a chemical modification of a nucleotide base within or proximal to a locus targeted by the endonuclease system. In this case, chemical modification of a nucleotide base generally refers to modification of a chemical moiety involved in base pairing, rather than modification of the sugar or phosphate portion of the nucleotide. The chemical modification may include deamination of an adenosine or cytosine nucleotide. In some cases, the endonuclease system configured to cause a chemical modification includes an endonuclease having a base editor linked or fused in-frame to the endonuclease. The endonuclease to which the base editor is fused or bound may contain an inactivating mutation in at least one catalytic residue of the endonuclease (e.g., in the RuvC domain). The base editor may be fused to the N-terminus or C-terminus of the endonuclease, or linked via chemical conjugation.Base editors may include any adenosine or cytosine deaminase, including, but not limited to, Adenosine Deaminase RNA Specific 1 (ADAR1), Adenosine Deaminase RNA Specific 2 (ADAR2), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 1 (APOBEC1), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 2 (APOBEC2), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 3A (APOBEC3A), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 3B (APOBEC3B), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 3C (APOBEC3C), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 3D (APOBEC3D), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit The base editors include Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 3F (APOBEC3F), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 3G (APOBEC3G), Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 3H (APOBEC3H), or Apolipoprotein B MRNA Editing Enzyme Catalytic Subunit 4 (APOBEC4), or functional fragments thereof. The base editor can include a yeast, eukaryotic, mammalian, or human base editor.
[0113] In some aspects, the present disclosure provides an endonuclease described herein configured to cause a chemical modification of a histone within or proximal to a locus targeted by the endonuclease system. In some cases, the endonuclease system configured to cause a chemical modification of a histone includes an endonuclease having a histone editor linked or fused in-frame to the endonuclease. The histone editor can be linked or fused to the N-terminus or C-terminus of the endonuclease. In some embodiments, the chemical modification can include methylation, acetylation, demethylation, or deacetylation. The endonuclease to which the histone editor is fused or bound can include an inactivating mutation in at least one catalytic residue of the endonuclease (e.g., in the RuvC domain). Histone editors are histone methyltransferases (e.g., ASH1L, DOT1L, EHMT1, EHMT2, EZH1, EZH2, MLL, MLL2, MLL3, MLL4, MLL5, NSD1, PRDM2, SET, SETBP1, SETD1A, SETD1B, SETD2, SETD3, SETD4, SETD5, SETD6, SETD7, SETD8, SETD9, SETDB1, SETDB2, SETMAR, SMYD1, S Histone editors may include histone demethylases (e.g., MYD2, SMYD3, SMYD4, SMYD5, SUV39H1, SUV39H2, SUV420H1, or SUV420H2), histone demethylases (e.g., KDM1, KDM2, KDM3, KDM4, KDM5, or KDM6 family), histone acetyltransferases (e.g., GNAT or HAT family acetyltransferases), or histone deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC4, HDAC5, HDAC6, HDAC7, HDAC8, HDAC9, HDAC10, HDAC11, SIRT1, SIRT2, SIRT3, SIRT4, SIRT5, SIRT6, or SIRT7). Histone editors may include yeast, eukaryotic, mammalian, or human histone editors.
[0114] In one aspect, the present disclosure provides a vector comprising a nucleic acid sequence described herein. Optionally, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure. The engineered guide ribonucleic acid structure may be configured to form a complex with an endonuclease. Optionally, the engineered guide ribonucleic acid structure comprises a guide ribonucleic acid sequence. Optionally, the guide ribonucleic acid sequence is configured to hybridize to a target deoxyribonucleic acid sequence. Optionally, the engineered guide ribonucleic acid structure comprises a tracr ribonucleic acid sequence. Optionally, the tracr ribonucleic acid sequence is configured to bind to an endonuclease. Optionally, the aforementioned vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.
[0115] In one aspect, the disclosure provides a cell comprising any of the vectors described herein.
[0116] In one aspect, the present disclosure provides a method of producing an endonuclease. The method may include culturing any of the cells described herein.
[0117] In one aspect, in some aspects, the present disclosure provides a method for binding, cleaving, labeling, or modifying a double-stranded deoxyribonucleic acid polynucleotide. The method may include contacting the double-stranded deoxyribonucleic acid polynucleotide with an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 endonuclease. In some cases, the endonuclease is a class 2 type II Cas endonuclease. The endonuclease may be complexed with an engineered guide ribonucleic acid structure. In some cases, the engineered guide ribonucleic acid structure is configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide. In some cases, the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM). In some cases, the endonuclease has a molecular weight of about 120 kDa or less, about 110 kDa or less, about 100 kDa or less, 90 kDa or less, about 80 kDa or less, about 70 kDa or less, about 60 kDa or less, about 50 kDa or less, about 40 kDa or less, about 30 kDa or less, about 20 kDa or less, or about 10 kDa or less. In some cases, the endonuclease includes a variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668.
[0118] In one aspect, in some aspects, the present disclosure provides a method for binding, cleaving, labeling, or modifying a double-stranded deoxyribonucleic acid polynucleotide. The method may include contacting the double-stranded deoxyribonucleic acid polynucleotide with an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 endonuclease. In some cases, the endonuclease is a class 2 type II Cas endonuclease. The endonuclease may be complexed with an engineered guide ribonucleic acid structure. In some cases, the engineered guide ribonucleic acid structure may be configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide. In some cases, the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM). In some cases, the PAM is NGG. In some cases, the endonuclease includes a variant having at least 50%, at least 55%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668.
[0119] In some cases, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some cases, the endonuclease is derived from a difficult-to-cultivate microorganism. In some cases, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, bacterial, eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. In some cases, the double-stranded deoxyribonucleic acid polynucleotide is a prokaryotic, archaeal, or bacterial double-stranded deoxyribonucleic acid polynucleotide from a species other than the species from which the endonuclease is derived.
[0120] In one aspect, the present disclosure provides a method for modifying a target nucleic acid locus. The method may include delivering an engineered nuclease system described herein to the target nucleic acid locus. In some cases, the endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure. In some cases, the complex is configured such that, upon binding to the target nucleic acid locus, the complex modifies the target nucleic acid locus. In some cases, modifying the target nucleic acid locus includes binding, nicking, cleaving, or labeling the target nucleic acid locus.
[0121] In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some cases, the target nucleic acid comprises genomic eukaryotic DNA, viral DNA, or bacterial DNA. In some cases, the target nucleic acid comprises bacterial DNA. The bacterial DNA may be from a bacterial species different from the species from which the endonuclease is derived. In some cases, the target nucleic acid locus is in vitro. In some cases, the nucleic acid locus is within a cell. In some cases, the endonuclease and the engineered guide nucleic acid structure are provided and encoded by separate nucleic acid molecules. In some cases, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some cases, the cell is from a species different from the species from which the endonuclease is derived.
[0122] In some cases, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid described herein or a vector described herein. In some cases, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease. In some cases, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. In some cases, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding the endonuclease. In some cases, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide.
[0123] In some cases, delivering the engineered nuclease system to the target nucleic acid locus includes delivering a deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid structure operably linked to a ribonucleic acid (RNA) pol III promoter. In some cases, the endonuclease creates a single-stranded or double-stranded break at or proximal to the target locus.
[0124] For example, the systems of the present disclosure can be used for a variety of applications, such as nucleic acid editing (e.g., gene editing), binding to nucleic acid molecules (e.g., sequence-specific binding), etc. Such systems may be used, for example, to inactivate viruses by targeting viral genomes or to prevent them from infecting host cells, to add genes or alter metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drivers for evolutionary selection, to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors, to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), such as inactivating enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), and to address (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, thereby inactivating the gene to confirm its function in cells. [Example]
[0125] Example 1. Discovery of new Cas effectors by metagenomics Metagenomic Mining Metagenomic samples were collected from sediments, soils, and animals. Deoxyribonucleic acid (DNA) was extracted using the Zymobiomics DNA mini-prep kit and analyzed by Illumina HiSeq.(登録商標) Samples were collected with the landowner's consent. DNA was extracted from the samples using the Qiagen DNeasy PowerSoil Kit or ZymoBIOMICS DNA Miniprep Kit. DNA was sent to the Vincent J. Coates Genomics Sequencing Laboratory at UC Berkeley for sequencing library construction (Illumina TruSeq) and sequencing on an Illumina HiSeq 4000 or Novaseq (150 base pair (bp) reads, target insert size 400–800 bp). In addition, publicly available high-temperature, soil, and marine metagenomic sequence data were downloaded from the NCBI SRA. Sequencing reads were trimmed using BBMap (Bushnell B., sourceforge.net / projects / bbmap / ) and assembled with Megahit (https: / / paperpile.com / c / QSZG6K / clMrh). Protein sequences were predicted using Progdigal (https: / / paperpile.com / c / QSZG6K / BJ6oW). HMM profiles of known type II CRISPR nucleases were constructed and searched against all predicted proteins using HMMER3 (hmmer.org). CRISPR arrays were predicted for assembled contigs using Minced (https: / / github.com / ctSkennerton / minced> or https: / / paperpile.com / c / QSZG6K / OPC44). Classification was assigned using Kaiju (https: / / paperpile.com / c / QSZG6K / nMi6k), and contig classification was determined by finding the consensus of all encoded proteins.
[0126] Predicted type II effector proteins were aligned with standards (e.g., SpCas9, SaCas9, AsCas9) using MAFFT (https: / / paperpile.com / c / QSZG6K / sVHNH), and phylogenetic trees were inferred using FastTree2 (https: / / paperpile.com / c / QSZG6K / osZNM). Novel families were identified from clades composed of sequences recovered in this study. From among the families, candidates were selected that contained all elements necessary for laboratory analysis (i.e., discovered using CRISPR arrays in fully assembled and annotated contigs). The selected representative sequences were aligned with the standard sequences using MUSCLE (https: / / paperpile.com / c / QSZG6K / ITOla), and catalytic and PAM-interacting residues were identified.
[0127] This metagenomic analysis workflow led to the delineation of the SMART (SMall ARchaeal-associated) endonuclease system described herein.
[0128] Discovery of SMART endonucleases with active residue signatures. By mining tens of thousands of high-quality CRISPR-Cas systems constructed from metagenomic data, we discovered novel effector nucleases that contain both RuvC and HNH domains but are unusually small (900 aa). These effector nucleases showed low sequence similarity (less than 20% amino acid identity) with archaeal Cas9 endonucleases. Phylogenetic analysis of the effector protein sequences showed that SMART systems are a divergent group compared to the well-studied type II systems of subtypes A, B, and C (Figure 1A).
[0129] These compact "SMART" effectors (~400-1000 amino acids, Figure 2) appeared in genomic loci adjacent to CRISPR arrays. Some of these adjacent SMART loci also contained sequences predicted to encode tracrRNA and CRISPR adaptation genes (e.g., genes involved in spacer acquisition) cas1, cas2, and / or cas4 within the same operon (Figure 3). Despite their compact size, SMART effectors encompass six putative HNH and RuvC catalytic residues when aligned with the reference SaCas9 sequence (Figure 4). Furthermore, 3D structure prediction identified residues involved in guide and target binding as well as PAM recognition, suggesting that SMART effectors are active dsDNA endonucleases.
[0130] Based on the location of key catalytic and binding residues, SMART nucleases contain three RuvC regions, an arginine-rich region that typically contains an RRxRR motif (e.g., a region with PF14239 homology), an HNH endonuclease domain, and a putative recognition region (Figures 5 and 6). These domains share low sequence similarity with reference sequences (Figure 7). In addition, SMART effectors, as well as reference archaeal sequences, contain the RRxRR motif and zinc-binding ribbon motif (CX) significantly more frequently than Cas9 nucleases. [2-4] C or CX [2-4] H) (Figure 8). Additionally, unlike Cas9 effector sequences, most SMART effectors contain significant hits to the Pfam domain PF14239, which is often associated with diverse endonucleases. Based on differences in SMART effector size, phylogenetic relatedness, and both operon and domain architecture, we classified these systems into two primary populations: SMART I and SMART II. The salient features of these groups are outlined below in Table 3, where differences are also illustrated compared to class 2 type II A / B / C Cas enzymes.
[0131] [Table 3]
[0132] SMART I endonuclease SMART I effectors range in size from approximately 700 to 1,050 amino acids. Common features in their genomic contexts include predicted tracrRNAs near adaptive module genes (e.g., genes involved in spacer acquisition) and CRISPR arrays, whose mechanisms resemble those of type II and type V CRISPR systems (Figures 3A, 3B, and 3C). The RRXRR motif-containing region in SMART I effectors is unique but may play a functional role similar to the arginine-rich bridge helix in Cas9 nuclease. When modeled against the SaCas9 crystal structure, the predicted 3D structures of SMART I effectors revealed unaligned regions within the recognition lobe (often encompassing the Pfam domain PF14239) and the RuvCII domain (Figure 5). The results indicated that these domains have a distinct origin from other type II effectors. Taken together with their branched placement in the type II effector phylogenetic tree and their low sequence similarity to known type II effectors (Figure 1A), these results indicate that SMART I endonucleases belong to a new group of type II CRISPR systems. In accordance with the accepted classification of CRISPR systems, these SMART I systems were classified as type II-D.
[0133] Putative single guide RNAs (sgRNAs) were engineered using environmental RNA expression data for the SMART I MG34-1 system. In addition, multiple sgRNAs designed from SMART I repeat and tracrRNA predictions were tested in vitro in PAM enrichment assays. For SMART I enzymes, optimal PAM sequence identification was achieved using end-repair and blunt-end ligation in this step, suggesting that these enzymes can produce staggered double-stranded DNA breaks. The assay confirmed dsDNA cleavage for MG34-1 (SEQ ID NO: 2), MG34-9 (SEQ ID NO: 9), and MG34-16 (SEQ ID NO: 17) with multiple sgRNA designs (Figure 7, representing the use of SEQ ID NOs: 612-615). MG34-1 demonstrated target recognition and cleavage preference for the NGGN PAM (Figure 8A). Cleavage site analysis showed selective cleavage at position 7 (Figure 8B). These results suggest a novel biochemical mechanism in comparison with the cleavage mechanisms of other type II enzymes that selectively cleave at positions 2–3 from the PAM and support a new classification of SMART I CRISPR systems.
[0134] Environmental expression data for several SMART I systems confirmed the in situ transcription of CRISPR arrays and intergenic regions encoding the predicted tracrRNA (Figures 3B and 3C). Furthermore, we evaluated instances of active CRISPR targeting by searching for spacer sequences matching other genome sequences assembled from the same or related metagenomes. Accordingly, we identified a phage genome targeted by one of the spacers encoded in the SMART I CRISPR array (Figures 3C and 3D). Analysis of the region flanking the target sequence suggested a 3' PAM sequence encompassing a GG motif (Figure 3D). These results suggest that SMART I CRISPR systems are active in their natural environment as RNA-guided effectors involved in phage defense, likely functioning as nucleases to cleave or degrade target DNA or RNA.
[0135] SMART I effectors are active, RNA-guided dsDNA CRISPR endonucleases. Environmental RNA expression data from the SMART I MG34-1 and MG34-16 systems (Figure 3B and Figure 3C, and Figure 9) were used to design putative single guide RNAs (sgRNAs). Furthermore, multiple sgRNAs designed from SMART I repeat and tracrRNA predictions were tested in an in vitro PAM enrichment assay (Figure 10). The assay confirmed programmable dsDNA cleavage for MG34-1, MG34-9, and MG34-16 with multiple sgRNA designs (Figure 10). MG34-1 and MG34-9 require the NGGN PAM for target recognition and cleavage (Figure 11A and Figure 11C). Cleavage site analysis showed selective cleavage at the 7 position (Figure 11B and Figure 11C). These results suggest a novel biochemical cleavage mechanism in comparison with that of the Cas9 enzyme, which selectively cleaves 3 positions from the PAM, and further support a new classification for the SMART I CRISPR system.
[0136] PAM enrichment assays without an end-repair step showed no activity for SMART I nucleases. The requirement for end-repair to generate blunt-ended fragments prior to ligation in the PAM enrichment protocol indicates that these enzymes generate staggered double-stranded DNA breaks.
[0137] Experiments performed in E. coli confirmed that the system possessed the necessary activity to function as a nuclease within the cells. E. coli expressing MG34-1 and sgRNA were transformed with a kanamycin resistance plasmid containing the sgRNA target. In the presence of antibiotic, successful targeting and cleavage of the antibiotic resistance plasmid resulted in growth defects. This assay confirmed approximately two-fold growth inhibition compared to a control experiment performed with a kanamycin resistance plasmid that did not contain the sgRNA target (Figure 12).
[0138] SMART II endonuclease SMART II effectors have a smaller size distribution (~400-600 amino acids) compared to SMART I effectors. Their genomic context suggested unusual repetitive regions or CRISPR arrays. Non-CRISPR repetitive regions encompass direct repeats ranging in size from approximately 10 to 30 bp. In some cases, they contain multiple distinct repeat units. Occasionally, common CRISPR identification algorithms would flag these regions as CRISPR systems; however, closer examination would reveal that regions identified as spacer sequences are repeated in the arrays. Although the arrays are not immediately adjacent to the effectors, they are located in the same genomic region (Figure 3A, MG35-236, and Figure 13A, e.g., >20 kb from the effector gene). SMART II system operons generally lacked adaptive module genes (e.g., genes involved in spacer acquisition).
[0139] The structural prediction identified all six RuvC and HNH nuclease catalytic residues frequently found in class 2 type II Cas effectors, as well as characteristic residues of Cas enzymes involved in guide RNA binding, target cleavage, and PAM recognition and interaction (Figure 6). SMART II effectors also contain multiple RRXRR and zinc-binding ribbon motifs (CX [2-4] C or CX [2-4] H), which may be involved in the recognition and binding of target nucleic acid motifs. Based on the locations of key residues, the predicted domain structure of SMART II nucleases consisted of three RuvC subdomains, an arginine-rich region containing an RRxRR motif (e.g., a domain with PF14239 homology), an HNH endonuclease domain, an unknown domain, and a recognition domain (REC) (Figure 6). The domain architecture of the SMART II effector differed from the known domain architecture of type II Cas9 nucleases (Figure 6 and Figure 14).
[0140] Environmental transcriptome data for several SMART II systems confirmed the expression of CRISPR arrays and other repeat regions in situ in their natural environments (Figure S13A). Transcription of the 5' untranslated regions (UTRs) of several SMART II effectors was also observed in the environmental expression data (Figure S13B), suggesting that this region may be important for either nuclease activity or regulation of the SMART system.
[0141] Preliminary in vitro experiments performed with SMART II effector proteins, repeat regions, and associated intergenic regions indicate that these enzymes may have the ability to cleave dsDNA, possibly in a programmable manner (see Figure 15). The results suggest that SMART II nuclease activity may be guided by RNA and / or DNA, using repeat regions such as CRISPR arrays, or requiring recognition of features encoded within genetic loci such as TIRs or 5'UTRs.
[0142] Several SMART II effectors were observed adjacent to putative insertion sequences (ISs) encoding the transposases TnpA and TnpB (Fig. 3A). The ends of the ISs were determined to encompass terminal inverted repeats (TIRs) with a predicted U-shaped structure, and target site duplications into which the ISs most likely integrate were also identified. Furthermore, several SMART II loci encoded putative TIRs flanking SMART II effectors (e.g., Fig. 3).
[0143] Example 2. Identification / confirmation of the PAM sequences of the endonucleases described herein Putative SMART endonucleases were expressed in an E. coli lysate-based expression system (PURExpress, New England Biolabs). In this system, the endonucleases were codon-optimized for E. coli and cloned into a vector with a T7 promoter and a C-terminal His tag. The genes were PCR-amplified using primer binding sites and terminator sequences 150 bp upstream and downstream of the T7 promoter, respectively. This PCR product was added to the NEB PURExpress and expressed at a final concentration of 5 nM at 37°C for 2 hours to produce the endonucleases for PAM assays.
[0144] Putative sgRNAs compatible with each of the SMART Cas enzymes described herein were identified from RNAseq reads assembled against contiguous CRISPR loci assembled from sequencing data. Secondary structures were determined for the tracr region from the RNAseq data and the repeat sequences from the CRISPR arrays using the Geneious software package (https: / / www.geneious.com). The final helix was trimmed and ligated to a GAAA tetraloop. Multiple lengths of repeat-antirepeat helix trimming were tested, as well as different spacer lengths and different tracr extension termination points (Figure 12, SEQ ID NOs: 612-615). The sgRNAs were then assembled via assembly PCR, purified using SPRI beads, and in vitro transcribed (IVT) according to the manufacturer's recommended protocol for short RNA transcripts (HiScribe T7 Kit, NEB). The RNA transcription reaction was cleaned up with the Monarch RNA kit and checked for purity via Tapestation (Agilent).
[0145] The PAM sequence was determined by sequencing a plasmid containing randomly generated candidate PAM sequences cleavable by the putative nuclease. In this system, a nucleotide sequence encoding the putative nuclease, optimized for E. coli codons, was transcribed and translated in vitro from a PCR fragment under the control of a T7 promoter. A second PCR fragment containing a minimal CRISPR array consisting of a T7 promoter followed by a repeat-spacer-repeat sequence was transcribed in the same reaction. Successful expression of the endonuclease and repeat-spacer-repeat sequence in the TXTL system, followed by CRISPR array processing, resulted in active in vitro CRISPR nuclease complexes.
[0146] A library of target plasmids containing spacer sequences matching sequences within a minimal array preceded by 8N mixed degenerate bases (potential PAM sequences) was incubated with the TXTL reaction products (10 mM Tris pH 7.5, 100 mM NaCl, and 10 mM MgCl2 with a 5-fold dilution of translated Cas enzyme, 5 nM of the 8N PAM plasmid library, and 50 nM of sgRNA targeting the PAM library). After 1–3 h, the reaction was stopped, and DNA was recovered using a DNA cleanup kit. Adapter sequences were ligated to DNA with active PAM sequences cleaved by an endonuclease; uncleaved DNA was blunt-ended and inaccessible for ligation. DNA segments containing active PAM sequences were then amplified by PCR using primers specific to the library and adapter sequences. PCR amplification products were resolved on a gel to identify amplicons corresponding to cleavage events. The amplified segments from the cleavage reaction were also used as templates for NGS library preparation or as substrates for Sanger sequencing. The resulting library, a subset of the starting 8N library, revealed sequences with PAM activity compatible with CRISPR complexes. For PAM testing using engineered RNA constructs, the same procedure was repeated, except that in vitro transcribed RNA was added along with the plasmid library and the minimal CRISPR array / tracr template was omitted. These assays used the following spacer sequence as the target: 5'-CGUGAGCCACCACGUCGCAAGCCUCGAC-3'.
[0147] After obtaining raw sequence reads from the PAM assay, reads were filtered for Phred quality scores >20. A 24-bp reference region representing known backbone DNA sequence adjacent to the PAM was used to locate the PAM-proximal region, and the adjacent 8-bp region was identified as a putative PAM. The distance between the PAM and the ligated adapter was also measured for each read. Reads that did not perfectly match the reference sequence or adapter sequence were excluded. PAM sequences were filtered by cleavage site frequency so that only PAMs with the most frequent cleavage site ±2 bp were included in the analysis. The filtered list of PAMs was used to generate sequence logos using Logomaker (Tareen A, Kinney JB. Logomaker: beautiful sequence logos in Python. Bioinformatics. 2020;36(7):2272-2274, incorporated herein by reference).
[0148] Example 3. Protocol for predicted RNA folding The predicted RNA folding of the active single RNA sequence was calculated at 37°C using the method of Andronescu 2007. The color of the base corresponds to the base pairing probability of that base, where red is high probability and blue is low probability.
[0149] Example 4. In vitro cleavage efficiency The endonuclease was expressed as a His-tagged fusion protein from an inducible T7 promoter in a protease-deficient Escherichia coli B strain. The endonuclease was fused to two nuclear localization signals (an N-terminal NLS nucleoplasmin duplex and a C-terminal simian virus 40 T antigen NLS PPKKKRK), a maltose-binding protein (MBP) tag, a tobacco etch virus (TEV) protease cleavage site, and a 6XHis tag in the following order from N- to C-terminus: 6XHis-MBP-TEV-NLS-gene-NLS-STOP. The protein was expressed under the pTac promoter in NEB Iq E. coli in autoinduction medium (MagicMedia ThermoFisher), grown at 30°C, and incubated at 16°C.
[0150] Cells expressing His-tagged proteins were lysed by sonication, and the His-tagged proteins were purified by Ni-NTA affinity chromatography on a HisTrap FF column (GE Lifescience) on an AKTA Avant FPLC (GE Lifescience). The eluate was analyzed by SDS-PAGE on an acrylamide gel (Bio-Rad) and stained with InstantBlue Ultrafast Coomassie (Sigma-Aldrich). Purity was determined using densitometry of the protein bands with ImageLab software (Bio-Rad). The purified endonuclease was dialyzed into a storage buffer consisting of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, and 5% glycerol, pH 7.5, and stored at -80°C.
[0151] Target DNA containing a spacer sequence and a PAM sequence (e.g., as determined in Example 2) was constructed by DNA synthesis. When the PAM had degenerate bases, a single representative PAM was selected. The target DNA consisted of a 2200-bp linear DNA fragment obtained by PCR amplification from a plasmid, with a PAM and spacer located 700 bp from one end. Successful cleavage yielded 700-bp and 1500-bp fragments. The target DNA, in vitro transcribed single RNA, and purified recombinant protein were combined in a cleavage buffer (10 mM Tris, 100 mM NaCl, 10 mM MgCl2) containing excess protein and RNA and incubated for 5 minutes to 3 hours, typically 1 hour. The reaction was stopped after 60 minutes of incubation by the addition of RNAse A. The reaction was then analyzed on a 1.2% TAE agarose gel, and the cleaved target DNA fragments were quantified using ImageLab software.
[0152] Example 5. Activity in E. coli E. coli lacks the ability to efficiently repair double-stranded DNA breaks. Therefore, genomic DNA breaks can be lethal. Taking advantage of this phenomenon, we test the activity of endonucleases in E. coli by recombinantly expressing the endonuclease and guide RNA in a target strain containing a spacer / target sequence and a PAM sequence integrated into the genomic DNA.
[0153] To test nuclease activity in bacterial cells, the BL21(DE3) strain (NEB) was transformed with plasmids containing the T7-driven effector and sgRNA (10 ng of each plasmid), inoculated onto plates, and grown overnight. Final colonies were grown overnight in triplicate and then subcultured in SOB and grown to an OD of 0.4–0.6. Cell cultures equivalent to OD 0.5 were transformed with 130 ng of kanamycin plasmids synthesized according to standard kit protocols (Zymo Mix and Go kit) with or without a spacer and PAM in the backbone. After heat shock, transformants were allowed to recover in SOB for 1 hour at 37°C and plated on induction medium (LB agar plates containing antibiotics and 0.05 mM IPTG) to determine nuclease efficiency. Colonies were quantified from the dilution series to measure overall suppression due to plasmid cleavage by the nuclease.
[0154] The results of such an assay are shown in Figure 12. In Figure 12, panel (A) shows replica plating of E. coli strains demonstrating plasmid cleavage. E. coli expressing MG34-1 and sgRNA were transformed with a kanamycin resistance plasmid containing a target for the sgRNA (+sp). The plate quadrant showing impaired growth (+sp) versus the negative control (no target and PAM (-sp)) indicates successful targeting and cleavage by the enzyme. The experiment was replicated twice and performed in triplicate. In Figure 12, panel (B) shows a graph of colony-forming unit (cfu) measurements from the replica plating experiment showing growth inhibition in the targeting condition (+sp) versus the non-targeting control (-sp) in (A), demonstrating that the plasmid was cleaved.
[0155] The engineered strain, which has a PAM sequence (e.g., as determined in Example 2) integrated into its genomic DNA, is transformed with DNA encoding the endonuclease. Transformants are then transformed with 50 ng of chemically synthesized guide RNA (e.g., crRNA) specific for the target sequence ("on-target") or non-specific for the target ("non-target"). After heat shock, transformants are recovered in SOC at 37°C for 2 hours. Nuclease efficiency is then determined by a 5-fold dilution series grown in induction medium. Colonies are quantified from a 3-fold dilution series.
[0156] Example 6. Verification of genome cleavage activity of MG CRISPR complex in mammalian cells To demonstrate targeting and cleavage activity in mammalian cells, the MG Cas effector protein sequence is tested in two mammalian expression vectors: (a) one with an SV40 NLS and a 2A-GFP tag at the C-terminus, and (b) one without a GFP tag and two SV40 NLS sequences at the N- and C-termini. The NLS sequences include any of the NLS sequences described herein. In some instances, the nucleotide sequence encoding the endonuclease is codon-optimized for expression in mammalian cells. The corresponding crRNA sequence with the targeting sequence is cloned into a second mammalian expression vector. The two plasmids are cotransfected into HEK293T cells. 72 hours after cotransfection of the expression plasmid and the gRNA targeting plasmid into HEK293T cells, DNA is extracted and used for next-generation sequencing (NGS) library preparation. To demonstrate the targeting efficiency of the enzyme in mammalian cells, the rate of NHEJ is measured via indels in the target site sequencing. At least 10 target sites were selected for testing the activity of each protein.
[0157] Example 7. Predicted activity of the MG family described herein In situ expression and protein sequence analysis indicate that these enzymes are active nucleases. They contain predicted endonuclease-associated domains (corresponding to the RRXRR and HNH endonuclease Pfam domains, Figures 2, 3A, and 3B) and predicted HNH and RuvC catalytic residues (e.g., Figures 2, 3A, and 3B, rectangles). Furthermore, the presence of the RRXRR motif, found in the RNase H-like protein family, indicates potential RNA targeting and nuclease activity (see Figure 2).
[0158] Expression data confirmed the in situ native activity of the MG34-1 nuclease candidate, tracrRNA, and CRISPR array ( Figure 4 ).
[0159] Example 8. Activity in mammalian cells following mRNA delivery For genome editing via mRNA-based cell transfection / transformation, the coding sequence is codon-optimized for mouse or human genes using the algorithms of Twist Bioscience or Thermo Fisher Scientific (GeneArt). Two nuclear localization signals (SV40 and nucleoplasmin) are added to the N- and C-termini of the coding endonuclease sequence. Additionally, untranslated regions derived from human complement 3 (C3) are added to both the 5' and 3' ends of the coding sequence within the cassette.
[0160] This cassette is then cloned into an mRNA production vector upstream of a long polyA stretch. The mRNA construct consists of the following: 5'UTR from C3 - SV40 NLS - codon-optimized SMART gene - nucleoplasmin NLS - 3'UTR from C3 - 107 polyA tail. Then, mRNA transcription is performed using an engineered T7 RNA polymerase (Hi-T7: New England Biolabs) driven by the T7 promoter. 5'-capping of the mRNA is co-transcriptionally induced using CleanCap AG (Trilink Biolabs). The mRNA is then purified using the MEGAclear Transcription Clean-Up kit (Thermo Fisher Scientific).
[0161] Using Lipofectamine Messenger Max (Thermo Fisher Scientific), mammalian cells are co-transfected with transcribed mRNA and a set of at least 10 guides targeting the genomic region of interest. After incubating the cells for a period of time (e.g., 48 hours), genomic DNA is isolated using the Purelink Genomic DNA extraction kit (Fisher Scientific). Specific primers are used to amplify the region of interest. Editing is then assessed by Sanger sequencing using Inference of CRISPR Edits, and the editing results are thoroughly analyzed by next-generation sequencing.
[0162] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited to the specific examples provided herein. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Many modifications, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It is to be understood that various alternatives to the embodiments of the present invention described herein may be utilized in practicing the invention. It is therefore contemplated that the present invention shall cover any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. 1. An in vitro method for modifying a target deoxyribonucleic acid locus, the method comprising: adding to the target deoxyribonucleic acid locus: (a) and (b): (a) an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease comprises a sequence having at least 90% sequence identity to SEQ ID NO: 2; (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, wherein the engineered guide ribonucleic acid structure comprises: (i) a guide ribonucleic acid sequence configured to hybridize to a portion of the target deoxyribonucleic acid locus; and (ii) a ribonucleic acid sequence configured to bind to the endonuclease, the ribonucleic acid sequence comprising a sequence having at least 90% sequence identity to a non-variable nucleotide of any one of SEQ ID NOs: 203, 202, or 613; a guide ribonucleic acid structure comprising an engineered guide ribonucleic acid structure comprising: and delivering wherein said complex modifies said target deoxyribonucleic acid locus.
2. The method of claim 1 , wherein the endonuclease is an archaeal endonuclease.
3. 3. The method of claim 1 or 2, wherein the endonuclease is a class 2 type II Cas endonuclease.
4. 4. The method of any one of claims 1 to 3, wherein the endonuclease further comprises one or more of an arginine-rich region containing an RRxRR motif, a domain with PF14239 homology, a recognition (REC) domain, a bridge-helix (BH) domain, a wedge (WED) domain, or a PAM-interacting (PI) domain.
5. 5. The method of claim 4, wherein the arginine-rich region, the domain having PF14239 homology, the recognition (REC) domain, the bridge helix (BH) domain, the wedge (WED) domain, or the PAM-interacting (PI) domain comprises a sequence having at least 85% sequence identity to an arginine-rich region comprising an RRxRR motif, a domain having PF14239 homology, a recognition (REC) domain, a bridge helix (BH) domain, a wedge (WED) domain, or a PAM-interacting (PI) domain of any one of SEQ ID NOs: 1-198, 221-459, 463-612, or 617-668, respectively.
6. The method of any one of claims 1 to 5, wherein the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease.
7. The method of any one of claims 1 to 6, wherein the endonuclease comprises a sequence having at least 95% sequence identity to SEQ ID NO:
2.
8. The method of any one of claims 1 to 7, wherein the endonuclease comprises the sequence of SEQ ID NO:
2.
9. 9. The method of any one of claims 1 to 8, wherein the endonuclease comprises a sequence having less than 80% sequence identity to SpCas9 endonuclease.
10. The method of any one of claims 1 to 9, wherein the guide ribonucleic acid sequence is complementary to a eukaryotic, fungal, plant, mammalian, or human genomic sequence.
11. The method of any one of claims 1 to 10, wherein the guide ribonucleic acid sequence is 15 to 24 nucleotides in length.
12. 12. The method of any one of claims 1-11, wherein the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having at least 95% sequence identity to a non-variable nucleotide of any one of SEQ ID NOs: 203, 202 or 613.
13. 13. The method of any one of claims 1 to 12, wherein the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having at least 95% sequence identity to nucleotides 23 to 157 of SEQ ID NO:203, nucleotides 23 to 93 of SEQ ID NO:202, or nucleotides 23 to 145 of SEQ ID NO:
613.
14. 14. The method of any one of claims 1 to 13, wherein the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having nucleotides 23 to 157 of SEQ ID NO:203, nucleotides 23 to 93 of SEQ ID NO:202, or nucleotides 23 to 145 of SEQ ID NO:
613.
15. The method of any one of claims 1 to 14, wherein the endonuclease and the ribonucleic acid sequence configured to bind to the endonuclease are derived from distinct bacterial species within the same phylum.
16. The method of any one of claims 1 to 15, wherein the engineered guide ribonucleic acid structure comprises any one of SEQ ID NOs: 203, 202 or 613.
17. 17. The method of any one of claims 1 to 16, wherein the engineered guide ribonucleic acid structure comprises a single ribonucleic acid polynucleotide comprising the guide ribonucleic acid sequence and the ribonucleic acid sequence configured to bind to the endonuclease.
18. 18. The method of any one of claims 1 to 17, wherein the sequence identity is determined by the BLASTP homology search algorithm using parameters of wordlength (W) = 3 and expectation (E) = 10, or the BLOSUM62 scoring matrix using a conditional composition score matrix adjustment with gap costs set to existence = 11 and extension = 1.
19. 19. The method of any one of claims 1 to 18, further comprising contacting the target deoxyribonucleic acid locus with a single-stranded or double-stranded deoxyribonucleic acid repair template comprising, from 5' to 3', a first homology arm comprising a sequence 5' to the target deoxyribonucleic acid locus, a synthetic deoxyribonucleic acid sequence, and a second homology arm comprising a sequence 3' to the target deoxyribonucleic acid locus.
20. 20. The method of any one of claims 1 to 19, wherein said modifying comprises binding, nicking, cleaving, or labeling said target deoxyribonucleic acid locus.
21. The method of any one of claims 1 to 20, wherein the target deoxyribonucleic acid locus is intracellular.
22. 22. The method of claim 21, wherein the cell is a eukaryotic cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell.
Citation Information
Patent Citations
RNA-guided nucleic acid modifying enzyme and method for using the same
JP2019534695A