Base editing enzyme
Patent Information
- Application Number
- JP2024519975
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-03
- Filing Date
- 2022-11-04
- Publication Date
- 2025-11-05
AI Technical Summary
Existing CRISPR systems for DNA manipulation and gene editing lack efficiency and specificity in targeting and modifying cytosine residues in eukaryotic nucleic acids, particularly in mammalian and human cells.
Development of engineered nucleic acid editing enzymes with cytosine deaminase activity, such as FAM72A-derived polypeptides, that can efficiently deaminate cytosine residues in eukaryotic nucleic acids, including mammalian and human cells, with high sequence identity and specificity.
The engineered enzymes achieve high fidelity and efficiency in converting cytosine to uracil in eukaryotic nucleic acids, enabling precise genetic modifications in mammalian and human cells.
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 276,461, filed November 5, 2021, No. 63 / 289,998, filed December 15, 2021, No. 63 / 342,824, filed May 17, 2022, No. 63 / 356,888, filed June 29, 2022, and No. 63 / 378,171, filed October 3, 2022, each of which is entitled "BASE EDITING ENZYMES" and is incorporated herein by reference in its entirety. This application is related to PCT Patent Application No. PCT / US2021 / 049962, which is incorporated herein by reference in its entirety. [Background technology]
[0002] Cas enzymes, along with their associated clustered regularly interspaced short palindromic repeats (CRISPR)-guided ribonucleic acid (RNA), are found to be widespread components of prokaryotic immune systems (approximately 45% of bacteria and 84% of archaea), where they function to protect these microorganisms from non-self nucleic acids, such as infectious viruses and plasmids, through CRISPR-RNA-guided nucleic acid cleavage. While deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid-interacting domains. While CRISPR DNA elements were observed as early as 1987, the programmable endonuclease cleavage capabilities of CRISPR complexes have only been recognized relatively recently, leading to the use of recombinant CRISPR systems in a variety of DNA manipulation and gene editing applications.
[0003] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format, and is incorporated herein by reference in its entirety. The XML copy created on November 4, 2022 is named 55921-742_601_SL.xml and is 2,274,288 KB in size. Summary of the Invention
[0004] In some aspects, the disclosure provides methods for deaminating cytosine residues in a eukaryotic nucleic acid sequence in a cell, the method comprising contacting the eukaryotic nucleic acid sequence with a peptide having cytosine deaminase activity, the peptide comprising a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, and 970-982, or a variant thereof. In some embodiments, the eukaryotic nucleic acid sequence is a mammalian, primate, or human nucleic acid sequence. In some embodiments, the cell is a mammalian, primate, or human cell. In some embodiments, the eukaryotic nucleic acid sequence comprises single-stranded DNA (ssDNA) or ribonucleic acid (RNA). In some embodiments, the polypeptide having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 809-811, 819, 826, 752, 777, 823, 668-671, 675, 650, 752, 774, 777, 806, 812, 816, 817, 818, 825, 827, 832, 970-982, or a variant thereof.In some embodiments, the polypeptide having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 808, 810-811, 819, 826, 752, 777, or 823, or a variant thereof. In some embodiments, the eukaryotic nucleic acid sequence comprises double-stranded DNA (dsDNA). In some embodiments, the polypeptide having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 810-811, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nucleic acid binding domain, an endonuclease, or a nickase.In some embodiments, the polypeptide having cytosine deaminase activity further comprises the endonuclease or the nickase, wherein the endonuclease or the nickase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597, 1120, 1122-1127, 1647, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nickase, wherein the nickase comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a uracil DNA glycosylase inhibitor sequence. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises the sequence FAM72A.In some embodiments, the FAM72A sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1121, or a variant thereof.
[0005] In some aspects, the disclosure provides a method of deaminating a cytosine residue in a primate nucleic acid sequence in a cell, the method comprising contacting the primate nucleic acid sequence with a polypeptide having cytosine deaminase activity, the polypeptide comprising a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 599-638, 660-675, and 828-835, or a variant thereof. In some embodiments, the eukaryotic nucleic acid sequence comprises double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), or ribonucleic acid (RNA). In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nucleic acid binding domain, an endonuclease, or a nickase. In some embodiments, the polypeptide having cytosine deaminase activity further comprises the endonuclease or the nickase, wherein the endonuclease or the nickase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597, 1120, 1122-1127, 1647, or a variant thereof.In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nickase, wherein the nickase comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a uracil DNA glycosylase inhibitor sequence. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises the sequence FAM72A. In some embodiments, the FAM72A sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1121, or a variant thereof.
[0006] In some aspects, the disclosure provides nucleic acids comprising an engineered nucleic acid sequence optimized for expression in a mammalian organism, wherein the nucleic acid encodes a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, 970-982, or a variant thereof. In some embodiments, the nucleic acid encodes a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 809-811, 819, 826, 752, 777, 823, 668-671, 675, 650, 752, 774, 777, 806, 812, 816, 817, 818, 825, 827, 832, 832, 970-982, or a variant thereof.
[0007] In some aspects, the disclosure provides nucleic acids encoding any of the polypeptides described herein.
[0008] In some aspects, the disclosure provides a vector comprising any of the nucleic acids described herein.
[0009] In some aspects, the present disclosure provides a fusion polypeptide comprising: (a) a domain having cytosine deaminase activity comprising a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, and 970-982, or a variant thereof; and (b) a nucleic acid binding domain, an endonuclease domain, or a nickase domain. In some embodiments, the domain having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 809-811, 819, 826, 752, 777, 823, 668-671, 675, 650, 752, 774, 777, 806, 812, 816, 817, 818, 825, 827, 832, 832, 970-982, or a variant thereof.In some embodiments, the domain having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 809-811, 819, 826, 752, 777, 823, or a variant thereof. In some embodiments, the fusion polypeptide comprises the endonuclease domain or the nickase domain, wherein the endonuclease domain or the nickase domain comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof. In some embodiments, the fusion protein comprises the nickase domain, wherein the nickase domain comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof.In some embodiments, the fusion protein comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 877-916 or 968-969, or a variant thereof.
[0010] In some aspects, the disclosure provides systems that include: (a) any of a fusion protein (e.g., an endonuclease base editor or an endonuclease-deaminase fusion); and (b) an engineered guide polynucleotide configured to form a complex with the endonuclease, the engineered guide polynucleotide comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease domain. In some embodiments, the engineered guide polynucleotide further comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 88-96, 917-931, 963-967, or 1099-1105, or a variant thereof.
[0011] In some aspects, the disclosure provides a polypeptide having adenosine deaminase activity, wherein the polypeptide has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, or at least 97% identity with any one of SEQ ID NOs: 50, 51, 385-443, 448-475. and a polypeptide comprising a sequence having at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 50, or a variant thereof, wherein the polypeptide comprises a substitution of at least one of residues T2, D7, E10, M13, W24, G32, K38, G45, G51, A63, E66, R75, C91, G93, H97, A107, E108, D109, P110, H124, A126, H129, F150, or S165, or any combination thereof, relative to SEQ ID NO: 50 when optimally aligned. In some embodiments, the substitutions are T2X1, D7X1, E10X1, M13X4, W24X1, G32X1, K38X2, G45X2, G51X5, A63X7, E66X5, E66X2, R75H, C91R, G93X6, H97X6, H97X5, A107X5, E108X2, D109N relative to SEQ ID NO: 50 or MG68-4 when optimally aligned. , P110H, H124X6, A126X2, H129R, H129N, F150P, F150S, S165X5, or any combination thereof, wherein X1 is A or G, X2 is D or E, X3 is N or Q, X4 is R or K, X5 is I, L, M, or V, X6 is F, Y, or W, and X7 is S or T.In some embodiments, the polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 836-860, or a variant thereof. In some embodiments, the polypeptide comprises any one of SEQ ID NOs: 839, 841, 843, 844, 847, 848, 849, 850, 851, 852, 859, or a variant thereof. In some embodiments, the substitution comprises W24G, G51V, E108D, P110H, F150P, D7G, E10G, or H129N, or any combination thereof, relative to SEQ ID NO: 50 or MG68-4 when optimally aligned. In some embodiments, the polypeptide further comprises a nucleic acid binding domain, an endonuclease domain, or a nickase domain. In some embodiments, the polypeptide comprises the endonuclease domain or the nickase domain, wherein the endonuclease domain or the nickase domain comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof.In some embodiments, the polypeptide comprises the nickase domain, wherein the nickase domain comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof.
[0012] In some aspects, the disclosure provides a system comprising: (a) any of the polypeptides or fusion polypeptides described herein; and (b) an engineered guide polynucleotide configured to form a complex with the endonuclease, the engineered guide polynucleotide comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease domain. In some embodiments, the engineered guide polynucleotide further comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 88-96, 917-931, 963-967, 1099-1105, or a variant thereof.
[0013] In some aspects, the disclosure provides a method for deaminating cytosine residues in a cell, the method comprising introducing into the cell (a) a vector encoding a polypeptide having cytosine deaminase activity and (b) a vector encoding a FAM72A protein. In some embodiments, the vector encoding the FAM72A protein has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:1115. or a variant thereof, or a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1121, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, or 970-982, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nucleic acid binding domain, an endonuclease domain, or a nickase domain.In some embodiments, the peptide having cytosine deaminase activity comprises the endonuclease domain or the nickase domain, wherein the endonuclease domain or the nickase domain comprises a sequence having at least at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity comprises the nickase domain, wherein the nickase domain comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof.
[0014] In some aspects, the disclosure provides an engineered nucleic acid editing polypeptide comprising (i) a sequence having cytosine deaminase activity and (ii) a sequence derived from a FAM72A protein. In some embodiments, the sequence having cytosine deaminase activity has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, 970-982, or a variant thereof. In some embodiments, the sequence derived from the FAM72A sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1121, or a variant thereof. In some embodiments, the polypeptide further comprises an endonuclease sequence comprising a RuvC domain and an HNH domain, wherein the endonuclease sequence is a class 2, type II endonuclease sequence. In some embodiments, the RuvC domain lacks nuclease activity. In some embodiments, the endonuclease comprises a nickase.In some embodiments, the Class 2, Type II endonuclease sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof. In some embodiments, the Class 2, Type II endonuclease comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597 when optimally aligned.
[0015] In some aspects, the present disclosure provides methods for editing a cytosine residue to a thymine residue in a cell, the method comprising contacting the cell with any of the cytosine deaminase fusion polypeptides described herein. In some embodiments, the cell is a prokaryotic cell, a eukaryotic cell, a mammalian cell, a primate cell, or a human cell.
[0016] In some aspects, the disclosure provides an engineered nucleic acid editing polypeptide comprising multiple domains derived from a Class 2, Type II endonuclease, the domains comprising RUVC-I, REC, HNH, RUVC-III, and WED domains; and a domain comprising a base editor sequence, wherein the base editor sequence is inserted (a) within the RUVC-I domain, (b) within the REC domain, (c) within the HNH domain, (d) within the RUV-CIII domain, (e) within the WED domain, (f) before the HNH domain, (g) before the RUV-CIII domain, or (h) between the RUVC-III and the WED domain. In some embodiments, the Class 2, Type II endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof. In some embodiments, the Class 2, Type II endonuclease comprises a sequence having at least 80% sequence identity to SEQ ID NO: 1647, or a variant thereof. In some embodiments, the base editor sequence comprises a deaminase sequence.In some embodiments, the deaminase sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, 970-982, 50, 51, 385-443, 448-475, or a variant thereof. In some embodiments, the cytosine deaminase sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, 970-982, or a variant thereof. In some embodiments, the cytosine deaminase sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 50, 51, 385-443, 448-475, or a variant thereof. In some embodiments, the deaminase has at least 80% sequence identity to SEQ ID NO: 386, or a variant thereof.In some embodiments, the deaminase sequence, when optimally aligned, comprises one of the following substitutions relative to SEQ ID NO: 50 or MG68-4: T2, D7, E10, M13, W24, G32, K38, G45, G51, A63, E66, R75, C91, G93, H97, A107, E108, D109, P110, H124, A126, H129, F150, or S165, or any combination thereof. In some embodiments, the engineered nucleic acid editing polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1128-1160, or a variant thereof. In some embodiments, the engineered nucleic acid editing polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1137, 1140, 1142, 1143, 1146, 1149, 1151-1158, or a variant thereof.In some embodiments, the engineered nucleic acid editing polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1139, 1152, 1158, or a variant thereof.
[0017] In some embodiments, the disclosure provides a polypeptide having adenosine deaminase activity, the polypeptide having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 129%. Provided are polypeptides comprising a sequence having at least 98%, or at least 99%, sequence identity, or variants thereof, wherein the polypeptides, when optimally aligned, comprise a substitution of a wild-type residue for a non-wild-type residue at residue 109 and one other residue, including any one of 24, 37, 49, 52, 83, 85, 107, 110, 112, 120, 123, 124, 147, 148, 150, 156, 157, 158, 166, 167, or 129, or any combination thereof, relative to SEQ ID NO: 386. In some embodiments, the sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:386. In some embodiments, the polypeptide, when optimally aligned, comprises a 109N substitution and at least one other substitution including any one of 24R, 37L, 49A, 52L, 83S, 85F, 107V, 110S, 112R, 120N, 123N, 124Y, 147C, 148Y, 148R, 150Y, 156V, 157F, 158N, 166I, or 129N, or any combination thereof, relative to SEQ ID NO: 386. In some embodiments, the peptide comprises any of the substitutions depicted in Figure 34B.The polypeptide has at least 80% sequence identity to any one of SEQ ID NOs: 1161-1183, or a variant thereof. In some embodiments, the polypeptide has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1170, 1179, or 1166, or a variant thereof. In some embodiments, the polypeptide further comprises an endonuclease or a nickase. In some embodiments, the polypeptide comprises the endonuclease or nickase, wherein the endonuclease or nickase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof. In some embodiments, the polypeptide comprises the nickase, wherein the nickase comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof.
[0018] In some aspects, the disclosure provides a polypeptide having cytosine deaminase activity, wherein the polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, 970-982, or a variant thereof, wherein the polypeptide comprises at least one of the modifications described in Table 12C. In some embodiments, the polypeptide optionally comprises one of the following amino acids: W90A, W90F, W90H, W90Y, Y120F, Y120H, Y121F, Y121H, Y121Q, Y121A, Y121D, Y121W, H122Y, H122F, H122I, H122A, H122W, H122D, Y 121T, R33A, R34A, R34K, H122A, R33A, R34A, R52A, N57G, H122A, E123A, E123Q, W127F, W127 H, W127Q, W127A, W127D, R39A, K40A, H128A, N63G, R58A, H121F, H121Y, H121Q, H121A, H121 D, H121W, R33A, K34A, H122A, H121A, R52A, P26R, P26A, N27R, N27A, W44A, W45A, K49G, S50G , R51G, R121A, I122A, N123A, Y88F, Y120F, P22R, P22A, K23A, K41R, K41A, E54A, E54A, E55A , K30A, K30R, M32A, M32K, Y117A, K118A, I119A, I119H, R120A, R121A, P46A, P46R, N29A, R27A, or N50G, or any combination thereof.In some embodiments, the polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1208-1315, or a variant thereof.
[0019] In some aspects, the present disclosure provides a polypeptide having cytosine deaminase activity, the polypeptide having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, or at least 85% similar to any one of SEQ ID NOs: 835, 1275, 668, 774, 818, 671, 667, 650, 827, 819, 823, 814, 813, 817, 628, 826, 1223, 834, 618, 621, 669, 833, 830. , at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, or a variant thereof, and an endonuclease or nickase. In some embodiments, the endonuclease comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, or 1122-1127, 1647, or a variant thereof. In some embodiments, the polypeptide comprises the nickase, wherein the nickase comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof. In some embodiments, the cytosine deaminase sequence has at least 80% sequence identity to any one of SEQ ID NOs:1275, 835, or 774, or a combination thereof.
[0020] In some aspects, the disclosure provides a polypeptide having adenosine deaminase activity, wherein the polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 50, 51, 385-443, 448-475, 1015-1098, or a variant thereof, wherein the polypeptide comprises any of the combinations of wild-type for non-wild-type residue substitutions listed in Table 12D. In some embodiments, the polypeptide has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1556-1638, or a variant thereof. In some embodiments, the polypeptide further comprises an endonuclease or a nickase. In some embodiments, the polypeptide comprises the endonuclease or nickase, wherein the endonuclease or nickase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, or 1122-1127, 1647, or a variant thereof.In some embodiments, the polypeptide comprises the nickase, wherein the nickase comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof.
[0021] In some aspects, the disclosure provides a polypeptide having adenosine deaminase activity, wherein the polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 50, 51, 385-443, 448-475, 1015-1098, or a variant thereof, wherein the polypeptide comprises any of the combinations of wild-type for non-wild-type residue substitutions listed in Table 13. In some embodiments, the sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 386, or a variant thereof. In some embodiments, the polypeptide further comprises an endonuclease or a nickase. In some embodiments, the polypeptide comprises the endonuclease or nickase, wherein the endonuclease or nickase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, or 1122-1127, 1647, or a variant thereof.In some embodiments, the polypeptide comprises the nickase, wherein the nickase comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof.
[0022] In some aspects, the disclosure provides a method of editing an APOA1 locus in a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide nucleic acid structure, wherein the engineered guide nucleic acid structure is configured to form a complex with the endonuclease, wherein the engineered guide nucleic acid structure comprises a spacer sequence configured to hybridize to a region of the APOA1 locus, wherein the engineered guide nucleic acid structure comprises at least one of 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 1109, 1111 The targeting sequence or its reverse complement has at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to 22, 23, 24, 25, or 26 contiguous nucleotides. In some embodiments, the engineered guide nucleic acid structure has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1431-1454. In some embodiments, the engineered guide nucleic acid structure comprises any of the nucleotide modifications listed in Table 13A. In some embodiments, the RNA-guided endonuclease is a Class 2, Type II endonuclease.In some embodiments, the RNA-guided endonuclease has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof.
[0023] In some aspects, the disclosure provides a method of editing an ANGPTL3 locus in a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide nucleic acid structure, wherein the engineered guide nucleic acid structure is configured to form a complex with the endonuclease, the engineered guide nucleic acid structure comprising a spacer sequence configured to hybridize to a region of the ANGPTL3 locus, the engineered guide nucleic acid structure comprising at least 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 110, The engineered guide nucleic acid structure comprises a targeting sequence or its reverse complement having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to 1, 22, 23, 24, 25, or 26 contiguous nucleotides. The engineered guide nucleic acid structure has at least 80% identity to any one of SEQ ID NOs: 1479-1483. In some embodiments, the engineered guide nucleic acid structure comprises any of the nucleotide modifications listed in Table 13A. In some embodiments, the RNA-guided endonuclease is a Class 2, Type II endonuclease. In some embodiments, the RNA-guided endonuclease has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof.
[0024] In some aspects, the disclosure provides a method of editing a TRAC locus in a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide nucleic acid structure, wherein the engineered guide nucleic acid structure is configured to form a complex with the endonuclease, the engineered guide nucleic acid structure comprising a spacer sequence configured to hybridize to a region of the TRAC locus, the engineered guide nucleic acid structure comprising at least 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 1109, 1109, 1111, 112 The targeting sequence or its reverse complement has at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with 2, 23, 24, 25, or 26 contiguous nucleotides. In some embodiments, the engineered guide nucleic acid structure has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1489-1490. In some embodiments, the auxiliary engineered guide nucleic acid structure comprises any of the nucleotide modifications listed in Table 13A. In some embodiments, the RNA-guided endonuclease is a Class 2, Type II endonuclease.In some embodiments, the RNA-guided endonuclease has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, 1122-1127, 1647, or a variant thereof.
[0025] In some aspects, the disclosure provides engineered adenosine base editor polypeptides, wherein the polypeptides comprise a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1647-1653.
[0026] In some aspects, the disclosure provides methods for deaminating cytosine residues in a eukaryotic nucleic acid sequence in a cell, the method comprising contacting the eukaryotic nucleic acid sequence with a peptide having cytosine deaminase activity, the peptide comprising a sequence having at least at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, and 970-982, or a variant thereof. In some embodiments, the eukaryotic nucleic acid sequence is a mammalian, primate, or human nucleic acid sequence. In some embodiments, the cell is a mammalian, primate, or human cell. In some embodiments, the eukaryotic nucleic acid sequence comprises single-stranded DNA (ssDNA) or ribonucleic acid (RNA). In some embodiments, the polypeptide having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 809-811, 819, 826, 752, 777, 823, 668-671, 675, 650, 752, 774, 777, 806, 812, 816, 817, 818, 825, 827, 832, 970-982, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 808, 810-811, 819, 826, 752, 777, or 823, or a variant thereof.In some embodiments, the eukaryotic nucleic acid sequence comprises double-stranded DNA (dsDNA). In some embodiments, the polypeptide having cytosine deaminase activity comprises a sequence having 80% identity to any one of SEQ ID NOs: 810-811. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nucleic acid binding domain, an endonuclease, or a nickase. In some embodiments, the polypeptide having cytosine deaminase activity further comprises the endonuclease or the nickase, wherein the endonuclease or the nickase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597, 1120, or 1122-1127, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nickase, wherein the nickase comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a uracil DNA glycosylase inhibitor sequence.In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises the sequence FAM72A. In some embodiments, the FAM72A sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1121, or a variant thereof.
[0027] In some aspects, the disclosure provides a method of deaminating a cytosine residue in a primate nucleic acid sequence in a cell, the method comprising contacting the primate nucleic acid sequence with a polypeptide having cytosine deaminase activity, the polypeptide comprising a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 599-638, 660-675, or 828-835, or a variant thereof. In some embodiments, the eukaryotic nucleic acid sequence comprises double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), or ribonucleic acid (RNA). In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nucleic acid binding domain, an endonuclease, or a nickase. In some embodiments, the polypeptide having cytosine deaminase activity further comprises the endonuclease or the nickase, wherein the endonuclease or the nickase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597, 1120, or 1122-1127, or a variant thereof.In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nickase, wherein the nickase comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a uracil DNA glycosylase inhibitor sequence. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises the sequence FAM72A. In some embodiments, the FAM72A sequence has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1121, or a variant thereof.
[0028] In some aspects, the disclosure provides nucleic acids comprising an engineered nucleic acid sequence optimized for expression in a mammalian organism, wherein the nucleic acid encodes a sequence having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, 970-982, or a variant thereof. In some embodiments, the nucleic acid comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 809-811, 819, 826, 752, 777, 823, 668-671, 675, 650, 752, 774, 777, 806, 812, 816, 817, 818, 825, 827, 832, 832, 970-982, or a variant thereof.
[0029] In some aspects, the present disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the viral vector is a non-viral vector or a viral vector. In some embodiments, the vector is a plasmid, minicircle, or plasmid vector. In some embodiments, the viral vector is an AAV vector.
[0030] In some aspects, the present disclosure provides a fusion polypeptide comprising: (a) a domain having cytosine deaminase activity comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, and 970-982, or a variant thereof; and (b) a nucleic acid binding domain, an endonuclease domain, or a nickase domain. In some embodiments, the domain having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 809-811, 819, 826, 752, 777, 823, 668-671, 675, 650, 752, 774, 777, 806, 812, 816, 817, 818, 825, 827, 832, 832, 970-982, or a variant thereof. In some embodiments, the domain having cytosine deaminase activity comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 809-811, 819, 826, 752, 777, 823, or a variant thereof.In some embodiments, the fusion polypeptide comprises the endonuclease domain or the nickase domain, wherein the endonuclease domain or the nickase domain comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, or 1122-1127, or a variant thereof. In some embodiments, the fusion protein comprises the nickase domain, wherein the nickase domain comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof. In some embodiments, the fusion protein comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs:877-916 or 968-969, or a variant thereof.
[0031] In some aspects, the disclosure provides a system comprising: (a) any of the fusion polypeptides described herein; and (b) an engineered guide polynucleotide configured to form a complex with the endonuclease, the engineered guide polynucleotide comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease domain. In some embodiments, the engineered guide polynucleotide further comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 88-96, 917-931, 963-967, or 1099-1105, or a variant thereof.
[0032] In some aspects, the disclosure provides a polypeptide having adenosine deaminase activity, wherein the polypeptide has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, or at least 97% identity with any one of SEQ ID NOs: 50, 51, 385-443, 448-475. and a polypeptide comprising a sequence having at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 50, or a variant thereof, wherein the polypeptide comprises a substitution of at least one of residues T2, D7, E10, M13, W24, G32, K38, G45, G51, A63, E66, R75, C91, G93, H97, A107, E108, D109, P110, H124, A126, H129, F150, or S165, or any combination thereof, relative to SEQ ID NO: 50 when optimally aligned. In some embodiments, the substitutions are T2X1, D7X1, E10X1, M13X4, W24X1, G32X1, K38X2, G45X2, G51X5, A63X7, E66X5, E66X2, R75H, C91R, G93X6, H97X6, H97X5, A107X5, E108X2, D109N, P11 OH, H124X6, A126X2, H129R, H129N, F150P, F150S, S165X5, or any combination thereof, wherein X1 is A or G, X2 is D or E, X3 is N or Q, X4 is R or K, X5 is I, L, M, or V, X6 is F, Y, or W, and X7 is S or T.In some embodiments, the polypeptide comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 836-860, or a variant thereof. In some embodiments, the polypeptide comprises any one of SEQ ID NOs: 839, 841, 843, 844, 847, 848, 849, 850, 851, 852, or 859. In some embodiments, the substitution, when optimally aligned, comprises W24G, G51V, E108D, P110H, F150P, D7G, E10G, or H129N, or any combination thereof, relative to SEQ ID NO: 50. In some embodiments, the polypeptide further comprises a nucleic acid binding domain, an endonuclease domain, or a nickase domain. In some embodiments, the polypeptide comprises the endonuclease domain or the nickase domain, wherein the endonuclease domain or the nickase domain comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 70-78, 596, 597-598, 1120, or 1122-1127, or a variant thereof.In some embodiments, the polypeptide comprises the nickase domain, wherein the nickase domain comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof.
[0033] In some aspects, the present disclosure provides a system comprising: (a) any of the polypeptides for a base editor fusion described herein (e.g., an endonuclease-deaminase fusion); and (b) an engineered guide polynucleotide configured to form a complex with the endonuclease, the engineered guide polynucleotide comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease domain. In some embodiments, the engineered guide polynucleotide further comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 88-96, 917-931, 963-967, or 1099-1105.
[0034] In some aspects, the disclosure provides methods for deaminating cytosine residues in a cell, the method comprising introducing into the cell (a) a vector encoding a polypeptide having cytosine deaminase activity and (b) a vector encoding a FAM72A protein. In some embodiments, the vector encoding the FAM72A protein comprises a sequence having at least 80% identity to SEQ ID NO: 1115 or encodes a sequence having at least 80% identity to SEQ ID NO: 1121. In some embodiments, the polypeptide having cytosine deaminase activity comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 1-49, 444-447, 599-675, 744-835, 970-982, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity further comprises a nucleic acid binding domain, an endonuclease domain, or a nickase domain. In some embodiments, the polypeptide having cytosine deaminase activity comprises the endonuclease domain or the nickase domain, wherein the endonuclease domain or the nickase domain comprises a sequence having at least 80% identity to any one of SEQ ID NOs:70-78, 596, 597-598, 1120, or 1122-1127, or a variant thereof. In some embodiments, the polypeptide having cytosine deaminase activity comprises the nickase domain, wherein the nickase domain comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NOs:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597, or any combination thereof.
[0035] In some aspects, the disclosure provides an engineered nucleic acid editing system comprising: an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, the endonuclease is a class 2, type II endonuclease, and the endonuclease is configured to lack nuclease activity; a base editor bound to the endonuclease; and an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the RuvC domain lacks nuclease activity. In some embodiments, the class 2, type II endonuclease comprises a nickase mutation. In some embodiments, the Class 2, Type II endonuclease, when optimally aligned, comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597. In some embodiments, the endonuclease, when optimally aligned, comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:72, or residue 17 relative to SEQ ID NO:75. In some embodiments, the endonuclease comprises a sequence having at least 95% sequence identity to any one of SEQ ID NOs:70-78, 597, or a variant thereof.In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: an endonuclease comprising a sequence having 95% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof; a base editor bound to the endonuclease; and an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a ribonucleic acid sequence configured to bind to the endonuclease. In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising any one of SEQ ID NOs: 360-368 or 598, or a variant thereof, wherein the endonuclease is a class 2, type II endonuclease, and the endonuclease is configured to lack nuclease activity; a base editor bound to the endonuclease; and an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the endonuclease comprises a nickase mutation. In some embodiments, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid. In some embodiments, the Class 2, Type II endonuclease comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, or residue 10 relative to SEQ ID NO:597 when optimally aligned.In some embodiments, the base editor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, or 599-675, or a variant thereof. In some embodiments, the base editor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 50-51 or 385-390. In some embodiments, the RuvC domain lacks nuclease activity. In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has less than 80% identity to Cas9 endonuclease. In some embodiments, the endonuclease further comprises an HNH domain. In some embodiments, the guide ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to a non-degenerate nucleotide in any one of SEQ ID NOs: 88-96, 488-489, or 679-680, or a variant thereof. In some aspects, the present disclosure provides an engineered nucleic acid editing system, the system comprising an engineered guide ribonucleic acid structure comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a ribonucleic acid sequence configured to bind to an endonuclease, the engineered ribonucleic acid sequence comprising a sequence having at least 80% sequence identity to a non-degenerate nucleotide in any one of SEQ ID NOs: 88-96, 488-489, or 679-680, or a variant thereof; a Class 2, Type II endonuclease configured to bind to the engineered guide ribonucleic acid; and a base editor bound to the endonuclease. In some embodiments, the base editor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 50-51 or 385-390. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence selected from the group consisting of SEQ ID NOs: 360-368 or 598.In some embodiments, the base editor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, or 599-675, or a variant thereof. In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenosine deaminase comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 50-51, 57, 385-443, 448-475, or 595, or a variant thereof. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 1-49, 444-447, 594, or 58-66, or a variant thereof. In some embodiments, the system further comprises a uracil DNA glycosylase inhibitor coupled to the endonuclease or the base editor. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67. In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some embodiments, the engineered guide ribonucleic acid structure comprises one ribonucleic acid polynucleotide comprising the guide ribonucleic acid sequence and the tracr ribonucleic acid sequence. In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence having at least 90% identity to one selected from SEQ ID NOs: 369-384, or a variant thereof.In some embodiments, the endonuclease is covalently linked directly to the base editor or covalently linked to the base editor via a linker. In some embodiments, the endonuclease comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:71, 72, or 74, residue 12 relative to SEQ ID NO:73 or 78, residue 17 relative to SEQ ID NO:75, residue 23 relative to SEQ ID NO:76, residue 8 relative to SEQ ID NO:77, or residue 10 relative to SEQ ID NO:597, when optimally aligned. In some embodiments, the endonuclease comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO:70, residue 13 relative to SEQ ID NO:72, or residue 17 relative to SEQ ID NO:75, when optimally aligned. In some embodiments, a polypeptide comprises the endonuclease and the base editor. In some embodiments, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid. In some embodiments, the system comprises Mg. 2+In some embodiments, (a) the endonuclease comprises a sequence at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 70, 71, 73, 74, 76, 78, 77, or 78, or a variant thereof, (b) the guide RNA structure comprises a sequence at least 70%, at least 80%, or at least 90% identical to the non-degenerate nucleotides of any one of SEQ ID NOs: 88, 89, 91, 92, 94, 96, 95, or 488, (c) the endonuclease is configured to bind to a PAM comprising any one of SEQ ID NOs: 360, 361, 363, 365, 367, or 368, or (d) the base editor comprises a sequence at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 58 or 595, or a variant thereof. In some embodiments, (a) the endonuclease comprises a sequence at least 70%, at least 80%, or at least 90% identical to any one of SEQ ID NOs: 70, 71, or 78, or a variant thereof, (b) the guide RNA structure comprises a sequence at least 70%, at least 80%, or at least 90% identical to at least one non-degenerate nucleotide of SEQ ID NOs: 88, 89, or 96, (c) the endonuclease is configured to bind to a PAM comprising any one of SEQ ID NOs: 360, 362, or 368, or (d) the base editor comprises a sequence at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 594, or a variant thereof. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or Smith-Waterman homology search algorithms. In some embodiments, the sequence identity is determined by the BLASTP homology search algorithm using a BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment. In some embodiments, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid.
[0036] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, the nucleic acid encoding a Class 2, Type II endonuclease linked to a base editor, wherein the endonuclease is derived from an uncultured microorganism.
[0037] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes an endonuclease having at least 70% sequence identity to any one of SEQ ID NOs: 70-78 linked to a base editor. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence having at least 90% identity to one selected from SEQ ID NOs: 369-384, or a variant thereof. In some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human.
[0038] In some aspects, the disclosure provides a vector comprising a nucleic acid sequence encoding a class 2, type II endonuclease linked to a base editor, wherein the endonuclease is derived from an uncultured microorganism.
[0039] In some aspects, the present disclosure provides a vector comprising a nucleic acid according to any of the aspects or embodiments described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with an endonuclease, the guide ribonucleic acid structure comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.
[0040] In some aspects, the present disclosure provides a cell comprising a vector according to any of the aspects or embodiments described herein.
[0041] In some aspects, the present disclosure provides a method of producing an endonuclease, comprising culturing a cell according to any of the aspects or embodiments described herein.
[0042] In some aspects, the disclosure provides methods of modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising contacting the double-stranded deoxyribonucleic acid polynucleotide with a complex comprising an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, the endonuclease is a class 2, type II endonuclease, and the endonuclease is configured to lack nuclease activity, a base editor bound to the endonuclease, and an engineered guide ribonucleic acid structure configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM). In some embodiments, the endonuclease comprising a RuvC domain and an HNH domain is covalently linked directly to the base editor or covalently linked to the base editor via a linker. In some embodiments, the endonuclease comprising a RuvC domain and an HNH domain comprises a sequence having at least 95% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof.
[0043] In some aspects, the disclosure provides methods of modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising contacting the double-stranded deoxyribonucleic acid polynucleotide with a complex comprising a class 2, type II endonuclease, a base editor bound to the endonuclease, and an engineered guide ribonucleic acid structure configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), wherein the PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 70-78 or 597. In some embodiments, the class 2, type II endonuclease is covalently linked to the base editor or is linked to the base editor via a linker. In some embodiments, the base editor comprises a sequence having at least 70%, at least 80%, at least 90%, or at least 95% identity to a sequence selected from SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, or 599-675, or a variant thereof. In some embodiments, the base editor comprises an adenine deaminase, the double-stranded deoxyribonucleic acid polynucleotide comprises an adenine, and modifying the double-stranded deoxyribonucleic acid polypeptide comprises converting the adenine to guanine. In some embodiments, the adenine deaminase comprises a sequence having at least 70%, 80%, 90%, or 95% sequence identity to any one of SEQ ID NOs: 50-51, 57, 385-443, 448-475, or 595, or a variant thereof. In some embodiments, the base editor comprises a cytosine deaminase, the double-stranded deoxyribonucleic acid polynucleotide comprises a cytosine, and modifying the double-stranded deoxyribonucleic acid polypeptide comprises converting the cytosine to a uracil. In some embodiments, the cytosine deaminase comprises a sequence having at least 70%, 80%, 90%, or 95% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 594, or 58-66, or a variant thereof.In some embodiments, the complex further comprises a uracil DNA glycosylase inhibitor bound to the endonuclease or the base editor. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67, or a variant thereof. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of the engineered guide ribonucleic acid structure and a second strand comprising the PAM. In some embodiments, the PAM is immediately adjacent to the 3' end of the sequence complementary to the sequence of the engineered guide ribonucleic acid structure. In some embodiments, the Class 2, Type II endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the Class 2, Type II endonuclease is derived from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0044] In some aspects, the present disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering to the target nucleic acid locus an engineered nucleic acid editing system of any of the aspects or embodiments described herein, wherein the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure, the complex being configured such that upon binding to the target nucleic acid locus, the complex modifies a nucleotide at the target nucleic acid locus. In some embodiments, the engineered nucleic acid editing system comprises adenine deaminase, the nucleotide is adenine, and modifying the target nucleic acid locus comprises converting the adenine to guanine. In some embodiments, the engineered nucleic acid editing system comprises cytidine deaminase and a uracil DNA glycosylase inhibitor, the nucleotide is cytosine, and modifying the target nucleic acid locus comprises converting the adenine to uracil. In some embodiments, the target nucleic acid locus comprises genomic DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, the cell is in an animal. In some embodiments, the cell is in a cochlea. In some embodiments, the cell is in an embryo. In some embodiments, the embryo is a two-cell stage embryo. In some embodiments, the embryo is a mouse embryo. In some embodiments, delivering an engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a nucleic acid of any of the aspects or embodiments described herein, or a vector of any of the aspects or embodiments described herein. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease.In some embodiments, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a capped mRNA containing the open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid structure operably linked to a ribonucleic acid (RNA) pol III promoter.
[0045] In some aspects, the disclosure provides an engineered nucleic acid editing polypeptide, the engineered nucleic acid editing polypeptide comprising: an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, the endonuclease is a Class 2, Type II endonuclease, and the endonuclease is configured to lack nuclease activity; and a base editor coupled to the endonuclease. In some embodiments, the endonuclease comprises a sequence having at least 95% sequence identity to any one of SEQ ID NOs: 70-78, 597, or a variant thereof.
[0046] In some aspects, the disclosure provides an engineered nucleic acid editing peptide comprising: an endonuclease having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof, wherein the endonuclease is configured to lack nuclease activity; and a base editor bound to the endonuclease. In some aspects, the disclosure provides an engineered nucleic acid editing polypeptide, the engineered nucleic acid editing polypeptide comprising: an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising any one of SEQ ID NOs: 360-368 or 598, wherein the endonuclease is a Class 2, Type II endonuclease, and the endonuclease is configured to lack nuclease activity; and a base editor coupled to the endonuclease. In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has less than 80% identity to Cas9 endonuclease. In some embodiments, the endonuclease further comprises an HNH domain. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to about 60-90 contiguous nucleotides selected from any one of SEQ ID NOs: 88-96, 488, 489, and 679-680. In some embodiments, the base editor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, or 599-675, or a variant thereof.In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenosine deaminase comprises a sequence having at least 70%, 80%, 90%, or 95% sequence identity to any one of SEQ ID NOs: 50-51, 57, 385-443, 448-475, or 595, or a variant thereof. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises a sequence having at least 70%, 80%, 90%, or 95% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 594, or 58-66, or a variant thereof.
[0047] In some aspects, the disclosure provides an engineered nucleic acid editing peptide, the engineered nucleic acid editing polypeptide comprising: an endonuclease (wherein the endonuclease is configured to lack activity); and a base editor bound to the endonuclease, wherein the base editor comprises a sequence having at least 70%, 80%, 90%, or 95% sequence identity to any one of SEQ ID NOs: 1-51, 385-386, 387-443, 444-447, 488-475, or 595, or a variant thereof. In some embodiments, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid. In some embodiments, the endonuclease is configured to be catalytically ineffective. In some embodiments, the endonuclease is a class II, type II endonuclease, or a class II, type V endonuclease. In some embodiments, the endonuclease comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 70-78, 597, or a variant thereof. In some embodiments, the endonuclease comprises a nickase mutation. In some embodiments, the endonuclease comprises an aspartic acid to alanine mutation at residue 9 relative to SEQ ID NO: 70, residue 13 relative to SEQ ID NOs: 71, 72, or 74, residue 12 relative to SEQ ID NO: 73, residue 17 relative to SEQ ID NO: 75, residue 23 relative to SEQ ID NO: 76, or residue 10 relative to SEQ ID NO: 597, when optimally aligned. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence selected from the group consisting of SEQ ID NOs: 360-368 or 598. In some embodiments, the base editor is an adenine deaminase.In some embodiments, the adenosine deaminase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 50-51, 385-443, or 448-475, or a variant thereof. In some embodiments, the adenosine deaminase comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 50-51, 385-390, or 595, or a variant thereof. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 1-49, 444-447, or a variant thereof. In some embodiments, the polypeptide further comprises a uracil DNA glycosylase inhibitor linked to the endonuclease or the base editor. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67, or a variant thereof. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence having at least 90% identity to one selected from SEQ ID NOs: 369-384, or a variant thereof. In some embodiments, the endonuclease is covalently linked directly to the base editor or covalently linked to the base editor via a linker.
[0048] In some aspects, the disclosure provides nucleic acids comprising engineered nucleic acid sequences optimized for expression in an organism, wherein the nucleic acid encodes a sequence having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-51, 385-386, 387-443, 444-447, or 488-475, or a variant thereof. In some embodiments, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human.
[0049] In some aspects, the present disclosure provides a vector comprising a nucleic acid according to any of the aspects or embodiments described herein, hi some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.
[0050] In some aspects, the present disclosure provides a cell comprising the vector of any one of the aspects or embodiments described herein.
[0051] In some aspects, the present disclosure provides a method of producing a base editor, comprising culturing the cell of any one of the aspects or embodiments described herein.
[0052] In some aspects, the present disclosure provides a system comprising: (a) a nucleic acid editing polypeptide according to any of the aspects or embodiments described herein; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the nucleic acid editing polypeptide, the engineered guide ribonucleic acid structure comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the guide ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 88-96, 488-489, or 679-680.
[0053] In some aspects, the present disclosure provides methods of modifying a target nucleic acid locus, the method comprising delivering the engineered nucleic acid editing polypeptide of any of the aspects or embodiments described herein or the system of any of the aspects or embodiments described herein to the target nucleic acid locus, wherein the complex is configured such that upon binding of the complex to the target nucleic acid locus, the complex modifies the target nucleic acid locus.
[0054] In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: (a) an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, the endonuclease is a class 2, type II endonuclease, and the RuvC domain lacks nuclease activity; (b) a base editor bound to the endonuclease; and (c) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the endonuclease comprises a sequence having at least 95% sequence identity to any one of SEQ ID NOs: 70-78.
[0055] In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: (a) an endonuclease having 95% sequence identity to any one of SEQ ID NOs:70-78, wherein the endonuclease comprises a RuvC domain that lacks nuclease activity; a base editor bound to the endonuclease; and an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease.
[0056] In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: (a) an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising SEQ ID NOs: 360-368 (wherein the endonuclease is a class 2, type II endonuclease, and the endonuclease comprises a RuvC domain lacking nuclease activity); (b) a base editor bound to the endonuclease; and (c) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease.
[0057] In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has less than 80% identity to a Cas9 endonuclease. In some embodiments, the endonuclease further comprises an HNH domain. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to about 60-90 contiguous nucleotides selected from any one of SEQ ID NOs: 88-96, 488, 489, and 679-680.
[0058] In some aspects, the present disclosure provides an engineered nucleic acid editing system, the system comprising: (a) an engineered guide ribonucleic acid structure comprising (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to an endonuclease, the tracr ribonucleic acid sequence comprising a sequence having at least 80% sequence identity to about 60-90 contiguous nucleotides selected from any one of SEQ ID NOs: 88-96, 488, 489, and 679-680; and a Class 2, Type II endonuclease configured to bind to the engineered guide ribonucleic acid.
[0059] In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence selected from the group consisting of SEQ ID NOs: 360-368. In some embodiments, the base editor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 1-51 and 385-475. In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenosine deaminase comprises a sequence having at least 95% identity to SEQ ID NO: 57. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises a sequence having at least 95% identity to SEQ ID NO: 58. In some embodiments, the cytosine deaminase comprises a sequence having at least 95% identity to any one of SEQ ID NOs: 59-66.
[0060] In some embodiments, the engineered nucleic acid editing system further comprises a uracil DNA glycosylase inhibitor, hi some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs:52-56 or SEQ ID NO:67.
[0061] In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some embodiments, the engineered guide nucleic acid structure comprises one ribonucleic acid polynucleotide comprising a guide ribonucleic acid sequence and a tracr ribonucleic acid sequence. In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the endonuclease is covalently linked directly to a base editor or covalently linked to a base editor via a linker. In some embodiments, the polypeptide comprises an endonuclease and a base editor. In some embodiments, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid. In some embodiments, the endonuclease comprises SEQ ID NO: 370. In some embodiments, the system comprises Mg 2+ The present invention further comprises a source of
[0062] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:70, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:88, and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO:360.
[0063] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:71, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:89, and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO:361.
[0064] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:73, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:91, and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO:363.
[0065] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:75, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:93, and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO:365.
[0066] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:76, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:94, and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO:366.
[0067] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:77, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:95, and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO:367.
[0068] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:78, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO:96, and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO:368.
[0069] In some embodiments, the base editor comprises an adenine deaminase. In some embodiments, the adenine deaminase comprises SEQ ID NO: 57. In some embodiments, the base editor comprises a cytosine deaminase. In some embodiments, the cytosine deaminase comprises SEQ ID NO: 58. In some embodiments, the engineered nucleic acid editing system described herein further comprises a uracil DNA glycosylation inhibitor. In some embodiments, the uracil DNA glycosylation inhibitor comprises SEQ ID NO: 67.
[0070] In some embodiments, sequence identity is determined by the BLASTP, CLUSTALW, MUSCLE, MAFFT, or Smith-Waterman homology search algorithm. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using a BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment.
[0071] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes a class 2, type II endonuclease linked to a base editor, and wherein the endonuclease is derived from an uncultured microorganism.
[0072] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes an endonuclease comprising a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 70-78, linked to a base editor. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human.
[0073] In some aspects, the present disclosure provides a vector comprising a nucleic acid sequence encoding a class 2, type II endonuclease linked to a base editor, wherein the endonuclease is derived from an uncultured microorganism. In some embodiments, the vector comprises a nucleic acid described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the nucleic acid comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus. In some aspects, the present disclosure provides a cell comprising the vector described herein. In some aspects, the present disclosure provides a method of producing an endonuclease, comprising culturing a cell described herein.
[0074] In some aspects, the disclosure provides methods of modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising contacting the double-stranded deoxyribonucleic acid polynucleotide with a complex comprising an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, and the endonuclease is a class 2, type II endonuclease, and the RuvC domain lacks nuclease activity, a base editor bound to the endonuclease, and an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM).
[0075] In some embodiments, the endonuclease comprising a RuvC domain and an HNH domain is covalently linked to the base editor directly or via a linker, hi some embodiments, the endonuclease comprising a RuvC domain and an HNH domain comprises a sequence having at least 95% sequence identity to any one of SEQ ID NOs: 70-78.
[0076] In some aspects, the disclosure provides a method of modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising contacting the double-stranded deoxyribonucleic acid polynucleotide with a complex comprising a class 2, type II endonuclease, a base editor bound to the endonuclease, and an engineered guide ribonucleic acid structure configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and the PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 360-368.
[0077] In some embodiments, the Class 2, Type II endonuclease is covalently linked to the base editor or is linked to the base editor via a linker. In some embodiments, the base editor comprises a sequence having at least 70%, at least 80%, at least 90%, or at least 95% identity to a sequence selected from SEQ ID NOs: 1-51 and 385-475. In some embodiments, the base editor comprises an adenine deaminase, the double-stranded deoxyribonucleic acid polynucleotide comprises an adenine, and modifying the double-stranded deoxyribonucleic acid polypeptide comprises converting the adenine to guanine. In some embodiments, the adenine deaminase comprises a sequence having at least 95% identity to SEQ ID NO: 57.
[0078] In some embodiments, the base editor comprises a cytosine deaminase, the double-stranded deoxyribonucleic acid polynucleotide comprises a cytosine, and modifying the double-stranded deoxyribonucleic acid polypeptide comprises converting the cytosine to a uracil. In some embodiments, the cytosine deaminase comprises a sequence having at least 95% identity to SEQ ID NO: 58. In some embodiments, the cytosine deaminase comprises a sequence having at least 95% identity to any one of SEQ ID NOs: 59-66.
[0079] In some embodiments, the complex further comprises a uracil DNA glycosylase inhibitor. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of the engineered guide ribonucleic acid structure and a second strand comprising the PAM. In some embodiments, the PAM is immediately adjacent to the 3' end of the sequence complementary to a sequence of the engineered guide ribonucleic acid structure.
[0080] In some embodiments, the Class 2, Type II endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the Class 2, Type II endonuclease is derived from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0081] In some aspects, the disclosure provides methods of modifying a target nucleic acid locus, the method comprising delivering an engineered nucleic acid editing system described herein to a target nucleic acid locus, wherein an endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure, the complex being configured such that upon binding of the complex to the target nucleic acid locus, the complex modifies a nucleotide at the target nucleic acid locus.
[0082] In some embodiments, the engineered nucleic acid editing system comprises adenine deaminase, the nucleotide is adenine, and modifying the target nucleic acid locus comprises converting the adenine to guanine. In some embodiments, the engineered nucleic acid editing system comprises cytidine deaminase and a uracil DNA glycosylase inhibitor, the nucleotide is cytosine, and modifying the target nucleic acid locus comprises converting the adenine to uracil. In some embodiments, the target nucleic acid locus comprises genomic DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, the cell is in an animal.
[0083] In some embodiments, the cell is in a cochlea. In some embodiments, the cell is in an embryo. In some embodiments, the embryo is a two-cell stage embryo. In some embodiments, the embryo is a mouse embryo. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a nucleic acid described herein or a vector described herein. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease.
[0084] In some embodiments, the nucleic acid comprises a promoter to which an open reading frame encoding an endonuclease is operably linked. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide ribonucleic acid structure operably linked to a ribonucleic acid (RNA) pol III promoter.
[0085] In some aspects, the disclosure provides an engineered nucleic acid editing polypeptide, comprising: an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, the endonuclease is a Class 2, Type II endonuclease, and the RuvC domain lacks nuclease activity; and a base editor bound to the endonuclease. In some embodiments, the endonuclease comprises a sequence having at least 95% sequence identity to any one of SEQ ID NOs: 70-78.
[0086] In some aspects, the disclosure provides an engineered nucleic acid editing polypeptide, the engineered nucleic acid editing polypeptide comprising: an endonuclease having at least 95% sequence identity to any one of SEQ ID NOs:70-78, the endonuclease comprising a RuvC domain lacking nuclease activity; and a base editor bound to the endonuclease.
[0087] In some aspects, the disclosure provides an engineered nucleic acid editing polypeptide, comprising: an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising SEQ ID NOs: 360-368, wherein the endonuclease is a Class 2, Type II endonuclease and comprises a RuvC domain that lacks nuclease activity; and a base editor coupled to the endonuclease.
[0088] In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has less than 80% identity to a Cas9 endonuclease. In some embodiments, the endonuclease further comprises an HNH domain. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to about 60-90 contiguous nucleotides selected from any one of SEQ ID NOs: 88-96, 488, 489, and 679-680. In some embodiments, the base editor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 1-51 and 385-475. In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenosine deaminase comprises a sequence having at least 95% identity to SEQ ID NO: 57. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises a sequence having at least 95% identity to SEQ ID NO: 58. In some embodiments, the adenosine deaminase comprises a sequence having at least 95% identity to any one of SEQ ID NOs: 59-66.
[0089]
[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0090] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]
[0091] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (herein referred to as "Figure" and "FIG.").
[0092] [Figure 1] 1 shows exemplary organizations of various classes and types of CRISPR loci. [Figure 2] 1 shows the structure of a base editor plasmid containing a T7 promoter driving expression of the system described herein. [Figure 3] Figure 1 shows the plasmid maps of the system described herein. MGA contains TadA* (from ABE8.17m)-SV40 NLS, and MGC contains APOBEC1 (from BE3) linked to a uracil glycosylase inhibitor and an SV40 NLS. [Figure 4]Shown are predicted catalytic residues in the RuvCI domains of selected endonucleases described herein, which have been mutated to disrupt nuclease activity and generate nickase enzymes. [Figure 5] An exemplary method for cloning a single guide RNA expression cassette into the system described herein is shown. One fragment contains a T7 promoter plus a spacer. The other fragment contains a spacer plus a single guide scaffold sequence plus a bidirectional terminator. The fragments are assembled into an expression plasmid, resulting in a functional construct capable of simultaneously expressing the sgRNA and base editor. [Figure 6A] Figure 1 shows the sgRNA design for targeting lacZ in E. coli.The spacer length used in the system described herein is 22 nucleotides.For the system selected herein, three sgRNAs are designed to target lacZ in E. coli to determine the editing window. [Figure 6B] Figure 1 shows the sgRNA design for targeting lacZ in E. coli.The spacer length used in the system described herein is 22 nucleotides.For the system selected herein, three sgRNAs are designed to target lacZ in E. coli to determine the editing window. [Figure 7] Nickase activity of selected mutant effectors is shown. 600-bp double-stranded DNA fragments labeled with fluorophores (6-FAM) at both 5' ends were incubated with purified enzyme supplemented with their cognate sgRNAs. Reaction products were resolved on a 10% TBE-urea denaturing gel. Double-stranded breaks yield bands at 400 and 200 bases. Nickase activity yields bands at 600 and 200 bases. [Figure 8A] 1 shows Sanger sequencing results demonstrating base editing by selected systems described herein. [Figure 8B] 1 shows Sanger sequencing results demonstrating base editing by selected systems described herein. [Figure 8C] 1 shows Sanger sequencing results demonstrating base editing by selected systems described herein. [Figure 9] The systems described herein demonstrate how base editing capabilities can be expanded with the endonucleases and base editors described herein. [Figure 10A] The base editing efficiency of adenine base editors (ABEs) containing TadA (ABE8.17m) and MG nickase is shown. TadA is a tRNA adenine deaminase, and TadA (ABE8.17m) is an engineered variant of E. coli TadA. Twelve MG nickases fused to TadA (ABE8.17m) were constructed and tested in E. coli. Three guides targeting lacZ were designed. The numbers in the boxes indicate the percentage of A to G conversion quantified by Edit R. ABE8.17m was used as a positive control in the experiment. [Figure 10B] The base editing efficiency of adenine base editors (ABEs) containing TadA (ABE8.17m) and MG nickase is shown. TadA is a tRNA adenine deaminase, and TadA (ABE8.17m) is an engineered variant of E. coli TadA. Twelve MG nickases fused to TadA (ABE8.17m) were constructed and tested in E. coli. Three guides targeting lacZ were designed. The numbers in the boxes indicate the percentage of A to G conversion quantified by Edit R. ABE8.17m was used as a positive control in the experiment. [Figure 11A]This figure shows the base editing efficiency of a cytosine base editor (CBE) containing rat APOBEC1, MG nickase, and Bacillus subtilis bacteriophage uracil glycosylase inhibitor (UGI(PBS1)). APOBEC1 is a cytosine deaminase. Twelve MG nickases fused to rAPOBEC1 at their N-terminus and UGI at their C-terminus were constructed and tested in E. coli. Three guides targeting lacZ were designed. The numbers in the boxes indicate the percentage of C-to-T conversion quantified by EditR. BE3 was used as a positive control in the experiment. [Figure 11B] This figure shows the base editing efficiency of a cytosine base editor (CBE) containing rat APOBEC1, MG nickase, and Bacillus subtilis bacteriophage uracil glycosylase inhibitor (UGI(PBS1)). APOBEC1 is a cytosine deaminase. Twelve MG nickases fused to rAPOBEC1 at their N-terminus and UGI at their C-terminus were constructed and tested in E. coli. Three guides targeting lacZ were designed. The numbers in the boxes indicate the percentage of C-to-T conversion quantified by EditR. BE3 was used as a positive control in the experiment. [Figure 12] The effect of MG uracil glycosylase inhibitors (UGIs) on the base editing activity of CBEs is shown. Panel A of Figure 12 shows a graph depicting the base editing activity of MG15-1 and variants containing N-terminal APOBEC1, MGC15-1, and C-terminal UGI. Three MG UGIs were tested for improved cytosine base editing activity in E. coli. Panel B of Figure 12 shows a graph depicting the base editing activity of BE3, which contains N-terminal rAPOBEC1, SpCas9 nickase, and C-terminal UGI. Two MG UGIs were tested for improved cytosine base editing activity in HEK293T cells. Editing efficiency was quantified by Edit R. [Figure 13A]Figure 1 shows a map of edited sites demonstrating the editing efficiency of cytosine base editors containing A0A2K5RDN7, MG nickase, and MG UGI. The construct contains N-terminal A0A2K5RDN7, MG nickase, and C-terminal MG69-1. For simplicity, the identity of the MG nickase is indicated in the figure. BE3 was used as a positive control for base editing. An empty vector was used as a negative control. Three independent experiments were performed on different days. Abbreviations: R: repeat, NEG: negative control. [Figure 13B] Figure 1 shows a map of edited sites demonstrating the editing efficiency of cytosine base editors containing A0A2K5RDN7, MG nickase, and MG UGI. The construct contains N-terminal A0A2K5RDN7, MG nickase, and C-terminal MG69-1. For simplicity, the identity of the MG nickase is indicated in the figure. BE3 was used as a positive control for base editing. An empty vector was used as a negative control. Three independent experiments were performed on different days. Abbreviations: R: repeat, NEG: negative control. [Figure 14] A positive selection method for TadA characterization in E. coli is shown. Figure 14A shows a map of one plasmid system used for TadA selection. The vector contains CAT(H193Y), a CAT-targeting sgRNA expression cassette, and an ABE expression cassette. The N-terminal TadA from E. coli and the C-terminal SpCas9(D10A) from Streptococcus pyogenes are shown in this figure. Figure 14B shows a sequencing trace demonstrating that when introduced / transformed into E. coli cells, it edits the A2 position of the template strand of CAT(H193Y), reverting the H193Y mutant to the wild type and restoring its activity. Abbreviations: CAT: chloramphenicol acetyltransferase. [Figure 15]This shows that the mutations caused by TadA enable high resistance to chloramphenicol (Cm). Figure 15A shows a photograph of a growth plate in which various concentrations of chloramphenicol were used to select for antibiotic resistance in E. coli. In this example, wild-type and two variants of TadA derived from E. coli (EcTadA) were tested. Figure 15B shows a summary table of the results demonstrating that ABEs carrying mutant TadA exhibit higher editing efficiency than wild-type. In these experiments, colonies were picked from plates with Cm concentrations of 0.5 μg / mL or higher. For simplicity, the identity of the deaminase is shown in the table. [Figure 16A] A photograph of a growth plate used to investigate MG68 TadA activity during positive selection is shown. Eight MG68 TadA candidates were tested against 0–2 μg / mL of chloramphenicol (ABEs contained an N-terminal TadA variant and a C-terminal SpCas9(D10A) nickase). For simplicity, the identity of the deaminase is shown. In this experiment, colonies were picked from plates with Cm ≥ 0.5 μg / mL. [Figure 16B] We summarize the editing efficiencies of MG TadA candidates and demonstrate that MG68-3 and MG68-4 drove adenine base editing. [Figure 17] This figure shows the improved base editing efficiency of MG68-4_nSpCas9 via the D109N mutation in MG68-4. Figure 17A shows a photograph of a growth plate in which wild-type MG68-4 and its variants were tested against 0 to 4 μg / mL of chloramphenicol. For simplicity, the identity of the deaminase is shown. The adenine base editors in this experiment include an N-terminal TadA variant and a C-terminal SpCas9(D10A) nickase. Panel (b) shows a summary table showing the editing efficiency of the MG TadA candidates. Figure 17B demonstrates that MG68-4 and MG68-4(D109N) exhibited adenine base editing, with the D109N mutant showing increased activity. In this experiment, colonies were picked from plates containing 0.5 μg / mL or higher Cm. [Figure 18]Base editing of MG68-4(D109N)_nMG34-1 is shown. Figure 18A shows a photograph of a growth plate from an experiment in which an ABE containing N-terminal MG68-4(D109N) and C-terminal SpCas9(D10A) nickase was tested against 0 to 2 μg / mL of chloramphenicol. Figure 18B shows a summary table showing editing efficiency with and without sgRNA. In this experiment, colonies were picked from plates with Cm of 1 μg / mL or higher. [Figure 19] MG68-4-nMG34-1 shows 28 MG68-4 variants designed for improved base editing activity (SEQ ID NOs: 448-475). Twelve residues were selected for targeted mutagenesis to improve editing of the enzyme. [Figure 20] Results of a gel-based deaminase assay demonstrating the activity of deaminases from several selected families (MG93, MG138, and MG139) are shown. Enzymes were expressed in an in vitro transcription-translation system from bacterial (E. coli codon-optimized) Purexpress cell lysates and incubated with 5' FAM-labeled ssDNA and USER enzyme (uracil DNA glycosylase and endonuclease VIII) for 2.5 hours at 37°C. The resulting DNA was resolved on a denaturing polyacrylamide gel and imaged. The positive control was a sequence with a synthetically incorporated U at the same position as the target C, and the negative control was a sequence containing neither U nor C. [Figure 21] Figure 1 shows the base editing efficiency of adenine base editing at specific nucleotide sites using MG68-4v1 fused to either nMG34-1 or nSpCas9. Nine guides were designed to target genomic loci in HEK293T cells. Abbreviations: MG68-4v1, MG68-4(D109N); nMG34-1, MG34-1 nickase; nSpCas9, SpCas9 nickase. [Figure 22]Figure 22 shows in vivo base editing using engineered MG34-1 and MG35-1 nickases. Panels (A) and (B) show base editing at four target loci in the E. coli genome. Figure 22A shows the ABE-MG34-1 base editor versus the reference ABE-SpCas9 (both using TadA*(8.8m) deaminase). Figure 22B shows the CBE-MG34-1 base editor versus the reference CBE-SpCas9 (both using rAPOBEC1 deaminase and PBS1 UGI). Figure 22C shows base editing in human HEK293T cells using ABE-MG34-1 nickase at three target loci. The target sequence for each locus in panels A, B, and C is shown above each heatmap. The predicted editing positions are represented by subscript numbers on the sequence and by their respective positions (boxes) on the heatmap. The heat maps in Figure 22A, B, and C represent the percentage of NGS reads that support editing. The values in Figure 22A and B represent the average of two independent experiments, and the values in panel (C) represent the average of three independent biological replicates. Figure 22D shows an E. coli survival assay. E. coli was transformed with a plasmid containing ABE, a nonfunctional chloramphenicol acetyltransferase (CAT H193Y) gene, and an sgRNA that either targets (targeted spacer) or does not target (non-targeted spacer) the CAT gene. E. coli survival under chloramphenicol selection depends on ABE base editing of the nonfunctional CAT gene to its wild-type sequence. The top panel of Figure 22E shows a diagram of the ABE construct with a C-terminal TadA*-(7.10) monomer and an engineered MG35-1 nickase containing an SV40 NLS fused to the C-terminus. Figure 22E, bottom panel: Transformed E. coli were grown on plates containing chloramphenicol concentrations of 0, 2, 3, 4, and 8 μg / mL. Plates also contained 100 μg / mL carbecillin and 0.1 mM IPTG. Colonies grown on plates containing chloramphenicol concentrations of 0, 2, 3, and 4 μg / mL were sequenced to assess reversion of the CAT gene. Experiments were performed in duplicate. [Figure 23]A gel-based deaminase assay showing the activity of a deaminase from one selected family (MG139) is shown. The enzyme was expressed in an in vitro transcription-translation system from bacterial (E. coli codon-optimized) Purexpress cell lysate and incubated with 5' FAM-labeled ssDNA and USER enzyme (uracil DNA glycosylase and endonuclease VIII) for 2.5 hours at 37°C. The resulting DNA was resolved on a denaturing polyacrylamide gel and imaged, as shown in Figure 23A. The positive control is a sequence with a synthetically incorporated U at the same position as the target C, and the negative control is a sequence containing neither U nor C. Figure 23B shows the percentage of deaminating activity of all active cytidine deaminases on ssDNA. The taxonomic classification of cytidine deaminases is shown. [Figure 24] Gel-based deaminase assays demonstrating the ssDNA and dsDNA activity of deaminases from several selected families (MG93, MG138, and MG139) are shown. Enzymes were expressed in an in vitro transcription-translation system from bacterial (E. coli codon-optimized) Purexpress cell lysates and incubated with 5' FAM-labeled ssDNA or dsDNA and USER enzyme (uracil DNA glycosylase and endonuclease VIII) for 2.5 hours at 37°C. The resulting DNA was resolved on a denaturing polyacrylamide gel and imaged. The positive control for ssDNA activity was a sequence with a synthetically incorporated U at the same position as the target C, and the negative control was a sequence containing neither U nor C. A positive control for dsDNA activity is the DddA toxin deaminase, which has been documented to be selective for dsDNA substrates (Mok, BY, de Moraes, MH, Zeng, J. et al. A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing. Nature 583, 631-637 (2020). https: / / doi.org / 10.1038 / s41586-020-2477-4). [Figure 25A]These data demonstrate that cytosine base editors (CBEs) containing novel cytidine deaminases with spCas9, MG3-6, or MG34-1 effectors exhibit varying editing levels in HEK293 cells. Each novel cytidine deaminase is fused to the N-terminus of the effector (spCas9, MG3-6, or MG34-1) via a linker. A uracil glycosylase inhibitor domain (UGI or MG69-1) is fused to the C-terminus of the effector, followed by a nuclear localization signal (NLS). Each CBE was transiently transfected into HEK293 cells, and the corresponding sgRNA was used to target five different genomic locations (spacer sequences are shown, and target cytosines are underlined). The editing levels (C to T (%)) of the spacer sequences and surrounding cytosines are shown for CBEs with each distinct cytidine deaminase effector (n=3). [Figure 25B] These data demonstrate that cytosine base editors (CBEs) containing novel cytidine deaminases with spCas9, MG3-6, or MG34-1 effectors exhibit varying editing levels in HEK293 cells. Each novel cytidine deaminase is fused to the N-terminus of the effector (spCas9, MG3-6, or MG34-1) via a linker. A uracil glycosylase inhibitor domain (UGI or MG69-1) is fused to the C-terminus of the effector, followed by a nuclear localization signal (NLS). Each CBE was transiently transfected into HEK293 cells, and the corresponding sgRNA was used to target five different genomic locations (spacer sequences are shown, and target cytosines are underlined). The editing levels (C to T (%)) of the spacer sequences and surrounding cytosines are shown for CBEs with each distinct cytidine deaminase effector (n=3). [Figure 25C]These data demonstrate that cytosine base editors (CBEs) containing novel cytidine deaminases with spCas9, MG3-6, or MG34-1 effectors exhibit varying editing levels in HEK293 cells. Each novel cytidine deaminase is fused to the N-terminus of the effector (spCas9, MG3-6, or MG34-1) via a linker. A uracil glycosylase inhibitor domain (UGI or MG69-1) is fused to the C-terminus of the effector, followed by a nuclear localization signal (NLS). Each CBE was transiently transfected into HEK293 cells, and the corresponding sgRNA was used to target five different genomic locations (spacer sequences are shown, and target cytosines are underlined). The editing levels (C to T (%)) of the spacer sequences and surrounding cytosines are shown for CBEs with each distinct cytidine deaminase effector (n=3). [Figure 26] Figure 26 shows the activity of cytidine deaminases (CDAs) fused to MG3-6. Cytidine deaminases were fused to MG3-6, and their activity was assessed by targeting engineered sites in reporter cell lines. Figure 26A shows the relative activity of various CDAs; the controls used were the literature A0A2K5RDN7 and the highly active CBE from rAPOBEC1. Figure 26B shows quantification of the activity of various CDAs compared to the highly active CDA A0A2K5RDN7. Figure 26C shows MG139-52 activity highlighting GA conversion, suggesting strand editing in the DNA / RNA heteroduplex at the opposite strand, i.e., the R-loop. [Figure 27] Figure 27 shows a toxicity assay in mammalian cells. The toxicity of CDA was measured by stable expression of CDA as CBE (fused to MG3-6). HEK293T cells stably expressing CBE were grown in puromycin for 3 days, and surviving cells were stained with crystal violet. The crystal violet dye was then solubilized with 1% SDS and quantified using a plate reader. Figure 27A shows a photograph of cells stained with crystal violet, and Figure 27B shows the quantification of Figure 27A. The absorbance was measured using a plate reader at 570 nm. [Figure 28] Mutations identified from chloramphenicol selection in E. coli are shown. The r1v1 variant was the starting variant for the evolution experiment. 24 variants were identified, and the associated mutations are shown in the table. [Figure 29] Beneficial mutations identified from variant screening in HEK293T are shown. The predicted structure of MG68-4 aligns with tRNAArg2 from S. aureus TadA (PDB 2B3J). The structure representation highlights key mutated residues. [Figure 30]
[0033] Figure 1 shows screening of MG68-4 variants in HEK293T cells. Four guides were used to screen the activity, editing window, and sequence preference of the engineered variants. [Figure 31] Figure 1 shows the sequencing results of the ABE-MG35-1 E. coli survival assay. Surviving colonies were picked from plates under chloramphenicol selection for the first experimental replicate and subjected to Sanger sequencing. Sequencing of four of the five selected colonies showed an A to G mutation on the minus strand, restoring CAT function on the plus strand from Y193 back to H (boxed nucleotides). Bystander base editing was observed in two of the five sequenced colonies. [Figure 32] Figure 1 shows an increase in cytosine base editing efficiency upon Fam72a expression. [Figure 33-1]Data are presented demonstrating that structurally optimized adenine base editors (ABEs) exhibit varying editing levels in HEK293 cells. Thirty-three ABEs were constructed by inserting the MG68-4 (D109N) deaminase upstream, downstream, or internally into the MG3-6_3-8 (D13A) nickase enzyme and cloned into the pCMV vector. These plasmids were co-transfected with plasmids containing one of eight sgRNAs targeting the HEK293 genome. Data shown are from an sgRNA targeting the ACAGACAAAACTGTGCTAGACA sequence. The editing levels (A to G (%)) of A5, A7, A8, A9, and A10 within the spacer sequence, as well as cell viability for each individual experiment (n=2) are shown. [Figure 33-2] Data are presented demonstrating that structurally optimized adenine base editors (ABEs) exhibit varying editing levels in HEK293 cells. Thirty-three ABEs were constructed by inserting the MG68-4 (D109N) deaminase upstream, downstream, or internally into the MG3-6_3-8 (D13A) nickase enzyme and cloned into the pCMV vector. These plasmids were co-transfected with plasmids containing one of eight sgRNAs targeting the HEK293 genome. Data shown are from an sgRNA targeting the ACAGACAAAACTGTGCTAGACA sequence. The editing levels (A to G (%)) of A5, A7, A8, A9, and A10 within the spacer sequence, as well as cell viability for each individual experiment (n=2) are shown. [Figure 33-3]Data are presented demonstrating that structurally optimized adenine base editors (ABEs) exhibit varying editing levels in HEK293 cells. Thirty-three ABEs were constructed by inserting the MG68-4 (D109N) deaminase upstream, downstream, or internally into the MG3-6_3-8 (D13A) nickase enzyme and cloned into the pCMV vector. These plasmids were co-transfected with plasmids containing one of eight sgRNAs targeting the HEK293 genome. Data shown are from an sgRNA targeting the ACAGACAAAACTGTGCTAGACA sequence. The editing levels (A to G (%)) of A5, A7, A8, A9, and A10 within the spacer sequence, as well as cell viability for each individual experiment (n=2) are shown. [Figure 33-4] Data are presented demonstrating that structurally optimized adenine base editors (ABEs) exhibit varying editing levels in HEK293 cells. Thirty-three ABEs were constructed by inserting the MG68-4 (D109N) deaminase upstream, downstream, or internally into the MG3-6_3-8 (D13A) nickase enzyme and cloned into the pCMV vector. These plasmids were co-transfected with plasmids containing one of eight sgRNAs targeting the HEK293 genome. Data shown are from an sgRNA targeting the ACAGACAAAACTGTGCTAGACA sequence. The editing levels (A to G (%)) of A5, A7, A8, A9, and A10 within the spacer sequence, as well as cell viability for each individual experiment (n=2) are shown. [Figure 33-5]Data are presented demonstrating that structurally optimized adenine base editors (ABEs) exhibit varying editing levels in HEK293 cells. Thirty-three ABEs were constructed by inserting the MG68-4 (D109N) deaminase upstream, downstream, or internally into the MG3-6_3-8 (D13A) nickase enzyme and cloned into the pCMV vector. These plasmids were co-transfected with plasmids containing one of eight sgRNAs targeting the HEK293 genome. Data shown are from an sgRNA targeting the ACAGACAAAACTGTGCTAGACA sequence. The editing levels (A to G (%)) of A5, A7, A8, A9, and A10 within the spacer sequence, as well as cell viability for each individual experiment (n=2) are shown. [Figure 34A] Figure 34A shows the rational design of MG68-4 variants. Figure 34A shows the structural alignment of E. coli TadA (PDB: 1z3a) and the predicted structure of MG68-4. The tRNA structure was taken from S. aureus TadA (PDB: 2b3j). [Figure 34B] Figure 34B shows the rational design of MG68-4 variants. Figure 34B shows the mutations identified from EcTadA for the development of adenine base editors (ABE7.10, ABE8.8m, ABE8.17m, and ABE8e) and the equivalent residues of EcTadA on MG68-4. The corresponding mutations in EcTadA were introduced into MG68-4. H129N was identified from bacterial selection in E. coli. Generally, a nuclear localization signal (SV40) was positioned on the C-terminus. For the two-NLS construct, one SV40 was used at the N-terminus and one SV40 was used at the C-terminus. For simplicity, the deaminase sequences of the adenine base editors are shown in the table. Abbreviations: MGA0.1, MG68-4; MGA1.1, MG68-4(D109N); MGA2.1, MG68-4(D109N / H129N); RD, rationally designed variant. [Figure 35]Screening of adenine base editors in HEK293T cells. The top three variants are highlighted. The starting variant is MGA1.1. For the 2NLS construct, one SV40 was used at the N-terminus and one SV40 was used at the C-terminus. Abbreviations: MGA0.1, MG68-4; MGA1.1, MG68-4(D109N); MGA2.1, MG68-4(D109N / H129N); RD, rationally designed variant. [Figure 36] 1 shows a table summarizing the base editing activity of the rationally designed ABE variants described herein. [Figure 37] A gel-based deaminase assay demonstrating the activity of variant deaminases from several selected families (MG93, MG139, and MG152) is shown. The enzymes were expressed in an in vitro transcription-translation system from bacterial (E. coli codon-optimized) Purexpress cell lysates and incubated with 5' FAM-labeled ssDNA and USER enzyme (uracil DNA glycosylase and endonuclease VIII) for 2.5 hours at 37°C. The resulting DNA was resolved on a denaturing polyacrylamide gel and imaged. The positive control was a sequence with a synthetically incorporated U at the same position as the target C, and the negative control was a sequence containing neither U nor C. [Figure 38-1] Figure 38 shows a gel-based deaminase assay using a dual fluorophore assay. Figure 38A shows a schematic of the substrate design. The substrate was designed to minimize overlap between the two fluorophores. The emission of Cy3 is approximately 560 nm, and the emission peak of Cy5.5 is approximately 700 nm. [Figure 38-2]Gel-based deaminase assays using dual fluorophore assays are shown. Figure 38B and Figure 38C show TBE-urea gel images captured using Cy3 and Cy5.5 filters, respectively. RF157 is a single nucleotide substrate with a FAM molecule and serves as a positive control to confirm that the user enzyme is cleaving in the reaction and that the filter is functional and can distinguish between either fluorophore. A master mix is used as a negative control to provide a baseline measurement of uncleaved substrate. Figure 38B: A deaminase that preferentially cleaves a substrate at T at position -1 gives a 65-nt fluorescent product. A substrate cleaved at C at position -1 gives a 45-nt product. A deaminase active at both C and T at position -1 gives a 30-nt product. Figure 38C: A deaminase that preferentially cleaves a substrate at G at position -1 gives a 65-nt fluorescent product. A substrate cleaved at C at position -1 gives a 45-nt product. Deaminases active at both A or G at position -1 give a 30 nt product. [Figure 39-1] For each variant tested in this study (MG93 family and MG152 family), the percentage of deamination at each −1 position relative to the target cytidine is shown. [Figure 39-2] For each variant tested in this study (MG93 family and MG152 family), the percentage of deamination at each −1 position relative to the target cytidine is shown. [Figure 39-3] For each variant tested in this study (MG93 family and MG152 family), the percentage of deamination at each −1 position relative to the target cytidine is shown. [Figure 39-4] For each variant tested in this study (MG93 family and MG152 family), the percentage of deamination at each −1 position relative to the target cytidine is shown. [Figure 40-1] The percentage of deamination at each −1 position relative to the target cytidine is shown for each variant (MG139 family) tested in this study. [Figure 40-2]The percentage of deamination at each −1 position relative to the target cytidine is shown for each variant (MG139 family) tested in this study. [Figure 40-3] The percentage of deamination at each −1 position relative to the target cytidine is shown for each variant (MG139 family) tested in this study. [Figure 41A] Figure 41A shows a summary of activity data for novel and engineered CDAs as CBEs in mammalian cells. Figure 41A shows the maximum editing efficiency detected for all CDAs tested across the five engineered spacers. [Figure 41B] Figure 41B shows a summary of activity data for novel and engineered CDAs as CBEs in mammalian cells. Figure 41B shows the maximum activity detected, normalized to an internal positive control across the five engineered spacers. The internal experimental positive control used for normalization was the highly active CDA "A0A2K5RDN7." [Figure 41C] Figure 41C shows a summary of activity data for novel and engineered CDAs as CBEs in mammalian cells. Figure 41C shows a side-by-side comparison of one of the leading candidates, "139-52-V6," versus the highly active positive control, "A0A2K5RDN7," for the two guides. 139-52-V6 exhibits similar editing efficiency compared to the highly active test CDAs. [Figure 42] Figure 1 shows the -1 nt preference of CDAs with editing activity of greater than 1% as CBEs in mammalian cells. Figure 2 shows a comparison of -1 nt preference in mammalian cells versus in vitro -1 nt preference. The -1 preference observed in mammalian cells as CBEs is roughly comparable to the in vitro preference. The in vitro preference shows a more gradual pattern than the CBE activity in mammalian cells. [Figure 43A] Examples of MG139-52wt and MG139-52v6 mutated to A at N27, showing differences in activity on ssDNA and / or RNA:DNA duplexes, are shown. Figure 43A shows the predicted structure of MG139-52 using A3H as a template (pdb:5W3V). The targeted mutation at N27 is indicated by an arrow and is located distal to the catalytic center and recognition loop 7. [Figure 43B] Examples of MG139-52wt and MG139-52v6 mutated at N27 to A are shown, demonstrating differences in activity on ssDNA and / or RNA:DNA duplexes. Figure 43B shows a diagram showing the DNA / RNA heteroduplex in the R-loop targeted by 139-52WT. CRISPResso output shows GA conversion, indicating deamination in the DNA strand forming the DNA / RNA heteroduplex. [Figure 43C] Examples of MG139-52wt and MG139-52v6 mutated at N27 to A are shown, demonstrating differences in activity at ssDNA and / or RNA:DNA duplexes. Figure 43C shows the CRISPREsso output showing that the GA change in the DNA / RNA heteroduplex is abolished in the N27A variant. Instead, such modification occurs outside the DNA / RNA heteroduplex, suggesting that deamination of the DNA / RNA heteroduplex is impaired. [Figure 44] The editing windows of the dominant CDAs are shown compared to the highly active CDA A0A2K5RDN7. The editing windows shown correspond to approximately 110 nt. The R-loop (Cas9 target) is shown as a box. The dominant candidates 152-6 and 139-52-V6 have smaller editing windows than A0A2K5RDN7, a favorable feature for avoiding off-target editing. The engineered CDA 139-52-V6 exhibits a smaller editing window than its WT counterpart 139-52. [Figure 45] Figure 1 shows the mammalian cytotoxicity of CDA stably expressed as a CBE. CDA expressed as a CBE was stably expressed in mammalian cells by lentiviral integration. Cytotoxicity was measured as a fold change relative to a low-activity, low-cytotoxicity CDA (rAPOBEC). Superior candidates (high editing efficiency) show moderate cytotoxic activity under these conditions. It is understood that cytotoxic activity is reduced when the system is transiently expressed. [Figure 46]The dimeric design of the MG68-4 variant is shown. Figure 46A shows the predicted structure of MG68-4 and a structural alignment of MG68-4 and SaTadA (PDB code: 2b3j). The distance between the N-terminus of the first monomer and the C-terminus of the second monomer is shown. Figure 46B shows the base editing efficiency comparing the monomeric and dimeric designs. TadA*8.8m was used as the benchmark. The target sequence is shown in the bar graph. A-to-G conversion was obtained from the highest editing position, A8. All deaminases were fused to the N-terminus of MG34-1(D10A). Editing was evaluated in HEK293T cells. [Figure 47] The effect of the D109Q mutation on C to G base substitutions is shown. A to G and C to G conversions were obtained from target sequences 633 and 634, respectively. The editing efficiencies of residue C6 of target sequence 633 and residue A8 of target sequence 634 are shown. All deaminases were fused to the N-terminus of MG34-1(D10A). Editing efficiencies were evaluated in HEK293T cells. [Figure 48] Base editing efficiency of a combinatorial library in HEK293T cells is shown. Beneficial mutations identified through rational design and directed evolution were introduced into MG68-4 to generate a combinatorial library. Variants were inserted into 3-68_DIV30_M_RDr1v1_B. Editing efficiency was evaluated in HEK293T cells. [Figure 49] 1 shows the effect of MG68-4 dimerization and / or MG68-4 amino acid sequence variants within the 3-68_DIV30 scaffold on A to G conversion rates in HEK293T cells. [Figure 50]Data demonstrating that MG35-1 nickase can function as a scaffold for an adenine base editor in E. coli cells are shown. Figure 50A shows a schematic diagram of the MG35-1 adenine base editor (ABE) containing a C-terminal TadA*-(7.10) monomer and an SV40 NLS fused to the C-terminus. Figure 50B shows a chloramphenicol selection experiment used to evaluate MG35-1 ABE base editing. Plasmids containing the MG35-1 ABE, a nonfunctional chloramphenicol acetyltransferase (CAT) gene, and an sgRNA that either targets the CAT gene (targeting sgRNA) or does not target the CAT gene (non-targeting sgRNA) are transformed into BL21(DE3) (Lucigen) E. coli cells. E. coli survival under chloramphenicol selection was dependent on the MG35-1 ABE base editing the nonfunctional CAT gene to its wild-type sequence. The transformed E. coli were plated on plates containing chloramphenicol at concentrations of 0, 2, 3, 4, and 8 μg / mL. The plates also contained 100 μg / mL carbecillin and 0.1 mM IPTG. Colonies grown on plates containing chloramphenicol at concentrations of 0, 2, 3, 4, and 8 μg / mL were sequenced to assess reversion of the CAT gene. Experiments were performed in duplicate. [Figure 51-1] The activity of 3-6 / 8 ABE in Apoa1 is shown. High A to G conversions were observed in 26 Apoa1 guides. For all spacers shown in the graph, base conversions at all A positions within the spacer region are shown. [Figure 51-2] The activity of 3-6 / 8 ABE in Apoa1 is shown. High A to G conversions were observed in 26 Apoa1 guides. For all spacers shown in the graph, base conversions at all A positions within the spacer region are shown. [Figure 51-3] The activity of 3-6 / 8 ABE in Apoa1 is shown. High A to G conversions were observed in 26 Apoa1 guides. For all spacers shown in the graph, base conversions at all A positions within the spacer region are shown. [Figure 52] Figure 1 shows the activity of 3-6 / 8 ABE with Angptl3. High A to G conversions were observed in the five Angptl3 guides. For all spacers shown in the graph, base conversions at all A positions within the spacer region are shown. [Figure 53] The activity of the 3-6 / 8 ABE in Trac is shown. High A to G conversions were observed for the two Trac guides. For all spacers shown in the graph, base conversions at all A positions within the spacer region are shown. [Figure 54-1] Background 3-6 / 8 ABE activity in Apoa1 is shown. Active guide primer pairs were tested in mock-nucleofected samples to assay background editing in the target region. Scale is 0-1%. [Figure 54-2] Background 3-6 / 8 ABE activity in Apoa1 is shown. Active guide primer pairs were tested in mock-nucleofected samples to assay background editing in the target region. Scale is 0-1%. [Figure 54-3] Background 3-6 / 8 ABE activity in Apoa1 is shown. Active guide primer pairs were tested in mock-nucleofected samples to assay background editing in the target region. Scale is 0-1%. [Figure 55]Figure 55 shows an E. coli survival assay using nMG35-1 ABE. E. coli was transformed with a plasmid containing nMG35-1-ABE, a nonfunctional chloramphenicol acetyltransferase (CAT Y193) gene, and an sgRNA that either targets (targeted spacer) or does not target (scrambled spacer) the CAT gene. Figure 55A shows a diagram depicting the target sequence with the predicted TAM. Cell growth depends on ABE base editing of the nonfunctional CAT gene (A at position 17 from the TAM / PAM, boxed) to restore activity. Figures 55B-55E show the base editing activity in E. coli of a base editor containing nMG35-1 fused to TadA deaminase using linkers of various lengths. The x-axis indicates the linkers listed in Table 14. [Figure 56-1] Figure 56A-B show evaluation of nMG35-1 ABE base editing in an E. coli survival assay under chloramphenicol selection, where cell growth is dependent on ABE base editing of the nonfunctional CAT gene stop codon and restoration of activity. Figures 56A-B show diagrams depicting the target sequence with predicted TAMs. An "A" base at position 11 (A) or 10 (B) from the TAM (boxed) is predicted to be edited to "G" to change the stop codon back to glutamine and restore chloramphenicol (cm) resistance. [Figure 56-2]Evaluation of nMG35-1 ABE base editing in an E. coli survival assay under chloramphenicol selection is shown, where cell growth is dependent on ABE base editing of a nonfunctional CAT gene stop codon and restoration of activity. Figure 56C: E. coli were transformed with a plasmid containing nMG35-1-ABE, a nonfunctional chloramphenicol acetyltransferase (CAT), and an sgRNA that either targets (targeting spacer) or does not target (no spacer) the CAT gene. Transformed E. coli were grown on plates containing chloramphenicol concentrations of 0, 2, 4, and 8 μg / mL. Plates also contained 100 μg / mL carbecillin and 0.1 mM IPTG. nMG35-1-ABE, which targets both STOP98Q and STOP122Q, contains both stop codons in the same gene that are required to restore CAT gene function. MIC: Minimum inhibitory concentration. [Figure 56-3] Figure 56D shows evaluation of nMG35-1 ABE base editing in an E. coli survival assay under chloramphenicol selection, where cell growth is dependent on ABE base editing of a nonfunctional CAT gene stop codon and restoration of activity. Figure 56D shows Sanger sequencing chromatograms of 5 of 18 colonies grown at 2 μg / mL chloramphenicol for the nMG35-1 ABE double reversion of STOP98Q and STOP122Q in the CAT gene. The chromatogram of a colony not showing reversion (colony 3) reveals a smaller peak of A to G conversion, which is likely obscured due to cotransformation with a non-edited plasmid. [Figure 57]Figure 1 shows data demonstrating that truncation of the predicted PLMP domain at the N-terminus of MG35-1 eliminates function of the MG35-1 ABE in E. coli. E. coli was transformed with a plasmid containing nMG35-1-ABE, a non-functional chloramphenicol acetyltransferase (CAT), and an sgRNA targeting either the CAT gene (WT (top) or PLMP domain truncated (bottom) MG35-1 ABE) or a non-target spacer (middle: WT MG35-1 ABE with a scrambled spacer). Transformed E. coli were grown on plates containing chloramphenicol concentrations of 0, 2, and 4 μg / mL. Plates also contained 100 μg / mL carbecillin and 0.1 mM IPTG. MIC: Minimum inhibitory concentration.
[0093] Brief description of the sequence listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.
[0094] SEQ ID NOs: 1-47 show the full-length peptide sequences of MG66 deaminases suitable for the engineered nucleic acid editing system described herein.
[0095] SEQ ID NOs: 48-49 show the full-length peptide sequences of MG67 deaminase suitable for the engineered nucleic acid editing system described herein.
[0096] SEQ ID NOs: 50-51 show the full-length peptide sequences of MG68 deaminase suitable for the engineered nucleic acid editing system described herein.
[0097] SEQ ID NOs: 52-56 show the sequences of uracil DNA glycosylase inhibitors suitable for the engineered nucleic acid editing system described herein.
[0098] SEQ ID NOs: 57 to 66 show the sequences of reference deaminases.
[0099] SEQ ID NO: 67 shows the sequence of a reference uracil DNA glycosylase inhibitor.
[0100] SEQ ID NO: 68 shows the sequence of an adenine base editor.
[0101] SEQ ID NO: 69 shows the sequence of a cytosine base editor.
[0102] SEQ ID NOs: 70-78 show full-length peptide sequences of MG nickases suitable for the engineered nucleic acid editing system described herein.
[0103] SEQ ID NOs: 79-87 show the protospacers and PAMs used in the in vitro nickase assays described herein.
[0104] SEQ ID NOs: 88-96 show the peptide sequences of single guide RNAs used in the in vitro nickase assays described herein.
[0105] SEQ ID NOs: 97 to 156 show the spacer sequences when E. coli lacZ is targeted.
[0106] SEQ ID NOs: 157 to 176 show the sequences of primers used for site-directed mutagenesis.
[0107] SEQ ID NOs: 177 to 178 show the sequences of primers for lacZ sequencing.
[0108] SEQ ID NOs: 179 to 342 show the sequences of the primers used during amplification.
[0109] SEQ ID NOs: 343 to 345 show the sequences of primers for lacZ sequencing.
[0110] SEQ ID NOs: 346 to 359 show the sequences of the primers used during amplification.
[0111] SEQ ID NOs: 360-368 show protospacer adjacent motifs suitable for the engineered nucleic acid editing system described herein.
[0112] SEQ ID NOs: 369-384 show nuclear localization sequences (NLS) suitable for the engineered nucleic acid editing system described herein.
[0113] SEQ ID NOs: 385-443 show the full-length peptide sequences of MG68 deaminases suitable for the engineered nucleic acid editing system described herein.
[0114] SEQ ID NOs: 444-447 show the full-length peptide sequences of MG121 deaminase suitable for the engineered nucleic acid editing system described herein.
[0115] SEQ ID NOs: 448-475 show the full-length peptide sequences of MG68 deaminases suitable for the engineered nucleic acid editing system described herein.
[0116] SEQ ID NOs: 476 and 477 show the sequences of adenine base editors.
[0117] SEQ ID NOs: 478 to 482 show the sequences of cytosine base editors.
[0118] SEQ ID NOs: 483-487 show the sequences of plasmids suitable for encoding the engineered nucleic acid editing system described herein.
[0119] SEQ ID NOs: 488 and 489 show the sgRNA scaffold sequences of MG15-1 and MG34-1.
[0120] SEQ ID NOs: 490 to 522 show the sequences of spacers used to target genomic loci in E. coli and HEK293T cells.
[0121] SEQ ID NOs: 523 to 585 show the sequences of primers used during amplification and Sanger sequencing.
[0122] SEQ ID NOs: 584-585 show the sequences of the primers used during amplification.
[0123] SEQ ID NO: 586 shows the sequence of an adenine base editor.
[0124] SEQ ID NO: 587 shows the sequence of a cytosine base editor.
[0125] SEQ ID NOs: 588 to 589 show the sequences of adenine base editors.
[0126] SEQ ID NOs: 590-593 show full-length peptide sequences of linkers suitable for the engineered nucleic acid editing system described herein.
[0127] SEQ ID NO: 594 shows the sequence of cytosine deaminase.
[0128] SEQ ID NO: 595 shows the sequence of adenosine deaminase.
[0129] SEQ ID NO: 596 shows the sequence of an MG34 active effector suitable for the engineered nucleic acid editing system described herein.
[0130] SEQ ID NO: 597 shows the sequence of an MG34 nickase suitable for the engineered nucleic acid editing system described herein.
[0131] SEQ ID NO: 598 shows the sequence of MG34 PAM.
[0132] SEQ ID NOs: 599-638 show the full-length peptide sequences of MG138 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0133] SEQ ID NOs: 639-659 show the full-length peptide sequences of MG139 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0134] SEQ ID NOs: 660-662 show the full-length peptide sequences of MG141 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0135] SEQ ID NOs: 663-664 show the full-length peptide sequence of MG142 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0136] SEQ ID NOs: 665-675 show the full-length peptide sequences of MG93 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0137] SEQ ID NOs: 676 to 678 show the sequences of adenine base editors.
[0138] SEQ ID NOs: 679 to 680 show the sgRNA scaffold sequences of MG34-1 and SpCas9.
[0139] SEQ ID NOs: 681 to 689 show spacer sequences used to target genomic loci in guide RNAs.
[0140] SEQ ID NOs: 690-707 show the sequences of primers used to amplify genomic targets of adenine base editors (ABEs) for next-generation sequencing (NGS) analysis.
[0141] SEQ ID NO: 708 shows the sequence of the blasticidin (BSD) resistance cassette.
[0142] SEQ ID NOs: 709-719 show spacer sequences used to target genomic loci in guide RNAs.
[0143] SEQ ID NOs: 720-726 show the sequences of plasmids suitable for encoding the engineered nucleic acid editing system described herein.
[0144] SEQ ID NOs: 728 to 729 show the sequences of adenine base editors.
[0145] SEQ ID NOs: 730 to 736 show spacer sequences used to target genomic loci in guide RNAs.
[0146] SEQ ID NOs: 737-738 show the sequences of plasmids suitable for encoding the engineered nucleic acid editing system described herein.
[0147] SEQ ID NOs: 739 to 740 show the sequences of cytidine base editors.
[0148] SEQ ID NO: 741 shows the sequence of a suitable plasmid encoding the A1CF gene.
[0149] SEQ ID NO: 742 shows the sequence of the RNA used to test CDA for RNA activity.
[0150] SEQ ID NO: 743 shows the sequence of a labeled primer for the poisoned primer extension assay used to test CDA for RNA activity.
[0151] SEQ ID NOs: 744-827 show the full-length peptide sequences of MG139 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0152] SEQ ID NO: 828 shows the full-length peptide sequence of MG93 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0153] SEQ ID NO: 829 shows the full-length peptide sequence of MG142 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0154] SEQ ID NOs: 830-835 show the full-length peptide sequences of MG152 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0155] SEQ ID NOs: 836 to 860 show the sequences of adenine base editors.
[0156] SEQ ID NOs: 861 to 864 show spacer sequences used to target genomic loci in guide RNAs.
[0157] SEQ ID NOs: 865-872 show the sequences of primers used to amplify genomic targets of adenine base editors (ABEs) for next-generation sequencing (NGS) analysis.
[0158] SEQ ID NOs: 873-875 show the sequences of plasmids suitable for encoding the engineered nucleic acid editing system described herein.
[0159] SEQ ID NO: 876 shows the sgRNA scaffold sequence of MG34-1.
[0160] SEQ ID NOs: 877 to 916 show the sequences of cytosine base editors.
[0161] SEQ ID NOs: 917-931 show the sequences of sgRNAs suitable for the engineered nucleic acid editing system described herein.
[0162] SEQ ID NOs: 932 to 961 show the sequences of primers used to amplify genomic targets of adenine base editors (ABEs) for next-generation sequencing (NGS) analysis.
[0163] SEQ ID NO: 962 shows sites engineered in mammalian cell lines with five PAMs compatible with Cas9 and MG3-6 editing.
[0164] SEQ ID NOs: 963-967 show the sequences of sgRNAs suitable for the engineered nucleic acid editing system described herein.
[0165] SEQ ID NOs: 968 to 969 show the sequences of cytosine base editors.
[0166] SEQ ID NO: 970 shows the full-length peptide sequence of MG139 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0167] SEQ ID NOs: 971-977 show the full-length peptide sequences of MG93 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0168] SEQ ID NOs: 978-981 show the full-length peptide sequences of MG138 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0169] SEQ ID NO: 982 shows the full-length peptide sequence of MG142 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0170] SEQ ID NOs: 983-1014 show the full-length peptide sequences of MG128 deaminases suitable for the engineered nucleic acid editing system described herein.
[0171] SEQ ID NOs: 1015-1026 show the full-length peptide sequences of MG129 deaminase suitable for the engineered nucleic acid editing system described herein.
[0172] SEQ ID NOs: 1027-1031 show the full-length peptide sequences of MG130 deaminase suitable for the engineered nucleic acid editing system described herein.
[0173] SEQ ID NOs: 1032-1040 show the full-length peptide sequences of MG131 deaminase suitable for the engineered nucleic acid editing system described herein.
[0174] SEQ ID NOs: 1041-1043 show the full-length peptide sequences of MG132 deaminase suitable for the engineered nucleic acid editing system described herein.
[0175] SEQ ID NOs: 1044-1057 show the full-length peptide sequences of MG133 deaminases suitable for the engineered nucleic acid editing system described herein.
[0176] SEQ ID NOs: 1058-1061 show the full-length peptide sequences of MG134 deaminase suitable for the engineered nucleic acid editing system described herein.
[0177] SEQ ID NOs: 1062-1069 show the full-length peptide sequences of MG135 deaminase suitable for the engineered nucleic acid editing system described herein.
[0178] SEQ ID NOs: 1070-1081 show the full-length peptide sequences of MG136 deaminase suitable for the engineered nucleic acid editing system described herein.
[0179] SEQ ID NOs: 1082-1098 show the full-length peptide sequences of MG137 deaminase suitable for the engineered nucleic acid editing system described herein.
[0180] SEQ ID NOs: 1099-1105 show the sequences of sgRNAs suitable for the engineered nucleic acid editing system described herein.
[0181] SEQ ID NOs: 1106 to 1111 show the sequences of MG35 PAM.
[0182] SEQ ID NO: 1112 shows the DNA sequence of the gene encoding the ABE-MG35-1 adenine base editor.
[0183] SEQ ID NO: 1113 shows the protein sequence of the ABE-MG35-1 adenine base editor.
[0184] SEQ ID NO: 1114 shows the nucleotide sequence of a plasmid encoding a Cas9-based cytosine base editor (CBE).
[0185] SEQ ID NO: 1115 shows the nucleotide sequence of a plasmid encoding Fam72a.
[0186] SEQ ID NOs: 1116-1117 show the sequences of the Cas9-CBE target sites.
[0187] SEQ ID NOs: 1118-1119 show the sequences of the NGS amplicons.
[0188] SEQ ID NO: 1120 shows the full-length peptide sequence of MG35 nuclease.
[0189] SEQ ID NO: 1121 shows the full-length peptide sequence of Fam72A.
[0190] SEQ ID NOs: 1121 to 1127 show the full-length peptide sequences of MG35 nuclease.
[0191] SEQ ID NOs: 1128 to 1160 show the full-length peptide sequences of the MG3-6 / 3-8 adenine base editors.
[0192] SEQ ID NOs: 1161 to 1186 show the full-length peptide sequences of the MG34-1 adenine base editors.
[0193] SEQ ID NOs: 1187-1195 show the sequences of sgRNAs suitable for the engineered nucleic acid editing system described herein.
[0194] SEQ ID NOs: 1196-1204 show spacer sequences used to target genomic loci in guide RNAs.
[0195] SEQ ID NO: 1205 shows the nucleotide sequence of a plasmid encoding the MG3-6 / 3-8 adenine base editor.
[0196] SEQ ID NO: 1206 shows the nucleotide sequence of a plasmid encoding a suitable sgRNA for the MG3-6 / 3-8 adenine base editor described herein.
[0197] SEQ ID NO: 1207 shows the nucleotide sequence of a plasmid encoding the MG34-1 adenine base editor.
[0198] SEQ ID NOs: 1208-1269 show the full-length peptide sequences of MG93 deaminases suitable for the engineered nucleic acid editing system described herein.
[0199] SEQ ID NOs: 1270-1296 show the full-length peptide sequences of MG139 deaminases suitable for the engineered nucleic acid editing system described herein.
[0200] SEQ ID NOs: 1297-1311 show the full-length peptide sequences of MG152 deaminase suitable for the engineered nucleic acid editing system described herein.
[0201] SEQ ID NOs: 1312-1313 show the full-length peptide sequence of MG138 deaminase suitable for the engineered nucleic acid editing system described herein.
[0202] SEQ ID NOs: 1314-1315 show the full-length peptide sequence of MG139 deaminase suitable for the engineered nucleic acid editing system described herein.
[0203] SEQ ID NOs: 1316 to 1319 show the nucleotide sequences of 5'-FAM-labeled ssDNA.
[0204] SEQ ID NOs: 1320 to 1321 show the nucleotide sequences of Cy5.5-labeled ssDNA.
[0205] SEQ ID NOs: 1322 to 1355 show the sequences of cytidine base editors.
[0206] SEQ ID NOs: 1356 to 1362 show the full-length peptide sequences of the MG34-1 adenine base editor.
[0207] SEQ ID NOs: 1363 to 1415 show the full-length peptide sequences of the MG3-6 / 3-8 adenine base editors.
[0208] SEQ ID NOs: 1416-1417 set forth the nucleotide sequences of sgRNAs suitable for use with the MG34-1 adenine base editor described herein.
[0209] SEQ ID NO: 1418 shows the nucleotide sequence of an sgRNA suitable for use in the MG3-6 / 3-8 adenine base editor described herein.
[0210] SEQ ID NOs: 1419-1420 show the DNA sequences of target sites suitable for targeting by the MG34-1 adenine base editor described herein.
[0211] SEQ ID NO: 1421 shows the DNA sequence of a target site suitable for targeting by the MG3-6 / 3-8 adenine base editor described herein.
[0212] SEQ ID NO: 1422 sets forth the nucleotide sequence of a plasmid suitable for expression of the MG34-1 adenine base editor described herein.
[0213] SEQ ID NO: 1423 sets forth the nucleotide sequence of a plasmid suitable for expression of the MG3-6 / 3-8 adenine base editor described herein.
[0214] SEQ ID NO: 1424 shows the full-length peptide sequence of the MG35-1 adenine base editor.
[0215] SEQ ID NOs: 1425-1426 set forth the nucleotide sequences of plasmids suitable for expression of the MG35-1 adenine base editor and sgRNA described herein.
[0216] SEQ ID NOs: 1427-1428 set forth the nucleotide sequences of sgRNAs suitable for use with the MG35-1 adenine base editor described herein.
[0217] SEQ ID NOs: 1429-1430 set forth the DNA sequences of target sites suitable for targeting by the MG35-1 adenine base editor described herein.
[0218] SEQ ID NOs: 1431-1454 show the nucleotide sequences of sgRNAs engineered to function with the MG3-6 / 3-8 adenine base editor to target APOA1.
[0219] SEQ ID NOs: 1455 to 1478 show the DNA sequences of the APOA1 target sites.
[0220] SEQ ID NOs: 1479-1483 show the nucleotide sequences of sgRNAs engineered to function with the MG3-6 / 3-8 adenine base editor to target ANGPTL3.
[0221] SEQ ID NOs: 1484 to 1488 show the DNA sequences of the ANGPTL3 target sites.
[0222] SEQ ID NOs: 1489-1490 show the nucleotide sequences of sgRNAs engineered to function with the MG3-6 / 3-8 adenine base editor to target TRAC.
[0223] SEQ ID NOs: 1491 to 1492 show the DNA sequences of the TRAC site.
[0224] SEQ ID NOs: 1493 to 1516 show the nucleotide sequences of NGS primers suitable for use in assessing base editing of APOA1.
[0225] SEQ ID NOs: 1517 to 1521 show the nucleotide sequences of NGS primers suitable for use in assessing base editing of ANGPTL3.
[0226] SEQ ID NOs: 1522-1523 show the nucleotide sequences of NGS primers suitable for use in assessing base editing of TRAC.
[0227] SEQ ID NOs: 1524 to 1547 show the nucleotide sequences of NGS primers suitable for use in assessing base editing of APOA1.
[0228] SEQ ID NOs: 1548 to 1552 show the nucleotide sequences of NGS primers suitable for use in assessing base editing of ANGPTL3.
[0229] SEQ ID NOs: 1553-1554 show the nucleotide sequences of NGS primers suitable for use in assessing base editing of TRAC.
[0230] SEQ ID NO: 1555 shows the nucleotide sequence of a plasmid suitable for use in mRNA production.
[0231] SEQ ID NOs: 1556 to 1562 show the full-length peptide sequences of MG131 adenine deaminase variants.
[0232] SEQ ID NOs: 1563 to 1566 show the full-length peptide sequences of MG134 adenine deaminase variants.
[0233] SEQ ID NOs: 1567 to 1574 show the full-length peptide sequences of MG135 adenine deaminase variants.
[0234] SEQ ID NOs: 1575 to 1589 show the full-length peptide sequences of MG137 adenine deaminase variants.
[0235] SEQ ID NOs: 1590 to 1599 show the full-length peptide sequences of MG68 adenine deaminase variants.
[0236] SEQ ID NOs: 1600 to 1602 show the full-length peptide sequences of MG132 adenine deaminase variants.
[0237] SEQ ID NOs: 1603 to 1616 show the full-length peptide sequences of MG133 adenine deaminase variants.
[0238] SEQ ID NOs: 1617 to 1624 show the full-length peptide sequences of MG136 adenine deaminase variants.
[0239] SEQ ID NOs: 1625 to 1633 show the full-length peptide sequences of MG129 adenine deaminase variants.
[0240] SEQ ID NOs: 1634 to 1638 show the full-length peptide sequences of MG130 adenine deaminase variants.
[0241] SEQ ID NOs: 1639 to 1644 show the full-length peptide sequences of the MG34-1 adenine base editor.
[0242] SEQ ID NOs: 1645-1646 show the nucleotide sequences of ssDNA substrates suitable for testing adenine deaminase activity in vitro. DETAILED DESCRIPTION OF THE INVENTION
[0243] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0244] The practice of some methods disclosed herein employs, unless otherwise indicated, techniques in immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, e.g., Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F.M.A.usubel, et al. eds.); the series Methods in Enzymology (Academic Press, Inc.), PCR2: A Practical Approach (M.J.MacPherson, B.D.Hames and G.R.Taylor eds. (1995)), Harlow and Lane, eds. (1988), Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)), which are incorporated herein by reference in their entireties.
[0245] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or variations thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0246] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0247] As used herein, "cell" generally refers to a biological cell. A cell can be the basic structural, functional, or biological unit of a living organism. A cell can originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-celled eukaryotic cells, protozoan cells, cells from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, bryophytes, mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.), and the like. Examples of cells include cells from various organisms, such as porcine algae (e.g., C. Agardh), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from vertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.), etc. In some cases, the cells are not derived from a naturally occurring organism (e.g., the cells may be synthetically produced and sometimes referred to as artificial cells).
[0248] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. A nucleotide may include synthetic nucleotides. A nucleotide may include synthetic nucleotide analogs. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term "nucleotide" may refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Examples of dideoxyribonucleoside triphosphates include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled, such as by using a moiety containing an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif.; fluoro-conjugated deoxynucleotides, fluoro-conjugated Cy3-dCTP, fluoro-conjugated Cy5-dCTP, fluoro-conjugated fluoroX-dCTP, fluoro-conjugated Cy3-dUTP, and fluoro-conjugated Cy5-dUTP available from Amersham, Arlington Heights, Ill.; and fluoro-conjugated Cy5-dUTP available from Boehringer Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Mannheim, Indianapolis, Ind.; and Molecular Examples of chromosomal labeled nucleotides include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, all available from Probes, Eugene, Oreg. Nucleotides may also be labeled or marked by chemical modification. The chemically modified single nucleotide may be a biotin-dNTP.Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0249] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are generally used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-, double-, or multiple-stranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may be present in a cell-free environment. A polynucleotide may be a gene or a fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may have any three-dimensional structure and may perform any function. A polynucleotide may contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acids, heterologous nucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, a loci (locus) defined from binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers.The sequence of nucleotides may be interrupted by non-nucleotide components.
[0250] The terms "transfection" or "transfected" generally refer to the introduction of nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule may be a gene sequence encoding an entire protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.
[0251] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein and generally refer to a polymer of at least two amino acid residues joined by a peptide bond. The term does not denote a specific length of the polymer, and is not intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary or tertiary structure (e.g., domains). The term also encompasses amino acid polymers modified by any other manipulation, such as disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. As used herein, the terms "amino acid" and "amino acids" generally refer to natural and unnatural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids may include natural amino acids and unnatural amino acids, which are chemically modified to include non-naturally occurring groups or chemical moieties on the amino acid. Amino acid analogs may refer to amino acid derivatives. The term "amino acid" includes both D- and L-amino acids.
[0252] As used herein, "non-naturally occurring" can generally refer to a nucleic acid or polypeptide sequence that is not found in a naturally occurring nucleic acid or protein. Non-naturally occurring can refer to an affinity tag. Non-naturally occurring can refer to a fusion. Non-naturally occurring can refer to a naturally occurring nucleic acid or polypeptide sequence that includes a mutation, insertion, or deletion. A non-naturally occurring sequence can exhibit or encode an activity (e.g., an enzymatic activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) that can also be exhibited by the nucleic acid or polypeptide sequence to which the non-naturally occurring sequence is fused. A non-naturally occurring nucleic acid or polypeptide sequence can be linked to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid or polypeptide sequence that encodes the chimeric nucleic acid or polypeptide.
[0253] As used herein, the term "promoter" generally refers to a regulatory DNA region that controls the transcription or expression of a gene and may be located adjacent to or overlapping the nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often called transcription factors, which promote the binding of RNA polymerase to DNA, thereby resulting in gene transcription. A "basal promoter," also called a "core promoter," may generally refer to a promoter that contains all the basic elements required to promote the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters may contain a TATA box or a CAAT box.
[0254] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) or by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0255] As used herein, "operably linked," "operably linked," "operably linked," or grammatical equivalents thereof generally refer to the juxtaposition of genetic elements, e.g., promoters, enhancers, polyadenylation sequences, etc., where the elements are in a relationship permitting them to operate in an expected manner. For example, a regulatory element, which may include a promoter sequence or an enhancer sequence, is operably linked to a coding region if the regulatory element helps initiate transcription of the coding sequence. There can be intervening residues between the regulatory element and the coding region so long as this functional relationship is maintained.
[0256] As used herein, a "vector" generally refers to a polymer or an association of polymers that contains or associates with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally contains genetic elements, such as regulatory elements, operably linked to a gene to facilitate expression of the gene in a target.
[0257] As used herein, "expression cassette" and "nucleic acid cassette" are generally used interchangeably to refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to a combination of a gene or genes with regulatory elements that are operably linked for expression.
[0258] A "functional fragment" of a DNA or protein sequence generally refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be the ability to affect expression in a manner attributable to the full-length sequence.
[0259] As used herein, an "engineered" entity generally refers to an entity that has been modified by human intervention. By way of non-limiting example, a nucleic acid may be modified by changing its sequence to one that does not occur in nature; a nucleic acid may be modified by ligating it to a nucleic acid with which it is not naturally associated, such that the ligated product has a function not present in the original nucleic acid; an engineered nucleic acid may be synthesized in vitro using a sequence that does not occur in nature; a protein may be modified by changing its amino acid sequence to a sequence that does not occur in nature; an engineered protein may acquire a new function or property. An "engineered" system includes at least one engineered component.
[0260] As used herein, "synthetic" and "artificial" are used interchangeably to refer to proteins or domains thereof that have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains.
[0261] As used herein, the term "tracrRNA" or "tracr sequence" generally refers to a nucleic acid having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity or similarity to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from S. pyogenes, S. aureus, etc.). tracrRNA can refer to a nucleic acid having up to about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity or similarity to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from S. pyogenes, S. aureus, etc.). tracrRNA can also refer to modified forms of tracrRNA, which may contain nucleotide alterations, such as deletions, insertions, or substitutions, variants, mutations, or chimeras. A tracrRNA may refer to a nucleic acid that may be at least about 60% identical to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from S. pyogenes, S. aureus, etc.) over a stretch of at least six consecutive nucleotides. For example, a tracrRNA sequence may be at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, or 100% identical to a wild-type exemplary tracrRNA sequence (e.g., a tracrRNA from S. pyogenes, S. aureus, etc.) over a stretch of at least six consecutive nucleotides. Type II tracrRNA sequences can be predicted on a genomic sequence by identifying regions that have complementarity to portions of repeat sequences in adjacent CRISPR arrays.
[0262] As used herein, "guide nucleic acid" can generally refer to a nucleic acid that can hybridize to another nucleic acid. A guide nucleic acid can be RNA. A guide nucleic acid can be DNA. A guide nucleic acid can be programmed to site-specifically bind to a nucleic acid sequence. The nucleic acid to be targeted, or the target nucleic acid, can comprise nucleotides. A guide nucleic acid can comprise nucleotides. A portion of the target nucleic acid can be complementary to a portion of the guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to and hybridizes with a guide nucleic acid can be referred to as a complementary strand. A strand of a double-stranded target polynucleotide that is complementary to a complementary strand and therefore not complementary to the guide nucleic acid can be referred to as a non-complementary strand. A guide nucleic acid can comprise a polynucleotide strand and can be referred to as a "single guide nucleic acid." A guide nucleic acid can comprise two polynucleotide strands and can be referred to as a "dual guide nucleic acid." Otherwise, the term "guide nucleic acid" can be inclusive, referring to both single and double guide nucleic acids. A guide nucleic acid may comprise a segment that may be referred to as a "nucleic acid targeting segment" or a "nucleic acid targeting sequence." The nucleic acid targeting segment may comprise a subsegment that may be referred to as a "protein binding segment" or a "protein binding sequence" or a "Cas protein binding segment."
[0263] The terms "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences generally refer to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical or have a certain percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and presence of 11, gap cost at an extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of word length (W) of 2, expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (default parameters for BLASTP are available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using the Smith-Waterman homology search algorithm with parameters of match of 2, mismatch of -1, and gap of -1; MUSCLE using default parameters; MAFFT using parameters retree of 2 and maximum iterations of 1,000; Novafold using default parameters; and HMMER hmmalign using default parameters.
[0264] As used herein, the term "RuvC_III domain" generally refers to the third, non-contiguous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three non-contiguous segments, RuvC_I, RuvC_II, and RuvC_III). RuvC domains or segments thereof can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF18541 for RuvC_III).
[0265] As used herein, the term "HNH domain" generally refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF01844 for domain HNH).
[0266] As used herein, the term "base editor" generally refers to an enzyme that catalyzes the conversion of one targeted base or base pair to another (e.g., A:T to G:C, C:G to T:A) without requiring the creation and repair of a double-stranded break. In some embodiments, the base editor is a deaminase.
[0267] As used herein, the term deaminase generally refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine deaminase, which catalyzes the hydrolytic deamination of adenine or adenosine (e.g., an engineered adenosine deaminase that deaminates adenosine in DNA). In some embodiments, the deaminase or deaminase domain is a cytidine (or cytosine) deaminase, which catalyzes the hydrolytic deamination of cytidine (or cytosine) or deoxycytidine to uridine (or uracil) or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine (or cytosine) deaminase domain, which catalyzes the hydrolytic deamination of cytosine (or cytosine) to uracil (or uridine). In some embodiments, the deaminase or deaminase domain is a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, mouse, or bacterium (e.g., E. coli). In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism that does not occur in nature.
[0268] In the context of two or more nucleic acid or polypeptide sequences, the term "optimally aligned" generally refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences aligned for maximum amino acid residue or nucleotide correspondence, as determined, for example, by the alignment producing the highest or "optimized" percent identity score.
[0269] The present disclosure also encompasses variants of any of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by finding amino acid residues (e.g., non-conserved residues) that vary between species without altering the basic function of the encoded protein. Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of the endonuclease protein sequences described herein. In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with substitutions such that the activity of one or more critical active site residues or guide RNA binding residues of the endonuclease are not destroyed.
[0270] The present disclosure also includes variants (e.g., reduced activity variants) of any of the enzymes described herein having substitutions of one or more catalytic residues to reduce or eliminate the activity of the enzyme. In some embodiments, reduced activity variants of the proteins described herein include disruptive substitutions of at least one, at least two, or all three catalytic residues. In some embodiments, any of the endonucleases described herein may include a nickase mutation. In some embodiments, any of the endonucleases described herein may include a RuvC domain that lacks nuclease activity. In some embodiments, any of the endonucleases described herein may be configured to cleave one strand of a double-stranded target deoxyribonucleic acid. In some embodiments, any of the endonucleases described herein may lack endonuclease activity or be configured to be catalytically ineffective.
[0271] Conservative substitution tables providing functionally similar amino acids are available in various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993)). The following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G), 2) Aspartic acid (D), glutamic acid (E), 3) Asparagine (N), Glutamine (Q), 4) Arginine (R), Lysine (K), 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V), 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W), 7) Serine (S), Threonine (T), and 8) Cysteine (C), methionine (M).
[0272] overview The discovery of new CRISPR enzymes with unique functionality and structure could further disrupt deoxyribonucleic acid (DNA) editing technology, offering the potential to improve speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeats (CRISPR) systems in microorganisms and the sheer diversity of microbial species, there are relatively few functionally characterized CRISPR enzymes in the literature. This is in part due to the inability to easily cultivate vast numbers of microbial species under laboratory conditions. Metagenomic sequencing from natural environmental niches representing a large number of microbial species could dramatically increase the number of documented new CRISPR systems and potentially expedite the discovery of novel oligonucleotide editing functions. A recent example of the utility of such an approach was demonstrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.
[0273] CRISPR systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences that encode an RNA-based targeting element; and (ii) an ORF encoding a nuclease polypeptide directed by the RNA-based targeting element along with accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (the PAM is typically a sequence not commonly represented in the host genome). Depending on the exact function and organization of the system, CRISPR systems are generally organized into two classes, five types, and 16 subtypes based on common functional characteristics and evolutionary similarities (see Figure 1).
[0274] Class I CRISPR systems have large multi-subunit effector complexes and include types I, III, and IV.
[0275] Type I CRISPR systems are considered to be of intermediate complexity in terms of components. In Type I CRISPR systems, an array of RNA targeting elements is transcribed as a long precursor crRNA (pre-crRNA), which is processed at the repeat element to release a short mature crRNA that, when followed by a suitable short consensus sequence called a protospacer adjacent motif (PAM), directs the nuclease complex to the nucleic acid target. This processing occurs via the endoribonuclease subunit (Cas6) of a larger endonuclease complex called Cascade, which also contains the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Type I nucleases function primarily as DNA nucleases.
[0276] Type III CRISPR systems can be characterized by the presence of a central nuclease known as Cas10, along with repeat-associated mysterious proteins (RAMPs) containing Csm or Cmr protein subunits. Similar to type I systems, mature crRNA is processed from pre-crRNA using a Cas6-like enzyme. Unlike type I and II systems, type III systems are thought to target and cleave DNA-RNA duplexes (such as the DNA strand used as a template for RNA polymerase).
[0277] Type IV CRISPR systems possess an effector complex containing a highly reduced large subunit nuclease (csf1), two genes for RAMP proteins of the Cas5 (csf3) and Cas7 (csf2) families, and, in some cases, a predicted small subunit gene; such systems are commonly found on endogenous plasmids.
[0278] Class II CRISPR systems generally have a single polypeptide multidomain nuclease effector and include types II, V, and VI.
[0279] The type II CRISPR system is considered the simplest in terms of components. In the type II CRISPR system, processing the CRISPR array into mature crRNA does not require the presence of a special endonuclease subunit, but rather a small trans-encoded crRNA (tracrRNA) with a region complementary to the array repeat sequence. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to generate a mature effector enzyme loaded with both tracrRNA and crRNA. Type II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure containing a RuvC-like endonuclease domain that fits into an RNase H fold with an unrelated HNH nuclease domain inserted into the RuvC-like nuclease domain fold. The RuvC-like domain is responsible for cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is responsible for cleavage of the displacement DNA strand.
[0280] Type V CRISPR systems are characterized by a nuclease effector (e.g., Cas12) structure similar to that of type II effectors, including a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use a tracrRNA to process the pre-crRNA into mature crRNA. However, unlike type II systems, which require RNAse III to cleave the pre-crRNA into multiple crRNAs, type V systems can cleave the pre-crRNA using the effector nuclease itself. Like type II CRISPR systems, type V CRISPR systems are also known as DNA nucleases. Unlike type II CRISPR systems, some type V enzymes (e.g., Cas12a) appear to have robust single-strand nonspecific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.
[0281] Type VI CRIPSR systems possess an RNA-guided RNA endonuclease. Instead of a RuvC-like domain, the single polypeptide effector of type VI systems (e.g., Cas13) contains two HEPN ribonuclease domains. Unlike both type II and type V systems, type VI systems may not require a tracrRNA to process pre-crRNA into crRNA. However, like type V systems, some type VI systems (e.g., C2C2) appear to possess robust single-strand nonspecific nuclease (ribonuclease) activity that is activated by cleavage of the target RNA by the initial crRNA.
[0282] Class II CRISPRs are simpler constructs and have therefore been the most widely adapted for engineering and development as engineered nuclease / genome editing applications.
[0283] One of the earliest adaptations of such a system for in vitro use can be found in Jinek et al. (Science. 2012 Aug 17;337(6096):816-21, incorporated herein by reference in its entirety). The Jinek study initially used (i) recombinantly expressed and purified full-length Cas9 (e.g., a Class II, Type II enzyme) isolated from S. pyogenes SF370, (ii) purified mature approximately 42-nt crRNA (the entire crRNA was transcribed in vitro from a synthetic DNA template bearing a T7 promoter sequence) with an approximately 20-nt 5' sequence complementary to the target DNA sequence to be cleaved followed by a 3' tracr binding sequence, (iii) purified tracrRNA transcribed in vitro from a synthetic DNA template bearing a T7 promoter sequence, and (iv) Mg 2+ Jinek later described an improved engineered system in which (ii) the crRNA is joined to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) that can itself guide Cas9 to the target.
[0284] Mali et al. (Science. 2013 Feb 15;339(6121):823-826), which is incorporated herein by reference in its entirety, later adapted this system for use in mammalian cells by providing a DNA vector encoding (i) an ORF encoding a codon-optimized Cas9 (e.g., a class II, type II enzyme) under a suitable mammalian promoter with a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal), and (ii) an ORF encoding an sgRNA (having a 5' sequence starting with G, followed by a 3' tracr binding sequence, a linker, and 20 nt of complementary targeting nucleic acid sequence attached to the tracrRNA sequence) under a suitable polymerase III promoter (e.g., a U6 promoter).
[0285] Base Editing Base editing is the conversion of one target base or base pair to another (e.g., A:T to G:C, C:G to T:A) without the need for double-strand break generation and repair. Base editing can be achieved using DNA and RNA base editors, which allow for the introduction of point mutations at specific sites in either DNA or RNA. Generally, DNA base editors can comprise a fusion of a catalytically inactive nuclease with a catalytically active base-modifying enzyme that acts on single-stranded DNA (ssDNA). RNA base editors can also be composed of similar RNA-specific enzymes. Base editing can increase the efficiency of gene modification while reducing off-target and random mutations in DNA.
[0286] DNA base editors are engineered ribonucleoprotein complexes that act as tools for single-base substitution in cells and organisms. They can be created by fusing an engineered base-modifying enzyme with a catalytically deficient CRISPR endonuclease variant that cannot cleave dsDNA but can unfold dsDNA in a protospacer adjacent motif (PAM) sequence-dependent manner, allowing the guide RNA to find its complementary target and indicate the ssDNA cut site. The guide RNA anneals to the complementary DNA, displacing the ssDNA fragment and directing the CRISPR "scissors" to the base modification site. The cellular repair machinery uses information from the complementary edited template to repair the nicked, unedited strand.
[0287] To date, two types of DNA editors have been developed: cytosine base editors (CBEs) and adenine base editors (ABEs). They have been shown to efficiently and precisely edit point mutations in DNA with minimal off-target DNA editing (see Nat Biotechnol. 2017;35:435-437, Nat Biotechnol. 2017;35:438-440, and Nat Biotechnol. 2017;35:475-480, the contents of each of which are incorporated herein by reference in their entirety). However, recent findings indicate that off-target modifications exist in DNA, and many off-target modifications are also introduced into RNA by DNA base editors.
[0288] MG Base Editor In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: (a) an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, wherein the endonuclease is a class 2, type II endonuclease, and wherein the endonuclease is configured to lack nuclease activity; (b) a base editor bound to the endonuclease; and (c) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof. Optionally, the RuvC domain lacks nuclease activity. Optionally, the endonuclease comprises a nickase mutation. Optionally, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid. Optionally, the ribonucleic acid sequence configured to bind to the endonuclease comprises a tracr sequence.
[0289] In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: (a) a nucleic acid sequence encoding any one of SEQ ID NOs: 70-78 or 597, or a variant thereof, which sequence is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least 135%, at least 136%, at least 137%, at least 138%, at least 139%, at least 140%, at least 141%, at least 142%, at least 143%, at least 144%, at least 145%, at least 146%, at least 147 The present invention includes an endonuclease having 8%, at least 99%, or 100% sequence identity with a target deoxyribonucleic acid sequence, the endonuclease configured to lack nuclease activity; a base editor bound to the endonuclease; and an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to the endonuclease. Optionally, the ribonucleic acid sequence configured to bind to the endonuclease comprises a tracr sequence. Optionally, the RuvC domain lacks nuclease activity. Optionally, the endonuclease comprises a nickase mutation. Optionally, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid.
[0290] In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: (a) an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising any one of SEQ ID NOs: 360-368 or 598, wherein the endonuclease is a class 2, type II endonuclease, and the endonuclease is configured to lack nuclease activity; (b) a base editor bound to the endonuclease; and (c) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to the endonuclease. Optionally, the ribonucleic acid sequence configured to bind to the endonuclease comprises a tracr sequence. Optionally, the endonuclease comprises a nickase mutation. Optionally, the RuvC domain lacks nuclease activity. In some cases, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid.
[0291] In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has less than 80% identity to a Cas9 endonuclease. In some embodiments, the endonuclease further comprises an HNH domain. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to about 60-90 contiguous nucleotides selected from any one of SEQ ID NOs: 88-96, 488-489, or 679-680, or a variant thereof. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 88-96, 488-489, or 679-680, or a variant thereof.
[0292] In some aspects, the disclosure provides an engineered nucleic acid editing system, the system comprising: (a) an engineered guide ribonucleic acid structure, comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a tracr ribonucleic acid sequence configured to bind to an endonuclease, the tracr ribonucleic acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 1109%, at least 1111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, at least 126%, at least 127%, at least 128%, at least 129%, at least 130%, at least 131%, at least 132%, at least 133%, at least 134%, at least 135%, at least 136%, at least 137%, at least 138%, at least 139%, at least and an engineered guide ribonucleic acid structure comprising a tracr ribonucleic acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the tracr ribonucleic acid sequence, or a variant thereof; and a Class 2, Type II endonuclease configured to bind to the engineered guide ribonucleic acid.
[0293] In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence comprising any one of 360, 362, or 368. In some embodiments, the base editor comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, or 599-675, or a variant thereof. In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenosine deaminase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NOs: 50-51, 57, 385-443, 448-475, or 595, or a variant thereof. In some embodiments, the base editor is a cytosine deaminase.In some embodiments, the cytosine deaminase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 594, or 58-66, or a variant thereof.
[0294] In some embodiments, the engineered nucleic acid editing system further comprises a uracil DNA glycosylase inhibitor. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 52-56, or SEQ ID NO: 67, or a variant thereof.
[0295] In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some embodiments, the engineered guide nucleic acid structure comprises one ribonucleic acid polynucleotide comprising a guide ribonucleic acid sequence and a tracr ribonucleic acid sequence. In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease.
[0296] The NLS may comprise any of the sequences in Table 1 below, or a combination thereof.
[0297] [Table 1]
[0298] In some embodiments, the endonuclease is covalently linked directly to the base editor or covalently linked to the base editor via a linker. In some embodiments, the linker joining any of the enzymes or domains described herein can comprise one or more copies of a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SGGSSGGSSGSETPGTSESATPESSGGSSGGS, SGSETPGTSESATPESA, GSGGS, SGSETPGTSESATPES, SGGSS, or GAAA, or any other linker sequence described herein. In some embodiments, the polypeptide comprises an endonuclease and a base editor. In some embodiments, the endonuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid. In some embodiments, the endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof. In some embodiments, the system comprises a Mg 2+ The present invention further comprises a source of
[0299] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 70, or a variant thereof; the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 88; and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 360.
[0300] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 71, or a variant thereof; the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 89; and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 361.
[0301] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 73, or a variant thereof; the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 91; and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 363.
[0302] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 75, or a variant thereof; the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 93; and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 365.
[0303] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 76, or a variant thereof; the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 94; and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 366.
[0304] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 77, or a variant thereof; the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 95; and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 367.
[0305] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 78, or a variant thereof; the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 96; and the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 368.
[0306] In some embodiments, the base editor comprises an adenine deaminase. In some embodiments, the adenine deaminase comprises SEQ ID NO: 57, or a variant thereof. In some embodiments, the base editor comprises a cytosine deaminase. In some embodiments, the cytosine deaminase comprises SEQ ID NO: 58, or a variant thereof. In some embodiments, the engineered nucleic acid editing system described herein further comprises a uracil DNA glycosylation inhibitor. In some embodiments, the uracil DNA glycosylation inhibitor comprises SEQ ID NO: 67, or a variant thereof.
[0307] In some embodiments, sequence identity is determined by the BLASTP, CLUSTALW, MUSCLE, MAFFT, or Smith-Waterman homology search algorithm. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using a BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment.
[0308] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes a class 2, type II endonuclease linked to a base editor, and wherein the endonuclease is derived from an uncultured microorganism.
[0309] In some aspects, the disclosure provides nucleic acids comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes an endonuclease having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof in combination with a base editor. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N- or C-terminus of the endonuclease. In some embodiments, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human.
[0310] In some aspects, the present disclosure provides a vector comprising a nucleic acid sequence encoding a class 2, type II endonuclease linked to a base editor, wherein the endonuclease is derived from an uncultured microorganism. In some embodiments, the vector comprises a nucleic acid described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the nucleic acid comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus. In some aspects, the present disclosure provides a cell comprising the vector described herein. In some aspects, the present disclosure provides a method of producing an endonuclease, comprising culturing a cell described herein.
[0311] In some aspects, the disclosure provides methods of modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising contacting the double-stranded deoxyribonucleic acid polynucleotide with a complex comprising an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, and the endonuclease is a class 2, type II endonuclease, and the RuvC domain lacks nuclease activity, a base editor bound to the endonuclease, and an engineered guide ribonucleic acid structure configured to bind to the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM).
[0312] In some embodiments, the endonuclease comprising a RuvC domain and an HNH domain is covalently linked to the base editor directly or via a linker. In some embodiments, the endonuclease comprising a RuvC domain and an HNH domain comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof.
[0313]
[0013] In some aspects, the disclosure provides methods of modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising contacting the double-stranded deoxyribonucleic acid polynucleotide with a complex comprising a class 2, type II endonuclease, a base editor bound to the endonuclease, and an engineered guide ribonucleic acid structure configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and the PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 360-368 or 598, or a variant thereof.
[0314] In some embodiments, the Class 2, Type II endonuclease is covalently linked to the base editor or is linked to the base editor via a linker. In some embodiments, the base editor comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a sequence selected from SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, or 599-675, or a variant thereof. In some embodiments, the base editor comprises an adenine deaminase, the double-stranded deoxyribonucleic acid polynucleotide comprises an adenine, and modifying the double-stranded deoxyribonucleic acid polypeptide comprises converting the adenine to guanine. In some embodiments, the adenine deaminase comprises a sequence having at least 95% identity to SEQ ID NO: 57, or a variant thereof.
[0315] In some embodiments, the base editor comprises a cytosine deaminase, the double-stranded deoxyribonucleic acid polynucleotide comprises a cytosine, and modifying the double-stranded deoxyribonucleic acid polypeptide comprises converting the cytosine to a uracil. In some embodiments, the cytosine deaminase comprises a sequence having at least 95% identity to SEQ ID NO: 58, or a variant thereof. In some embodiments, the cytosine deaminase comprises a sequence having at least 95% identity to any one of SEQ ID NOs: 59-66, or a variant thereof.
[0316] In some embodiments, the complex further comprises a uracil DNA glycosylase inhibitor. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70%, 80%, 90%, or 95% identity to any one of SEQ ID NOs: 52-56 or SEQ ID NO: 67, or a variant thereof. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of the engineered guide ribonucleic acid structure and a second strand comprising the PAM. In some embodiments, the PAM is immediately adjacent to the 3' end of the sequence complementary to a sequence of the engineered guide ribonucleic acid structure.
[0317] In some embodiments, the Class 2, Type II endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the Class 2, Type II endonuclease is derived from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0318] In some aspects, the disclosure provides methods of modifying a target nucleic acid locus, the method comprising delivering an engineered nucleic acid editing system described herein to a target nucleic acid locus, wherein an endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure, the complex being configured such that upon binding of the complex to the target nucleic acid locus, the complex modifies a nucleotide at the target nucleic acid locus.
[0319] In some embodiments, the engineered nucleic acid editing system comprises adenine deaminase, the nucleotide is adenine, and modifying the target nucleic acid locus comprises converting the adenine to guanine. In some embodiments, the engineered nucleic acid editing system comprises cytidine deaminase and a uracil DNA glycosylase inhibitor, the nucleotide is cytosine, and modifying the target nucleic acid locus comprises converting the adenine to uracil. In some embodiments, the target nucleic acid locus comprises genomic DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, the cell is in an animal.
[0320] In some embodiments, the cell is in a cochlea. In some embodiments, the cell is in an embryo. In some embodiments, the embryo is a two-cell stage embryo. In some embodiments, the embryo is a mouse embryo. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a nucleic acid described herein or a vector described herein. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease.
[0321] In some embodiments, the nucleic acid comprises a promoter to which an open reading frame encoding an endonuclease is operably linked. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivering the engineered nucleic acid editing system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide ribonucleic acid structure operably linked to a ribonucleic acid (RNA) pol III promoter.
[0322] In some aspects, the disclosure provides engineered nucleic acid editing polypeptides comprising an endonuclease comprising a RuvC domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism, the endonuclease is a Class 2, Type II endonuclease, and the endonuclease is configured to lack nuclease activity. In some embodiments, the endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof.
[0323] In some aspects, the disclosure provides an engineered nucleic acid editing peptide comprising: an endonuclease having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof, wherein the endonuclease is configured to lack nuclease activity; and a base editor linked to the endonuclease.
[0324] In some aspects, the disclosure provides an engineered nucleic acid editing polypeptide, comprising: an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising any one of SEQ ID NOs: 360-368 or 598, wherein the endonuclease is a Class 2, Type II endonuclease, and wherein the endonuclease is configured to lack nuclease activity; and a base editor coupled to the endonuclease.
[0325] In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has less than 80% identity to a Cas9 endonuclease. In some embodiments, the endonuclease further comprises an HNH domain. In some embodiments, the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to about 60-90 contiguous nucleotides selected from any one of SEQ ID NOs: 88-96, 488-489, or 679-680, or a variant thereof. In some embodiments, the ribonucleic acid sequence configured to bind to the endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 88-96, 488-489, or 679-680, or a variant thereof.In some embodiments, the base editor comprises a sequence having at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 70-78 or 597, or a variant thereof. In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenosine deaminase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 50-51, 57, 385-443, 448-475, or 595, or a variant thereof. In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-49, 444-447, 594, or 58-66, or a variant thereof.
[0326] The systems of the present disclosure can be used for a variety of applications, such as nucleic acid editing (e.g., gene editing), binding to nucleic acid molecules (e.g., sequence-specific binding), etc. Such systems can be used, for example, to address (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivating enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to prevent them from infecting host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors.
[0327] [Table 2-1]
[0328] [Table 2-2]
[0329] [Table 2-3]
[0330] [Table 2-4]
[0331] [Table 2-5]
[0332]
Table 2-6
[0333]
Table 2-7
[0334]
Table 2-8
[0335]
Table 2-9
[0336]
Table 2-10
[0337]
Table 2-11
[0338]
Table 2-12
[0339]
Table 2-13
[0340]
Table 2-14
[0341]
Table 2-15
[0342]
Table 2-16
[0343]
Table 2-17
[0344]
Table 2-18
[0345]
Table 2-19
[0346]
Table 2-20
[0347]
Table 2-21
[0348]
Table 2-22
[0349]
Table 2-23
[0350]
Table 2-24
[0351]
Table 2-25
[0352]
Table 2-26
[0353]
Table 2-27
[0354]
Table 2-28
[0355]
Table 2-29
[0356]
Table 2-30
[0357]
Table 2-31
[0358]
Table 2-32
[0359]
Table 2-33
[0360]
Table 2-34
[0361]
Table 2-35
[0362]
Table 2-36
[0363]
Table 2-37
[0364]
Table 2-38
[0365]
Table 2-39
[0366]
Table 2-40
[0367]
Table 2-41
[0368]
Table 2-42
[0369]
Table 2-43
[0370]
Table 2-44
[0371]
Table 2-45
[0372]
Table 2-46
[0373]
Table 2-47
[0374]
Table 2-48
[0375]
Table 2-49
[0376]
Table 2-50
[0377]
Table 2-51
[0378]
Table 2-52
[0379]
Table 2-53
[0380]
Table 2-54
[0381]
Table 2-55
[0382]
Table 2-56
[0383]
Table 2-57
[0384]
Table 2-58
[0385]
Table 2-59
[0386]
Table 2-60
[0387]
Table 2-61
[0388]
Table 2-62
[0389]
Table 2-63
[0390]
Table 2-64
[0391]
Table 2-65
[0392]
Table 2-66
[0393]
Table 2-67
[0394]
Table 2-68
[0395]
Table 2-69
[0396]
Table 2-70
[0397]
Table 2-71
[0398]
Table 2-72
[0399]
Table 2-73
[0400]
Table 2-74
[0401]
Table 2-75
[0402]
Table 2-76
[0403]
Table 2-77
[0404]
Table 2-78
[0405]
Table 2-79
[0406]
Table 2-80
[0407]
Table 2-81
[0408]
Table 2-82
[0409]
Table 2-83
[0410]
Table 2-84
[0411]
Table 2-85
[0412]
Table 2-86
[0413] [Table 2-87] [Example]
[0414] Example 1 - Plasmid construction of base editors To generate base editing enzymes that utilize CRISPR function to target these base edits, effector enzymes were fused to the exemplary deaminases described herein in various configurations. This process involved first constructing a vector suitable for generating the fusion enzyme. Two entry plasmid vectors, MGA and MGC, were first constructed.
[0415] To construct the MGA (Metagenomi adenine base editor) entry plasmid containing the T7 promoter-His tag-TadA*(ABE8.17m)-SV40 NLS, three DNA fragments were amplified from pAL6. To construct the MGC (Metagenomi cytosine base editor) entry plasmid containing the T7 promoter-His tag-APOBEC1(BE3)-UGI-SV40 NLS, APOBEC1 and UGI-SV40 NLS were amplified from pAL9, and two pieces of the vector backbone were amplified from pAL6 (see Figure 3).
[0416] To introduce mutations into the effectors, source plasmids containing the MG1-4, MG1-6, MG3-6, MG3-7, MG3-8, MG4-5, MG14-1, MG15-1, or MG18-1 effector gene sequences were amplified with Q5 DNA polymerase using forward and reverse primers incorporating the appropriate mutations. The linear DNA fragments were then phosphorylated and ligated. The DNA templates were digested with DpnI using KLD Enzyme Mix (New England Biolabs) according to the manufacturer's instructions.
[0417] To generate the pMGA and pMGC expression plasmids, genes were amplified from the plasmids carrying the mutated effectors and cloned into the MGA and MGC entry plasmids via the XhoI and SacII sites, respectively. To clone the sgRNA expression cassette containing the T7 promoter-sgRNA-bidirectional terminator into the BE expression plasmid, one set of primers (P366 as the forward primer) was used to amplify the T7 promoter-spacer sequence, while another set of primers (P367 as the reverse primer) was used to amplify the spacer sequence-sgRNA scaffold-bidirectional terminator, using the pTCM plasmid as a template (see Figure 2). The two fragments were assembled into pMGA and pMGC via the XbaI site, yielding pMGA-sgRNA and pMGC-sgRNA, respectively.
[0418] [Table 3-1]
[0419] [Table 3-2]
[0420] All amplified DNA fragments were purified using a QIAquick Gel Extraction Kit (Qiagen) and assembled using NEBuilder HiFi DNA Assembly (New England Biolabs). The resulting assemblies were propagated in Endura electrocompetent cells (Lucergen) according to the manufacturer's instructions (Figures 4 and 5). The DNA sequences of all cloned genes were confirmed at ELIM BIOPHARM.
[0421] [Table 4]
[0422] Example 2 - Protein Expression and Purification The T7 promoter-driven mutated effector genes in the pMGA and pMGC plasmids were expressed in E. coli BL21(DE3) cells in Magic Media by transformation with each of the respective plasmids described in Example 1 above, according to the manufacturer's instructions (Thermo). After incubation at 16°C for 40 hours, the transformed cells were harvested and suspended in lysis buffer (HisTrap equilibration buffer: 20 mM Tris (Sigma T2319-100 ml), 300 mM sodium chloride (VWR VWRVE529-500 ml), 5% glycerol, 10 mM MgCl2, and 10 mM imidazole (Sigma 68268-100 ml-F), pH 7.5), and EDTA-free protease inhibitors (Pierce), and frozen in a -80°C freezer. The cells were then thawed on ice, sonicated, clarified, filtered, and affinity purified. According to the manufacturer's specifications, the protein was applied to a Cytiva 5 ml HisTrap FF column on an Äkta Avant FPLC and eluted with an isocratic elution of 20 mM Tris (Sigma T2319-100 ml), 300 mM sodium chloride (VWR VWRVE529-500 ml), 5% glycerol, 10 mM MgCl2, 250 mM imidazole (Sigma 68268-100 ml-F), pH 7.5. The elution fraction containing the His-tagged effector protein was concentrated and buffer exchanged into 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 5% glycerol, pH 7.5. Protein concentration was determined by bicinchoninic acid assay (Thermo) and adjusted after determining relative purity by SDS-PAGE densitometry in an Image Lab (Bio-Rad) (see Figure 7).
[0423] Example 3 - In vitro nickase assay A linear fragment of LacZ containing the effector targeting sequence was amplified using Q5 DNA polymerase with 6-carboxyfluorescein (6-FAM)-labeled primers P141 and P146 (SEQ ID NOs: 179 and 180) synthesized by IDT. DNA fragments containing sgRNAs containing a T7 promoter followed by a 20-bp or 22-bp spacer sequence were transcribed in vitro using the HiScribe T7 High Yield RNA Synthesis Kit (New England Biolabs) according to the manufacturer's instructions. Synthetic sgRNAs with sequences corresponding to the designated sgRNAs in the Sequence Listing were purified using the Monarch RNA Cleanup Kit (New England Biolabs) according to the user's manual, and the concentration was measured using a Nanodrop™.
[0424] To determine DNA nickase activity, each purified mutant effector was first supplemented with its cognate sgRNA. The reaction was initiated by adding the linear DNA substrate to a 15 μL reaction mixture containing 10 mM Tris pH 7.5, 10 mM MgCl2, and 100 mM NaCl, 150 nM enzyme, 150 nM RNA, and 15 nM DNA. The reaction was incubated at 37°C for 2 hours. The digested DNA was purified using AMPure XP SPRI paramagnetic beads (Beckman Coulter) and eluted with 6 μL of TE buffer (10 mM Tris, 1 mM EDTA, pH 8.0). Nicked DNA was resolved on a 10% TBE-urea denaturing gel (Biorad) and imaged by ChemiDoc (Bio-Rad) (see Figure 7, where the depicted enzyme exhibits nickase activity by producing bands at 600 bases and 200 bases, compared to 400 bases and 200 bases for the wild-type enzyme). The results showed that all nickase mutants tested in Figure 7 exhibited their expected nickase activity but not wild-type cleavage activity, except for MG4-5(D17A), which was inconclusive.
[0425] Example 4 - Introduction of a base editor into E. coli Plasmids were transformed into electrocompetent BL21(DE3) cells (Lucergen) according to the manufacturer's instructions. After electroporation, cells were recovered in expression recovery medium at 37°C for 1 hour and then spread onto LB plates containing 100 μg / ml ampicillin and 0.1 mM IPTG. After overnight growth at 37°C, colonies were picked and the lacZ gene was amplified using Q5 DNA polymerase (New England Biolabs) with primers P137 and P360. The resulting PCR products were purified and sequenced by Sanger sequencing at ELIM BIOPHARM. Base editing was determined by examining the presence of C-to-T or A-to-G transversions in the target protospacer region for cytosine or adenine base editors, respectively.
[0426] To evaluate editing efficiency in E. coli, the plasmids were transformed into electrocompetent BL21(DE3) (Lucergen), and the electroporated cells were allowed to recover for 1 hour at 37°C using expression recovery medium. Then, 10 μL of recovered cells were inoculated into 990 μL of SOB containing 100 μL / mg ampicillin and 0.1 mM IPTG in a 96-well deep-well plate and grown at 37°C for 20 hours. One μL of cells induced for base editor expression was used to amplify the lacZ gene in a 20 μL PCR reaction (Q5 DNA polymerase) using primers P137 and P360. The resulting PCR product was purified and sequenced by Sanger sequencing at ELIM BIOPHARM. Quantification of editing efficiency was performed using Edit R as described in Example 12.
[0427] [Table 5]
[0428] Example 5 - Protein nucleofection and amplicon sequencing in mammalian cells (prophetic) Nucleofection is performed on mammalian cells (e.g., K-562, Neuro-2A, or RAW264.7) using the Lonza 4D Nucleofector and the Lonza SF Cell Line 4D-Nucleofector X Kit S (catalog number V4XC-2032) according to the manufacturer's recommendations. After preparing the SF nucleofection buffer, 200,000 cells per nucleofection are resuspended in 5 μl of buffer. In the remaining 15 μl of buffer per nucleofection, 20 pmol of chemically modified sgRNA from Synthego is combined with 18 pmol of base editor enzyme (e.g., ABE8e) and incubated at room temperature for 5 minutes to form a complex. The cells are added to a 20 μl nucleofection cuvette, followed by the addition of the protein solution, and the mixture is triturated to mix. Cells were nucleofected using program CM-130, and then immediately added to each well with 80 μl of warmed medium to recover. After 5 minutes, 25 μl of each sample was added to 250 μl of fresh medium in a 48-well poly-d-lysine plate (Corning). After an additional 3 days of culture, the cells were treated for genomic DNA extraction in the same manner as the lipofected cells described above.
[0429] After Illumina barcoding, PCR products are pooled and purified by electrophoresis on a 2% agarose gel using the Monarch DNA Gel Extraction Kit (New England Biolabs) and eluted in 30 μl of HO. DNA concentrations are quantified with a Qubit dsDNA High Sensitivity Assay Kit (Thermo Fisher Scientific) and sequenced on an Illumina MiSeq instrument (paired-end reads, R1: 250–280 cycles, R2: 0 cycles) according to the manufacturer's protocol.
[0430] Sequencing reads were demultiplexed using MiSeq Reporter (Illumina), and FASTQ files were analyzed using CRISPResso2. Dual edits in individual alleles were analyzed by a Python script. Base editing values represent n = 3 independent biological replicates collected by different researchers, and mean ± SD is shown. Base editing values are reported as the percentage of reads with adenine mutagenesis relative to the total number of aligned reads.
[0431] Example 6 - Plasmid nucleofection and whole genome sequencing (predictive) in mammalian cells All plasmids were assembled using the uracil-specific excision reagent (USER) cloning method. Guide RNA plasmids for SpCas9, SaCas9, and all engineered variants were assembled. Plasmids for mammalian cell transfection were prepared using the ZymoPURE Plasmid Midiprep Kit (Zymo Research Corporation). HEK293T cells (ATCC CRL-3216) were cultured in Dulbecco's modified Eagle's medium (Corning) supplemented with 10% fetal bovine serum (ThermoFisher Scientific) and maintained at 37°C with 5% CO2.
[0432] HEK293T cells were seeded onto 48-well poly-d-lysine plates (Corning) in the same culture medium. 12–16 h after plating, cells were transfected with 1.5 μl of Lipofectamine 2000 (ThermoFisher Scientific) using 750 ng of base editor plasmid, 250 ng of guide RNA plasmid, and 10 ng of green fluorescent protein as a transfection control. Cells were cultured for 3 days with the medium replaced after the first day, then washed with approximately 1 Å of PBS (ThermoFisher Scientific). Genomic DNA was then extracted by adding 100 μl of freshly prepared lysis buffer (10 mM Tris-HCl, pH 7.5, 0.05% SDS, 25 μg ml of proteinase K (ThermoFisher Scientific)) directly to each transfected well. The mixture was incubated at 37°C for 1 h and then heat-inactivated at 80°C for 30 min. The genomic DNA lysate is then immediately used for high-throughput sequencing (HTS).
[0433] We performed HTS on genomic DNA from HEK293T cells. After Illumina barcoding, PCR products were pooled and purified by electrophoresis on a 2% agarose gel using the Monarch DNA Gel Extraction Kit (NEB) and eluted in 30 μl of HO. DNA concentration was quantified using the Qubit dsDNA High Sensitivity Assay Kit (ThermoFisher Scientific) and sequenced on an Illumina MiSeq instrument (paired-end reads, R1: 250–280 cycles, R2: 0 cycles) according to the manufacturer's protocol.
[0434] Example 7 - Determining Edit Window (Predictive) To examine the editing window region, the cytosine with the highest CT conversion frequency in a given sgRNA is normalized to 1, and other cytosines at positions ranging from 30 nt upstream to 10 nt downstream of the PAM sequence of the same sgRNA (43 bp in total) are then normalized. The normalized CT conversion frequencies are then sorted and compared according to their positions for all tested sgRNAs for the given base editor. The comprehensive editing window (CEW) is defined to span positions with an average CT conversion efficiency greater than 0.6 after normalization.
[0435] To examine the substrate preference for each cytidine deaminase, C-sites were first classified according to their position in the sgRNA-targeted region, and positions containing at least one C-site with a normalized C-to-C conversion frequency of ≥ 0.8 were included in the subsequent analysis. Selected C-sites were then compared according to the base type (NC or CN) upstream or downstream of the edited cytosine. For cytidine deaminases that exhibit efficient C-to-C conversion at both the N- and C-termini of the endonuclease, substrate preference was assessed by integrating the respective NT-CBE and CT-CBE. Statistical analysis was performed using one-way ANOVA, with p<0.05 considered significant.
[0436] Example 8a - Testing off-target analysis using whole genome sequencing and transcriptomics in mammalian cells (predictive) HEK293T cells were seeded onto 48-well poly-d-lysine-coated plates at a density of 3.104 cells per well in antibiotic-free DMEM + GlutaMAX medium (Thermo Fisher Scientific) 16–20 h prior to lipofection. 750 ng of nickase or base editor expression plasmid DNA was combined with 250 ng of sgRNA expression plasmid DNA in 15 μl of Opti-MEM + GlutaMAX. This was then combined with 10 μl of lipid mixture containing 1.5 μl of Lipofectamine 2000 and 8.5 μl of Opti-MEM + GlutaMAX per well. Three days after transfection, cells were harvested and either DNA or RNA was collected. For DNA analysis, cells were washed once in PBS and then lysed in 100 μl of QuickExtract Buffer (Lucigen) according to the manufacturer's instructions. For RNA collection, the MagMAX mirVana Total RNA Isolation Kit (Thermo Fisher Scientific) is used with KingFisher Flex.
[0437] Genomic DNA from mammalian cells was fragmented and adapter-ligated using the Nextera DNA Flex Library Prep Kit (Illumina) in a 96-well plate using Nextera Index Primers (Illumina) according to the manufacturer's instructions. Library size and concentration were confirmed using a Fragment Analyzer (Agilent), and the DNA was sent to Novogene for WGS using an Illumina HiSeq system.
[0438] All targeted NGS data are analyzed by performing four general operations: (1) alignment, (2) duplicate marking, (3) variant calling, and (4) variant background filtering to remove artifacts and germline mutations. Mutant reference alleles and alternative alleles are reported relative to the plus strand of the reference genome.
[0439] For whole-transcriptome sequencing, mRNA selection is performed using the NEBNext Poly(A) mRNA Magnetic Isolation Module (New England BioLabs). RNA library preparation is performed using the NEBNext Ultra II RNA Library Prep Kit for Illumina (New England BioLabs). Based on the RNA input amount, cycle number 12 is used for PCR enrichment of adapter-ligated DNA. NEBNext Sample Purification Beads (New England BioLabs) are used throughout all size selections performed by this method. NEBNext Multiplex Oligos for Illumina (New England BioLabs) are used for multiplex indexing according to the PCR recipe outlined in the protocol. Prior to sequencing, sample quality is confirmed using the high-sensitivity D1000 ScreenTape on a 4200 TapeStation System (Agilent). Libraries are pooled and sequenced using NovaSeq (Novogene). Targeted RNA sequencing is then performed. Complementary DNA is generated by PCR with reverse transcription (RT-PCR) from isolated RNA using the SuperScript IV One-Step RT-PCR system with EZDnase (Thermo Fisher Scientific) according to the manufacturer's instructions.
[0440] The following program was used: 58°C for 12 minutes, 98°C for 2 minutes, followed by PCR cycles that varied depending on the amplicon: 32 cycles for CTNNB1 and IP90 (98°C for 10 seconds, 60°C for 10 seconds, 72°C for 30 seconds). After combined RT-PCR, the amplicons were barcoded and sequenced using the Illumina MiSeq sequencer described above. The first 125 nucleotides in each amplicon, starting with the first base after the end of the forward primer in each amplicon, were aligned to a reference sequence and used to analyze the maximum A-to-I frequency in each amplicon. Off-target DNA sequencing was performed using a two-step PCR and barcoding method using primers to prepare samples for sequencing using the Illumina MiSeq sequencer described above.
[0441] Example 8b - Analysis of off-target editing by whole genome sequencing and transcriptomics (predictive) Transfected cells prepared as in Example 8a were harvested after 3 days, and genomic DNA was isolated using the Agencourt DNAdvance Genomic DNA Isolation Kit (Beckman Coulter) according to the manufacturer's instructions. On-target and off-target genomic regions of interest were amplified by PCR using flanking HTS primer pairs. PCR amplification was performed using 5 ng of genomic DNA as a template with Phusion High-Fidelity DNA Polymerase (ThermoFisher) according to the manufacturer's instructions. The number of cycles was determined separately for each primer pair to ensure that the reaction was stopped in the linear range of amplification (30, 28, 28, 28, 32, and 32 cycles for EMX1, FANCF, HEK293 site 2, HEK293 site 3, HEK293 site 4, and RNF2 primers, respectively). PCR products were purified using RapidTips (Diffinity Genomics). The purified DNA was amplified by PCR using primers containing sequencing adapters. Products are gel-purified and quantified using the Quant-iT™ PicoGreen dsDNA Assay Kit (ThermoFisher) and the KAPA Library Quantification Kit-Illumina (KAPA Biosystems). Samples are sequenced on an Illumina MiSeq as previously described.
[0442] Sequencing reads were automatically demultiplexed using MiSeq Reporter (Illumina), and individual FASTQ files were analyzed with a custom Matlab script. Each read was pairwise aligned to the appropriate reference sequence using the Smith-Waterman algorithm. Base calls with a Q score below 31 were replaced with N and therefore excluded from nucleotide frequency calculations. This process resulted in an expected MiSeq base calling error rate of approximately 1 in 1,000. Aligned sequences that did not contain gaps in the read and reference sequences were saved in an alignment table, from which base frequencies were tabulated for each locus. Indel frequencies were quantified with a custom Matlab script.
[0443] Sequencing reads are scanned for exact matches to two 10-bp sequences flanking a window in which indels can occur. If an exact match is not found, the read is excluded from analysis. If the length of this indel window exactly matches the reference sequence, the read is classified as not containing an indel. If the indel window is more than two bases longer or shorter than the reference sequence, the sequencing read is classified as an insertion or deletion, respectively.
[0444] Example 9 - Mouse editing experiments (predictive) It is envisioned that base editors comprising novel DNA-targeting nuclease domains fused to novel deaminase domains can be validated as therapeutic candidates by testing in appropriate mouse models of disease.
[0445] One example of a suitable model includes mice engineered to express human PCSK9 protein, as described, for example, in Herbert et al. (10.1161 / ATVBAHA.110.204040). PCSK9 protein regulates LDL receptor (LDLR) levels and affects serum cholesterol levels. Mice expressing human PCSK9 protein exhibit elevated cholesterol levels and more rapid progression of atherosclerosis. PCSK9 is a validated drug target for reducing lipid levels in people at high risk for cardiovascular disease due to abnormally high plasma lipid levels (https: / / doi.org / 10.1038 / s41569-018-0107-8). Reducing PCSK9 levels through genome editing is expected to permanently lower lipid levels throughout an individual's lifetime, thus resulting in a lifelong reduction in cardiovascular disease risk. One genome editing approach can involve targeting the coding sequence of the PCSK9 gene and editing the sequence to generate a premature stop codon, thereby preventing translation of PCSK9 mRNA into a functional protein. Targeting a region near the 5' end of the coding sequence is useful for blocking translation of most of the protein. To generate a stop codon (TGA, TAA, TAG) with high efficiency and specificity, it is necessary to target a region of the PCSK9 coding sequence, and the editing window is positioned on the appropriate sequence so that the most frequent editing event results in a stop codon. Therefore, the availability of a multiple base editing system with a wide range of PAMs or a base editing system with a degenerate PAM is useful for accessing more potential target sites within the PCSK9 gene. Furthermore, additional editing systems with a low frequency of off-target editing (e.g., within 1% or less of on-target editing events) are also useful for performing gene editing in this situation.
[0446] To achieve significant reductions in plasma lipid levels, the base editing efficiency required for therapeutic efficacy is in the range of 50% or higher. An example of the use of a base editor to create a stop codon in the PCSK9 gene is that of Carreras et al. (https: / / doi.org / 10.1186 / s12915-018-0624-2), in which 10%–34% of PCSK9 alleles were edited to create a stop codon. While this level of editing was sufficient to result in measurable reductions in plasma lipid levels in mice, higher editing efficiencies are required for therapeutic use in humans.
[0447] To identify the optimal base editing (BE) system and guide for introducing a stop codon into the PCSK9 gene, screening can be performed in a mouse liver cell line, such as Hepa1-6 cells. In silico screening can be used to initially identify guides that target the PCSK9 gene using various available BE systems. To select from a large number of potential guides, in silico analysis can be performed to determine which guides have an editing window that encompasses sequences that may generate a stop codon upon editing. Preference can then be given to guides that are close to the 5' end of the coding sequence. The resulting set of guides and BE proteins can be combined to form a ribonucleoprotein complex (RNP) and nucleofected into Hepa1-6 cells. After 72 hours, the efficiency of editing at the target site can be determined by next-generation sequencing (NGS). Based on these in vitro results, one or more BE / guide combinations that resulted in the most frequent stop codon formation can be selected for in vivo testing.
[0448] Application in human therapeutic settings requires safe and effective methods for delivering base-editing components, including base editors and guide RNAs. In vivo delivery methods can be divided into viral and non-viral methods. Among viral vectors, adeno-associated virus (AAV) is the virus of choice for clinical use due to its safety record, efficient delivery to multiple tissues and cell types, and established manufacturing process. Large base editors (BEs) exceed the packaging capacity of AAVs, which precludes packaging them into a single AAV. An approach using split intein technology to package BEs into two AAVs has been demonstrated to be successful in mice (https: / / doi.org / 10.1038 / s41551-019-0501-5), but the need for two viruses can complicate development and manufacturing. Another drawback of AAV is that the virus lacks mechanisms for promoting integration into the host cell genome. While most of the AAV genome remains episomal, portions of the AAV genome integrate at random double-strand breaks that naturally occur within the cell (Curr Opin Mol Ther. 2009 August;11(4):442-447). This can result in persistence of the gene sequence expressing the BE throughout the life of the organism. Furthermore, the AAV genome can persist as an episome in the nucleus of transduced cells and be maintained for years, resulting in long-term expression of the BE in these cells. This potentially increases the risk of off-target effects, as the risk of off-target events is a function of the time the editing enzyme is active. Adenoviruses (Ads), such as Ad5, can efficiently deliver DNA payloads to the mammalian liver and can package up to 45 kb of DNA. However, adenoviruses are known to induce strong immune responses in mammals, including patients (http: / / dx.doi.org / 10.1136 / gut.48.5.733), potentially leading to serious adverse events, including death (https: / / doi.org / 10.1016 / j.ymthe.2020.02.010).
[0449] Nonviral delivery vectors, including lipid nanoparticles and polymeric nanoparticles (reviewed in doi:10.1038 / mt.2012.79), offer several advantages over viral delivery vectors, including lower immunogenicity and transient expression of nucleic acid transporters. The transient expression induced by nonviral delivery vectors is expected to minimize off-target events, making them particularly suitable for genome editing applications. Additionally, unlike viral vectors, nonviral delivery offers the potential for repeated administration to achieve therapeutic efficacy. While there is no theoretical limit to the size of nucleic acid molecules that can be packaged into nonviral vectors, in practice, packaging efficiency decreases as the size of the nucleic acid increases and particle size increases.
[0450] BEs can be delivered in vivo using non-viral vectors, such as lipid nanoparticles (LNPs), by encapsulating synthetic mRNA encoding the BE together with a guide RNA into the LNP. This can be done using any suitable method, such as those described in Finn et al. (DOI: 10.1016 / j.celrep.2018.02.014) or Yin et al. (DOI: 10.1038 / nbt.3471). LNPs can deliver their cargo preferentially to hepatocytes in the liver, which is also the target organ / cell type when attempting to disrupt PCSK9 gene expression. To demonstrate proof-of-concept for this approach, we envision that a BE consisting of a novel genome-editing protein fused to a deaminase domain can be encoded in synthetic mRNA and packaged into LNPs together with an appropriate guide RNA targeting a selected site in the mouse PCSK9 gene. In the case of mice engineered to express the human PCSK9 gene, the guide can be designed to selectively target the human PCSK9 gene or both the human and mouse PCSK9 genes. After injection of these LNPs, editing efficiency at on-target sites in the genome of hepatocytes can be analyzed by amplicon sequencing or other methods such as tracking indels by degradation (doi:10.1093 / nar / gku936). Physiological effects can be determined by measuring lipid levels in the blood of mice, including total cholesterol and triglyceride levels, using standard methods.
[0451] Another example of a disease that can be modeled in mice to evaluate novel BEs is primary hyperoxaluria type I. Primary hyperoxaluria type I (PH1) is a rare autosomal recessive disease caused by a deficiency in the AGXT gene, which encodes the enzyme alanine-glyoxylate aminotransferase. This results in defective glyoxylate metabolism and the accumulation of the toxic metabolite oxalate. One approach to treating this disease is to reduce the expression of the enzyme glycolate oxidase (GO), which produces glyoxylate from glycolate, thereby reducing the amount of substrate (glyoxylate) available for oxalate formation. PH1 can be modeled in mice in which both copies of the AGXT gene are knocked out (agxt - / - mice), resulting in a significant three-fold increase in urinary oxalate levels compared to wild-type controls. Therefore, agxt - / - mice can be used to evaluate the efficacy of novel base editors designed to generate stop codons in the coding sequences of endogenous mouse GO genes. To identify the optimal BE system and guide for introducing a stop codon into a GO gene, a screen can be performed in a mouse liver cell line, such as Hepa1-6 cells. In silico screening can be used to initially identify guides that target GO genes using various available BE systems. To select from a large number of potential guides, in silico analysis can be performed to determine which guides have editing windows that encompass sequences that may generate a stop codon upon editing. In some cases, guides near the 5' end of the coding sequence can be utilized. The resulting set of guides and BE proteins can be combined to form a ribonucleoprotein complex (RNP) and nucleofected into Hepa1-6 cells. After 72 hours, editing efficiency at the target site can be determined by next-generation sequencing (NGS). Based on these in vitro results, one or more BE / guide combinations that resulted in the most frequent stop codon formation can be selected for in vivo testing in mice.
[0452] The BE and guide can be delivered to mice using an AAV virus with a split intein system to express the BE and a third AAV to deliver the guide. Alternatively, due to the packaging capacity of adenoviruses >40 kb, adenovirus type 5 can be used to deliver the BE and guide in a single virus. Furthermore, the BE can be delivered as mRNA with the guide RNA packaged in an appropriate LNP. After intravenous injection of LNP into agxt- / - mice, urinary oxalate levels can be monitored over time to determine whether oxalate levels are reduced, which may indicate that the BE is active and has the expected therapeutic effect. To determine whether the BE has introduced a stop codon, the appropriate region of the GO gene can be PCR amplified from genomic DNA extracted from the livers of treated and control mice. The resulting PCR products can be sequenced using next-generation sequencing to determine the frequency of sequence variations.
[0453] Example 10 - Discovery of a new deaminase gene We discovered novel deaminases by mining 4 Tbp (terabase pairs) of proprietary and publicly collected metagenomic sequencing data from diverse environments (soil, sediment, groundwater, thermophilic, human, and non-human microbiomes). We constructed HMM profiles of documented deaminases and searched against all predicted proteins using HMMER3 (hmmer.org) to identify deaminases from our database. Predicted and reference (e.g., eukaryotic APOBEC1, bacterial TadA) deaminases were aligned with MAFFT, and phylogenetic trees were inferred using FastTree2. Novel families and subfamilies were defined by identifying clades composed of the sequences disclosed herein. Candidates were selected based on the presence of key catalytic residues indicative of enzymatic function (see, for example, SEQ ID NOs: 1-51, 385-386, 387-443, 444-447, 488-475, 599-675, 744-835, or 970-982).
[0454] Example 11 - Plasmid construction Gene DNA fragments were synthesized by either Twist Bioscience or Integrated DNA Technologies (IDT). Plasmid DNA was amplified in Endura electrocompetent cells (Lucigen) and isolated using the QIAprep Spin Miniprep Kit (Qiagen). Vector backbones were prepared by restriction enzyme digestion of the plasmids. Inserts were amplified with Q5 High-Enrichment DNA Polymerase (New England Biolabs) using primers (SEQ ID NOs: 690-707) ordered from either Elim BIOPHARM or IDT. Both the vector backbone and insert were purified by gel extraction using a Gel DNA Recovery Kit (Zymo Research). One or more DNA fragments were assembled into vectors (SEQ ID NOs: 483-487, 720-726, or 737-738) using NEBuilder HiFi DNA assembly (New England Biolabs).
[0455] Example 12 - Assessing base editing efficiency in E. coli by sequencing Five nanograms of extracted DNA prepared as in Example 4 was used as a template, and primers (P137 and P360) were used for PCR amplification. The resulting product was submitted for Sanger sequencing at ELIM BIOPHARM. The primers used for sequencing are shown in Tables 6 and 7 (SEQ ID NOs: 523 to 531).
[0456] [Table 6-1]
[0457] [Table 6-2]
[0458] [Table 6-3]
[0459] [Table 6-4]
[0460] [Table 6-5]
[0461] [Table 7]
[0462] Figures 8A-8C show examples of base editing by the enzymes investigated by this experiment, as assessed by Sanger sequencing.
[0463] Figures 10A-10B show the base editing efficiency of adenine base editors (ABEs) using TadA(ABE8.17m) (SEQ ID NO: 596) and MG nickase according to Table 3. TadA is a tRNA adenine deaminase, and TadA(ABE8.17m) is an engineered variant of E. coli TadA. Twelve MG nickases fused to TadA(ABE8.17m) were constructed and tested in E. coli. Three guides targeting lacZ were designed. The numbers in the boxes indicate the percentage of A to G conversion quantified by Edit R at each position. ABE8.17m was used as a positive control in the experiment.
[0464] Figures 11A-11B show the base editing efficiency of a cytosine base editor (CBE) containing rat APOBEC1, MG nickase, and Bacillus subtilis bacteriophage uracil glycosylase inhibitor (UGI(PBS1)). APOBEC1 is a cytosine deaminase. Twelve MG nickases fused to rAPOBEC1 at the N-terminus and UGI at the C-terminus were constructed and tested in E. coli. Three guides targeting lacZ were designed. The numbers in the boxes indicate the percentage of C-to-T conversion quantified by EditR. BE3 was used as a positive control in the experiment.
[0465] Figure 12 shows the effect of MG uracil glycosylase inhibitors (UGIs) on base editing activity when added to CBE. (a) MGC15-1 contains N-terminal APOBEC1, MG15-1 nickase, and C-terminal UGI. Three MG UGIs were tested for improved cytosine base editing activity in E. coli. (b) BE3 contains N-terminal rAPOBEC1, SpCas9 nickase, and C-terminal UGI. Two MG UGIs were tested for improved cytosine base editing activity in HEK293T cells. Editing efficiency was quantified by Edit R.
[0466] Example 13 - Cell culture, transfection, next generation sequencing, and base editing analysis HEK293T cells were grown and passaged at 37°C and 5% CO in Dulbecco's modified Eagle's medium + GlutaMAX (Gibco) supplemented with 10% (v / v) fetal bovine serum (Gibco). 4Cells were seeded onto 96-well cell culture plates (Costar) treated for cell attachment and grown for 20–24 h. The spent medium was refreshed with fresh medium immediately before transfection. 200 ng of expression plasmid and 1 μL of Lipofectamine 2000 (ThermoFisher Scientific) per well were used for transfection, according to the manufacturer's instructions. Transfected cells were grown for 3 days, harvested, and gDNA was extracted using QuickExtract (Lucigen) according to the manufacturer's instructions. The target region for base editing was amplified using Q5 high-fidelity DNA polymerase (New England Biolabs) with the primers listed in Tables 8 and 9 (SEQ ID NOs: 538–585), and DNA was extracted as a template.
[0467] [Table 8]
[0468] [Table 9-1]
[0469] [Table 9-2]
[0470] [Table 9-3]
[0471] [Table 9-4]
[0472] PCR products were purified using the HighPrep PCR Clean-up System (MAGBIO) according to the manufacturer's instructions. The effect of uracil glycosylase inhibitors (UGIs) on base editing of candidate enzymes was analyzed by submitting the PCR products to Elim BIOPHARM for Sanger sequencing, and efficiency was quantified using EditR. To analyze base editing of A0A2K5RND7-MG nickase-MG69-1, adapters used for next-generation sequencing (NGS) were added to the PCR products in a subsequent PCR reaction using primers compatible with the KAPA HiFi HotStart ReadyMix PCR Kit (Roche) and TruSeq DNA Library Prep Kits (Illumina). The DNA concentration of the resulting products was quantified using TapeStation (Agilent), and samples were pooled together to prepare libraries for NGS analysis. The resulting libraries were quantified by qPCR using an Aria Real-Time PCR System (Agilent), and high-throughput sequencing was performed using an Illumina Miseq instrument according to the manufacturer's instructions. Sequencing data were analyzed for base editing by Cripresso2.
[0473] Figures 13A-13B show maps of sites targeted by base editors, illustrating the base editing efficiency of cytosine base editors containing a CMP / dCMP-type deaminase domain-containing protein (uniprot accession A0A2K5RDN7), MG nickase, and MG UGI. The construct contains N-terminal A0A2K5RDN7, MG nickase, and C-terminal MG69-1. For simplicity, the identity of the MG nickase is indicated in the figure. BE3 (APOBEC1) was used as a positive control for base editing. An empty vector was used as a negative control. Three independent experiments were performed on different days. Abbreviations: R: repeat, NEG: negative control.
[0474] [Table 10]
[0475] Example 14 - Positive selection of base-edited mutants in E. coli Figure 14 shows the positive selection method for TadA characterization in E. coli. Panel (a) shows a map of one plasmid system used for TadA selection. The vector contains CAT(H193Y), a CAT-targeting sgRNA expression cassette, and an ABE expression cassette. The N-terminal TadA from E. coli and the C-terminal SpCas9(D10A) from Streptococcus pyogenes are shown in this figure. Panel (b) shows a sequencing trace demonstrating that when introduced / transformed into E. coli cells, it edits the A2 position of the template strand of CAT(H193Y), reverting the H193Y mutant to the wild type and restoring its activity. Abbreviations: CAT: chloramphenicol acetyltransferase.
[0476] One microliter of the plasmid solution at a concentration of 10 ng / μL was transformed into 25 μL of BL21(DE3) electrocompetent cells (Lucigen) and recovered in 975 μL of expression recovery medium at 37°C for 1 hour. 50 μL of the resulting cells were spread onto LB agar plates containing 100 μg / mL carbenicillin, 0.1 mM IPTG, and an appropriate amount of chloramphenicol. The plates were incubated at 37°C until colonies were pickable. Colony PCR was used to amplify the genomic region containing the base edits, and the resulting products were submitted for Sanger sequencing at ELIM BIOPHARM. The primers used for PCR and sequencing are listed in Table 10 (SEQ ID NOs: 532-537).
[0477] [Table 11]
[0478] Figure 15 shows that mutations induced by TadA confer high resistance to chloramphenicol (Cm). Panel (a) shows a photograph of a growth plate in which various concentrations of chloramphenicol were used to select for antibiotic resistance in E. coli. In this example, wild-type and two variants of E. coli-derived TadA (EcTadA) were tested. Panel (b) shows a summary table of results demonstrating that ABEs carrying mutant TadA exhibit higher editing efficiency than wild-type. In these experiments, colonies were picked from plates with Cm concentrations of 0.5 μg / mL or higher. For simplicity, the identity of the deaminase is shown in the table, while the effector (SpCas9) and construct configuration are shown in the figure above.
[0479] Figures 16A-16B show an investigation of MG68 TadA activity during positive selection. Figure 16A shows photographs of growth plates from an experiment in which eight MG68 TadA candidates were tested against 0-2 μg / mL chloramphenicol (ABE contained an N-terminal TadA variant and a C-terminal SpCas9(D10A) nickase). For simplicity, the identity of the deaminase is shown. Panel (b) shows a summary table showing the editing efficiency of the MG68 TadA candidates. Figure 16B demonstrates that MG68-3 and MG68-4 drove adenine base editing. In this experiment, colonies were picked from plates with Cm at 0.5 μg / mL or higher.
[0480] Figure 17 shows the improved base editing efficiency of MG68-4_nSpCas9 via the D109N mutation in MG68-4. Panel (a) shows a photograph of a growth plate in which wild-type MG68-4 and its variants were tested against 0 to 4 μg / mL of chloramphenicol. For simplicity, the identity of the deaminase is shown. The adenine base editors in this experiment include an N-terminal TadA variant and a C-terminal SpCas9(D10A) nickase. Panel (b) shows a summary table showing the editing efficiency of the MG TadA candidates. Panel (b) demonstrates that MG68-4 and MG68-4(D109N) exhibited adenine base editing, with the D109N mutant showing increased activity. In this experiment, colonies were picked from plates containing 0.5 μg / mL or higher Cm.
[0481] Figure 18 shows base editing of MG68-4(D109N)_nMG34-1. Panel (a) shows a photograph of a growth plate from an experiment in which an ABE containing an N-terminal MG68-4(D109N) and a C-terminal SpCas9(D10A) nickase was tested against 0 to 2 μg / mL of chloramphenicol. Panel (b) shows a summary table showing editing efficiency with and without sgRNA. In this experiment, colonies were picked from plates with Cm of 1 μg / mL or higher.
[0482] Figure 19 shows 28 MG68-4 variants designed to improve MG68-4-nMG34-1 base editing activity. Twelve residues were selected for targeted mutagenesis to improve editing of the enzyme.
[0483] Example 15 - Plasmid construction of E. coli optimized constructs All plasmids for cytidine deaminase expression were prepared by Twist Biosciences. Each construct was codon-optimized for E. coli expression and inserted into the XhoI and BamHI restriction sites of the pET-21(+) vector. The sequence was designed to exclude a BsaI restriction site. The following sequence was added to the beginning of each construct: 5'-GAAATAATTTTGTTTAACTTTAAGAAGGAGATATACATATGGGCAGCAGTCATCATCATCACCATCAC-3'. This sequence encodes a ribosome binding site and an N-terminal hexahistidine tag. A stop codon was added to the end of each CDA sequence to prevent incorporation of the C-terminal HisTag encoded by pET-21(+).
[0484] Example 16 - Plasmid construction of mammalian optimized constructs All plasmids for cytidine deaminase expression in mammalian cells were codon-optimized and ordered from Twist Biosciences. Each construct was codon-optimized for H. sapiens expression. The following restriction sites were avoided: BsaI, SphI, EcoRI, BmtI, BstX, BlpI, and BamHI. The following sequence was added 5' to the codon-optimized sequence: ACCGGTGCTAGCCCACC. This sequence contains a BmtI restriction site used for downstream cloning and a Kozak sequence for maximum translation. The following sequence was added 3' to the codon-optimized CDA: AGCGCATGC. This sequence contains a SphI restriction site to allow downstream cloning, and the stop codon was removed in all constructs.
[0485] Example 17 - Cell culture, transfection, next generation sequencing, and base editing analysis HEK293T cells were grown and passaged in Dulbecco's modified Eagle's medium + GlutaMAX (Gibco) supplemented with 10% (v / v) fetal bovine serum (Gibco) at 37°C and 5% CO2. 4Cells were seeded onto 96-well cell culture plates (Costar) treated for cell attachment and grown for 20–24 h. The spent medium was refreshed with fresh medium immediately before transfection. 300 ng of expression plasmid and 1 μL of Lipofectamine 2000 (ThermoFisher Scientific) per well were used for transfection, according to the manufacturer's instructions. Transfected cells were grown for 3 days, harvested, and gDNA was extracted using QuickExtract (Lucigen) according to the manufacturer's instructions. The target region for base editing was amplified using Q5 High-Fidelity DNA Polymerase (New England Biolabs) with primers (SEQ ID NOs: 690–707, 865–872, and 932–961), and DNA was extracted as a template. PCR products were purified using the HighPrep PCR Clean-up System (MAGBIO) according to the manufacturer's instructions. To analyze base substitutions in adenine base editors, adapters used for next-generation sequencing (NGS) were added to the PCR products in subsequent PCR reactions using primers compatible with the KAPA HiFi HotStart ReadyMix PCR Kit (Roche) and the TruSeq DNA Library Prep Kit (Illumina). The DNA concentration of the resulting products was quantified using a TapeStation (Agilent), and the samples were pooled together to prepare libraries for NGS analysis. The resulting libraries were quantified by qPCR using an Aria real-time PCR system (Agilent), and high-throughput sequencing was performed using an Illumina Miseq instrument according to the manufacturer's instructions. Sequencing data were analyzed for base editing by Crispresso2.
[0486] Example 18 - In vitro deaminase in-gel assay A linear DNA construct containing cytidine deaminase was amplified from the aforementioned plasmid in Twist via PCR. All constructs were washed with SPRI Cleanup (Lucigen) and eluted with 10 mM Tris buffer. The enzyme was expressed from the PCR template in an in vitro transcription-translation system, PURExpress (NEB), for 2 hours at 37°C. The deamination reaction was prepared by mixing 2 μL of PURExpress reaction with 2 μM of 5'-FAM-labeled ssDNA (IDT) and 1 U of User Enzyme (NEB) in 1× Cutsmart buffer (NEB). The reaction was incubated at 37°C for 2 hours and then quenched by adding 4 units of Proteinase K (NEB) and incubating at 55°C for 10 minutes. The reaction was further processed by adding 11 μL of 2× RNA loading dye and incubating at 75°C for 10 minutes. All reaction conditions were analyzed by gel electrophoresis on a 10% denaturing gel (Biorad). DNA bands were visualized using a Chemi-Doc imager (Biorad), and band intensity was quantified using BioRad Image Lab v6.0. Successful deamination was observed by visualization of a 10-bp fluorescently labeled band in the gel (Figure 20). Results showed that MG93-3 to MG93-7, MG93-11, MG138-17, MG138-20, MG138-23, MG139-12, and MG139-19 to MG139-21 were able to deaminate cytidine-containing substrates.
[0487] The in vitro activity of over 90 novel cytidine deaminases on cytosine-containing ssDNA substrates was measured in all four possible 5'-NC contexts (Figure 23). Thirty-eight of these cytidine deaminases exhibited ssDNA deamination activity, five of which were capable of substantially complete deamination of the target cytidine (MG139-84 / SEQ ID NO: 808, MG139-86 / SEQ ID NO: 810, MG139-87 / SEQ ID NO: 811, MG139-95 / SEQ ID NO: 819, and MG139-102 / SEQ ID NO: 826; see, e.g., Figure 23). Furthermore, some of the deaminases also exhibited greater than 50% deamination of the target cytosine (MG139-30 / SEQ ID NO: 752, MG139-55 / SEQ ID NO: 777, MG139-99 / SEQ ID NO: 823). Although most reported DNA cytidine deaminases operate primarily on ssDNA, often with a preference for the base immediately 5' to the substrate C, related dsDNA substrates were also included as controls (Figure 24), validating that MG139-86 and MG139-87 can also deaminate dsDNA substrates.
[0488] Example 19 - NGS-based deep deamination in vitro assay To determine cytosine deaminase activity and binding site preference, we generated an ssDNA library with a single target C. Briefly, we synthesized the ssDNA substrate oligonucleotide 5′-NNNCNNN flanked by a 21-nt and 21-nt region containing an adenine, a 20-nt upstream randomized barcode, and two conserved primer binding sites (Integrated DNA Technologies).
[0489] This resulted in an oligonucleotide pool with 4096 unique substrate sequences. For non-target C deamination events, each oligo contained a unique barcode to determine the original variable region after sequencing. First, deaminases were expressed from PCR templates in an in vitro transcription-translation system, PURExpress (NEB), at 37°C for 2 hours. PURExpress was then incubated with 0.5 pmol of the substrate oligonucleotide pool in 50 mM Tris, pH 7.5, 75 mM NaCl at 37°C for 1 hour.
[0490] A. Half of the processed pool was amplified using the Accel-NGS 1S Plus Kit (Swift) to generate a dsDNA pool, which was then further amplified with unique dual indexes and sequenced on a MiSeq for >15,000 reads per sample.
[0491] B. Half of the treated pool was annealed to the appropriate 3' barcoded adapter (IDT) and treated with T4 DNA polymerase for 20 minutes at 12°C to generate a dsDNA pool. This pool was uniquely dual-index amplified using conserved regions and sequenced on a MiSeq for >15,000 sequences per sample.
[0492] Example 20 - Lentivirus production and transduction HEK293T cells were grown and passaged in Dulbecco's modified Eagle's medium + GlutaMAX (Gibco) supplemented with 10% (v / v) fetal bovine serum (Gibco) at 37°C and 5% CO. The day before transfection, cells were plated at 5 × 10 per dish. 6The cells were seeded at 1000 x g. On the day of transfection, 8 μg of PsPax, 1 μg of pMD2-G, and 9 μg of a plasmid containing cytidine deaminase fused to MG3-6 or Cas9 were mixed together and packaged in Mirus LT1 transfection reagent (Mirus Bio). The mixture was transfected into HEK293T cells. Lentivirus was harvested 3 days after transfection, filtered through a 0.4 μM filter, and immediately used to transduce cells. Transduction occurred by adding half the volume of the virus-containing supernatant to the cells along with 8 μg / mL polybrene.
[0493] Example 21 - Adenine and cytidine base editors in E. coli and mammalian cells To demonstrate that the small type II CRISPR nuclease, MG34-1, can be used as a base editor, we generated constructs containing TadA*(8.17m)-nMG34-1 (ABE-MG34-1, SEQ ID NO: 727), where TadA*(8.17m) is an engineered TadA from E. coli, and rAPOBEC1-nMG34-1-UGI(PBS) (CBE-MG34-1, SEQ ID NO: 739), where rAPOBEC1 is rat APOBEC1 and UGI(PBS) is a uracil glycosylase inhibitor from Bacillus subtilis bacteriophage. TadA*(8.17m)-nSpCas9 (SEQ ID NO: 728) and rAPOBEC1-nSpCas9-UGI(PBS) (SEQ ID NO: 740) were generated as positive controls for editing profile analysis. Four guides (SEQ ID NOs: 729-736) targeting the lacZ gene in E. coli were designed and prepared for each base editor construct. The plasmids were transformed into BL21(DE3) and recovered in recovery medium at 37°C for 1 hour. Cell plates were plated on LB agar plates containing 100 μg / mL carbenicillin and 0.1 mM IPTG. After growing the cells at 37°C for 16-20 hours, colony PCR was used to amplify the target region in the E. coli genome, and the resulting products were analyzed by Sanger sequencing at Elim BIOPHARM (Figures 22A-22C). Sequencing results showed that both ABE-MG34-1 and CBE-MG34-1 edited the target locus in the E. coli genome at levels and within editing windows comparable to those of the positive control SpCas9 base editor (Figures 22A and 22B). Furthermore, TadA*(8.17m)-nMG34-1 showed higher base substitutions at two target loci. ABE-MG34-1 also demonstrated base editing in human cells across three different genomic targets with an editing efficiency of up to 22% (Figure 22C).
[0494] To determine whether the SMART HNH endonuclease-associated RNA enzymes and ORF (HEARO) enzymes could be used as base editors, we constructed an ABE by fusing the TadA*-(7.10) deaminase monomer to the C-terminus of an engineered MG35-1 gene containing the D59A mutation (Figure 22E). The A-to-G editing of this ABE was tested in a positive selection single-plasmid E. coli system, where the ABE required the chloramphenicol acetyltransferase (CAT) gene containing the Y193 mutation to be reverted to H193 to survive chloramphenicol selection (Figure 22D). The plasmid contained an sgRNA with either a spacer targeting the mutant CAT gene or a scrambled non-targeting spacer region (control). Colony enrichment was detected in E. coli transformed with ABE-MG35-1 targeting the CAT gene when grown on plates containing 2, 3, and 4 μg / mL chloramphenicol, but no colonies grew on plates containing 8 μg / mL chloramphenicol (Figure 22E). Sanger sequencing confirmed that 26 of 30 colonies picked from the 2, 3, and 4 μg / mL plates transformed with the targeted spacer contained the expected Y193H reversion (Table 11 and Figure 31).
[0495] [Table 12]
[0496] Because a single revertant CAT gene is sufficient to confer colony survival, the four colonies lacking the revertant CAT sequence are understood to contain more unedited than edited ...
Claims
1. An engineered nucleic acid editing system, comprising: (a) a cytidine deaminase or a nucleic acid encoding said cytidine deaminase, wherein said cytidine deaminase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 774 or 835; (b) an endonuclease domain or a nucleic acid encoding the endonuclease domain; An engineered nucleic acid editing system comprising:
2. The engineered nucleic acid editing system described in claim 1, wherein the cytidine deaminase comprises a sequence having at least 90% sequence identity to either one of SEQ ID NOs: 774 or 835.
3. The engineered nucleic acid editing system described in claim 2, wherein the cytidine deaminase comprises the sequence of sequence number 774.
4. The engineered nucleic acid editing system described in claim 2, wherein the cytidine deaminase comprises the sequence of SEQ ID NO:
835.
5. An engineered nucleic acid editing system described in any one of claims 1 to 4, wherein the endonuclease domain includes a nickase.
6. The engineered nucleic acid editing system of claim 1, further comprising one or more of a uracil DNA glycosylase inhibitor or a nucleic acid encoding the uracil DNA glycosylase inhibitor, or a FAM72A protein or a nucleic acid encoding the FAM72A protein.
7. The engineered nucleic acid editing system further comprises an engineered guide polynucleotide or a nucleic acid encoding the engineered guide polynucleotide, wherein the engineered guide polynucleotide is i) a guide ribonucleic acid sequence configured to hybridize to a target nucleic acid sequence; ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease domain; and 2. The engineered nucleic acid editing system of claim 1, comprising:
8. The engineered nucleic acid editing system described in claim 7, wherein the guide ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 1491 to 1492.
9. The engineered nucleic acid editing system of claim 2, wherein the engineered nucleic acid editing system comprises the cytidine deaminase and the endonuclease domain.
10. The engineered nucleic acid editing system of claim 9, wherein the engineered nucleic acid editing system comprises a fusion protein comprising the cytidine deaminase and the endonuclease domain.
11. An engineered nucleic acid editing system described in any one of claims 6 to 10, wherein the cytidine deaminase comprises the sequence of SEQ ID NO:
774.
12. An engineered nucleic acid editing system described in any one of claims 6 to 10, wherein the cytidine deaminase comprises the sequence of SEQ ID NO:
835.
13. Use of a polypeptide comprising a cytidine deaminase in the manufacture of a pharmaceutical for deaminating cytosine residues in a nucleic acid sequence, said use comprising contacting said polypeptide with said nucleic acid sequence, and said cytidine deaminase comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 774 or 835.
14. The use described in claim 13, wherein the cytidine deaminase comprises a sequence having at least 90% sequence identity with any one of SEQ ID NOs: 774 or 835.
15. The use described in claim 14, wherein the cytidine deaminase comprises the sequence of SEQ ID NO:
774.
16. The use described in claim 14, wherein the cytidine deaminase comprises the sequence of SEQ ID NO:
835.
17. The use of any one of claims 13 to 16, wherein the polypeptide is associated with an endonuclease.
18. The use described in claim 17, wherein the polypeptide is a fusion protein.
19. The use of claim 17, wherein the endonuclease comprises a nickase.
20. The use described in claim 17, wherein the nucleic acid sequence comprises a sequence having at least 80% sequence identity with at least 18 consecutive nucleotides of any one of SEQ ID NOs: 1491 to 1492.
21. The use described in claim 17, wherein the nucleic acid sequence is located in the TRAC locus.
22. The use described in claim 17, wherein the nucleic acid sequence is within a eukaryotic cell.