Wide-ranging base editor mutagenesis
The development of base editor fusion proteins with wider editing windows addresses the limitations of traditional genome editing technologies, enabling efficient mutagenesis and mutation screening in non-coding regions, facilitating the study of protein function and disease-associated mutations.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-02
- Publication Date
- 2026-04-09
AI Technical Summary
Existing genome editing technologies, such as Cas9-directed base editors, are limited by a protospacer adjacent motif (PAM) and narrow editing windows, restricting the saturation mutagenesis of non-coding regions, which are complex and vast, making it difficult to study protein function and disease-associated mutations effectively.
Development of base editor fusion proteins and systems that enable multiple edits across large genomic windows, utilizing split deaminases linked to DNA binding proteins and epitope domains, allowing for wider editing windows up to 450 bps, overcoming the limitations of traditional base editors.
Enables greater scaled screening and editing saturation, facilitating the study of protein function and disease-associated mutations with sustained mutagenesis, particularly in non-coding regions, by allowing multiple edits across large genomic windows.
Smart Images

Figure US2025049245_09042026_PF_FP_ABST
Abstract
Description
WIDE-RANGING BASE EDITOR MUTAGENESIS RELATED APPLICATIONS
[0001] This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Application, U.S.S.N.63 / 702,555, filed October 2, 2024, which is incorporated herein by reference. REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0002] The contents of the electronic sequence listing (B119570204WO00-SEQ- GJM.xml; Size: 535,985 bytes; and Date of Creation: October 2, 2025) are herein incorporated by reference in its entirety. BACKGROUND OF INVENTION
[0003] Transcription factors (TFs) and cis-regulatory elements (CREs), including promoters, enhancers, and insulators, have a pivotal role in orchestrating gene expression to ultimately dictate cellular identity, function, and state. Understanding the precise function of these regulatory elements and proteins that bind to them is thus key to unraveling how specific cell states and disease are established, and how such states can be deliberately guided. Further, most of the genetic variation that underlies common disease is found in non- coding regulatory regions. Numerous studies have utilized CRISPR interference (CRISPRi), which relies upon a catalytically dead Cas9 (dCas9) fused to a KRAB repressor domain, to elucidate regulatory roles of specific sets of CREs. Recent breakthroughs in genome editing technologies now enable the systematic mutation of bases directly in cells, in their native genomic contexts. While base editor tiling presents an exceptional opportunity to elucidate functional contribution of bases in CREs and disease, challenges remain.
[0004] The complexity and vastness of non-coding space poses a significant challenge. Enhancers can be as large as 1 kilobases, contain multiple TF motifs, function through co-binding events, and work in combination with other enhancers to control gene expression. Prior methods have generally relied on massively parallel reporter assays for saturation mutagenesis, where an enhancer sequence is systematically altered and placed in the cells with a lentivirus (Kinney JB, McCandlish DM. Massively Parallel Assays and Quantitative Sequence-Function Relationships. Annu Rev Genomics Hum Genet.2019 Aug 31;20:99-127. doi: 10.1146 / annurev-genom-083118-014845. Epub 2019 May 15. PMID: 31091417.). While these experiments can be informative, they lack the endogenous locus and 1 / 169#14463983v1surrounding elements that can contribute to control. Therefore, groups have relied on Cas9- directed base editors to target endogenous loci for perturbations. However, saturation mutagenesis of the target region with Cas9-directed base editors is restricted by a protospacer adjacent motif (PAM) and narrow editing windows (4 bps). Further, the number of potential single base edits to test in non-coding regions is incredibly high. SUMMARY OF INVENTION
[0005] The present disclosure provides base editor fusion proteins, complexes, and systems capable of installation of multiple edits across large genomic windows (e.g., between 4bps and 500 bps). In some embodiments, the disclosure provides a first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain, a second fusion protein comprising a N-terminal portion of a split deaminase and a second epitope binding domain, and a third fusion protein comprising a DNA binding protein domain and one or more epitope domains. Such systems may be useful for studying protein function at amino acid-level resolution, for example, to catalogue disease-associated mutations, profile protein-drug interactions, and / or to map protein sequence-function landscapes.
[0006] Figure 1 illustrates an exemplary base editor system (e.g., complex) as contemplated herein. In some embodiments, the base editor system comprises a third fusion protein comprising a DNA binding protein domain that is capable of binding a target nucleic acid sequence (e.g., Cas9 proteins that bind DNA sequence via gRNAs). As shown in Figure 1, the DNA binding protein domain is fused to one or more epitope domains. These “epitope domains” comprise amino acid sequences (e.g., analogous to antigen sequences) recognized by a first epitope binding domain (e.g., an antibody or scFv, etc.) of a first fusion protein comprising a C-terminal portion of a split deaminase (e.g., DddA) and / or second epitope binding domain (e.g., an antibody or scFv, etc.) of a second fusion protein comprising a N- terminal portion of the split deaminase (e.g., DddA). As can be seen in Figure 1, when a first fusion protein binds the epitope domain adjacent a second fusion protein the split deaminase becomes a functional deaminase. In this way, the “epitope domains” act as a linker between the DNA binding protein domain and the deaminase domain. In some embodiments, the functional deaminase is capable of editing double stranded DNA.
[0007] Those of skill in the art will understand and appreciate that the DNA binding protein domain may be programmed to target a nucleic acid of interest, and the number of the “one or more epitope domains” tailored to enable editing over a narrow (e.g., ~4 -40 bps) or wide (e.g., ~130-450 bps). 2 / 169#14463983v1
[0008] In various embodiments, the editing window is from 4–10 bps, 5–20 bps, 6–30 bps, 7–40 bps, 8–50 bps, 9–60 bps, 10–70 bps, 11–80 bps, 12–90 bps, 13–100 bps, 14–110 bps, 15–120 bps, 16–130 bps, 17–140 bps, 18–150 bps, 19–160 bps, 20–170 bps, 21–180 bps, 22–190 bps, 23–200 bps, 24–210 bps, 25–220 bps, 26–230 bps, 27–240 bps, 28–250 bps, 29–260 bps, 30–270 bps, 31–280 bps, 32–290 bps, 33–300 bps, 34–310 bps, 35–320 bps, 36–330 bps, 37–340 bps, 38–350 bps, 39–360 bps, 40–370 bps, 41–380 bps, 42–390 bps, 43–400 bps, 44–410 bps, 45–420 bps, 46–430 bps, 47–440 bps, 48–450 bps, 49–460 bps, or 50–500 bps.
[0009] In some embodiments, the base editors and complexes disclosed herein comprise a DNA binding protein domain linked to one or more epitope domains. The one or more epitope domains comprise binding motifs to which a split deaminase (e.g., split DddA) can bind. This configuration allows multiple split deaminases to bind to a single DNA binding protein comprising multiple epitope domains (e.g., 4x, 5, 10x, etc.), which enables multiple edits across large genomic windows (e.g., between 4 bps and 450 bps), relative to a 4 bp window offered by traditional base editors. Other aspects of the disclosure relate to methods, for example, methods of mutating one or more nucleotides in a target nucleic acid, and methods of performing mutational screens. The disclosure also provides compositions, polynucleotides, vectors, pharmaceutical compositions, cells, kits, and systems comprising the base editor fusion proteins and complexes contemplated herein.
[0010] The inventors of the present disclosure have developed a new base editor system with wider base editing (WBE) windows. It is believed that the newly discovered base editor systems enable greater scaled screening and editing saturation. For example, it is believed that the new base editor system enables a vast array of cytidine-to-thymine (or adenosine to guanine) edits across a genomic window of about 130 bp or greater (e.g., 450 bps), relative to a ~4 bp window in conventional base editors. Additionally, the new base editor system does not edit in the guide binding region, allowing for sustained mutagenesis, which is a distinct advantage over conventional systems in the art. In some aspects, the present disclosure relates to a complex comprising a (i) a first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain; (ii) a second fusion protein comprising a N-terminal portion of the split deaminase and a second epitope binding domain; and (iii) a third fusion protein comprising a DNA binding protein and one or more epitope domains. In some embodiments, the complex further comprises a gRNA configured to guide the DNA binding protein to a target nucleic acid sequence. 3 / 169#14463983v1
[0011] In some cases, a deaminase may be derived from a larger enzyme comprising a cytidine deaminase domain (e.g., the deaminase is truncated). For example, DddA is a cytidine deaminase derived from the type VI secretion system (T6SS)-associated deaminases (e.g., the T6SS-associated deaminase comprises the DddAdeaminase). Functional full-length enzymes contemplated herein include, but are not limited to, the amino acid sequences listed in SEQ ID NOs: 1-181 (Table 1). In some embodiments, the full-length enzymes comprise an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 100% identical to the amino acid sequences of any one of SEQ ID NOs: 1-181. In some embodiments, the full-length enzymes comprise an amino acid sequence that is identical to the amino acid sequence of any one of SEQ ID NOs: 1-181.
[0012] Additionally, functional truncated deaminases contemplated herein include, but are not limited to, the amino acid sequences listed in SEQ ID NOs: 3-4,14- 15,42,51,55,64,68-69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,194- 195,198-206,209-234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302- 337,339-343,345-351,354-364,366,368,370,375-380,382-386 (Table 2). In some embodiments, the full-length enzymes comprise an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acid sequences of any one of SEQ ID NOs: 3-4,14- 15,42,51,55,64,68-69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,194- 195,198-206,209-234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302- 337,339-343,345-351,354-364,366,368,370,375-380,382-386. In some embodiments, the full-length enzymes comprise an amino acid sequence that is identical to the amino acid sequence of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71- 73,90,95,108,142,145,151,159-160,172,174,176,178-181,194-195,198-206,209-234,236- 243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345-351,354- 364,366,368,370,375-380,382-386 (Table 2).
[0013] In some embodiments, the full-length enzyme comprises a cytidine deaminase derived from the type VI secretion system (T6SS)-associated deaminases (e.g., the T6SS- associated deaminase comprises the DddAdeaminase, DddA, SEQ ID NO: 1). In some embodiments, the C-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 1322-1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397- 4 / 169#14463983v1comprises amino acid sequence corresponding to amino acid positions 1322-1426, 1333- 1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO:1.
[0014] However, any one of the deaminases disclosed in Table 1 (SEQ ID NOs: 1- 181) may be split into a C-terminal portion. For example, in some embodiments, C-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of any one of SEQ ID NOs: 2-193 that correspond to amino acids 1322-1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO :1 or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 2-193 that corresponds to amino acids 1322-1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO :1.
[0015] Truncated versions of the full-length enzymes may also be used. For example, the DddA toxinis a truncated version of the T6SS-associated deaminase comprising an amino acid sequence identical to SEQ ID NO: 194, as shown in Table 2. In some embodiments, the DddA toxin is split into a C-terminal portion comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% , or 100% identical to the amino acids corresponding to amino acid positions 34- 138, 55-138, 69-138, 83-138, 99-138, or 109-138 of SEQ ID NO: 194 (e.g., DddA toxin) or SEQ ID NO: 201 (e.g., DddA11toxin).
[0016] Additional functional truncated enzymes (e.g., deaminases) recited in Table 1 are shown in Table 2 (SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71- 73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386). In some embodiments, a C-terminal portion of a split deaminase comprises an amino acid sequence that is at least sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68- 69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that correspond to amino acids 34-138, 55-138, 69-138, 83-138, 99-138, or 109-138 of SEQ ID NO: 194 or SEQ ID NO :201, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 3-4,14- 15,42,51,55,64,68-69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195- 5 / 169#14463983v1195,198-200, 202,209-234,236-243,245-247,249-256,258-260,263,267-282,284-287,289- 300,302-337,339-343,345-351,354-364,366,368,370,375-380,382-386 that corresponds to amino acids 34-138, 55-138, 69-138, 83-138, 99-138, or 109-138 of SEQ ID NO: 194 or SEQ ID NO: 201.
[0017] In some embodiments, the N-terminal portion of the split deaminase comprises an amino acid sequence that is that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 1-1321, 1-1332, 1-1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO: 1. In some embodiments, the N-terminal portion of the split deaminase comprises amino acid sequence corresponding to amino acid positions 1-1321, 1- 1332, 1-1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO: 1.
[0018] As mentioned above, any one of the deaminases disclosed in Table 1 (SEQ ID NOs: 1-181) may be split into a N-terminal portion and fused to an epitope binding domain. For example, in some embodiments, N-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of any one of SEQ ID NOs: 2-193 that correspond to amino acids 1-1321, 1-1332, 1-1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO: 1, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 2-193 that corresponds to amino acids 1-1321, 1-1332, 1-1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO :1.
[0019] Truncated versions of the full-length enzymes may also be used in the base editors contemplated herein. Thus, in some embodiments, a N-terminal portion of the split deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 194-385.
[0020] In some embodiments, a N-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 1-33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO: 194 (e.g., DddA toxin) or SEQ ID NO: 201 (e.g., DddA11toxin).
[0021] Just as any one of the full-length enzymes may be spilt into a C-terminal portion and an N-terminal portion, so too can any one of the truncated deaminases shown in Table 2 (SEQ ID NOs: 194-385). Thus, in some embodiments, the N-terminal portion of the split deaminase comprises an amino acid sequence that is at least sequence at least 80%, at 6 / 169#14463983v1least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68- 69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that correspond to amino acids 1-33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO: 194 or SEQ ID NO: 201, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68- 69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that corresponds to amino acids 1-33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO:194 or SEQ ID NO: 201.
[0022] The first fusion protein and second fusion protein further comprise an epitope binding domain (e.g., a first fusion protein comprises a first epitope binding domain, and a second fusion protein comprises a second epitope binding domain). In some embodiments, the first epitope binding domain and / or second binding epitope domain comprises an antibody (e.g., a bivalent antibody), an antibody fragment (e.g., a bivalent single chain variable fragment, scFv), a peptide, or a chromobody / nanobody (e.g., a bivalent chromobody / nanobody). In some embodiments, the first epitope binding domain and second epitope binding domain are configured to bind the same target domain. In other embodiments, the first epitope binding domain and second epitope binding domain bind different target domains (see Figure 1).
[0023] The first fusion protein and / or the second fusion protein may further comprise one or more additional domains. In some cases, the one or more additional domains comprise a uracil glycosylase inhibitor domain. In some embodiments, the UGI domain comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acid sequence of SEQ ID NO: 387.
[0024] In some embodiments, the one or more additional domains comprises a chromatin modifying enzyme domain. In some embodiments, the chromatin modifying enzyme domain is selected from the group consisting of acetylases, deacetylases, methyltransferases, demethylases, ligases, helicases, ubiquitin, deubiquitinases, and chromatin remodelers.
[0025] Additionally, according to some embodiments, the base editing system (e.g., complexes as shown in Figure 1) comprises a third fusion protein comprising one or more 7 / 169#14463983v1epitope domains (e.g., target binding sites for the first and / or second epitope binding domains of the first and / or second fusion proteins, respectively). In some embodiments, the third fusion protein comprises a DNA binding protein domain, such as a dead Cas9 (dCas9, e.g., comprises a D10A and H840A mutation) or a nickase Cas9 (nCas9, e.g., comprises a D10A or H840A mutation), which in association with a gRNA, mediates binding of the third fusion protein to a target nucleic acid sequence. Other DNA binding proteins are also contemplated and are discussed elsewhere herein.
[0026] In some embodiments, the one or more epitope domains are fused to a DNA binding protein domain. In some embodiments, the one or more epitope domains comprises any one of the following: a multimeric GCN4 peptide comprising an amino acid sequence EELLSKNYHLENEVARLKK (SEQ ID NO: 388), an ALFA-tag peptide comprising an amino acid sequence SRLEEELRRRLTE (SEQ ID NO: 389), a BC2 peptide comprising an amino acid sequence PDRKAAVSHWQQ (SEQ ID NO: 390) or PDRVRAVSHWSS (SEQ ID NO: 391), a peptide comprising an amino acid sequence AVERYLKDQQLLGIW (SEQ ID NO: 392), gp41 peptide comprising an amino acid sequence KNEQELLELDKWASL (SEQ ID NO: 393), a haloalkane dehalogenase, a benzylguanine ligand, a benzyluganine ligand, a FLAG peptide comprising an amino acid sequence of DYKDDDDK (SEQ ID NO: 394), a myc peptide comprising an amino acid sequence of EQKLISEEDL (SEQ ID NO: 395), an HA peptide comprising an amino acid sequence of YPYDVPDYA (SEQ ID NO: 396), a polyHistidine tag having between 6 and 10 histidine residues, a C-tag having amino acid sequence of EPEA (SEQ ID NO: 409), a Twin-Step tag having streptavidin conjugated to an amino acid sequence WSHPQFEK-(GGGS)2-GGSA-WSHPQFEK (SEQ ID NO: 397), biotin, streptavidin, or protein A / G.
[0027] In some embodiments, a third fusion protein comprises one or more repeats of the one or more epitope domains. In some embodiments, the third fusion protein comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten repeats of the one or more epitope domains.
[0028] In some embodiments, a first epitope binding domain of a first fusion protein is configured to bind to any one of the one or more epitope domains recited above. Similarly, a second epitope binding domain of a second fusion protein is configured to bind to any one of the one or more epitopes domains recited above.
[0029] Other aspects of the disclosure relate to complexes comprising a first fusion protein, a second fusion protein, and a third fusion protein. In some embodiments, the first fusion protein comprises a C-terminal portion of a split deaminase and a first epitope binding 8 / 169#14463983v1domain, the second fusion protein comprises a N-terminal portion of the split deaminase and a second epitope binding domain, and the third fusion protein comprises a DNA binding protein and one or more epitope domains. In some embodiments, the first epitope domain (e.g., antibody, antibody fragment, nanobody, etc.) of the first fusion protein binds to one or more epitope domains of the third fusion protein. In some embodiments, the second epitope domain (e.g., antibody, antibody fragment, nanobody, etc.) of the second fusion protein binds to one or more epitope domains of the third fusion protein. The first and / or second fusion proteins may bind to the same epitope domain or to different epitope domains.
[0030] Upon binding of the first epitope binding domain and second epitope binding domain to the one or more epitope domains of the third fusion protein, a C-terminal portion of a split deaminase and a N-terminal portion of the split deaminase form a functional deaminase (e.g., SEQ ID NOs: 1-386).
[0031] Other aspects of the disclosure relate to (1) compositions comprising any one or more of the first fusion proteins, second fusion proteins, and / or third fusion proteins disclosed herein; (2) polynucleotides encoding any one or more of the first fusion proteins, second fusion proteins, and / or third fusion proteins disclosed herein; (3) vectors comprising any of the polynucleotides disclosed herein; (4) pharmaceutical compositions comprising any one or more of the first fusion proteins, second fusion proteins, third fusion proteins, compositions, complexes, polynucleotides, and / or vectors disclosed herein; (5) cells comprising any one or more of the first fusion proteins, second fusion proteins, third fusion proteins, compositions, complexes, polynucleotides, vectors, and / or pharmaceutical compositions disclosed herein; and / or (6) kits comprising any one of the first fusion proteins, second fusion proteins, third fusion proteins, compositions, complexes, polynucleotides, vectors, pharmaceutical compositions, and / or cells disclosed herein.
[0032] Other aspects of the disclosure relate to one or more methods and / or uses, for example, of any one of the first fusion proteins, second fusion proteins, third fusion proteins, compositions, complexes, polynucleotides, vectors, pharmaceutical compositions, and / or cells disclosed herein.
[0033] In some embodiments, the methods relate to mutating one or more nucleotides in a target nucleic acid (e.g., genome of a cell). In some embodiments, the method comprises contacting a cell, or the target nucleic acid, with (i) one or more of any one of the first fusion proteins disclosed herein, any one of the second fusion proteins disclosed herein, and any one of the third fusion proteins disclosed herein, or (ii) one or more of any one of the compositions disclosed herein, or (iii) one or more of any one of the complexes disclosed 9 / 169#14463983v1herein, or (iv) one or more of any one of the polynucleotides disclosed herein, or (v) one or more of any one of the vectors disclosed herein, or (vi) one or more of any one of the pharmaceutical compositions disclosed herein, or (vii) one or more of any one of the cells disclosed herein, or (viii) one or more of any one of the kits disclosed herein.
[0034] In some embodiments, the step of contacting results in the binding of the first epitope binding domain of the first fusion protein to the first epitope domain of the third fusion protein, and the binding of the second epitope binding domain of the second fusion protein to the second epitope domain of the third fusion protein (e.g., resulting in a functional base editor complex). In some embodiments, upon the binding of the first and / or second epitope binding domains to their respective epitope domains, the N-terminal portion of the split deaminase and the C-terminal portion of the split deaminase form a functional deaminase. In some embodiments, the DNA binding protein directs the functional deaminase to install one or more mutations at one or more target sites of the target double stranded nucleic acid. In some embodiments, the one or more target sites is in a coding region and / or a non-coding region of the target gene.
[0035] In other embodiments, the methods relate to performing a mutational screen in a target nucleic acid. In some embodiments, the methods comprise contacting the target nucleic acid, or a cell comprising said target nucleic acid, with (i) one or more of any one of the first fusion proteins disclosed herein, any one of the second fusion proteins disclosed herein, and any one of the third fusion proteins disclosed herein, or (ii) one or more of any one of the compositions disclosed herein, or (iii) one or more of any one of the complexes disclosed herein, or (iv) one or more of any one of the polynucleotides disclosed herein or (v) one or more of any one of the vectors disclosed herein, or (vi) one or more of any one of the pharmaceutical compositions disclosed herein, or (vii) one or more of any one of the cells disclosed herein, or (viii) one or more of any one of the kits disclosed herein. In some embodiments, the methods further comprise applying a selection condition to the cells, and determining the mutations caused by step (i) that enable selection. Exemplary selection conditions include antibiotic selection, fluorescent reporter assays, and / or presence or absence of cell surface antigens. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 illustrates an exemplary base editing system as contemplated herein. In some embodiments, the base editing system comprises a first fusion protein comprising a 10 / 169#14463983v1first epitope binding domain, a second fusion protein comprising a second epitope binding domain, and a third fusion protein comprising one or more epitope domains.
[0037] Figure 2 illustrates exemplary configurations of a first fusion protein and a second fusion protein as contemplated herein. In some embodiments, the first fusion protein encodes a C-terminal portion of a split DddA system having the following configurations: a- GCN4 scFv-DddAC-UGI (denoted as CU) or a-GCN4 scFv-UGI-DddAC (denoted as UC). In some embodiments, the second fusion protein encodes a N-terminal portion of the split DddA system having the following configurations: a-GCN4 scFv-DddAN-UGI (denoted as NU) or a-GCN4-scFv-UGI-DddAN (denoted as UN).
[0038] Figure 3A-3D illustrates a split DddA system for wide range base editing. Schematic of base editor system components and configuration (Figure 3A). Constructs for the expression of fusions of N-terminus DddA (N), C-terminus DddA (C), and uracil DNA glycosylase inhibitor (UGI, also denoted U) delivered by lentiviral vectors (Figure 3B). Occurrence of base edits in K562 cells base edited with the system at the VEGFA locus analyzed by amplicon sequencing of a 300 bp region. WT, wildtype (Figure 3C). Co- occurrence ratios of base edits in selected system combinations (Figure 3D). Edited cytidines that met a threshold of 3% editing and above are displayed in Figures 3C and 3D.
[0039] Figure 4 illustrates the tunability of exemplary split DddA systems. Lentiviral constructions of the dCas9 system with various GCN4 repeats tested (Figure 4A). Lentiviral constructs of the split DddA system components delivered in Figure 4C (Figure 4B). Occurrence of base edits in Jurkat cells transduced with the split DddA system directed toward the VEGFA site (Figure 4C). Lentiviral constructs of the split DddA11 system components delivered in Figure 4E (Figure 4D). Occurrence of base edits in Jurkat cells transduced with the split DddA system directed toward the VEGFA site (Figure 4E). Edited cytidines meeting a threshold of 3% editing and above are displayed in the figure.
[0040] Figure 5 illustrates the workflow for analyzing edited loci with PacBio long- read sequencing. PCR1 captures 0.8-3.5 kb genomic loci of interest. A nested PCR (PCR2) uses tagged primers to create dual-barcoded amplicons primed for Kinnex library preparation based on MAS-ISO-seq methods (see Al’Khafaji et al., Nat. Biotech 2023, incorporated herein by reference in its entirety). During Kinnex library prep, amplicons are stitched together to form 18-20 kb Revio-compatible library molecules. After sequencing, bam files are demultiplexed with custom code and analyzed with CRISPRLungo.
[0041] Figure 6A shows Jurkat cells base-edited with the split DddA11 GCN4x4 system at the VEGFA locus with various sgRNAs. Cells were sorted 5 days post-lentiviral 11 / 169#14463983v1transduction and harvested 5 days post-sort. Edited genomic DNA was analyzed by PacBio sequencing of a 2.6 kb region. Edited cytosines that meet a threshold of at least 3% editing are displayed; all cytosines and guanines within the sgRNA binding regions are also included. The numbering of cytosines and guanines in the reference are based on their position in the sequencing amplicon. The frequency of cytosine editing with a 5′ neighboring thymine is included for each guide (right column).
[0042] Figure 6B shows sequences of VEGFA targeting guides used in Figure 6a, including the protospacer sequence and PAM. ‘VEGFA site 4’ sgRNA was previously deployed in pilot experiments using Illumina sequencing (from top to bottom, SEQ ID NOs: 455-459). DEFINITIONS
[0043] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.
[0044] As used herein and in the claims, the singular forms “a,” “an,” and “the” include the singular and the plural reference unless the context clearly indicates otherwise. Thus, for example, a reference to “an agent” includes a single agent and a plurality of such agents.
[0045] An “adeno-associated virus” or “AAV” is a virus which infects humans and some other primate species. The wild-type AAV genome is a single-stranded deoxyribonucleic acid (ssDNA), either positive- or negative-sensed. The genome comprises two inverted terminal repeats (ITRs), one at each end of the DNA strand, and two open reading frames (ORFs): rep and cap between the ITRs. The rep ORF comprises four overlapping genes encoding Rep proteins required for the AAV life cycle. The cap ORF comprises overlapping genes encoding capsid proteins: VP1, VP2 and VP3, which interact together to form the viral capsid. VP1, VP2 and VP3 are translated from one mRNA transcript, which can be spliced in two different manners: either a longer or shorter intron can be excised resulting in the formation of two isoforms of mRNAs: a ~2.3 kb- and a ~2.6 kb-long mRNA isoform. The 12 / 169#14463983v1capsid forms a supramolecular assembly of approximately 60 individual capsid protein subunits into a non-enveloped, T-1 icosahedral lattice capable of protecting the AAV genome. The mature capsid is composed of VP1, VP2, and VP3 (molecular masses of approximately 87, 73, and 62 kDa respectively) in a ratio of about 1:1:10.
[0046] rAAV particles may comprise a nucleic acid vector (e.g., a recombinant genome), which may comprise at a minimum: (a) one or more heterologous nucleic acid regions comprising a sequence encoding a protein or polypeptide of interest (e.g., a split Cas9 or split base) or an RNA of interest (e.g., a gRNA), or one or more nucleic acid regions comprising a sequence encoding a Rep protein; and (b) one or more regions comprising inverted terminal repeat (ITR) sequences (e.g., wild-type ITR sequences or engineered ITR sequences) flanking the one or more nucleic acid regions (e.g., heterologous nucleic acid regions). In some embodiments, the nucleic acid vector is between 4 kb and 5 kb in size (e.g., 4.2 to 4.7 kb in size). In some embodiments, the nucleic acid vector further comprises a region encoding a Rep protein. In some embodiments, the nucleic acid vector is circular. In some embodiments, the nucleic acid vector is single-stranded. In some embodiments, the nucleic acid vector is double-stranded. In some embodiments, a double-stranded nucleic acid vector may be, for example, a self-complimentary vector that contains a region of the nucleic acid vector that is complementary to another region of the nucleic acid vector, initiating the formation of the double-strandedness of the nucleic acid vector.
[0047] In genetics, the “antisense” strand of a segment within double-stranded DNA is the template strand, and which is considered to run in the 3' to 5' orientation. By contrast, the “sense” strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'. In the case of a DNA segment that encodes a protein, the sense strand is the strand of DNA that has the same sequence as the mRNA, which takes the antisense strand as its template during transcription, and eventually undergoes (typically, not always) translation into a protein. The antisense strand is thus responsible for the RNA that is later translated to protein, while the sense strand possesses a nearly identical makeup to that of the mRNA. Note that for each segment of dsDNA, there will possibly be two sets of sense and antisense, depending on which direction one reads (since sense and antisense is relative to perspective). It is ultimately the gene product, or mRNA, that dictates which strand of one segment of dsDNA is referred to as sense or antisense.
[0048] “Base editing” refers to genome editing technology that involves the conversion of a specific nucleic acid base into another at a targeted genomic locus. 13 / 169#14463983v1
[0049] The terms “base editor (BE)” refers to an agent comprising a polypeptide that is capable of making a modification to a base (e.g., A, T, C, G, or U) within a nucleic acid sequence (e.g., DNA or RNA) that converts one base to another (e.g., A to G, A to C, A to T, C to T, C to G, C to A, G to A, G to C, G to T, T to A, T to C, T to G). In some embodiments, the base editor is capable of deaminating a base within a nucleic acid such as a base within a DNA molecule. Such base editors may include a nucleic acid programmable DNA binding protein (napDNAbp) fused to a deaminase. Some base editors include CRISPR-mediated fusion proteins that are utilized in the base editing methods described herein. In some embodiments, the base editor comprises a nuclease-inactive Cas9 (dCas9) fused to a deaminase which binds a nucleic acid in a guide RNA-programmed manner via the formation of an R-loop, but does not cleave the nucleic acid. For example, the dCas9 domain of the fusion protein may include a D10A and a H840A mutation (which renders Cas9 capable of cleaving only one strand of a nucleic acid duplex), as described in PCT / US2016 / 058344, which published as WO 2017 / 070632 on April 27, 2017, and is incorporated herein by reference in its entirety. The DNA cleavage domain of S. pyogenes Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA (the “targeted strand”, or the strand in which editing or deamination occurs), whereas the RuvC1 subdomain cleaves the non-complementary strand containing the PAM sequence (the “non-edited strand”). The RuvC1 mutant D10A generates a nick in the targeted strand, while the HNH mutant H840A generates a nick on the non-edited strand (see Jinek et al., Science, 337:816-821(2012); Qi et al., Cell.28;152(5):1173-83 (2013)).
[0050] In some embodiments, the base editor comprises a DNA binding domain (e.g., a programmable DNA binding domain such as a dCas9 or nCas9) that directs it to a target sequence. In some embodiments, the base editor comprises a base modification domain fused to a programmable DNA binding domain (e.g., a dCas9 or nCas9). The terms “base modifying enzyme” and “base modification domain,” which are used interchangeably herein, refer to an enzyme that can modify a base and convert one base to another (e.g., a deaminase such as a cytidine deaminase). The base modifying enzyme of the base editor may target cytosine (C) bases in a nucleic acid sequence and convert the C to thymine (T) base. In some embodiments, C to T editing is carried out by a deaminase, e.g., a cytidine deaminase. Base editors that can carry out other types of base conversions (e.g., A to G, C to G, etc.) are also contemplated. 14 / 169#14463983v1
[0051] The term “deaminase” or “deaminase domain” refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is a cytidine (or cytosine) deaminase, which catalyzes the hydrolytic deamination of cytidine or cytosine. In some embodiments, the deaminase is configured to carry out the deamination reaction on a double stranded nucleic acid (e.g., double stranded DNA).
[0052] The deaminases provided herein may be from any organism, such as a bacterium. In some embodiments, the deaminase or deaminase domain is a variant of a naturally- occurring deaminase from an organism. In some embodiments, the deaminase or deaminase domain does not occur in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a naturally- occurring deaminase.
[0053] As used herein, the term “DNA binding protein” or “DNA binding protein domain” refers to any protein that localizes to and binds a specific target DNA nucleotide sequence (e.g. a gene locus of a genome). This term embraces RNA-programmable proteins, which associate (e.g. form a complex) with one or more nucleic acid molecules (i.e., which includes, for example, guide RNA in the case of Cas systems) that direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., DNA sequence) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof) associated with the protein. Exemplary RNA-programmable proteins are CRISPR- Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g. engineered or modified), and may include a Cas9 equivalent from any type of CRISPR system (e.g. type II, V, VI), including Cpf1 (a type-V CRISPR-Cas systems) (now known as Cas12a), C2c1 (a type V CRISPR-Cas system), C2c2 (a type VI CRISPR-Cas system), C2c3 (a type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, Cas13d, Cas14, Argonaute, and nCas9. Further Cas-equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299), the contents of which are incorporated herein by reference.
[0054] The term “DNA editing efficiency,” as used herein, refers to the number or proportion of intended base pairs that are edited. For example, if a base editor edits 10% of the base pairs that it is intended to target (e.g., within a cell or within a population of cells), then the base editor can be described as being 10% efficient. Some aspects of editing 15 / 169#14463983v1efficiency embrace the modification (e.g. deamination) of a specific nucleotide within DNA, without generating a large number or percentage of insertions or deletions (i.e., indels). It is generally accepted that editing while generating less than 5% indels (as measured over total target nucleotide substrates) is high editing efficiency. The generation of more than 20% indels is generally accepted as poor or low editing efficiency. Indel formation may be measured by techniques known in the art, including high-throughput screening of sequencing reads.
[0055] As used herein, the terms “upstream” and “downstream” are terms of relativity that define the linear position of at least two elements located in a nucleic acid molecule (whether single or double-stranded) that is orientated in a 5ʹ-to-3ʹ direction. In particular, a first element is upstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 5ʹ to the second element. For example, a SNP is upstream of a Cas9-induced nick site if the SNP is on the 5ʹ side of the nick site. Conversely, a first element is downstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 3ʹ to the second element. For example, a SNP is downstream of a Cas9-induced nick site if the SNP is on the 3 ʹ side of the nick site. The nucleic acid molecule can be a DNA (double or single stranded). RNA (double or single stranded), or a hybrid of DNA and RNA. The analysis is the same for single strand nucleic acid molecule and a double strand molecule since the terms upstream and downstream are in reference to only a single strand of a nucleic acid molecule, except that one needs to select which strand of the double stranded molecule is being considered. Often, the strand of a double stranded DNA which can be used to determine the positional relativity of at least two elements is the “sense” or “coding” strand. In genetics, a “sense” strand is the segment within double-stranded DNA that runs from 5ʹ to 3ʹ, and which is complementary to the antisense strand of DNA, or template strand, which runs from 3ʹ to 5ʹ. Thus, as an example, a SNP base is “downstream” of a promoter sequence in a genomic DNA (which is double- stranded) if the SNP base is on the 3ʹ side of the promoter on the sense or coding strand.
[0056] The term “effective amount,” as used herein, refers to an amount of a biologically active agent that is sufficient to elicit a desired biological response. For example, in some embodiments, an effective amount of a base editor may refer to the amount of the editor that is sufficient to edit a target site nucleotide sequence, e.g., a genome. In some embodiments, an effective amount of a base editor provided herein, e.g., of a fusion protein comprising a dCas9 domain and a guide RNA may refer to the amount of the fusion protein that is sufficient to induce editing of a target site specifically bound and edited by the fusion protein. 16 / 169#14463983v1As will be appreciated by the skilled artisan, the effective amount of an agent, e.g., a fusion protein, a nuclease, a hybrid protein, a protein dimer, a complex of a protein (or protein dimer) and a polynucleotide, or a polynucleotide, may vary depending on various factors as, for example, on the desired biological response, e.g., on the specific allele, genome, or target site to be edited, on the cell or tissue being targeted, and on the agent being used.
[0057] The term “fusion protein,” as used herein, refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C- terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein,” respectively. A protein may comprise different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the binding of the protein to a target site) and a nucleic acid cleavage domain or a catalytic domain of a nucleic-acid editing protein. Another example includes a Cas9 or equivalent thereof fused to an adenosine deaminase. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.
[0058] Two proteins or protein domains are considered to be “fused” when a peptide bond is formed linking the two proteins or two protein domains. In some embodiments, a linker (e.g., a peptide linker) is present between the two proteins or two protein domains. The term “linker,” as used herein, refers to a chemical group or a molecule linking two molecules or moieties, e.g., two domains of a fusion protein, such as, for example, a nuclease-inactive Cas9 domain and a nucleic acid editing domain (e.g., a deaminase domain). Typically, the linker is positioned between, or flanked by, two groups, molecules, or other moieties and connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. 17 / 169#14463983v1
[0059] The term “guide nucleic acid” or “guide sequence” refers the one or more nucleic acid molecules which associate with and direct or otherwise program a napDNAbp protein to localize to a specific target nucleotide sequence (e.g., a gene locus of a genome) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof) associated with the protein, thereby causing the napDNAbp protein to bind to the nucleotide sequence at the specific target site. A non-limiting example is a guide RNA of a Cas protein of a CRISPR-Cas genome editing system. Chemically, guide nucleic acids can be all RNA, all DNA, or a chimeric of RNA and DNA. The guide nucleic acids may also include nucleotide analogs. Guide nucleic acids can be expressed as transcription products or can be synthesized.
[0060] As used herein, a “spacer sequence” is the sequence of the guide RNA (~20 nts in length) which has the same sequence (with the exception of uridine bases in place of thymine bases) as the protospacer of the PAM strand of the target (DNA) sequence, and which is complementary to the target strand (or non-PAM strand) of the target sequence.
[0061] As used herein, the “target sequence” refers to the ~20 nucleotides in the target DNA sequence that have complementarity to the protospacer sequence in the PAM strand. The target sequence is the sequence that anneals to or is targeted by the spacer sequence of the guide RNA. The spacer sequence of the guide RNA and the protospacer have the same sequence (except the spacer sequence is RNA, and the protospacer is DNA).
[0062] As used herein, the terms “guide RNA core,” “guide RNA scaffold sequence,” and “backbone sequence,” which are used interchangeably, refer to the region (or sequence) within the gRNA that is responsible for Cas9 binding. It does not include the 20 bp spacer sequence that is used to guide Cas9 to target DNA. This region also known as the crRNA / tracrRNA. The guide RNA backbone sequence is separate from the guide sequence, or spacer, region of the guide RNA, which has complementarity to a protospacer of a nucleic acid molecule.
[0063] As used herein, the term “protospacer” refers to the sequence (e.g., a ~20 bp sequence) in DNA adjacent to the PAM (protospacer adjacent motif) sequence which shares the same sequence as the spacer sequence of the guide RNA, and which is complementary to the target sequence of the non-PAM strand. The spacer sequence of the guide RNA anneals to the target sequence located on the non-PAM strand. In order for Cas9 to function it also requires a specific protospacer adjacent motif (PAM) that varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease, derived from S. pyogenes, recognizes a PAM sequence of NGG that is found directly downstream of the 18 / 169#14463983v1protospacer sequence in the genomic DNA, on the non-target strand. The skilled person will appreciate that the literature in the state of the art sometimes refers to the “protospacer” as the ~20-nt target-specific guide sequence on the guide RNA itself, rather than referring to it as a “spacer” (and that the protospacer (DNA) and the spacer (RNA) have the same sequence). Thus, the term “protospacer” as used herein may be used interchangeably with the term “spacer.” The context of the description surrounding the appearance of either “protospacer” or “spacer” will help inform the reader as to whether the term is refence to the gRNA or the DNA sequence. Both usages of these terms are acceptable since the state of the art uses both terms in each of these ways.
[0064] A “protospacer adjacent motif” (PAM) is typically a sequence of nucleotides located adjacent to (e.g., within 10, 9, 8, 7, 6, 5, 4, 3, 3, or 1 nucleotide(s) of a target sequence). A PAM sequence is “immediately adjacent to” a target sequence if the PAM sequence is contiguous with the target sequence (that is, if there are no nucleotides located between the PAM sequence and the target sequence). In some embodiments, a PAM sequence is a wild-type PAM sequence. Examples of PAM sequences include, without limitation, NGG, NGR, NNGRR(T / N), NNNNGATT, NNAGAAW, NGGAG, NAAAAC, AWG, and CC. In some embodiments, a PAM sequence is obtained from Streptococcus pyogenes (e.g., NGG or NGR). In some embodiments, a PAM sequence is obtained from Staphylococcus aureus (e.g., NNGRR(T / N)). In some embodiments, a PAM sequence is obtained from Neisseria meningitidis (e.g., NNNNGATT,). In some embodiments, a PAM sequence is obtained from Streptococcus thermophilus (e.g., NNAGAAW or NGGAG) . In some embodiments, a PAM sequence is obtained from Treponema denticola (e.g., NAAAAC). In some embodiments, a PAM sequence is obtained from Escherichia coli (e.g., AWG). In some embodiments, a PAM sequence is obtained from Pseudomonas aeruginosa (e.g., CC). Other PAM sequences are contemplated. A PAM sequence is typically located downstream (i.e., 3′) from the target sequence, although in some embodiments a PAM sequence may be located upstream (i.e., 5′) from the target sequence.
[0065] The term “host cell,” as used herein, refers to a cell that can host, replicate, and transfer a phage vector useful for a continuous evolution process as provided herein. In embodiments where the vector is a viral vector, a suitable host cell is a cell that may be infected by the viral vector, can replicate it, and can package it into viral particles that can infect fresh host cells. A cell can host a viral vector if it supports expression of genes of viral vector, replication of the viral genome, and / or the generation of viral particles. One criterion to determine whether a cell is a suitable host cell for a given viral vector is to determine 19 / 169#14463983v1whether the cell can support the viral life cycle of a wild-type viral genome that the viral vector is derived from. In some embodiments, the viral vector is a phage and the host cell is a bacterial cell. In some embodiments, the host cell is an E. coli cell. Suitable E. coli host strains will be apparent to those of skill in the art, and include, but are not limited to, New England Biolabs (NEB) Turbo, Top10F’, DH12S, ER2738, ER2267, and XL1-Blue MRF’. These strain names are art recognized and the genotype of these strains has been well characterized. It should be understood that the above strains are exemplary only and that the invention is not limited in this respect. The term “fresh,” as used herein interchangeably with the terms “non-infected” or “uninfected” in the context of host cells, refers to a host cell that has not been infected by a viral vector comprising a gene of interest as used in a continuous evolution process provided herein. A fresh host cell can, however, have been infected by a viral vector unrelated to the vector to be evolved or by a vector of the same or a similar type but not carrying the gene of interest.
[0066] In some embodiments, the host cell is a prokaryotic cell, for example, a bacterial cell. In some embodiments, the host cell is an E. coli cell. In some embodiments, the host cell is a eukaryotic cell, for example, a yeast cell, a plant cell, an insect cell, or a mammalian cell. In some embodiments, the cell is a human cell. The type of host cell will, of course, depend on the viral vector employed, and suitable host cell / viral vector combinations will be readily apparent to those of skill in the art.
[0067] The term “mutation,” as used herein, refers to a substitution of a residue within a sequence, e.g., a nucleic acid or amino acid sequence, with another residue; a deletion or insertion of one or more residues within a sequence; or a substitution of a residue within a sequence of a genome in a subject to be corrected. Mutations are typically described herein by identifying the original residue followed by the position of the residue within the sequence and by the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art, and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). Mutations can include a variety of categories, such as single base polymorphisms, microduplication regions, indel, and inversions, and is not meant to be limiting in any way. Mutations can include “loss-of- function” mutations which are mutations that reduce or abolish a protein activity. Most loss- of-function mutations are recessive, because in a heterozygote the second chromosome copy carries an unmutated version of the gene coding for a fully functional protein whose presence compensates for the effect of the mutation. There are some exceptions where a loss-of- 20 / 169#14463983v1function mutation is dominant, one example being haploinsufficiency, where the organism is unable to tolerate the approximately 50% reduction in protein activity suffered by the heterozygote. This is the explanation for a few genetic diseases in humans, including Marfan syndrome, which results from a mutation in the gene for the connective tissue protein called fibrillin. Mutations also embrace “gain-of-function” mutations, which is one which confers an abnormal activity on a protein or cell that is otherwise not present in a normal condition. Many gain-of-function mutations are in regulatory sequences rather than in coding regions and can therefore have a number of consequences. Because of their nature, gain-of-function mutations are usually dominant. Many loss-of-function mutations are recessive, such as autosomal recessive.
[0068] The term “napDNAbp” which stand for “nucleic acid programmable DNA binding protein” refers to any protein that may associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., which may broadly be referred to as a “napDNAbp- programming nucleic acid molecule” and includes, for example, guide RNA in the case of Cas systems) which direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a gene locus of a genome) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof) associated with the protein, thereby causing the protein to bind to the nucleotide sequence at the specific target site. This term napDNAbp embraces CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and may include a Cas9 equivalent from any type of CRISPR system (e.g., type II, V, VI).
[0069] A “uracil glycosylase inhibitor (UGI)” refers to a protein that inhibits the activity of uracil-DNA glycosylase. Suitable UGI proteins for use in accordance with the present disclosure include, for example, those published in Wang et al., J. Biol. Chem.264:1163- 1171(1989); Lundquist et al., J. Biol. Chem.272:21408-21419(1997); Ravishankar et al., Nucleic Acids Res.26:4880-4887(1998); and Putnam et al., J. Mol. Biol.287:331-346(1999), each of which is incorporated herein by reference. Non-limiting, exemplary proteins that may be used as a UGI of the present disclosure and their respective sequences are provided below. In some embodiments, the UGI is a variant of a naturally-occurring deaminase from an organism, and the variants do not occur in nature. For example, in some embodiments, the UGI is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a naturally-occurring UGI 21 / 169#14463983v1from an organism or any UGIs provided herein. In some embodiments, the UGI comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 30%, no more than 25%, no more than 20%, no more than 15%, no more than 10%, no more than 5%, no more than 1% longer or shorter) than any of the UGIs provided herein. In some embodiments, the UGI comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 20 amino acids, no more than 15 amino acids, no more than 10 amino acids, no more than 5 amino acids, no more than 2 amino acids longer or shorter) than any of the UGIs provided herein.
[0070] A “nuclear localization signal” or “NLS” refers to as an amino acid sequence that “tags” a protein for import into the cell nucleus by nuclear transport. Typically, this signal consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface. One or more NLS may be added to the N- or C-terminus of a protein, or internally (e.g., between two protein domains).
[0071] Nucleic acids of the present disclosure may include one or more genetic elements. A “genetic element” refers to a particular nucleotide sequence that has a role in nucleic acid expression (e.g., promoter, enhancer, terminator) or encodes a discrete product of an engineered nucleic acid (e.g., a nucleotide sequence encoding a guide RNA, a protein and / or an RNA interference molecule).
[0072] A “promoter” refers to a control region of a nucleic acid sequence at which initiation and rate of transcription of the remainder of a nucleic acid sequence are controlled. A promoter may also contain sub-regions at which regulatory proteins and molecules may bind, such as RNA polymerase and other transcription factors. Promoters may be constitutive, inducible, activatable, repressible, tissue-specific, or any combination thereof. A promoter drives expression or drives transcription of the nucleic acid sequence that it regulates. Herein, a promoter is considered to be “operably linked” when it is in a correct functional location and orientation in relation to a nucleic acid sequence it regulates to control (“drive”) transcriptional initiation and / or expression of that sequence.
[0073] A promoter may be one naturally associated with a gene or sequence, as may be obtained by isolating the 5′ non-coding sequences located upstream of the coding segment of a given gene or sequence. Such a promoter is referred to as an “endogenous promoter.” In some embodiments, a coding nucleic acid sequence may be positioned under the control of a recombinant or heterologous promoter, which refers to a promoter that is not normally associated with the encoded sequence in its natural environment. Such promoters may include promoters of other genes; promoters isolated from any other cell; and synthetic 22 / 169#14463983v1promoters or enhancers that are not “naturally occurring” such as, for example, those that contain different elements of different transcriptional regulatory regions and / or mutations that alter expression through methods of genetic engineering that are known in the art. In addition to producing nucleic acid sequences of promoters and enhancers synthetically, sequences may be produced using recombinant cloning and / or nucleic acid amplification technology, including polymerase chain reaction (PCR).
[0074] In some embodiments, promoters used in accordance with the present disclosure are “inducible promoters,” which are promoters that are characterized by regulating (e.g., initiating or activating) transcriptional activity when in the presence of, influenced by or contacted by an inducer signal. An inducer signal may be endogenous or a normally exogenous condition (e.g., light), compound (e.g., chemical or non-chemical compound) or protein that contacts an inducible promoter in such a way as to be active in regulating transcriptional activity from the inducible promoter. Thus, a “signal that regulates transcription” of a nucleic acid refers to an inducer signal that acts on an inducible promoter. A signal that regulates transcription may activate or inactivate transcription, depending on the regulatory system used. Activation of transcription may involve directly acting on a promoter to drive transcription or indirectly acting on a promoter by inactivation a repressor that is preventing the promoter from driving transcription. Conversely, deactivation of transcription may involve directly acting on a promoter to prevent transcription or indirectly acting on a promoter by activating a repressor that then acts on the promoter.
[0075] In genetics, a “sense” strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'. In the case of a DNA segment that encodes a protein, the sense strand is the strand of DNA that has the same sequence as the mRNA, which takes the antisense strand as its template during transcription, and eventually undergoes (typically, not always) translation into a protein. The antisense strand is thus responsible for the RNA that is later translated to protein, while the sense strand possesses a nearly identical makeup to that of the mRNA. Note that for each segment of dsDNA, there will possibly be two sets of sense and antisense, depending on which direction one reads (since sense and antisense is relative to perspective). It is ultimately the gene product, or mRNA, that dictates which strand of one segment of dsDNA is referred to as sense or antisense.
[0076] The term “subject,” as used herein, refers to an individual organism, for example, an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human 23 / 169#14463983v1primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, a goat, a cow, a cat, or a dog. In some embodiments, the subject is a vertebrate, an amphibian, a reptile, a fish, an insect, a fly, or a nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject may be of either sex and at any stage of development.
[0077] A “subject in need thereof” refers to an individual who has a disease, a sign and / or symptom of a disease, or a predisposition toward a disease, with the purpose to cure, heal, alleviate, relieve, alter, remedy, ameliorate, improve, or affect the disease, the symptom of the disease, or the predisposition toward the disease. In some embodiments, the subject is a mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is human. In some embodiments, the mammal is a rodent. In some embodiments, the rodent is a mouse. In some embodiments, the rodent is a rat. In some embodiments, the mammal is a companion animal. A “companion animal” refers to pets and other domestic animals. Non-limiting examples of companion animals include dogs and cats; livestock, such as horses, cattle, pigs, sheep, goats, and chickens; and other animals, such as mice, rats, guinea pigs, and hamsters.
[0078] The term “target site” refers to a sequence within a nucleic acid molecule that is edited by a base editor (BE) or base editor disclosed herein. The term “target site,” in the context of a single strand, also can refer to the “target strand” which anneals or binds to the spacer sequence of the guide RNA. The target site can refer, in certain embodiments, to a segment of double-stranded DNA that includes the protospacer (i.e., the strand of the target site that has the same nucleotide sequence as the spacer sequence of the guide RNA) on the PAM-strand (or non-target strand) and target strand, which is complementary to the protospacer and the spacer alike, and which anneals to the spacer of the guide RNA, thereby targeting or programming a Cas9 base editor to target the target site.
[0079] A “transcriptional terminator” is a nucleic acid sequence that causes transcription to stop. A transcriptional terminator may be unidirectional or bidirectional. It is comprised of a DNA sequence involved in specific termination of an RNA transcript by an RNA polymerase. A transcriptional terminator sequence prevents transcriptional activation of downstream nucleic acid sequences by upstream promoters. A transcriptional terminator may be necessary in vivo to achieve desirable expression levels or to avoid transcription of certain sequences. A transcriptional terminator is considered to be “operably linked to” a nucleotide sequence when it is able to terminate the transcription of the sequence it is linked to. 24 / 169#14463983v1
[0080] The most commonly used type of terminator is a forward terminator. When placed downstream of a nucleic acid sequence that is usually transcribed, a forward transcriptional terminator will cause transcription to abort. In some embodiments, bidirectional transcriptional terminators are provided, which usually cause transcription to terminate on both the forward and reverse strand. In some embodiments, reverse transcriptional terminators are provided, which usually terminate transcription on the reverse strand only.
[0081] In prokaryotic systems, terminators usually fall into two categories (1) rho- independent terminators and (2) rho-dependent terminators. Rho-independent terminators are generally composed of palindromic sequence that forms a stem loop rich in G-C base pairs followed by several T bases. Without wishing to be bound by theory, the conventional model of transcriptional termination is that the stem loop causes RNA polymerase to pause, and transcription of the poly-A tail causes the RNA:DNA duplex to unwind and dissociate from RNA polymerase.
[0082] In eukaryotic systems, the terminator region may comprise specific DNA sequences that permit site-specific cleavage of the new transcript so as to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch of about 200 A residues (polyA) to the 3′ end of the transcript. RNA molecules modified with this polyA tail appear to more stable and are translated more efficiently. Thus, in some embodiments involving eukaryotes, a terminator may comprise a signal for the cleavage of the RNA. In some embodiments, the terminator signal promotes polyadenylation of the message. The terminator and / or polyadenylation site elements may serve to enhance output nucleic acid levels and / or to minimize read through between nucleic acids.
[0083] Terminators for use in accordance with the present disclosure include any terminator of transcription described herein or known to one of ordinary skill in the art. Examples of terminators include, without limitation, the termination sequences of genes such as, for example, the bovine growth hormone terminator, and viral termination sequences, such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA, and arcA terminator. In some embodiments, the termination signal may be a sequence that cannot be transcribed or translated, such as those resulting from a sequence truncation.
[0084] A “Woodchuck Hepatitis Virus (WHP) Posttranscriptional Regulatory Element (WPRE)” is a DNA sequence that, when transcribed creates a tertiary structure enhancing expression. Commonly used in molecular biology to increase expression of genes delivered 25 / 169#14463983v1by viral vectors. WPRE is a tripartite regulatory element with gamma, alpha, and beta components.
[0085] The full WPRE sequence is 609 bp long:
[0086] GCTTATCGATAATCAACCTCTGGATTACAAAATTTGTGAAAGATTGAC TGGTATTCTTAACTATGTTGCTCCTTTTACGCTATGTGGATACGCTGCTTTAATGC CTTTGTATCATGCTATTGCTTCCCGTATGGCTTTCATTTTCTCCTCCTTGTATAAAT CCTGGTTGCTGTCTCTTTATGAGGAGTTGTGGCCCGTTGTCAGGCAACGTGGCGT GGTGTGCACTGTGTTTGCTGACGCAACCCCCACTGGTTGGGGCATTGCCACCACC TGTCAGCTCCTTTCCGGGACTTTCGCTTTCCCCCTCCCTATTGCCACGGCGGAACT CATCGCCGCCTGCCTTGCCCGCTGCTGGACAGGGGCTCGGCTGTTGGGCACTGAC AATTCCGTGGTGTTGTCGGGGAAATCATCGTCCTTTCCTTGGCTGCTCGCCTATGT TGCCACCTGGATTCTGCGCGGGACGTCCTTCTGCTACGTCCCTTCGGCCCTCAATC CAGCGGACCTTCCTTCCCGCGGCCTGCTGCCGGCTCTGCGGCCTCTTCCGCGTCTT CGCCTTCGCCCTCAGACGAGTCGGATCTCCCTTTGGGCCGCCTCCCCGCATCGAT ACCG (SEQ ID NO: 408).
[0087] The terms “nucleic acid” and “polynucleotide,” as used herein, refer to a compound comprising a base and an acidic moiety, e.g., a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides are linear molecules, in which adjacent nucleotides are linked to each other via a phosphodiester linkage. In some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g. nucleotides and / or nucleosides). In some embodiments, “nucleic acid” refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms “oligonucleotide” and “polynucleotide” can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, “nucleic acid” encompasses RNA as well as single and / or double-stranded DNA. Nucleic acids may be naturally occurring, for example, in the context of a genome, a transcript, an mRNA, tRNA, rRNA, siRNA, snRNA, a plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule may be a non-naturally occurring molecule, e.g., a recombinant DNA or RNA, an artificial chromosome, an engineered genome (e.g., an engineered viral vector), an engineered vector, or fragment thereof, or a synthetic DNA, RNA, or DNA / RNA hybrid, optionally including non-naturally occurring nucleotides or nucleosides. Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. 26 / 169#14463983v1Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can comprise nucleoside analogs such as analogs having chemically modified bases or sugars, and backbone modifications. A nucleic acid sequence is presented in the 5′ to 3′ direction unless otherwise indicated. In some embodiments, a nucleic acid is or comprises natural nucleosides (e.g. adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, 5-methylcytidine, 2-aminoadenosine, C5- bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8- oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2′-fluororibose, ribose, 2′-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5′-N-phosphoramidite linkages).
[0088] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein, and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to a protein, peptide, or polypeptide of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, etc. A protein, peptide, or polypeptide may also be a single molecule or may be a multi-molecular complex. A protein, peptide, or polypeptide may be just a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof. The term “fusion protein” as used herein refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy- terminal (C-terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy- terminal fusion protein,” respectively. A protein may comprise different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the binding of the protein to a target site) and a nucleic acid cleavage domain or a catalytic 27 / 169#14463983v1domain of a nucleic-acid editing protein. In some embodiments, a protein is in a complex with, or is in association with, a nucleic acid, e.g., RNA or DNA. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), which are incorporated herein by reference.
[0089] The term “recombinant,” as used herein, in the context of proteins or nucleic acids refers to proteins or nucleic acids that do not occur in nature, but are the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that comprises at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations as compared to any naturally occurring sequence. The fusion proteins (e.g., base editors) described herein are made by recombinant technology. Recombinant technology is familiar to those skilled in the art.
[0090] The term “pharmaceutically-acceptable carrier” means a pharmaceutically- acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc magnesium, calcium or zinc stearate, or steric acid), or solvent encapsulating material, involved in carrying or transporting the compound from one site (e.g., the delivery site) of the body, to another site (e.g., organ, tissue or portion of the body). A pharmaceutically acceptable carrier is “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the tissue of the subject (e.g., physiologically compatible, sterile, physiologic pH, etc.).
[0091] A “therapeutically effective amount,” as used herein, refers to the amount of each therapeutic agent (e.g., base editor, rAAV) described in the present disclosure required to confer therapeutic effect on the subject, either alone or in combination with one or more other therapeutic agents. Effective amounts vary, as recognized by those skilled in the art, depending on the particular condition being treated, the severity of the condition, the individual subject parameters including age, physical condition, size, gender, and weight, the duration of the treatment, the nature of concurrent therapy (if any), the specific route of administration and like factors within the knowledge and expertise of the health practitioner. These factors are well known to those of ordinary skill in the art and can be addressed with 28 / 169#14463983v1no more than routine experimentation. It is generally preferred that a maximum dose of the individual components or combinations thereof be used, that is, the highest safe dose according to sound medical judgment. It will be understood by those of ordinary skill in the art, however, that a subject may insist upon a lower dose or tolerable dose for medical reasons, psychological reasons or for virtually any other reasons. Empirical considerations, such as the half-life, generally will contribute to the determination of the dosage. For example, therapeutic agents that are compatible with the human immune system, such as polypeptides comprising regions from humanized antibodies or fully human antibodies, may be used to prolong half-life of the polypeptide and to prevent the polypeptide being attacked by the host's immune system.
[0092] The terms “treatment,” “treat,” and “treating,” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. As used herein, the terms “treatment,” “treat,” and “treating” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay onset of a symptom or inhibit onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, for example, to prevent or delay their recurrence.
[0093] As used herein, the term “variant” refers to a protein having characteristics that deviate from what occurs in nature that retains at least one functional i.e. binding, interaction, or enzymatic ability and / or therapeutic property thereof. A “variant” is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, at least about 99.9% or 100% identical to the wild type protein. For instance, a variant of Cas9 may comprise a Cas9 that has one or more changes in amino acid residues as compared to a wild type Cas9 amino acid sequence. As another example, a variant of a deaminase may comprise a deaminase that has one or more changes in amino acid residues as compared to a wild type deaminase amino acid sequence, e.g. following ancestral sequence 29 / 169#14463983v1reconstruction of the deaminase. These changes include chemical modifications, including substitutions of different amino acid residues truncations, covalent additions (e.g. of a tag), and any other mutations. The term also encompasses circular permutants, mutants, truncations, or domains of a reference sequence, and which display the same or substantially the same functional activity or activities as the reference sequence. This term also embraces fragments of a wild type protein.
[0094] The level or degree of which the property is retained may be reduced relative to the wild type protein but is typically the same or similar in kind. Generally, variants are overall very similar, and in many regions, identical to the amino acid sequence of the protein described herein. A skilled artisan will appreciate how to make and use variants that maintain all, or at least some, of a functional ability or property.
[0095] The variant proteins may comprise, or alternatively consist of, an amino acid sequence which is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%, identical to, for example, the amino acid sequence of a wild-type protein, or any protein provided herein.
[0096] By a polypeptide having an amino acid sequence at least, for example, 95% “identical” to a query amino acid sequence, it is intended that the amino acid sequence of the subject polypeptide is identical to the query sequence except that the subject polypeptide sequence may include up to five amino acid alterations per each 100 amino acids of the query amino acid sequence. In other words, to obtain a polypeptide having an amino acid sequence at least 95% identical to a query amino acid sequence, up to 5% of the amino acid residues in the subject sequence may be inserted, deleted, or substituted with another amino acid. These alterations of the reference sequence may occur at the amino- or carboxy-terminal positions of the reference amino acid sequence or anywhere between those terminal positions, interspersed either individually among residues in the reference sequence or in one or more contiguous groups within the reference sequence.
[0097] As a practical matter, whether any particular polypeptide is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to, for instance, the amino acid sequence of a protein such as a PrP protein, can be determined conventionally using known computer programs. A preferred method for determining the best overall match between a query sequence (a sequence of the present invention) and a subject sequence, also referred to as a global sequence alignment, can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci.6:237-245 (1990)). In a sequence alignment the query and subject sequences are either both nucleotide sequences or both 30 / 169#14463983v1amino acid sequences. The result of said global sequence alignment is expressed as percent identity. Preferred parameters used in a FASTDB amino acid alignment are: Matrix=PAM 0, k-tuple=2, Mismatch Penalty=1, Joining Penalty=20, Randomization Group Length=0, Cutoff Score=1, Window Size=sequence length, Gap Penalty=5, Gap Size Penalty=0.05, Window Size=500 or the length of the subject amino acid sequence, whichever is shorter.
[0098] If the subject sequence is shorter than the query sequence due to N- or C-terminal deletions, not because of internal deletions, a manual correction must be made to the results. This is because the FASTDB program does not account for N- and C-terminal truncations of the subject sequence when calculating global percent identity. For subject sequences truncated at the N- and C-termini, relative to the query sequence, the percent identity is corrected by calculating the number of residues of the query sequence that are N- and C- terminal of the subject sequence, which are not matched / aligned with a corresponding subject residue, as a percent of the total bases of the query sequence. Whether a residue is matched / aligned is determined by results of the FASTDB sequence alignment. This percentage is then subtracted from the percent identity, calculated by the above FASTDB program using the specified parameters, to arrive at a final percent identity score. This final percent identity score is what is used for the purposes of the present invention. Only residues to the N- and C-termini of the subject sequence, which are not matched / aligned with the query sequence, are considered for the purposes of manually adjusting the percent identity score. That is, only query residue positions outside the farthest N- and C-terminal residues of the subject sequence.
[0099] The term “vector,” as used herein, refers to a nucleic acid that can be modified to encode a gene of interest and that is able to enter into a host cell and replicate within the host cell, and then transfer a replicated form of the vector into another host cell. Exemplary suitable vectors include viral vectors, such as AAV vectors or bacteriophages and filamentous phage, and conjugative plasmids. Additional suitable vectors will be apparent to those of skill in the art based on the instant disclosure.
[0100] As used herein, the term “wild type” is a term of the art understood by skilled persons and means the typical form of an organism, strain, gene or characteristic as it occurs in nature as distinguished from mutant or variant forms. DETAILED DESCRIPTION
[0101] The present disclosure provides base editor fusion proteins, complexes, and systems that are capable of introducing multiple nucleotide edits across extended genomic 31 / 169#14463983v1windows (e.g., between about 4 base pairs (bps) and about 500 bps). Such systems may be useful, for example, for mutating one or more nucleotides within a target nucleic acid. In certain embodiments, the gene editing system comprises: (a) a first fusion protein including a C-terminal portion of a split deaminase operably linked to a first epitope-binding domain; (b) a second fusion protein including an N-terminal portion of a split deaminase operably linked to a second epitope-binding domain; and (c) a third fusion protein including a DNA-binding protein domain and one or more epitope domains capable of binding the first and second fusion proteins. When assembled, the interaction of the epitope domains with the first and second epitope-binding domains reconstitutes an active deaminase domain, thereby generating a functional base editor. The disclosed base editor fusion proteins, complexes, and systems can be applied, in some embodiments, to perform mutational screening of a gene of interest. For example, the editors described herein enable cytidine-to-thymine conversions across a wide genomic window of about 4–500 bps (e.g., about 130 bps), representing a significant expansion over the narrow ~4 bp editing windows characteristic of traditional base editors. Additional aspects of the disclosure relate to methods of use, including methods for mutating one or more nucleotides in a target nucleic acid and methods for conducting mutational screens. Further embodiments provide compositions, polynucleotides, vectors, pharmaceutical formulations, cells, kits, and systems comprising or employing the disclosed base editor fusion proteins and complexes.
[0102] In certain aspects, the present disclosure provides base editor fusion proteins, complexes, and systems useful, for example, for performing base editor screens (e.g., to catalogue disease-associated mutations, profile protein-drug interactions, and / or to map protein sequence-function landscapes). In some embodiments, the base editors disclosed herein enable cytidine-to-thymine edits across a genomic editing window of about 130 bp or greater (e.g., ~450 bps), relative to a 4 bp window offered by traditional base editors.
[0103] In various embodiments, the editing window is from 4–10 bps, 5–20 bps, 6–30 bps, 7–40 bps, 8–50 bps, 9–60 bps, 10–70 bps, 11–80 bps, 12–90 bps, 13–100 bps, 14–110 bps, 15–120 bps, 16–130 bps, 17–140 bps, 18–150 bps, 19–160 bps, 20–170 bps, 21–180 bps, 22–190 bps, 23–200 bps, 24–210 bps, 25–220 bps, 26–230 bps, 27–240 bps, 28–250 bps, 29–260 bps, 30–270 bps, 31–280 bps, 32–290 bps, 33–300 bps, 34–310 bps, 35–320 bps, 36–330 bps, 37–340 bps, 38–350 bps, 39–360 bps, 40–370 bps, 41–380 bps, 42–390 bps, 43–400 bps, 44–410 bps, 45–420 bps, 46–430 bps, 47–440 bps, 48–450 bps, 49–460 bps, or 50–500 bps. 32 / 169#14463983v1
[0104] Other aspects of the disclosure relate to methods, for example, methods of mutating one or more nucleotides in a target nucleic acid, and methods of performing mutational screens. The disclosure also provides compositions, polynucleotides, vectors, pharmaceutical compositions, cells, kits, and systems comprising the base editor fusion proteins and complexes contemplated herein.
[0105] Non-coding regulator regions, such as transcription factors (TFs) and cis- regulator elements (CREs) include promoters, enhancers, and insulators. These elements have a pivotal role in orchestrating gene expression that ultimately dictate cellular identity, function, and state. Additionally, most genetic variation that underlies common disease is found in non-coding regulatory regions. However, the complexity and vastness of non- coding space poses a significant challenge to understanding these processes. For example, enhancers can be as large as 1 kilobases, contain multiple transcription factor (TF) motifs, function through co-binding events, and work in combination with other enhancers to control gene expression. Prior studies in the art have generally relied on massively parallel reporter assays for saturation mutagenesis, where an enhancer sequence is systematically altered and placed in the cells with a lentivirus. While these experiments can be informative, they lack the endogenous locus and surrounding elements that can contribute to control. More recently, the field has relied on Cas9-directed base editors to target endogenous loci for perturbations. However, saturation mutagenesis of the target region with Cas9-directed base editors is restricted by a protospacer adjacent motif (PAM) and narrow editing windows (4 bps). Further, the number of potential single base edits to test in non-coding regions is incredibly high.
[0106] Without being bound by theory, the inventors of the present disclosure have developed a new base editor system with wider base editing (WBE) windows. The newly discovered base editor systems enable greater scaled screening and editing saturation. For example, the new base editor system enables a vast array of cytidine-to-thymine edits across a genomic window of about 180 bp, relative to a ~4 bp window with conventional base editors. Additionally, the new base editor system does not edit in the guide binding region, allowing for sustained mutagenesis, which is a distinct advantage over conventional systems in the art.
[0107] An exemplary base editing system is shown in Figure 1. In some embodiments, the base editing system comprises a deaminase split into an N-terminal portion and a C-terminal portion. In some embodiments, the base editing system comprises a first fusion protein comprising the C-terminal portion of the split deaminase and a second fusion 33 / 169#14463983v1protein comprising the N-terminal portion of the split deaminase. In some embodiments, the first fusion protein further comprises a first epitope binding domain, and the second fusion protein further comprises a second epitope binding domain.
[0108] Additionally, the base editing system comprises a third fusion protein comprising one or more epitope domains. In some embodiments, the third fusion protein comprises a DNA binding protein domain, such as a dead Cas9 (dCas9), which in association with a gRNA, mediates binding of the third fusion protein to a target gene.
[0109] In some embodiments, a base editing system comprises a complex comprising the first, second, and third fusion proteins. For example, in some embodiments, the first epitope binding domain of the first fusion protein is configured to bind to one or more epitope domains of the third fusion protein, and the second epitope binding domain of the second fusion protein is configured to bind to one or more epitope domains of the third fusion protein. In some embodiments, upon binding of the second and third fusion proteins to the first fusion protein, the N-terminal portion of the split deaminase and the C-terminal portion of the split deaminase form a functional deaminase. In some embodiments, the functional deaminase is configured to install cytidine-to-thymine base edits on double-stranded DNA on both sides of the third fusion protein of the complex.
[0110] As discussed above, in some embodiments, the base editor systems disclosed herein comprise a first fusion protein that when placed adjacent a second fusion protein form a functional deaminase that can instill cytidine-to-thymine edits in double-stranded DNA. Any suitable deaminase known to the skilled artisan that can edit double-stranded DNA may be “split” and used in any of the base editor systems disclosed herein. In some cases, the deaminase may be derived from a larger enzyme comprising a cytidine deaminase domain (e.g., the deaminase is truncated). For example, DddAis a cytidine deaminase derived from the type VI secretion system (T6SS)-associated deaminases (e.g., the T6SS-associated deaminase comprises the DddA deaminase). Thus, in some embodiments, a first fusion protein and a second fusion protein may comprise a C-terminal and an N-terminal portion, respectively, of a split T6SS-associated deaminase (SEQ ID NO: 1) or a split DddA deaminase.
[0111] Functional full-length enzymes contemplated herein include, but are not limited to, the amino acid sequences listed in SEQ ID NOs: 1-181 (Table 1). In some embodiments, the full-length enzymes comprise an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or 100% identical to the amino acid sequences of any one of SEQ ID NOs: 1-181. In 34 / 169#14463983v1some embodiments, the full-length enzymes comprise an amino acid sequence that is identical to the amino acid sequence of any one of SEQ ID NOs: 1-181.
[0112] Functional truncated deaminase contemplated herein include, but are not limited to, the amino acid sequences listed in SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71- 73,90,95,108,142,145,151,159-160,172,174,176,178-181,194-195,198-206,209-234,236- 243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345-351,354- 364,366,368,370,375-380,382-386 (Table 2). In some embodiments, the full-length enzymes comprise an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acid sequences of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71- 73,90,95,108,142,145,151,159-160,172,174,176,178-181,194-195,198-206,209-234,236- 243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345-351,354- 364,366,368,370,375-380,382-386. In some embodiments, the full-length enzymes comprise an amino acid sequence that is identical to the amino acid sequence of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71-73,90,95,108,142,145,151,159-160,172,174,176,178- 181,194-195,198-206,209-234,236-243,245-247,249-256,258-260,263,267-282,284-287,289- 300,302-337,339-343,345-351,354-364,366,368,370,375-380,382-386. First fusion protein
[0113] In some embodiments, a base editor system disclosed herein comprises a first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain. In some embodiments, the C-terminal portion of the split deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or is 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 1-181.
[0114] In some embodiments, the C-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or 100% identical to the amino acids corresponding to positions 1322-1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387- 1426, or 1397-1426 of SEQ ID NO: 1. In some embodiments, the C-terminal portion of the split deaminase comprises amino acid sequence corresponding to amino acid positions 1322- 1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO: 1. 35 / 169#14463983v1
[0115] However, any one of the deaminases disclosed in Table 1 (SEQ ID NOs: 1- 181) may be split into a C-terminal portion and an N-terminal portion. For example, in some embodiments, C-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of any one of SEQ ID NOs: 2-193 that correspond to amino acids 1322-1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO:1, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 2-193 that corresponds to amino acids 1322-1426, 1333- 1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO:1.
[0116] As stated earlier, truncated versions of the full-length enzymes may also be used in the base editors contemplated herein. Thus, in some embodiments, a C-terminal portion of the split deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or is 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 3-4,14- 15,42,51,55,64,68-69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,194- 195,198-206,209-234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302- 337,339-343,345-351,354-364,366,368,370,375-380,382-386.
[0117] In some embodiments, a C-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 34-138, 45-138, 55-138, 69-138, 83-138, 99-138, 109-138 of SEQ ID NO: 194 (e.g., DddA toxin) or SEQ ID NO :201 (e.g., DddA11toxin).
[0118] Just as any one of the full-length enzyme may be spilt into a C-terminal portion and an N-terminal portion, so too can any one of the truncated deaminases shown in Table 2 (SEQ ID NOs: 194-385). Thus, in some embodiments, the C-terminal portion of the split deaminase comprises an amino acid sequence that is at least sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68- 69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that correspond to amino acids 34-138, 45-138, 55-138, 69-138, 83-138, 99-138, 109-138 of SEQ ID NO:194 or SEQ ID NO :201, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 3-4,14- 15,42,51,55,64,68-69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195- 36 / 169#14463983v1195,198-200, 202,209-234,236-243,245-247,249-256,258-260,263,267-282,284-287,289- 300,302-337,339-343,345-351,354-364,366,368,370,375-380,382-386 that corresponds to amino acids34-138, 45-138, 55-138, 69-138, 83-138, 99-138, 109-138 of SEQ ID NO:194 or SEQ ID NO:201.
[0119] In some embodiments, a first fusion protein comprises a portion of a DddA toxin. In some embodiments, the first fusion protein comprises a portion of a mutated version of the DddA toxin, such as the DddA11toxin. In some embodiments, the DddA11toxin(SEQ ID NO: 201) comprises the following mutations relative to the DddAtoxin sequence (SEQ ID NO: 194): S1330I, A1341V, N1342S, E1370K, T1380I and T1413I.
[0120] In some embodiments, the C-terminal portion of a split deaminase comprises any one of the following amino acid sequences:
[0121] G1333 DddAtoxin-C
[0122] PTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMT ETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 446).
[0123] G1397 DddAtoxin-C
[0124] AIPVKRGATGETKVFTGNSNSPKSPTKGGC (SEQ ID NO: 447).
[0125] DddA11 toxin-C
[0126] AIPVKRGATGETKVFIGNSNSPKSPTKGGC (SEQ ID NO: 448).
[0127] In some embodiments, the C-terminal portion of a split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 446-448
[0128] In some embodiments, a C-terminal portion of a split deaminase comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the nucleic acid sequence of SEQ ID NO: 449
[0129] DddA11 toxin-C 37 / 169#14463983v1
[0130] GCCATTCCAGTGAAGCGCGGCGCTACCGGTGAAACCAAAGTGTTTA TCGGTAACAGCAACAGCCCGAAGAGCCCGACCAAAGGCGGTTGC (SEQ ID NO: 449).
[0131] In some embodiments, a first fusion protein comprises a first epitope binding domain. In some embodiments, the first epitope binding domain comprises an antibody, an antibody fragment (e.g., a bivalent single chain variable fragment, scFv), a peptide, or a nanobody. In some embodiments, the first epitope binding domain is configured to bind to a first target epitope domain (e.g., fused to a DNA binding protein domain).
[0132] In some embodiments, the first epitope binding domain comprises a bivalent scFv configured to bind to a multimeric GCN4 peptide comprising an amino acid sequence EELLSKNYHLENEVARLKK (SEQ ID NO: 388). In some embodiments, the first epitope binding domain comprises a bivalent nanobody configured to bind to an ALFA-tag peptide comprising an amino acid sequence SRLEEELRRRLTE (SEQ ID NO: 389). In some embodiments, the first epitope binding domain comprises a bivalent nanobody configured to bind to BC2 peptide comprising an amino acid sequence PDRKAAVSHWQQ (SEQ ID NO: 390) or PDRVRAVSHWSS (SEQ ID NO: 391). In some embodiments, the first epitope binding domain comprises a bivalent VHH 2E7-derived chromobody / nanobody configured to bind a peptide comprising an amino acid sequence AVERYLKDQQLLGIW (SEQ ID NO: 392). In some embodiments, the first epitope binding domain comprises a bivalent gp41 nanobody configured to bind to a gp41 peptide comprising an amino acid sequence KNEQELLELDKWASL (SEQ ID NO: 393).
[0133] In some embodiments, a first epitope binding domain comprises a bivalent chloroalkane linker configured to covalently bind to haloalkane dehalogenase. In some embodiments, the first epitope binding domain comprises a bivalent O2-alkylguanine-DNA alkyltransferase configured to form a covalent bond with a benzylguanine ligand. In some embodiments, the first epitope binding domain comprises a bivalent O6-alkylguanine-DNA alkyltransferase configured to form a covalent bond with a benzyluganine ligand.
[0134] In some embodiments, the first epitope binding domain comprises an antibody, an antibody fragment, a peptide, or a nanobody configured to bind to a FLAG peptide comprising an amino acid sequence of DYKDDDDK (SEQ ID NO: 394), a myc peptide comprising an amino acid sequence of EQKLISEEDL (SEQ ID NO: 395), an HA peptide comprising an amino acid sequence of YPYDVPDYA (SEQ ID NO: 396), a polyHistidine tag having between 6 and 10 histidine residues, a C-tag having amino acid sequence of EPEA (SEQ ID NO: 409), a Twin-Step tag having streptavidin conjugated to an 38 / 169#14463983v1amino acid sequence WSHPQFEK-(GGGS)2-GGSA-WSHPQFEK (SEQ ID NO: 397), biotin, streptavidin, or protein A / G.
[0135] In some embodiments, the first fusion protein further comprises a uracil glycosylase inhibitor domain (UGI). In some embodiments, the UGI domain comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acid sequence of SEQ ID NO: 387.
[0136] UGI
[0137] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDES TDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 387)
[0138] In some embodiments, the UGI domain is positioned adjacent to and on either side of the deaminase domain. For example, in some embodiments, the first fusion protein comprises one or more of the following configurations: [epitope binding domain]-[C-terminal portion of split deaminase domain]-[UGI domain] [epitope binding domain]-[UGI domain]-[C-terminal portion of split deaminase domain]
[0139] The skilled artisan will understand, however, that the above recited configurations are not intended to be limiting in any way and that the first fusion protein may comprise additional domains, as appropriate, in other embodiments.
[0140] In some embodiments, the first fusion protein further comprises a chromatin modifying enzyme domain. Chromatin modifying enzymes, and other chromatin-modifying proteins, fall into three broad categories: writers, readers, and erasers. The function of these proteins is to dynamically maintain cell identity and regulate processes such as differentiation, development, proliferation and genome integrity via recognition of specific 'marks' (covalent post-translational modifications) on histone proteins and DNA. In normal cells, tissues and organs, precise co-ordination of these proteins ensures expression of only those genes required to specify phenotype or which are required at specific times, for specific functions. Chromatin modifications allow DNA modifications not coded by the DNA sequence to be passed on through the genome and underlies heritable phenomena such as X chromosome inactivation, aging, heterochromatin formation, reprogramming, and gene silencing (epigenetic control).
[0141] To date at least eight distinct types of modifications are found on histones. These include small covalent modifications, such as acetylation, methylation, and phosphorylation, the attachment of larger modifiers, such as ubiquitination or sumoylation, ADP ribosylation, proline isomerization, and deimination. 39 / 169#14463983v1
[0142] In some embodiments, a chromatin modifying enzyme domain comprises writer proteins, such as histone methyltransferases, histone acetyltransferases, kinases, and ubiquitin ligases.
[0143] In some embodiments, a chromatin modifying enzyme domain comprises readers such as methyl-lysine-recognition motifs such as bromodomains, chromodomains, tudor domains, PHD zinc fingers, PWWP domains and MBT domains.
[0144] In some embodiments, a chromatin modifying enzyme domain comprises erasers such as histone demethylases and histone deacetylases (HDACs and sirtuins).
[0145] In some embodiments, a chromatin modifying enzyme domain is selected from the group consisting of acetylase, deacetylase, methyltransferase, demethylase, ligase, helicase, deubiquitinase, and chromatin remodeler.
[0146] In some embodiments, a first fusion protein comprising a epitope binding domain, a C-terminal portion of split deaminase domain, a UGI domain, and a chromatin modifying enzyme domain can be arranged in 24 unique different combinations (e.g., 4 *3 *2* 1 = 24 permutations), all of which are contemplated herein. Second Fusion Protein
[0147] In some embodiments, a base editor system disclosed herein comprises a second fusion protein comprising a N-terminal portion of a split deaminase and a second epitope binding domain. In some embodiments, the N-terminal portion of the split deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or is 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 1-181.
[0148] In some embodiments, the N-terminal portion of the split deaminase comprises an amino acid sequence that is that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to 1-1321, 1-1332, 1-1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO:1. In some embodiments, the N-terminal portion of the split deaminase comprises amino acid sequence corresponding to amino acid positions 1-1321, 1-1332, 1- 1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO :1.
[0149] As mentioned above, any one of the deaminases disclosed in Table 1 (SEQ ID NOs: 1-181) may be split into a C-terminal portion and an N-terminal portion. For example, in some embodiments, N-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 40 / 169#14463983v199%, at least 99.5%, at least 99.8%, or 100% identical to a portion of any one of SEQ ID NOs: 2-193 that correspond to amino acids 1-1321, 1-1332, 1-1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO :1, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 2-193 that corresponds to amino acids 1-1321, 1-1332, 1-1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO :1.
[0150] As stated earlier, truncated versions of the full-length enzymes may also be used in the base editors contemplated herein. Thus, in some embodiments, a N-terminal portion of the split deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or is 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 194-385.
[0151] In some embodiments, a N-terminal portion of the split deaminase comprises an amino acid sequence that is that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 1-33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO :194 (e.g., DddA toxin) or SEQ ID NO :201 (e.g., DddA11 toxin).
[0152] Just as any one of the full-length enzymes may be spilt into a C-terminal portion and an N-terminal portion, so too can any one of the truncated deaminases shown in Table 2 (SEQ ID NOs: 194-385). Thus, in some embodiments, the N-terminal portion of the split deaminase comprises an amino acid sequence that is at least sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68- 69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that correspond to amino acids 1-33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO:194 or SEQ ID NO:201 or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71- 73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that corresponds to amino acids 1-33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO:194 or SEQ ID NO:201.
[0153] In some embodiments, a second fusion protein comprises a portion of a DddA toxin. In some embodiments, the second fusion protein comprises a portion of a mutated version of the DddA toxin, such as the DddA11 toxin. In some embodiments, the DddA11toxin 41 / 169#14463983v1(SEQ ID NO: 201) comprises the following mutations relative to the DddAtoxsequence (SEQ ID NO: 194): S1330I, A1341V, N1342S, E1370K, T1380I and T1413I.
[0154] In some embodiments, the N-terminal portion of a split deaminase comprises any one of the following amino acid sequences:
[0155] G1333 DddAtox-N
[0156] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGG (SEQ ID NO: 450.
[0157] G1397 DddAtox-N
[0158] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPY PNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTV VPPEG SEQ ID NO: 451.
[0159] DddA11toxin-N
[0160] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFISGGPTPY PNYVSAGHVEGQSALFMRDNGISEGLVFHNNPKGTCGFCVNMIETLLPENAKMTVV PPEG 452
[0161] In some embodiments, the N-terminal portion of a split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs:450-452.
[0162] In some embodiments, a N-terminal portion of a split deaminase comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the nucleic acid sequence of SEQ ID NO: 453.
[0163] DddA11toxin-N
[0164] ggcagctacgccctgggtccgtatcagattagcgccccgcagctgccagcttacaatggtcagaccgtgggta ccttctactatgtgaacgacgcgggcggtctggagagcaaggtgtttatcagcggcggtccaaccccgtacccaaactatgtcagtgc cggtcatgtggagggtcagagcgccctgttcatgcgtgataacggcatcagcgagggtctggtgttccacaacaacccgaaaggcac ctgcggtttttgcgtgaacatgatcgagaccctgctgccggaaaacgcgaaaatgaccgtggtgccgccggaaggt (SEQ ID NO: 453).
[0165] In some embodiments, a second fusion protein comprises a second epitope binding domain. In some embodiments, the second epitope binding domain comprises an 42 / 169#14463983v1antibody, an antibody fragment (e.g., a bivalent single chain variable fragment, scFv), a peptide, or a nanobody. In some embodiments, the second epitope binding domain is configured to bind to a second target epitope domain (e.g., fused to a DNA binding protein domain).
[0166] In some embodiments, the second epitope binding domain comprises a bivalent scFv configured to bind to a multimeric GCN4 peptide comprising an amino acid sequence EELLSKNYHLENEVARLKK (SEQ ID NO: 388). In some embodiments, the second epitope binding domain comprises a bivalent nanobody configured to bind to an ALFA-tag peptide comprising an amino acid sequence SRLEEELRRRLTE (SEQ ID NO: 389). In some embodiments, the second epitope binding domain comprises a bivalent nanobody configured to bind to BC2 peptide comprising an amino acid sequence PDRKAAVSHWQQ (SEQ ID NO: 390) or PDRVRAVSHWSS (SEQ ID NO: 391). In some embodiments, the second epitope binding domain comprises a bivalent VHH 2E7-derived chromobody / nanobody configured to bind a peptide comprising an amino acid sequence AVERYLKDQQLLGIW (SEQ ID NO: 392). In some embodiments, the second epitope binding domain comprises a bivalent gp41 nanobody configured to bind to a gp41 peptide comprising an amino acid sequence KNEQELLELDKWASL (SEQ ID NO: 393).
[0167] In some embodiments, a second epitope binding domain comprises a bivalent chloroalkane linker configured to covalently bind to haloalkane dehalogenase. In some embodiments, the second epitope binding domain comprises a bivalent O2-alkylguanine- DNA alkyltransferase configured to form a covalent bond with a benzylguanine ligand. In some embodiments, the second epitope binding domain comprises a bivalent O6- alkylguanine-DNA alkyltransferase configured to form a covalent bond with a benzyluganine ligand.
[0168] In some embodiments, the second epitope binding domain comprises an antibody, an antibody fragment, a peptide, or a nanobody configured to bind to a FLAG peptide comprising an amino acid sequence of DYKDDDDK (SEQ ID NO: 394), a myc peptide comprising an amino acid sequence of EQKLISEEDL (SEQ ID NO: 395), an HA peptide comprising an amino acid sequence of YPYDVPDYA (SEQ ID NO: 396), a polyHistidine tag having between 6 and 10 histidine residues, a C-tag having amino acid sequence of EPEA (SEQ ID NO: 409), a Twin-Step tag having streptavidin conjugated to an amino acid sequence WSHPQFEK-(GGGS)2-GGSA-WSHPQFEK (SEQ ID NO: 397), biotin, streptavidin, or protein A / G. 43 / 169#14463983v1
[0169] In some embodiments, the second fusion protein further comprises a uracil glycosylase inhibitor domain (UGI). In some embodiments, the UGI domain comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acid sequence of SEQ ID NO: 387.
[0170] UGI
[0171] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDES TDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 387)
[0172] In some embodiments, the UGI domain is positioned adjacent to and on either side of the deaminase domain. For example, in some embodiments, the first fusion protein comprises one or more of the following configurations: [epitope binding domain]-[N-terminal portion of split deaminase domain]-[UGI domain] [epitope binding domain]-[UGI domain ]-[N-terminal portion of split deaminase domain]
[0173] The skilled artisan will understand, however, that the above recited configurations are not intended to be limiting in any way and that the first fusion protein may comprise additional domains, as appropriate, in other embodiments.
[0174] In some embodiments, the second fusion protein further comprises a chromatin modifying enzyme domain. Chromatin modifying enzymes, and other chromatin- modifying proteins, fall into three broad categories: writers, readers, and erasers. The function of these proteins is to dynamically maintain cell identity and regulate processes such as differentiation, development, proliferation and genome integrity via recognition of specific 'marks' (covalent post-translational modifications) on histone proteins and DNA. In normal cells, tissues and organs, precise co-ordination of these proteins ensures expression of only those genes required to specify phenotype or which are required at specific times, for specific functions. Chromatin modifications allow DNA modifications not coded by the DNA sequence to be passed on through the genome and underlies heritable phenomena such as X chromosome inactivation, aging, heterochromatin formation, reprogramming, and gene silencing (epigenetic control).
[0175] To date at least eight distinct types of modifications are found on histones. These include small covalent modifications such as acetylation, methylation, and phosphorylation, the attachment of larger modifiers such as ubiquitination or sumoylation, and ADP ribosylation, proline isomerization and deimination. 44 / 169#14463983v1
[0176] In some embodiments, a chromatin modifying enzyme domain comprise writer proteins such as histone methyltransferases, histone acetyltransferases, some kinases and ubiquitin ligases.
[0177] In some embodiments, a chromatin modifying enzyme domain comprises readers such as methyl-lysine-recognition motifs such as bromodomains, chromodomains, tudor domains, PHD zinc fingers, PWWP domains and MBT domains.
[0178] In some embodiments, a chromatin modifying enzyme domain comprises erasers such as histone demethylases and histone deacetylases (HDACs and sirtuins).
[0179] In some embodiments, a chromatin modifying enzyme domain is selected from the group consisting of acetylase, deacetylase, methyltransferase, demethylase, ligase, helicase, deubiquitinase, and chromatin remodeler.
[0180] In some embodiments, a second fusion protein comprising a epitope binding domain, a N-terminal portion of split deaminase domain, a UGI domain, and a chromatin modifying enzyme domain can be arranged in 24 unique different combinations (e.g., 4 *3 *2* 1 = 24 permutations), all of which are contemplated herein. Third Fusion Protein
[0181] In some embodiments, a base editing system disclosed herein comprises a third fusion protein. In some embodiments, the third fusion protein comprises a DNA binding protein domain and one or more epitope domains.
[0182] In some embodiments, a DNA binding protein domain comprises a nucleic acid programmable DNA binding protein (napDNAbp). Any suitable napDNAbp domain known in the art may be used in the base editors described herein, such as those described in detail in “A Review on the Mechanisms and Applications of CRISPR / Cas9 / Cas12 / Cas13 / Cas14 Proteins Utilized for Genome Engineering” by V.E. Hillary and S.A. Ceasar, in Mol. Biotechnol.2023; 65(3): 311-325, which is incorporated herein by reference in its entirety. For example, in various embodiments, the napDNAbp may be any Class 2 CRISPR-Cas system, including any type II, type V, or type VI CRISPR- Cas enzyme. Given the rapid development of CRISPR-Cas as a tool for genome editing, there have been constant developments in the nomenclature used to describe and / or identify CRISPR-Cas enzymes, such as Cas9 and Cas9 orthologs. This application references CRISPR-Cas enzymes with nomenclature that may be old and / or new as described in United State Patent Application 63 / 136,194 (described elsewhere herein) or Makarova et al., The 45 / 169#14463983v1CRISPR Journal, Vol.1, No.5, 2018, which is incorporated herein by reference in its entirety.
[0183] Other napDNAbps are also possible in other embodiments. For example, in some embodiments, the napDNAbp comprises the canonical SpCas9, or any ortholog Cas9 protein, or any variant Cas9 protein —including any naturally occurring variant, mutant, or otherwise engineered version of Cas9 — that is known or that may be made or evolved through a directed evolutionary or otherwise mutagenic process. In various embodiments, the Cas9 or Cas9 variants have a nickase activity, i.e., only cleave one strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variants have inactive nucleases, i.e., are “dead” Cas9 proteins. Other variant Cas9 proteins that may be used are those having a smaller molecular weight than the canonical SpCas9 (e.g., for easier delivery) or having modified or rearranged primary amino acid structure (e.g., the circular permutant formats).
[0184] In various embodiments described herein, the base editors comprise a napDNAbp, such as a Cas9 protein. These proteins are “programmable” by way of their becoming complexed with a guide RNA (or a pegRNA, as the case may be), which guides the Cas9 protein to a target site on the DNA which possess a sequence that is complementary to the spacer portion of the gRNA (or pegRNA) and also which possesses the required PAM sequence. However, in certain embodiment envisioned here, the napDNAbp may be substituted with a different type of programmable protein, such as a zinc finger nuclease or a transcription activator-like effector nuclease (TALEN). See U.S. Ser. No.12 / 965,590; U.S. Ser. No.13 / 426,991 (U.S. Pat. No.8,450,471); U.S. Ser. No.13 / 427,040 (U.S. Pat. No. 8,440,431); U.S. Ser. No.13 / 427,137 (U.S. Pat. No.8,440,432); and U.S. Ser. No. 13 / 738,381, all of which are incorporated by reference herein in their entirety. In addition, TALENS are described in WO 2015 / 027134, US 9,181,535, Boch et al., “Breaking the Code of DNA Binding Specificity of TAL-Type III Effectors,” Science, vol.326, pp.1509-1512 (2009), Bogdanove et al., TAL Effectors: Customizable Proteins for DNA Targeting, Science, vol.333, pp.1843-1846 (2011), Cade et al., “Highly efficient generation of heritable zebrafish gene mutations using homo- and heterodimeric TALENs,” Nucleic Acids Research, vol.40, pp.8001-8010 (2012), and Cermak et al., “Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting,” Nucleic Acids Research, vol.39, No.17, e82 (2011), each of which are incorporated herein by reference. See also, for example, in Carroll et al., “Genome Engineering with Zinc-Finger Nucleases,” Genetics, Aug 2011, Vol.188: 773-782; Durai et al., “Zinc finger nucleases: custom- designed molecular scissors for genome engineering of plant and mammalian cells,” Nucleic 46 / 169#14463983v1Acids Res, 2005, Vol.33: 5978-90; and Gaj et al., “ZFN, TALEN, and CRISPR / Cas-based methods for genome engineering,” Trends Biotechnol.2013, Vol.31: 397-405, each of which is incorporated herein by reference in their entireties.
[0185] In some embodiments, the DNA binding protein comprises a napDNAbp, wherein the napDNAbp comprises a dead Cas9 (dCas9).
[0186] In some embodiments, a third fusion protein comprises one or more epitope domains. Any suitable epitope domain known to the skilled artisan may be used herein. In some embodiments, the one or more epitope domains comprises a multimeric GCN4 peptide comprising an amino acid sequence EELLSKNYHLENEVARLKK (SEQ ID NO: 388). In some embodiments, the one or more epitope domains comprises an ALFA-tag peptide comprising an amino acid sequence SRLEEELRRRLTE (SEQ ID NO: 389). In some embodiments, the one or more epitope domains comprises a BC2 peptide comprising an amino acid sequence PDRKAAVSHWQQ (SEQ ID NO: 390) or PDRVRAVSHWSS (SEQ ID NO: 391). In some embodiments, the one or more epitope domains comprises a peptide comprising an amino acid sequence AVERYLKDQQLLGIW (SEQ ID NO: 392). In some embodiments, the one or more epitope domains comprises a gp41 peptide comprising an amino acid sequence KNEQELLELDKWASL (SEQ ID NO: 393).
[0187] In some embodiments, the one or more epitope domains comprises a haloalkane dehalogenase. In some embodiments, the one or more epitope binding domains comprises a benzylguanine ligand. In some embodiments, the one or more epitope domains comprises a benzyluganine ligand.
[0188] In some embodiments, the one or more epitope domains comprises a FLAG peptide comprising an amino acid sequence of DYKDDDDK (SEQ ID NO: 394), a myc peptide comprising an amino acid sequence of EQKLISEEDL (SEQ ID NO: 395), an HA peptide comprising an amino acid sequence of YPYDVPDYA (SEQ ID NO: 396), a polyHistidine tag having between 6 and 10 histidine residues, a C-tag having amino acid sequence of EPEA, a Twin-Step tag having streptavidin conjugated to an amino acid sequence WSHPQFEK-(GGGS)2-GGSA-WSHPQFEK (SEQ ID NO: 397), biotin, streptavidin, or a protein A / G.
[0189] In some embodiments, the one or more epitope domains comprises a first epitope domain, a second epitope domain, a third epitope domain, a fourth epitope domain and so forth. In some embodiments, the one or more epitope domains are the same. In other embodiments, the one or more epitope domains are different. For example, in some embodiments, the third fusion protein has any one of the following exemplary configurations: 47 / 169#14463983v1[DNA binding protein]-[first epitope domain]n[DNA binding protein]-[first epitope domain]n-[second epitope domain]n [DNA binding protein]-[first epitope domain]n-[second epitope domain]n-[ first epitope domain]n-[second epitope domain]n, where n represent the number of repeats of the epitope domain.
[0190] In some embodiments, a third fusion protein comprises one or more repeats of the one or more epitope domains. For example, in some embodiments, the third fusion protein comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten repeats of each epitope domain of the one or more epitope domains.
[0191] As discussed elsewhere herein, aspects of the disclosure relate to split sites for deaminases recited in Tables 1 and 2, relative to split sites of DddAtoxin. Those of skill in the art will understand that various alignment programs and algorithms known in the art can be used to identify amino acid sequences and amino acid residues that are homologous to a reference deaminase amino acid sequence or amino acid residue (e.g., a reference DddAtoxin). Exemplary alignment programs include, but are not limited to, Clustal Omega (sequence) or Alphafold (structure). Additionally, split protein designs may also be determined using the methods described in Dagliyan, O., Krokhotin, A., Ozkan-Dagliyan, I. et al. Computational design of chemogenetic and optogenetic split proteins. Nat Commun 9, 4042 (2018). https: / / doi.org / 10.1038 / s41467-018-06531-4, which is hereby incorporated by reference in its entirety.
[0192] Compositions and Complexes
[0193] The present disclosure also provides compositions comprising any one of the first fusion proteins, second fusion proteins, and / or third fusion proteins disclosed herein. For example, in some embodiments, the composition comprises a first fusion protein, a second fusion protein, or a third fusion protein. In other embodiments, the composition comprises a first fusion protein and a second fusion protein. In some embodiments, the composition comprises a first fusion protein and a third fusion protein. In certain embodiments, the composition comprises a second fusion protein and a third fusion protein. In other embodiments, the composition comprises a first fusion protein, a second fusion protein, and a third fusion protein.
[0194] In some embodiments, a composition comprises a first fusion protein comprising any one of the C-terminal portions of a split deaminase (e.g., Table 1 and 2) and 48 / 169#14463983v1any one of the first epitope binding domains disclosed herein. In some embodiments, the composition comprises a second fusion protein comprising any one of the N-terminal portions of a split deaminase (e.g., Table 1 and 2) and any one of the second epitope binding domains disclosed herein. In some embodiments, the compositions comprise a third fusion protein comprising any one of the DNA binding proteins and any one of the one or more epitope domains as disclosed herein.
[0195] The present disclosure also provides for complexes comprising the first, second, and / or third fusion proteins disclosed herein. For example, in some embodiments, the complex comprises a first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain, a second fusion protein comprising a N- terminal portion of the split deaminase and a second epitope binding domain; and a third fusion protein comprising a DNA binding protein and one or more epitope domains.
[0196] In some embodiments, a first epitope binding domain (e.g. of the first fusion protein) and a second epitope binding domain (e.g., of the second fusion protein) bind to the same epitope domain (e.g., of the third fusion protein). However, in certain other embodiments, the first epitope binding domain (e.g. of the first fusion protein) and the second epitope binding domain (e.g., of the second fusion protein) bind to different epitope domains (e.g., of the third fusion protein).
[0197] In some embodiments, a first epitope binding domain of a first fusion protein is configured to bind to a first epitope domain of the third fusion protein. In some embodiments, a second epitope binding domain of a second fusion protein is configured to bind a second epitope domain of the third fusion protein. In some embodiments, the first epitope domain of the third fusion protein is the same as the second epitope domain of the second fusion protein (e.g., the first epitope binding domain and second epitope binding domain bind to the same epitope domain on the third fusion protein). In some embodiments, the first epitope domain of the third fusion protein is different than the second epitope domain of the second fusion protein (e.g., the first epitope binding domain and second epitope binding domain bind to different epitope domains on the third fusion protein).
[0198] In some embodiments, upon binding of a first fusion protein and a second fusion protein to a third fusion protein (e.g., via the epitope binding domain-epitope domain interaction), a N-terminal portion of a split deaminase and a C-terminal portion of a split deaminase form a functional deaminase. In some embodiments, a functional deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at 49 / 169#14463983v1least 98%, at least 99%, at least 99.5%, at least 99.8%, or is 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 1-386.
[0199] In some embodiments, a functional deaminase is configured to install one or more cytidine-to-thymine edits in a double stranded DNA. In some embodiments, the edit is installed on the coding DNA strand (e.g., 3′-5′ strand). In other embodiments, the edit is installed on the non-coding DNA strand (e.g., 5′-3′). 50 / 169#14463983v151 / 169#14463983v152 / 169#14463983v153 / 169#14463983v154 / 169#14463983v155 / 169#14463983v156 / 169#14463983v157 / 169#14463983v158 / 169#14463983v159 / 169#14463983v160 / 169#14463983v161 / 169#14463983v162 / 169#14463983v163 / 169#14463983v164 / 169#14463983v165 / 169#14463983v166 / 169#14463983v167 / 169#14463983v168 / 169#14463983v169 / 169#14463983v170 / 169#14463983v171 / 169#14463983v172 / 169#14463983v173 / 169#14463983v174 / 169#14463983v175 / 169#14463983v176 / 169#14463983v177 / 169#14463983v178 / 169#14463983v179 / 169#14463983v180 / 169#14463983v181 / 169#14463983v182 / 169#14463983v183 / 169#14463983v184 / 169#14463983v185 / 169#14463983v186 / 169#14463983v187 / 169#14463983v188 / 169#14463983v189 / 169#14463983v190 / 169#14463983v191 / 169#14463983v192 / 169#14463983v193 / 169#14463983v194 / 169#14463983v195 / 169#14463983v196 / 169#14463983v197 / 169#14463983v198 / 169#14463983v199 / 169#14463983v1100 / 169#14463983v1101 / 169#14463983v1102 / 169#14463983v1103 / 169#14463983v1104 / 169#14463983v1105 / 169#14463983v1106 / 169#14463983v1107 / 169#14463983v1108 / 169#14463983v1109 / 169#14463983v1110 / 169#14463983v1111 / 169#14463983v1112 / 169#14463983v1113 / 169#14463983v1114 / 169#14463983v1115 / 169#14463983v1116 / 169#14463983v1117 / 169#14463983v1118 / 169#14463983v1119 / 169#14463983v1120 / 169#14463983v1121 / 169#14463983v1122 / 169#14463983v1123 / 169#14463983v1124 / 169#14463983v1125 / 169#14463983v1126 / 169#14463983v1Nuclear Localization Signal 127 / 169#14463983v1
[0200] In various embodiments, the base editors disclosed herein further comprise one or more, preferably, at least two nuclear localization signals. In certain embodiments, the base editors comprise at least two NLSs. In embodiments with at least two NLSs, the NLSs can be the same NLSs, or they can be different NLSs. In addition, the NLSs may be expressed as part of a fusion protein with the remaining portions of the base editors. The location of the NLS fusion can be at the N-terminus, the C-terminus, or within a sequence of a base editor (e.g., inserted between the encoded napDNAbp domain (e.g., Cas9) and the one or more epitope domains (e.g., GCN4).
[0201] The NLSs may be any known NLS sequence in the art. The NLSs may also be any future-discovered NLSs for nuclear localization. The NLSs also may be any naturally-occurring NLS, or any non-naturally occurring NLS (e.g., an NLS with one or more desired mutations).
[0202] A nuclear localization signal or sequence (NLS) is an amino acid sequence that tags, designates, or otherwise marks a protein for import into the cell nucleus by nuclear transport. Typically, this signal consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface. Different nuclear localized proteins may share the same NLS. An NLS has the opposite function of a nuclear export signal (NES), which targets proteins out of the nucleus. A nuclear localization signal can also target the exterior surface of a cell. Thus, a single nuclear localization signal can direct the entity with which it is associated to the exterior of a cell and to the nucleus of a cell. Such sequences can be of any size and composition, for example, more than 25, 25, 15, 12, 10, 8, 7, 6, 5, or 4 amino acids, but will preferably comprise at least a four to eight amino acid sequence known to function as a nuclear localization signal (NLS).
[0203] The term “nuclear localization sequence” or “NLS” refers to an amino acid sequence that promotes import of a protein into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and would be apparent to the skilled artisan. For example, NLS sequences are described in Plank et al., International PCT application PCT / EP2000 / 011690, filed November 23, 2000, published as WO / 2001 / 038547 on May 31, 2001, the contents of which are incorporated herein by reference. In some embodiments, an NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 410), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 411), KRTADGSEFESPKKKRKV (SEQ ID NO: 412), or KRTADGSEFEPKKKRKV (SEQ ID NO: 413). In other embodiments, NLS comprises the amino acid sequences NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 414), PAAKRVKLD (SEQ ID NO: 415), 128 / 169#14463983v1RQRRNELKRSF (SEQ ID NO: 416), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 417). Other protein domains
[0204] In some embodiments, the base editor described herein may comprise one or more additional protein domains (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more domains in addition to the base editor components). A base editor may comprise any additional protein sequence, and optionally a linker sequence between any two domains. Examples of protein domains that may be fused to a base editor or component thereof (e.g., the napDNAbp domain, the base modification domain, or the NLS domain) include, without limitation, epitope tags and reporter gene sequences. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins including blue fluorescent protein (BFP). A base editor may be fused to a gene sequence encoding a protein or a fragment of a protein that bind DNA molecules or bind other cellular molecules, including, but not limited to, maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD) fusions, GAL4 DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Additional domains that may form part of a base editor are described in US Patent Publication No.2011 / 0059502, published March 10, 2011, and incorporated herein by reference in its entirety.
[0205] In an aspect of the invention, a reporter gene which includes, but is not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT) beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins including blue fluorescent protein (BFP), may be introduced into a cell to encode a gene product which serves as a marker by which to measure the alteration or modification of expression of the gene product. In a further embodiment of the invention, the DNA molecule encoding the gene product may be introduced into the cell via a vector. In certain embodiments of the invention the gene product is luciferase. In a further embodiment of the invention the expression of the gene product is decreased. 129 / 169#14463983v1Guide sequence (e.g., a guide RNA)
[0206] In various embodiments, the base editors may be complexed, bound, or otherwise associated with (e.g., via any type of covalent or non-covalent bond) one or more guide sequences, i.e., the sequence which becomes associated or bound to the base editor and directs its localization to a specific target sequence having complementarity to the guide sequence or a portion thereof. The particular design embodiments of a guide sequence will depend upon the nucleotide sequence of a genomic target site of interest (i.e., the desired site to be edited) and the type of napDNAbp (e.g., type of Cas protein) present in the base editor, among other factors, such as PAM sequence locations, percent G / C content in the target sequence, the degree of microhomology regions, secondary structures, etc.
[0207] In general, a guide sequence is any polynucleotide sequence having sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a napDNAbp (e.g., a Cas9, Cas9 homolog, or Cas9 variant) to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting example of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, a guide sequence is about or more than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length.
[0208] In some embodiments, a guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. The ability of a guide sequence to direct sequence-specific binding of a base editor to a target sequence may be assessed by any suitable assay. For example, the components of a base editor, including the guide sequence to be tested, may be provided to a host cell having the corresponding target sequence, such as by transfection with vectors encoding the components of a base editor disclosed herein, followed by an assessment of preferential cleavage within the target sequence, such as by Surveyor assay as described herein. Similarly, cleavage of a target polynucleotide sequence may be evaluated in a test tube by providing the target sequence, components of a base editor, 130 / 169#14463983v1including the guide sequence to be tested and a control guide sequence different from the test guide sequence, and comparing binding or rate of cleavage at the target sequence between the test and control guide sequence reactions. Other assays are possible, and will occur to those skilled in the art.
[0209] A guide sequence may be selected to target any target sequence. In some embodiments, the target sequence is a sequence within a genome of a cell. Exemplary target sequences include those that are unique in the target genome. For example, for the S. pyogenes Cas9, a unique target sequence in a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGG, where NNNNNNNNNNNNXGG (N is A, G, T, or C; and X can be anything) has a single occurrence in the genome. A unique target sequence in a genome may include an S. pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGG where NNNNNNNNNNNXGG (N is A, G, T, or C; and X can be anything) has a single occurrence in the genome. For the S. thermophilus CRISPR1Cas9, a unique target sequence in a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXXAGAAW where NNNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; and W is A or T) has a single occurrence in the genome. A unique target sequence in a genome may include an S. thermophilus CRISPR 1 Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXXAGAAW where NNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; and W is A or T) (SEQ ID NO: 425) has a single occurrence in the genome. For the S. pyogenes Cas9, a unique target sequence in a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGGXG where NNNNNNNNNNNNXGGXG (N is A, G, T, or C; and X can be anything) has a single occurrence in the genome. A unique target sequence in a genome may include an S. pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGGXG where NNNNNNNNNNNXGGXG (N is A, G, T, or C; and X can be anything) has a single occurrence in the genome. In each of these sequences “M” may be A, G, T, or C, and need not be considered in identifying a sequence as unique.
[0210] In some embodiments, a guide sequence is selected to reduce the degree of secondary structure within the guide sequence. Secondary structure may be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimal Gibbs free energy. An example of one such algorithm is mFold, as described by Zuker & Stiegler (Nucleic Acids Res.9 (1981), 133-148). Another example folding algorithm is the online webserver RNAfold, developed at Institute for Theoretical Chemistry at the 131 / 169#14463983v1University of Vienna, using the centroid structure prediction algorithm (see, e.g., A. R. Gruber et al., 2008, Cell 106(1): 23-24; and PA Carr & GM Church, 2009, Nature Biotechnology 27(12): 1151-62). Additional algorithms may be found in Chuai, G. et al., DeepCRISPR: optimized CRISPR guide RNA design by deep learning, Genome Biol.19:80 (2018), and U.S. application Ser. No.61 / 836,080 and U.S. Patent No.8,871,445, issued October 28, 2014, the entireties of each of which are incorporated herein by reference.
[0211] In some embodiments, a gRNA comprises a tracr mate sequence comprising a sequence that has sufficient complementarity with a tracr sequence to promote one or more of: (1) excision of a guide sequence flanked by tracr mate sequences in a cell containing the corresponding tracr sequence; and (2) formation of a complex at a target sequence, wherein the complex comprises the tracr mate sequence hybridized to the tracr sequence. In general, degree of complementarity is with reference to the optimal alignment of the tracr mate sequence and tracr sequence, along the length of the shorter of the two sequences. Optimal alignment may be determined by any suitable alignment algorithm, and may further account for secondary structures, such as self-complementarity within either the tracr sequence or tracr mate sequence. In some embodiments, the degree of complementarity between the tracr sequence and tracr mate sequence along the length of the shorter of the two when optimally aligned is about or more than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or higher. In some embodiments, the tracr sequence is about or more than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and tracr mate sequence are contained within a single transcript, such that hybridization between the two produces a transcript having a secondary structure, such as a hairpin. Preferred loop forming sequences for use in hairpin structures are four nucleotides in length, and most preferably have the sequence GAAA. However, longer or shorter loop sequences may be used, as may alternative sequences. The sequences preferably include a nucleotide triplet (for example, AAA), and an additional nucleotide (for example C or G). Examples of loop forming sequences include CAAA and AAAG. In an embodiment of the invention, the transcript or transcribed polynucleotide sequence has at least two or more hairpins. In certain embodiments, the transcript has two, three, four or five hairpins. In a further embodiment of the invention, the transcript has at most five hairpins.
[0212] In some embodiments, the single transcript further includes a transcription termination sequence; preferably this is a polyT sequence, for example six T nucleotides. Further non-limiting examples of single polynucleotides comprising a guide sequence, a tracr 132 / 169#14463983v1mate sequence, and a tracr sequence are as follows (listed 5′ to 3′), where “N” represents a base of a guide sequence, the first block of lower case letters represent the tracr mate sequence, and the second block of lower case letters represent the tracr sequence, and the final poly-T sequence represents the transcription terminator: (1) NNNNNNNNgtttttgtactctcaagatttaGAAAtaaatcttgcagaagctacaaagataaggctt catgccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT (SEQ ID NO: 430); (2) NNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaatca acaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT (SEQ ID NO: 431); (3) NNNNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaa atca acaccctgtcattttatggcagggtgtTTTTT (SEQ ID NO: 432); (4) NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAAtagcaagttaaaataaggctagtccgttatcaacttga aaa agtggcaccgagtcggtgcTTTTTT (SEQ ID NO: 433); (5) NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAATAGcaagttaaaataaggctagtccgttatcaactt gaa aaagtgTTTTTTT (SEQ ID NO: 434); and (6) NNNNNNNNNNNNNNNNNNNNgttttagagctagAAATAGcaagttaaaataaggctagtccgttatcaTTT TT TTT (SEQ ID NO: 435). In some embodiments, sequences (1) to (3) are used in combination with Cas9 from S. thermophilus CRISPR1. In some embodiments, sequences (4) to (6) are used in combination with Cas9 from S. pyogenes. In some embodiments, the tracr sequence is a separate transcript from a transcript comprising the tracr mate sequence. Linkers
[0213] In certain embodiments, linkers may be used to link any of the peptides or peptide domains or domains of the base editor (e.g., a C-terminal portion of a split deaminase with a first epitope binding domain).
[0214] As defined above, the term “linker,” as used herein, refers to a chemical group or a molecule linking two molecules or domains, e.g., a binding domain and a cleavage domain of a nuclease. In some embodiments, a linker joins a gRNA binding domain of a napDNAbp and the catalytic domain of a recombinase. In some embodiments, a linker joins a dCas9 and base editor domain (e.g., an adenine oxidase). Typically, the linker is positioned between, or flanked by, two groups, molecules, or other domains and connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical domain. Chemical domains 133 / 169#14463983v1include, but are not limited to, disulfide, hydrazone, thiol and azo domains. In some embodiments, the linker is 5-100 amino acids in length, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45- 50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. In some embodiments, the linker is a molecule in length. Longer or shorter linkers are also contemplated.
[0215] The linker may be as simple as a covalent bond, or it may be a polymeric linker many atoms in length. In certain embodiments, the linker is a polypeptide or based on amino acids. In other embodiments, the linker is not peptide-like. In certain embodiments, the linker is a covalent bond (e.g., a carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.). In certain embodiments, the linker is a carbon-nitrogen bond of an amide linkage. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic linker. In certain embodiments, the linker is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminoalkanoic acid. In certain embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta-alanine, 3-aminopropanoic acid, 4-aminobutanoic acid, 5- pentanoic acid, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In certain embodiments, the linker is based on a carbocyclic domain (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol domain (PEG). In other embodiments, the linker comprises amino acids. In certain embodiments, the linker comprises a peptide. In certain embodiments, the linker comprises an aryl or heteroaryl domain. In certain embodiments, the linker is based on a phenyl ring. The linker may included functionalized domains to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.
[0216] In some other embodiments, the linker comprises the amino acid sequence (GGGGS)n(SEQ ID NO: 436), (G)n(SEQ ID NO: 437), (EAAAK)n(SEQ ID NO: 438), (GGS)n(SEQ ID NO: 439), (SGGS)n(SEQ ID NO: 440), (XP)n(SEQ ID NO: 441), or any combination thereof, wherein n is independently an integer between 1 and 30, and wherein X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS)n(SEQ ID NO: 454), wherein n is 1, 3, or 7. In some embodiments, the linker 134 / 169#14463983v1comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 442). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 443). In some embodiments, the linker comprises the amino acid sequence SGGSGGSGGS (SEQ ID NO: 444). In some embodiments, the linker comprises the amino acid sequence SGGS (SEQ ID NO: 445).
[0217] In some embodiments, a linker is used to fuse one or more domains of a first fusion protein, a second fusion protein, or a third fusion protein. For example, in some embodiments, a linker fuses a C-terminal portion of a split deaminase to a first epitope binding domain (e.g. first fusion protein). In some embodiments, a linker fuses a N-terminal portion of a split deaminase to a second epitope binding domain (e.g., a second fusion protein). In other embodiments, a linker fuses a DNA binding protein domain to one or more epitope domains (e.g., a third fusion protein). A linker may also be used to fuse one or more epitope domains together (e.g., to fuse a first epitope domain to a second epitope domain). Polynucleotides
[0218] The terms “nucleic acid”, “nucleic acid molecule”, “ribonucleotide”, “polynucleotide”, “nucleotide sequence”, “nucleic acid sequence”, and “oligonucleotide” refer to a single nucleotide or a series of nucleotide bases (also called “nucleotides”) in DNA and RNA. The polynucleotides can be chimeric mixtures or derivatives or modified versions thereof, single-stranded or double-stranded. The oligonucleotide can be modified at the base moiety, sugar moiety, or phosphate backbone, for example, to improve stability of the molecule, its hybridization parameters, etc. The antisense oligonucleotide may comprise a modified base moiety which is selected from the group including, but not limited to, 5- fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4- acetylcytosine, 5-(carboxyhydroxylmethyl) uracil, 5-carboxymethylaminomethyl-2- thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2- methyladenine, 2-methylguanine, 3-methylcytosine, 5- methylcytosine, N6-adenine, 7- methylguanine, 5-methylaminomethyluracil, 5- methoxyaminomethyl-2-thiouracil, beta-D- mannosylqueosine, 5’-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6- isopentenyladenine, wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2- thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil- 5-oxyacetic acid methylester, uracil-5-oxyacetic acid, 5-methyl-2- thiouracil, 3-(3-amino-3-N-2-carboxypropyl) uracil, a 135 / 169#14463983v1thio-guanine, and 2,6-diaminopurine. A nucleotide sequence typically carries genetic information, including the information used by cellular machinery to make proteins and enzymes. These terms include double- or single-stranded genomic and cDNA, RNA, any synthetic and genetically manipulated polynucleotide, and both sense and antisense polynucleotides. This includes single- and double-stranded molecules, i.e., DNA-DNA, DNA-RNA and RNA-RNA hybrids, as well as “protein nucleic acids” (PNAs) formed by conjugating bases to an amino acid backbone. This also includes nucleic acids containing carbohydrate or lipids. Exemplary DNAs include single-stranded DNA (ssDNA), double- stranded DNA (dsDNA), plasmid DNA (pDNA), genomic DNA (gDNA), complementary DNA (cDNA), antisense DNA, chloroplast DNA (ctDNA or cpDNA), microsatellite DNA, mitochondrial DNA (mtDNA or mDNA), kinetoplast DNA (kDNA), provirus, lysogen, repetitive DNA, satellite DNA, and viral DNA. Exemplary RNAs include single-stranded RNA (ssRNA), double-stranded RNA (dsRNA), small interfering RNA (siRNA), messenger RNA (mRNA), precursor messenger RNA (pre-mRNA), small hairpin RNA or short hairpin RNA (shRNA), microRNA (miRNA), guide RNA (gRNA), transfer RNA (tRNA), antisense RNA (asRNA), heterogeneous nuclear RNA (hnRNA), coding RNA, non-coding RNA (ncRNA), long non-coding RNA (long ncRNA or lncRNA), satellite RNA, viral satellite RNA, signal recognition particle RNA, small cytoplasmic RNA, small nuclear RNA (snRNA), ribosomal RNA (rRNA), Piwi-interacting RNA (piRNA), polyinosinic acid, ribozyme, flexizyme, small nucleolar RNA (snoRNA), spliced leader RNA, viral RNA, and viral satellite RNA.
[0219] Polynucleotides described herein may be synthesized by standard methods known in the art, (e.g., by use of an automated DNA synthesizer etc.). As examples, phosphorothioate oligonucleotides may be synthesized by the method of Stein et al., Nucl. Acids Res., 16, 3209, (1988), methylphosphonate oligonucleotides can be prepared by use of controlled pore glass polymer supports (Sarin et al., Proc. Natl. Acad. Sci. U.S.A.85, 7448- 7451, (1988)). A number of methods have been developed for delivering antisense DNA or RNA to cells, e.g., antisense molecules can be injected directly into the tissue site, or modified antisense molecules, designed to target the desired cells (antisense linked to peptides or antibodies that specifically bind receptors or antigens expressed on the target cell surface) can be administered systemically. Alternatively, RNA molecules may be generated by in vitro and in vivo transcription of DNA sequences encoding the antisense RNA molecule. Such DNA sequences may be incorporated into a wide variety of vectors that incorporate suitable RNA polymerase promoters such as the T7 or SP6 polymerase 136 / 169#14463983v1promoters. Alternatively, antisense cDNA constructs that synthesize antisense RNA constitutively or inducibly, depending on the promoter used, can be introduced stably into cell lines. However, it is often difficult to achieve intracellular concentrations of the antisense sufficient to suppress translation of endogenous mRNAs. Therefore, a preferred approach utilizes a recombinant DNA construct in which the antisense oligonucleotide is placed under the control of a strong promoter. The use of such a construct to transfect target cells in the patient will result in the transcription of sufficient amounts of single stranded RNAs that will form complementary base pairs with the endogenous target gene transcripts and thereby prevent translation of the target gene mRNA. For example, a vector can be introduced in vivo such that it is taken up by a cell and directs the transcription of an antisense RNA. Such a vector can remain episomal or become chromosomally integrated, as long as it can be transcribed to produce the desired antisense RNA. Such vectors can be constructed by recombinant DNA technology methods standard in the art. Vectors can be plasmid, viral, or others known in the art, used for replication and expression in mammalian cells. Expression of the sequence encoding the antisense RNA can be by any promoter known in the art to act in mammalian, preferably human, cells. Such promoters can be inducible or constitutive. Any type of plasmid, cosmid, yeast artificial chromosome, or viral vector can be used to prepare the recombinant DNA construct that can be introduced directly into the tissue site.
[0220] The polynucleotides may be flanked by natural regulatory (expression control) sequences or may be associated with heterologous sequences, including promoters, internal ribosome entry sites (IRES) and other ribosome binding site sequences, enhancers, response elements, suppressors, signal sequences, polyadenylation sequences, introns, 5´- and 3´-non- coding regions, and the like. The nucleic acids may also be modified by many means known in the art. Non-limiting examples of such modifications include methylation, “caps”, substitution of one or more of the naturally occurring nucleotides with an analog, and internucleotide modifications, such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoroamidates, carbamates, etc.) and with charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.). Polynucleotides may contain one or more additional covalently linked moieties, such as, for example, proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), intercalators (e.g., acridine, psoralen, etc.), chelators (e.g., metals, radioactive metals, iron, oxidative metals, etc.), and alkylators. The polynucleotides may be derivatized by formation of a methyl or ethyl phosphotriester or an alkyl phosphoramidate linkage. Furthermore, the polynucleotides herein may also be modified with a label capable of providing a detectable signal, either 137 / 169#14463983v1directly or indirectly. Exemplary labels include radioisotopes, fluorescent molecules, isotopes (e.g., radioactive isotopes), biotin, and the like.
[0221] In some embodiments, a polynucleotide encodes any one of the first fusion proteins, second fusion proteins, and / or a third fusion proteins as disclosed herein. Vectors and Reagents
[0222] Several embodiments of the making and using of the base editors of the invention relate to vector systems comprising one or more vectors, or vectors as such. Vectors may be designed to clone and / or express the base editors as disclosed herein. Vectors may also be designed to clone and / or express one more gRNAs having complementarity to the target sequence, as disclosed herein. Vectors may also be designed to transfect the base editors and gRNAs of the disclosure into one or more cells, e.g., a target diseased eukaryotic cell for treatment with the base editor systems and methods disclosed herein.
[0223] Vectors can be designed for expression of base editor transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells. For example, base editor transcripts can be expressed in bacterial cells such as Escherichia coli, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells. Suitable host cells are discussed further in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press. San Diego, Calif. (1990). Alternatively, expression vectors encoding one or more base editors described herein can be transcribed and translated in vitro, for example using T7 promoter regulatory sequences and T7 polymerase.
[0224] Vectors may be introduced and propagated in any suitable prokaryotic cell. In some embodiments, a prokaryote is used to amplify copies of a vector to be introduced into a eukaryotic cell or as an intermediate vector in the production of a vector to be introduced into a eukaryotic cell (e.g., amplifying a plasmid as part of a viral vector packaging system). In some embodiments, a prokaryote is used to amplify copies of a vector and express one or more nucleic acids, such as to provide a source of one or more proteins for delivery to a host cell or host organism. Expression of proteins in prokaryotes is most often carried out in Escherichia coli with vectors containing constitutive or inducible promoters directing the expression of either fusion or non-fusion proteins.
[0225] Fusion expression vectors also may be used to express the base editors of the disclosure. Such vectors generally add a number of amino acids to a protein encoded therein, such as to the amino terminus of the recombinant protein. Such fusion vectors may serve one 138 / 169#14463983v1or more purposes, such as: (i) to increase expression of a recombinant protein; (ii) to increase the solubility of a recombinant protein; and (iii) to aid in the purification of a recombinant protein by acting as a ligand in affinity purification. Often, in fusion expression vectors, a proteolytic cleavage site is introduced at the junction of the fusion domain and the recombinant protein to enable separation of the recombinant protein from the fusion domain subsequent to purification of the fusion protein. Such enzymes, and their cognate recognition sequences, include Factor Xa, thrombin and enterokinase. Example fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith and Johnson, 1988. Gene 67: 31-40), pMAL (New England Biolabs, Beverly, Mass.), and pRIT5 (Pharmacia, Piscataway, N.J.) that fuse glutathione S-transferase (GST), maltose E binding protein, or protein A, respectively, to the target recombinant protein.
[0226] Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (Amrann et al., (1988) Gene 69:301-315) and pET 11d (Studier et al., Gene Expression Technology: Methods In Enzymology 185, Academic Press, San Diego, Calif. (1990) 60-89).
[0227] In some embodiments, a vector is a yeast expression vector for expressing the base editors described herein. Examples of vectors for expression in yeast Saccharomyces cerivisae include pYepSec1 (Baldari, et al., 1987. EMBO J.6: 229-234), pMFa (Kuijan and Herskowitz, 1982. Cell 30: 933-943), pJRY88 (Schultz et al., 1987. Gene 54: 113-123), pYES2 (Invitrogen Corporation, San Diego, Calif.), and picZ (InVitrogen Corp, San Diego, Calif.).
[0228] In some embodiments, a vector drives protein expression in insect cells using baculovirus expression vectors. Baculovirus vectors available for expression of proteins in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith, et al., 1983. Mol. Cell. Biol.3: 2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170: 31-39).
[0229] In some embodiments, a vector is capable of driving expression of one or more sequences in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329: 840) and pMT2PC (Kaufman, et al., 1987. EMBO J.6: 187-195). When used in mammalian cells, the expression vector's control functions are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. For other suitable expression systems for both prokaryotic and eukaryotic cells see, e.g., Chapters 16 and 17 of Sambrook, et al., Molecular Cloning: A Laboratory Manual.2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989. 139 / 169#14463983v1
[0230] In some embodiments, the recombinant mammalian expression vector is capable of directing expression of the nucleic acid preferentially in a particular cell type (e.g., tissue-specific regulatory elements are used to express the nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include the albumin promoter (liver-specific; Pinkert, et al., 1987. Genes Dev.1: 268-277), lymphoid-specific promoters (Calame and Eaton, 1988. Adv. Immunol.43: 235- 275), in particular promoters of T cell receptors (Winoto and Baltimore, 1989. EMBO J.8: 729-733) and immunoglobulins (Baneiji, et al., 1983. Cell 33: 729-740; Queen and Baltimore, 1983. Cell 33: 741-748), neuron-specific promoters (e.g., the neurofilament promoter; Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA 86: 5473-5477), pancreas- specific promoters (Edlund, et al., 1985. Science 230: 912-916), and mammary gland-specific promoters (e.g., milk whey promoter, U.S. Pat. No.4,873,316 and European Application Publication No.264,166). Developmentally-regulated promoters are also encompassed, e.g., the murine hox promoters (Kessel and Gruss, 1990. Science 249: 374-379) and the α- fetoprotein promoter (Campes and Tilghman, 1989. Genes Dev.3: 537-546).
[0231] In some embodiments, a vector comprises a polynucleotide encoding a first fusion protein. In some embodiments, the vector comprises a polynucleotide encoding a second fusion protein. In some embodiments, the vector comprises a polynucleotide encoding a third fusion protein. The vector may also, according to some embodiments, comprise more than one polynucleotide. For example, in some embodiments, the vector may comprise a first polynucleotide encoding the first fusion protein and a second polynucleotide encoding a second fusion protein. Similarly, in some embodiments, the vector may comprise a first polynucleotide encoding the first fusion protein and a third polynucleotide encoding a third fusion protein. In other embodiments, the vector may comprise a second polynucleotide encoding the second fusion protein and a third polynucleotide encoding a third fusion protein.
[0232] In some embodiments, a vector comprises a viral vector. In some cases, the viral vector is a viral transfer vector. In some cases, the viral vector is a viral envelope vector. In some embodiments, the viral vector is selected from the group consisting of adenoviruses, adeno-associated viruses, and lentiviruses.
[0233] In some embodiments, a vector comprises a mammalian expression vector. In some embodiments, the mammalian expression vector is delivered by lipid nanoparticles, liposomes, or polymers. Any suitable lipid nanoparticle, liposome, or polymer for delivering the mammalian expression vector known to the skilled artisan may be used herein. 140 / 169#14463983v1Pharmaceutical compositions
[0234] Other embodiments of the present disclosure relate to pharmaceutical compositions comprising any of the fusion proteins or the fusion protein-gRNA complexes described herein. The term “pharmaceutical composition”, as used herein, refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition comprises additional agents (e.g., for specific delivery, increasing half-life, or other therapeutic compounds).
[0235] In some embodiments, any of the fusion proteins, gRNAs, and / or complexes described herein are provided as part of a pharmaceutical composition. In some embodiments, the pharmaceutical composition comprises any of the fusion proteins provided herein. In some embodiments, the pharmaceutical composition comprises any of the complexes provided herein. In some embodiments pharmaceutical composition comprises a gRNA, a napDNAbp-dCas9 fusion protein, and a pharmaceutically acceptable excipient. In some embodiments pharmaceutical composition comprises a gRNA, a napDNAbp-nCas9 fusion protein, and a pharmaceutically acceptable excipient. Pharmaceutical compositions may optionally comprise one or more additional therapeutically active substances.
[0236] In some embodiments, compositions provided herein are administered to a subject, for example, to a human subject, in order to effect a targeted genomic modification within the subject. In some embodiments, cells are obtained from the subject and contacted with a any of the pharmaceutical compositions provided herein. In some embodiments, cells removed from a subject and contacted ex vivo with a pharmaceutical composition are re- introduced into the subject, optionally after the desired genomic modification has been effected or detected in the cells. Methods of delivering pharmaceutical compositions comprising nucleases are known, and are described, for example, in U.S. Pat. Nos.6,453,242; 6,503,717; 6,534,261; 6,599,692; 6,607,882; 6,689,558; 6,824,978; 6,933,113; 6,979,539; 7,013,219; and 7,163,824, the disclosures of all of which are incorporated by reference herein in their entireties. Although the descriptions of pharmaceutical compositions provided herein are principally directed to pharmaceutical compositions which are suitable for administration to humans, it will be understood by the skilled artisan that such compositions are generally suitable for administration to animals or organisms of all sorts. Modification of pharmaceutical compositions suitable for administration to humans in order to render the compositions suitable for administration to various animals is well understood, and the ordinarily skilled veterinary pharmacologist can design and / or perform such modification 141 / 169#14463983v1with merely ordinary, if any, experimentation. Subjects to which administration of the pharmaceutical compositions is contemplated include, but are not limited to, humans and / or other primates; mammals, domesticated animals, pets, and commercially relevant mammals such as cattle, pigs, horses, sheep, cats, dogs, mice, and / or rats; and / or birds, including commercially relevant birds such as chickens, ducks, geese, and / or turkeys.
[0237] Formulations of the pharmaceutical compositions described herein may be prepared by any method known or hereafter developed in the art of pharmacology. In general, such preparatory methods include the step of bringing the active ingredient(s) into association with an excipient and / or one or more other accessory ingredients, and then, if necessary and / or desirable, shaping and / or packaging the product into a desired single- or multi-dose unit.
[0238] Pharmaceutical formulations may additionally comprise a pharmaceutically acceptable excipient, which, as used herein, includes any and all solvents, dispersion media, diluents, or other liquid vehicles, dispersion or suspension aids, surface active agents, isotonic agents, thickening or emulsifying agents, preservatives, solid binders, lubricants and the like, as suited to the particular dosage form desired. Remington’s The Science and Practice of Pharmacy, 21stEdition, A. R. Gennaro (Lippincott, Williams & Wilkins, Baltimore, MD, 2006; incorporated in its entirety herein by reference) discloses various excipients used in formulating pharmaceutical compositions and known techniques for the preparation thereof. See also PCT application PCT / US2010 / 055131 (Publication No. WO / 2011053982), filed Nov.2, 2010, incorporated in its entirety herein by reference, for additional suitable methods, reagents, excipients and solvents for producing pharmaceutical compositions comprising a nuclease. Except insofar as any conventional excipient medium is incompatible with a substance or its derivatives, such as by producing any undesirable biological effect or otherwise interacting in a deleterious manner with any other component(s) of the pharmaceutical composition, its use is contemplated to be within the scope of this disclosure.
[0239] As used here, the term “pharmaceutically acceptable carrier” means a pharmaceutically acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc magnesium, calcium or zinc stearate, or steric acid), or solvent encapsulating material, involved in carrying or transporting the compound from one site (e.g., the delivery site) of the body, to another site (e.g., organ, tissue or portion of the body). A pharmaceutically acceptable carrier is “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the tissue of the subject (e.g., physiologically compatible, sterile, physiologic pH, etc.). 142 / 169#14463983v1Some examples of materials which can serve as pharmaceutically acceptable carriers include: (1) sugars, such as lactose, glucose and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose, and its derivatives, such as sodium carboxymethyl cellulose, methylcellulose, ethyl cellulose, microcrystalline cellulose and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricating agents, such as magnesium stearate, sodium lauryl sulfate and talc; (8) excipients, such as cocoa butter and suppository waxes; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil and soybean oil; (10) glycols, such as propylene glycol; (11) polyols, such as glycerin, sorbitol, mannitol and polyethylene glycol (PEG); (12) esters, such as ethyl oleate and ethyl laurate; (13) agar; (14) buffering agents, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer's solution; (19) ethyl alcohol; (20) pH buffered solutions; (21) polyesters, polycarbonates and / or polyanhydrides; (22) bulking agents, such as polypeptides and amino acids (23) serum component, such as serum albumin, HDL and LDL; (22) C2-C12 alcohols, such as ethanol; and (23) other non-toxic compatible substances employed in pharmaceutical formulations. Wetting agents, coloring agents, release agents, coating agents, sweetening agents, flavoring agents, perfuming agents, preservative and antioxidants may also be present in the formulation. The terms such as “excipient”, “carrier”, “pharmaceutically acceptable carrier” or the like are used interchangeably herein.
[0240] In some embodiments, the pharmaceutical composition is formulated for delivery to a subject, e.g., for gene editing. Suitable routes of administrating the pharmaceutical composition described herein include, without limitation: topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, gingival, intradental, intracochlear, transtympanic, intraorgan, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseus, periocular, intratumoral, intracerebral, and intracerebroventricular administration.
[0241] In some embodiments, the pharmaceutical composition described herein is administered locally to a diseased site. In some embodiments, the pharmaceutical composition described herein is administered to a subject by injection, by means of a catheter, by means of a suppository, or by means of an implant, the implant being of a porous, non-porous, or gelatinous material, including a membrane, such as a sialastic membrane, or a fiber.
[0242] In some embodiments, the pharmaceutical composition is formulated in accordance with routine procedures as a composition adapted for intravenous or 143 / 169#14463983v1subcutaneous administration to a subject, e.g., a human. In some embodiments, pharmaceutical composition for administration by injection are solutions in sterile isotonic aqueous buffer. Where necessary, the pharmaceutical can also include a solubilizing agent and a local anesthetic such as lignocaine to ease pain at the site of the injection. Generally, the ingredients are supplied either separately or mixed together in unit dosage form, for example, as a dry lyophilized powder or water free concentrate in a hermetically sealed container such as an ampoule or sachette indicating the quantity of active agent. Where the pharmaceutical is to be administered by infusion, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline. Where the pharmaceutical composition is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed prior to administration.
[0243] The pharmaceutical composition can be contained within a lipid particle or vesicle, such as a liposome or microcrystal, which is also suitable for parenteral administration. The particles can be of any suitable structure, such as unilamellar or plurilamellar, so long as compositions are contained therein. Compounds can be entrapped in “stabilized plasmid-lipid particles” (SPLP) containing the fusogenic lipid dioleoylphosphatidylethanolamine (DOPE), low levels (5-10 mol%) of cationic lipid, and stabilized by a polyethyleneglycol (PEG) coating (Zhang Y. P. et al., Gene Ther.1999, 6:1438-47). Positively charged lipids such as N-[1-(2,3-dioleoyloxi)propyl]-N,N,N- trimethyl-amoniummethylsulfate, or “DOTAP,” are particularly preferred for such particles and vesicles. The preparation of such lipid particles is well known. See, e.g., U.S. Patent Nos. 4,880,635; 4,906,477; 4,911,928; 4,917,951; 4,920,016; and 4,921,757; each of which is incorporated herein by reference.
[0244] The pharmaceutical composition described herein may be administered or packaged as a unit dose, for example. The term “unit dose” when used in reference to a pharmaceutical composition of the present disclosure refers to physically discrete units suitable as unitary dosage for the subject, each unit containing a predetermined quantity of active material calculated to produce the desired therapeutic effect in association with the required diluent; i.e., carrier, or vehicle.
[0245] Further, the pharmaceutical composition can be provided as a pharmaceutical kit comprising (a) a container containing a compound of the invention in lyophilized form and (b) a second container containing a pharmaceutically acceptable diluent (e.g., sterile water) for injection. The pharmaceutically acceptable diluent can be used for reconstitution or dilution of the lyophilized compound of the invention. Optionally associated with such 144 / 169#14463983v1container(s) can be a notice in the form prescribed by a governmental agency regulating the manufacture, use or sale of pharmaceuticals or biological products, which notice reflects approval by the agency of manufacture, use or sale for human administration.
[0246] In another aspect, an article of manufacture containing materials useful for the treatment of the diseases described above is included. In some embodiments, the article of manufacture comprises a container and a label. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The containers may be formed from a variety of materials such as glass or plastic. In some embodiments, the container holds a composition that is effective for treating a disease described herein and may have a sterile access port. For example, the container may be an intravenous solution bag or a vial having a stopper pierceable by a hypodermic injection needle. The active agent in the composition is a compound of the invention. In some embodiments, the label on or associated with the container indicates that the composition is used for treating the disease of choice. The article of manufacture may further comprise a second container comprising a pharmaceutically acceptable buffer, such as phosphate-buffered saline, Ringer’s solution, or dextrose solution. It may further include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, syringes, and package inserts with instructions for use.
[0247] In some embodiments, a pharmaceutical composition comprises any one of the first fusion proteins, second fusion proteins, third fusion proteins, compositions, complexes, polynucleotides, or vectors disclosed herein, or any combination thereof.
[0248] In some embodiments, a pharmaceutical composition comprises a lentiviral packaging vector. In some embodiments, a pharmaceutical composition comprises a lentiviral envelope vector. Any suitable lentiviral packaging and / or envelope vectors known to the skilled artisan may be used herein. Kits and cells
[0249] This disclosure provides kits comprising a nucleic acid construct comprising nucleotide sequences encoding the fusion proteins, gRNAs, and / or complexes described herein. Some embodiments of this disclosure provide kits comprising a nucleic acid construct comprising a first fusion protein, a second fusion, or a third fusion protein. In some embodiments, the nucleotide sequence encodes any of the fusion proteins provided herein. In some embodiments, the nucleotide sequence comprises a heterologous promoter that drives expression of the fusion protein. The nucleotide sequence may further comprise a 145 / 169#14463983v1heterologous promoter that drives expression of the gRNA, or a heterologous promoter that drives expression of the fusion protein and the gRNA.
[0250] In some embodiments, the kit further comprises an expression construct encoding a guide nucleic acid backbone, e.g., a guide RNA backbone, wherein the construct comprises a cloning site positioned to allow the cloning of a nucleic acid sequence identical or complementary to a target sequence into the guide nucleic acid, e.g., guide RNA backbone. In some embodiments, the kit further comprises an expression construct comprising a nucleotide sequence encoding a UGI domain.
[0251] The disclosure further provides kits comprising a fusion protein as provided herein, a gRNA having complementarity to a target sequence, and one or more of the following: cofactor proteins, buffers, media, and target cells (e.g. human cells). Kits may comprise combinations of several or all of the aforementioned components.
[0252] Some embodiments of this disclosure provide cells comprising any of the fusion proteins or complexes provided herein. In some embodiments, the cells comprise a nucleotide that encodes any of the fusion proteins provided herein. In some embodiments, the cells comprise any of the nucleotides or vectors provided herein.
[0253] In some embodiments, a host cell is transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, a cell is transfected as it naturally occurs in a subject. In some embodiments, a cell that is transfected is taken from a subject. In some embodiments, the cell is derived from cells taken from a subject, such as a cell line. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLa- S3, Huh1, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panc1, PC-3, TF1, CTLL-2, C1R, Rat6, CV1, RPTE, A10, T24, J82, A375, ARH-77, Calu1, SW480, SW620, SKOV3, SK-UT, CaCo2, P388D1, SEM-K2, WEHI-231, HB56, TIB55, Jurkat, J45.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelial, BALB / 3T3 mouse embryo fibroblast, 3T3 Swiss, 3T3-L1, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A 172, A20, A253, A431, A-549, ALC, B16, B35, BCP-1 cells, BEAS-2B, bEnd.3, BHK-21, BR 293. BxPC3. C3H-10T1 / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO Dhfr − / −, COR-L23, COR-L23 / CPR, COR-L23 / 5010, COR-L23 / R23, COS-7, COV- 434, CML T1, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepa1c1c7, HL- 146 / 169#14463983v160, HMEC, HT-29, Jurkat, JY cells, K562 cells, Ku812, KCL22, KG1, KYO1, LNCap, Ma- Mel 1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCK II, MDCK 11, MOR / 0.2R, MONO-MAC 6, MTD-1A, MyEnd, NCI-H69 / CPR, NCI- H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell line, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic varieties thereof. Cell lines are available from a variety of sources known to those with skill in the art (see, e.g., the American Type Culture Collection (ATCC) (Manassus, Va.)). In some embodiments, a cell transfected with one or more vectors described herein is used to establish a new cell line comprising one or more vector-derived sequences. In some embodiments, a cell transiently transfected with the components of a CRISPR system as described herein (such as by transient transfection of one or more vectors, or transfection with RNA), and modified through the activity of a CRISPR complex, is used to establish a new cell line comprising cells containing the modification but lacking any other exogenous sequence. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells are used in assessing one or more test compounds.
[0254] In some embodiments, a cell comprises one or more of any one of a first fusion proteins, second fusion proteins, third fusion proteins, compositions, complexes, polynucleotides, vectors, or pharmaceutical compositions disclosed herein, or any combination thereof.
[0255] In some embodiments, a kit comprises one or more of any one of a first fusion proteins, second fusion proteins, third fusion proteins, compositions, complexes, polynucleotides, vectors, pharmaceutical compositions, or cells disclosed herein, or any combination thereof. Methods and uses
[0256] The present disclosure also provides methods of using the fusion proteins, complexes, systems, compositions, polynucleotides, vectors, cells, and kits disclosed herein. For example, in some embodiments, the methods relate to mutating one or more nucleotides in a target nucleic acid. In some embodiments, the methods comprise contacting a cell, or a target nucleic acid, with any one of the first fusion proteins comprising a C-terminal portion of a split deaminase and a first epitope binding domain, a second fusion proteins comprising a N-terminal portion of a split deaminase and a second epitope binding domain, and a third 147 / 169#14463983v1fusion proteins comprising a DNA binding protein domain and one or more epitope domains disclosed herein. In some embodiments, the contacting results in the binding of the first epitope binding domain of the first fusion protein to a first epitope domain of the third fusion protein and the binding of a second epitope binding domain of the second fusion protein to a second epitope domain of the third fusion protein. Upon binding of the first and second epitope binding domains to the first and second epitope domains, respectively, the N-terminal portion of the split deaminase and the C-terminal portion of the split deaminase form a functional deaminase. In some embodiments, the DNA binding protein directs the functional deaminase to install one or more mutations at one or more target sites of the target double stranded nucleic acid. In some embodiments, the one or more target sites is in a coding region and / or a non-coding region of the target gene.
[0257] In some embodiments, the methods relate to mutagenesis screens at non- coding (or coding) elements of a locus of interest. For example, in some embodiments, the methods relate to performing a mutagenesis screen at non-coding elements of the gene locus (e.g. CD28). In some embodiments, the methods use libraries of gRNAs to tile across a target gene region (e.g. to tile across the entire gene locus). In some embodiments, the methods the libraries comprise between 50 and 500 gRNAs. In some embodiments, the base editing window per locus is between 4 bps and 450 bps. In some embodiments, the base editing window per locus is about 200 bps.
[0258] In some embodiments, the methods comprise contacting a cell, or a target nucleic acid, with any one of the compositions comprising the first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain, a second fusion protein comprising a N-terminal portion of a split deaminase and a second epitope binding domain, and a third fusion proteins comprising a DNA binding protein domain and one or more epitope domains disclosed herein. In some embodiments, the contacting results in the binding of the first epitope binding domain of the first fusion protein to a first epitope domain of the third fusion protein and the binding of a second epitope binding domain of the second fusion protein to a second epitope domain of the third fusion protein. Upon binding of the first and second epitope binding domains to the first and second epitope domains, respectively, the N-terminal portion of the split deaminase and the C-terminal portion of the split deaminase form a functional deaminase. In some embodiments, the DNA binding protein directs the functional deaminase to install one or more mutations at one or more target sites of the target double stranded nucleic acid. In some embodiments, the one or more target sites is in a coding region and / or a non-coding region of the target gene. 148 / 169#14463983v1
[0259] In some embodiments, the methods comprise contacting a cell, or a target nucleic acid, with any one of the complexes comprising a first fusion protein, a second fusion protein, and a third fusion protein disclosed herein. In some embodiments, the contacting results in the binding of the first epitope binding domain of the first fusion protein to a first epitope domain of the third fusion protein and the binding of a second epitope binding domain of the second fusion protein to a second epitope domain of the third fusion protein. Upon binding of the first and second epitope binding domains to the first and second epitope domains, respectively, the N-terminal portion of the split deaminase and the C-terminal portion of the split deaminase form a functional deaminase. In some embodiments, the DNA binding protein directs the functional deaminase to install one or more mutations at one or more target sites of the target double stranded nucleic acid. In some embodiments, the one or more target sites is in a coding region and / or a non-coding region of the target gene.
[0260] In some embodiments, the methods comprise contacting a cell, or a target nucleic acid, with one or more of any one of the polynucleotides, vectors, pharmaceutical compositions, cells, and / or kits disclosed herein. In some embodiments, the contacting results in the binding of the first epitope binding domain of the first fusion protein to a first epitope domain of the third fusion protein and the binding of a second epitope binding domain of the second fusion protein to a second epitope domain of the third fusion protein. Upon binding of the first and second epitope binding domains to the first and second epitope domains, respectively, the N-terminal portion of the split deaminase and the C-terminal portion of the split deaminase form a functional deaminase. In some embodiments, the DNA binding protein directs the functional deaminase to install one or more mutations at one or more target sites of the target double stranded nucleic acid. In some embodiments, the one or more target sites is in a coding region and / or a non-coding region of the target gene.
[0261] In other embodiments, the methods relate to performing a mutational screen in a target nucleic acid. In some embodiments, the methods comprise contacting the target nucleic acid, or a cell comprising said target nucleic acid, with any one of the first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain, a second fusion protein comprising a N-terminal portion of a split deaminase and a second epitope binding domain, and a third fusion protein comprising a DNA binding protein domain and one or more epitope domains disclosed herein. In some embodiments, the methods further comprise applying a selection condition (e.g., addition of antibody such as Trametinib) to the cells and determining the mutations caused by contacting the target nucleic 149 / 169#14463983v1acid with the fusion proteins that enable selection (e.g., cells that are resistant to Trametinib identified).
[0262] In some embodiments, the methods comprise contacting a cell with a first nucleic acid sequence encoding a third fusion protein (e.g., dCas9-GCN4 (10x)) fused to a first fluorescent protein (e.g., blue florescent protein), a second nucleic acid sequence encoding a first fusion protein (e.g., DddA-C-term) fused to a second fluorescent protein (e.g., mCherry), and a third nucleic acid sequence encoding second fusion protein (DddA-N- term), and an sgRNA library. In some embodiments, the third nucleic acid sequence further encodes an antibiotic resistance gene. In some embodiments, the antibiotic resistance gene is puromycin.
[0263] In some embodiments, upon contacting the cells with the first, second, and third nucleic acid sequences, the first fusion protein, second fusion protein, and third fusion proteins are expressed and assemble as described elsewhere herein. The library of sgRNAs is also expressed and each sgRNA guides a third fusion protein, via the DNA binding protein domain, to a different target nucleic acid sequences across the gene of interest (e.g., tiling across the gene of interest).
[0264] In some embodiments, a selection is subsequently applied to the cells. For example, CD28 is highly expressed in Jurkat cells and is involved in T cell activation and cytokine production. The base editors and complexes disclosed herein may be used to introduce mutations throughout non-coding elements in the CD28 locus and said Jurkat cells screened to determine which mutations reduce expression of the CD28. In this way, the skilled artisan can determine the effect non-coding elements have on protein expression, which may represent new drug targets for disease treatment.
[0265] As another example, MEK1 is a protein known to stimulate the RAS-RAF- MEK-ERK pathway in cancer. Trametinib targets MEK1 signaling and is used as a cancer therapy. However, resistance to MEK1 can occur. The base editors and complexes disclosed herein can be used to determine mutations that lead to this resistance. This information can be used to develop improved therapies.
[0266] In some embodiments, the methods comprise contacting the target nucleic acid, or a cell comprising said target nucleic acid, with any one of the compositions comprising the first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain, a second fusion protein comprising a N-terminal portion of a split deaminase and a second epitope binding domain and a third fusion proteins comprising a DNA binding protein domain and one or more epitope domains disclosed herein. In some 150 / 169#14463983v1embodiments, the methods further comprise applying a selection condition to the cells and determining the mutations caused by contacting the target nucleic acid with the fusion proteins that enable selection.
[0267] In some embodiments, the methods comprise contacting the target nucleic acid, or a cell comprising said target nucleic acid, with any one of the complexes comprising a first fusion protein, a second fusion protein, and a third fusion protein disclosed herein. In some embodiments, the methods further comprise applying a selection condition to the cells and determining the mutations caused by contacting the target nucleic acid with the fusion proteins that enable selection.
[0268] In some embodiments, the methods comprise contacting the target nucleic acid, or a cell comprising said target nucleic acid, with one or more of any one of the polynucleotides, vectors, pharmaceutical compositions, cells, and / or kits disclosed herein. In some embodiments, the methods further comprise applying a selection condition to the cells and determining the mutations caused by contacting the target nucleic acid with the fusion proteins that enable selection.
[0269] Any suitable selection condition may be used herein. Exemplary selection conditions are described in detail in Examples 1 and 2 and in Figure 3 and Figure 4.
[0270] The present disclosure further provides for use of any one of the fusion proteins, compositions, complexes, polynucleotides, vectors, pharmaceutical compositions, cells, or kits disclosed herein. For example, in some embodiments, the disclosure provides for the use of any one of the first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain, a second fusion protein comprising a N- terminal portion of a split deaminase and a second epitope binding domain, and a third fusion protein comprising a DNA binding protein domain and one or more epitope domains, as disclosed herein, for mutating one or more nucleotides in a target double stranded nucleic acid.
[0271] In other embodiments, the disclosure provides for the use of any one of the compositions comprising first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain, a second fusion protein comprising a N- terminal portion of a split deaminase and a second epitope binding domain, and a third fusion protein comprising a DNA binding protein domain and one or more epitope domains. as disclosed herein, for mutating one or more nucleotides in a target double stranded nucleic acid. 151 / 169#14463983v1
[0272] In other embodiments, the disclosure provides for the use of any one of the complexes comprising a first fusion protein, a second fusion protein, and a third fusion protein, as disclosed herein, for mutating one or more nucleotides in a target double stranded nucleic acid.
[0273] In other embodiments, the disclosure provides for the use of any one of the polynucleotides, vectors, pharmaceutical compositions, cells, and / or kits, as disclosed herein, for mutating one or more nucleotides in a target double stranded nucleic acid and / or performing a mutagenesis assay.
[0274] In some embodiments, the fusion proteins, complexes, compositions, polynucleotides, vectors, pharmaceutical compositions, cells, kits, and methods provided herein provide wider editing windows than traditional base editors. For example, while traditional base editors have a genomic editing window of about 4 bps, the base editors contemplated herein enable editing across a much wider ( ~180 bp – 450 bp) genomic window. Thus, in some embodiments, the base editors disclosed herein enable C>T (or G>A edits) edits on double stranded DNA across a genomic window of greater than or equal to 4 bp, greater than or equal to 10 bp, greater than or equal to 20 bp, greater than or equal to 40 bp, greater than or equal to 60 bps, greater than or equal to 80 bps, greater than or equal to 100 bps, greater than or equal to 120 bps, greater than or equal to 140 bps, greater than or equal to 160 bps, or greater than or equal to 180 bps. In other embodiments, the base editors disclosed herein enable C>T (or G>A edits) edits on double stranded DNA across a genomic window of less than or equal to 180 bps, less than or equal to 160 bps, less than or equal to 140 bps, less than or equal to 120 bps, less than or equal to 100 bps, less than or equal to 80 bps, less than or equal to 60 bps, less than or equal to 40 bps, less than or equal to 20 bps, less than or equal to 10 bps, or less than or equal to 4 bps. Combinations of the above recited ranges are also possible in other embodiments. For example, in some embodiments, the base editors disclosed herein enable C>T (or G>A edits) edits on double stranded DNA across a genomic window of greater than or equal to 4 bps and less than or equal to 180 bps.
[0275] In other embodiments, the base editors disclosed herein enable C>T (or G>A edits) edits on double stranded DNA across a genomic window of between 4 bps and 180 bps, of between 4 bps and 200 bps, of between 4 bps and 225 bps, of between 4 bps and 250 bps, of between 4 bps and 300 bps.
[0276] In some embodiments, the base editors contemplated herein enable editing across a much wider ( ~5 bp – 450 bp) genomic window. In some embodiments, base editors enable C>T (or G>A) edits on double-stranded DNA across a genomic window greater than 152 / 169#14463983v1or equal to 5 bps, greater than or equal to 55 bps, greater than or equal to 105 bps, greater than or equal to 155 bps, greater than or equal to 205 bps, greater than or equal to 255 bps, greater than or equal to 305 bps, greater than or equal to 355 bps, or greater than or equal to 405 bps. In some embodiments, base editors enable C>T (or G>A) edits on double-stranded DNA across a genomic window of less than or equal to 450 bps, less than or equal to 400 bps, less than or equal to 350 bps, less than or equal to 300 bps, less than or equal to 250 bps, less than or equal to 200 bps, less than or equal to 150 bps, less than or equal to 100 bps, or less than or equal to 50 bps. Combinations are also possible in some embodiments. For example, in some embodiments, base editors enable C>T (or G>A) edits on double stranded DNA across a genomic window greater than or equal to 5 bps and less than or equal to 450 bps. Combinations of other ranges are also possible (e.g., greater than or equal to 5 bps and less than or equal to 450 bps). Other ranges are also possible. EXAMPLES
[0277] Example 1. An exemplary DddA system for wide range base editing.
[0278] Base editor components and configurations tested in this example were delivered to K562 cells using lentiviral vectors (shown in Figure 3B). Cells were treated with 3 vectors. One of the three vectors encoded a dCas9-10xGCN4-BFP construct. A second vector encoded the N-terminal domain of the split DddA deaminase, and the third vector included the C-terminal domain of the split DddA deaminase. The N-terminal domain further comprised an sgRNA targeting the VEGFA locus. Figures 3C and 3D illustrate the editing efficiencies for the various combinations and configurations tested. UN+CU+sgRNA and UN+UC+sgRNA exhibited the greatest editing efficiencies and the widest base editing windows. No editing was observed in the sgRNA binding region, suggesting that single stranded DNA was not being editing.
[0279] Example 2. Demonstration of the tunability of exemplary DddA system for wide range base editing.
[0280] In Example 1, the dCas9 fusion protein comprised 10 repeats of the GCN4 domain. In this example, the effect of varying the number of repeat epitope domains fused to the dCas9 domain on the base editing window was explored. Lentiviral vectors encoding UN and CU (Figure 4B) were delivered to Jurkat cells in addition to vectors encoding dCas9 constructs comprising 1x, 2x, 4x, 5x, or 10x GCN4 repeats. As shown in Figure 4C, constructs having fewer repeat epitope domains had smaller base editing windows relative to 153 / 169#14463983v1constructs having more repeat epitope domains. In some cases, increasing the number of epitope domains increased the editing efficiency at a specific base pair position, however this was not true for all edited positions.
[0281] Example 3. Evaluation of edited loci using PacBio long-read sequencing.
[0282] As shown in Figure 5, edited loci was analyzed using PacBio long-read sequencing. PCR1 was used to capture 0.8-3.5 kb genomic loci of interest. A nested PCR (PCR2) was then used to tag primers to create dual-barcoded amplicons primed for Kinnex library preparation based on a MAS-ISO-seq methods described in Al’Khafaji et al., Nat. Biotech 2023. During Kinnex library prep, amplicons were stitched together to form 18-20 kb Revio-compatible library molecules. After sequencing, bam files were demultiplexed with custom code and analyzed with CRISPRLungo.
[0283] Figures 6A and 6B show the analysis described in Figure 5 applied to Jurkat cells base-edited with a split DddA11 GCN4x4 system, as described elsewhere herein, at the VEGFA locus with various sgRNAs. Cells were sorted 5 days post-lentiviral transduction and harvested 5 days post-sort. Edited genomic DNA was analyzed by PacBio sequencing of a 2.6 kb region. Edited cytosines that meet a threshold of at least 3% editing are displayed in Figure 6A; all cytosines and guanines within the sgRNA binding regions are also included. The numbering of cytosines and guanines in the reference are based on their position in the sequencing amplicon. The frequency of cytosine editing with a 5′ neighboring thymine is included for each guide and ranged between 88% and 100% (right column).
[0284] The data show that, in some cases, editing occurs as far as 450 bp away from the end of the sgRNA (PAM-proximal end) (e.g., see sgRNA ID# sg1). REFERENCES
[0285] Fiaz et al. Int. J. Mol. Sci.2021, 22, 5585.
[0286] Komor et al., Nature 2016, 533, 420-424.
[0287] Lue and Lia., Molecular Cell 2023, 83, 2167-2187.
[0288] Mok et al., Nature 2020, 583, 631-637.
[0289] Tanenbaum et al., Cell 2014, 159, 635-646.
[0290] Vaisvila et al., Mol. Cell 2024, 84, 854-866.
[0291] He et al., Biorxiv 2024
[0292] Iyer et al., NAR 2011, 39, 9473-9497.
[0293] Cho et al., Biorxiv 2024 154 / 169#14463983v1
[0294] Kinney et al., Annu Rev Genomics Hum Genet. Annual Reviews; 2019, 20, 99- 127. EQUIVALENTS AND SCOPE
[0295] In the claims articles such as “a,” “an,” and “the” may mean one or more than one unless indicated to the contrary or otherwise evident from the context. Claims or descriptions that include “or” between one or more members of a group are considered satisfied if one, more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process unless indicated to the contrary or otherwise evident from the context. The invention includes embodiments in which exactly one member of the group is present in, employed in, or otherwise relevant to a given product or process. The invention includes embodiments in which more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process.
[0296] Furthermore, the invention encompasses all variations, combinations, and permutations in which one or more limitations, elements, clauses, and descriptive terms from one or more of the listed claims is introduced into another claim. For example, any claim that is dependent on another claim may be modified to include one or more limitations found in any other claim that is dependent on the same base claim. Where elements are presented as lists, e.g., in Markush group format, each subgroup of the elements is also disclosed, and any element(s) may be removed from the group. It should it be understood that, in general, where the invention, or aspects of the invention, is / are referred to as comprising particular elements and / or features, certain embodiments of the invention or aspects of the invention consist, or consist essentially of, such elements and / or features. For purposes of simplicity, those embodiments have not been specifically set forth in haec verba herein. It is also noted that the terms “comprising” and “containing” are intended to be open and permits the inclusion of additional elements or steps. Where ranges are given, endpoints are included. Furthermore, unless otherwise indicated or otherwise evident from the context and understanding of one of ordinary skill in the art, values that are expressed as ranges can assume any specific value or sub–range within the stated ranges in different embodiments of the invention, to the tenth of the unit of the lower limit of the range, unless the context clearly dictates otherwise.
[0297] This application refers to various issued patents, published patent applications, journal articles, and other publications, all of which are incorporated herein by reference. If there is a conflict between any of the incorporated references and the instant specification, the 155 / 169#14463983v1specification shall control. In addition, any particular embodiment of the present invention that falls within the prior art may be explicitly excluded from any one or more of the claims. Because such embodiments are deemed to be known to one of ordinary skill in the art, they may be excluded even if the exclusion is not set forth explicitly herein. Any particular embodiment of the invention may be excluded from any claim, for any reason, whether or not related to the existence of prior art.
[0298] Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific embodiments described herein. The scope of the present embodiments described herein is not intended to be limited to the above Description, but rather is as set forth in the appended claims. Those of ordinary skill in the art will appreciate that various changes and modifications to this description may be made without departing from the spirit or scope of the present invention, as defined in the following claims. 156 / 169#14463983v1
Claims
CLAIMS What is claimed is:
1. A first fusion protein, comprising a C-terminal portion of a split deaminase and a first epitope binding domain.
2. The first fusion protein of claim 1, wherein the C-terminal portion of the split deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 1-181.
3. The first fusion protein of claim 1, wherein the C-terminal portion of the split deaminase comprises an amino acid sequence that is that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 1322-1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO:
1.
4. The first fusion protein of claim 1, wherein the C-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% identical to a portion of any one of SEQ ID NOs: 2-193 that correspond to amino acids 1322-1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO:1, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 2-193 that corresponds to amino acids 1322-1426, 1333-1426, 1343-1426, 1357-1426, 1371-1426, 1387-1426, or 1397-1426 of SEQ ID NO:
1.
5. The first fusion protein of claim 1, wherein the C-terminal portion of the split deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71- 73,90,95,108,142,145,151,159-160,172,174,176,178-181,194-195,198-206,209-234,236- 157 / 169#14463983v1243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345-351,354- 364,366,368,370,375-380,382-386.
6. The first fusion protein of claim 1, wherein the C-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 34-138, 45-138, 55-138, 69-138, 83-138, 99-138, or 109-138 of SEQ ID NO: 194 or SEQ ID NO :
201.
7. The first fusion protein of claim 1, wherein the C-terminal portion of the split deaminase comprises an amino acid sequence that is at least sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71- 73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that correspond to amino acids 34-138, 45-138, 55-138, 69-138, 83-138, 99-138, or 109-138 of SEQ ID NO:194 or SEQ ID NO:201, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 3-4,14- 15,42,51,55,64,68-69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195- 195,198-200, 202,209-234,236-243,245-247,249-256,258-260,263,267-282,284-287,289- 300,302-337,339-343,345-351,354-364,366,368,370,375-380,382-386 that structurally corresponds to amino acids 34-138, 45-138, 55-138, 69-138, 83-138, 99-138, or 109-138 of SEQ ID NO: 194 or SEQ ID NO:
201.
8. The first fusion protein of any one of claims 1-7, wherein the first epitope binding domain comprises an antibody, an antibody fragment, a peptide, or a nanobody.
9. The first fusion protein of any one of claims 1-8, wherein the first epitope binding domain is configured to bind to a first target epitope domain.
10. The first fusion protein of claim 9, wherein the first epitope binding domain comprises a bivalent scFv antibody fragment. 158 / 169#14463983v111. The first fusion protein of claim 9, wherein the first target epitope domain comprises a multimeric GCN4 peptide comprising an amino acid sequence EELLSKNYHLENEVARLKK (SEQ ID NO: 388).
12. The first fusion protein of any one of claims 1-11, further comprising an uracil glycosylase inhibitor (UGI) domain.
13. The first fusion protein of claim 12, wherein the UGI domain is positioned adjacent to and on either side of the deaminase domain.
14. The first fusion protein of any one of claims 1-13, further comprising chromatin modifying enzyme domain.
15. The first fusion protein of claim 14, wherein the chromatin modifying enzyme domain is selected from the group consisting of acetylase, deacetylase, methyltransferase, demethylase, ligase, helicase, ubiquitin, deubiquitinase, and chromatin remodeler.
16. A second fusion protein, comprising a N-terminal portion of the split deaminase and a second epitope binding domain.
17. The second fusion protein of claim 16, wherein the N-terminal portion of the split deaminase comprises an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to a portion of an amino acid sequence of any one of SEQ ID NOs: 1-181.
18. The second fusion protein of claim 16, wherein the N-terminal portion of the split deaminase comprises an amino acid sequence that is that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 1-1323, 1-1334, 1-1344, 1-1358, 1- 1372, 1-1388, or 1-1398 of SEQ ID NO:
1.
19. The second fusion protein of claim 16, wherein the N-terminal portion of the split deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% identical to a portion 159 / 169#14463983v1of any one of SEQ ID NOs: 2-193 that correspond to amino acids 1-1321, 1-1332, 1-1342, 1- 1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO: 1, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 2-193 that structurally corresponds to amino acids 1-1321, 1-1332, 1-1342, 1-1356, 1-1370, 1-1386, or 1-1396 of SEQ ID NO:
1.
20. The second fusion protein of claim 16, wherein the N-terminal portion of the split deaminase comprises an amino acid sequence that is that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or 100% identical to the amino acids corresponding to positions 1-33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO: 194 or SEQ ID NO:
201.
21. The second fusion protein of claim 16, wherein the N-terminal portion of the split deaminase comprises an amino acid sequence that is at least sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68-69,71- 73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that correspond to amino acids 1-33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO: 194 or SEQ ID NO: 201, or an amino acid sequence that is identical to a portion of any one of SEQ ID NOs: 3-4,14-15,42,51,55,64,68- 69,71-73,90,95,108,142,145,151,159-160,172,174,176,178-181,195-195,198-200, 202,209- 234,236-243,245-247,249-256,258-260,263,267-282,284-287,289-300,302-337,339-343,345- 351,354-364,366,368,370,375-380,382-386 that structurally corresponds to amino acids 1- 33, 1-44, 1-54, 1-68, 1-82, 1-98, or 1-108 of SEQ ID NO: 194 or SEQ ID NO:
201.
22. The second fusion protein of any one of claims 16-21, wherein the second epitope binding domain comprises an antibody, an antibody fragment, a peptide, or a nanobody.
23. The second fusion protein of any one of claims 16-22, wherein the second epitope binding domain is configured to bind to the second target epitope and comprises a bivalent scFv antibody fragment. 160 / 169#14463983v124. The second fusion protein of claim 23, wherein the second target epitope domain comprises a multimeric GCN4 peptide comprising an amino acid sequence EELLSKNYHLENEVARLKK (SEQ ID NO: 388).
25. The second fusion protein of any one of claims 16-24, further comprising an uracil glycosylase inhibitor domain.
26. The second fusion protein of claim 25, wherein the UGI domain is positioned adjacent to and on either side of the deaminase domain.
27. The second fusion protein of any one of claims 16-26, further comprising chromatin modifying enzyme domain.
28. The second fusion protein of claim 27, wherein the chromatin modifying enzyme domain is selected from the group consisting of acetylase, deacetylase, methyltransferase, demethylase, ligase, helicase, ubiquitanse, deubiquitinase, and chromatin remodeler.
29. A third fusion protein comprising a DNA binding protein domain and one or more epitope domains.
30. The third fusion protein of claim 29, wherein the DNA binding protein domain comprises a programmable DNA binding protein, transcription factor, or chromatin binding protein.
31. The third fusion protein of claim 30, wherein the programmable DNA binding protein comprises dead Cas9 (dCas9) or nickase Cas9 (nCas9).
32. The third fusion protein of any one of claims 29-31, wherein the one or more epitope domains is selected from the group consisting of streptavidin, EELLSKNYHLENEVARLKK (SEQ ID NO: 388) (GCN4 peptide), SRLEEELRRRLTE (SEQ ID NO: 389) (ALFA peptide), PDRKAAVSHWQQ (SEQ ID NO: 390) (BC2 peptide), PDRVRAVSHWSS (SEQ ID NO: 391) (Spot peptide), AVERYLKDQQLLGIW (SEQ ID NO: 392) (PepTag peptide) and KNEQELLELDKWASL (SEQ ID NO: 393) (gp41 peptide). 161 / 169#14463983v133. The third fusion protein of any one of claims 29-32, wherein the third fusion protein comprises one or more repeats of the one or epitope domains.
34. The third fusion protein of any one of claims 29-33, wherein the third fusion protein comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten repeats of the one or more epitope domains.
35. The third fusion protein of any one of claims 29-34, further comprising a linker configured to tether one or more domains of the third fusion protein.
36. A composition comprising (i) the first fusion protein of any one of claims 1-15, (ii) the second fusion protein of any one of claims 16-28, and / or (iii) the third fusion protein of any one of claims 29-35.
37. A complex comprising: (i) a first fusion protein comprising a C-terminal portion of a split deaminase and a first epitope binding domain; (ii) a second fusion protein comprising a N-terminal portion of the split deaminase and a second epitope binding domain; and (iii) a third fusion protein comprising a DNA binding protein and one or more epitope domains.
38. The complex of claim 37, wherein the first epitope binding domain and the second epitope binding domain bind to the same epitope domain.
39. The complex of claim 37 or 38, wherein the first epitope domain and the second epitope domain bind to different epitope domains.
40. The complex of any one of claims 37-39, wherein the first epitope binding domain of the first fusion protein is configured to bind to a first epitope domain of the third fusion protein. 162 / 169#14463983v141. The complex of any one of claims 37-40, wherein the second epitope binding domain of the second fusion protein is configured to bind a second epitope domain of the third fusion protein.
42. The complex of claim 41, wherein upon binding of the first and second fusion proteins to the third fusion protein, the N-terminal portion of the split deaminase and the C- terminal portion of the split deaminase form a functional deaminase.
43. A polynucleotide encoding the first fusion protein of any one of claims 1-15.
44. A polynucleotide encoding the second fusion protein of any one of claims 16-28.
45. A polynucleotide encoding the third fusion protein of any one of clams 29-35.
46. A vector comprising the polynucleotide of claim 43.
47. The vector of claim 46, wherein the vector comprises a viral transfer vector.
48. The vector of claim 47, wherein the viral transfer vector is selected from the group consisting of adenoviruses, adeno-associated viruses, and lentiviruses.
49. The vector of claim 46, wherein the vector comprises a mammalian expression vector.
50. The vector of claim 46, wherein the mammalian expression system is delivered by lipid nanoparticles, liposomes, or polymers.
51. A vector comprising the polynucleotide of claim 44.
52. The vector of claim 51, wherein the vector comprises a viral transfer vector.
53. The vector of claim 52, wherein the viral transfer vector is selected from the group consisting of adenoviruses, adeno-associated viruses, and lentiviruses.
54. The vector of claim 51, wherein the vector comprises a mammalian expression vector. 163 / 169#14463983v155. The vector of claim 54, wherein the mammalian expression vector is delivered by lipid nanoparticles, liposomes, or polymers.
56. A vector comprising the polynucleotide of claim 45.
57. The vector of claim 56, wherein the vector comprises a viral transfer vector.
58. The vector of claim 57, wherein the viral transfer vector is selected from the group consisting of adenoviruses, adeno-associated viruses, and lentiviruses.
59. The vector of claim 56, wherein the vector comprises a mammalian expression vector.
60. The vector of claim 59, wherein the mammalian expression system is delivered by lipid nanoparticles, liposomes, or polymers.
61. A pharmaceutical composition comprising the first fusion protein of any one of claims 1-15, the second fusion protein of any one of claims 16-28, the third fusion protein of any one of claims 29-35, the composition of claim 36, the complex of any one of claims 37-42, the polynucleotide of any one of claims 43-45, the vector of any one of claims 46-50, the vector of any one of claims 51-55, the vector of any one of claims 56-60, or any combination thereof.
62. The pharmaceutical composition of claim 61, further comprising a lentiviral packaging vector.
63. The pharmaceutical composition of claim 61 or 62, further comprising an envelope vector.
64. The pharmaceutical composition of any one of claims 61-63, further comprising a pharmaceutically acceptable excipient.
65. A cell comprising the first fusion protein of any one of claims 1-15, the second fusion protein of any one of claims 16-28, the third fusion protein of any one of claims 29-35, the 164 / 169#14463983v1composition of claim 36, the complex of any one of claims 37-42, the polynucleotide of any one of claims 43-45, the vector of any one of claims 46-50, the vector of any one of claims 51-55, the vector of any one of claims 56-60, the pharmaceutical composition of any one of claims 61-64, or any combination thereof.
66. A kit comprising the first fusion protein of any one of claims 1-15, the second fusion protein of any one of claims 16-28, the third fusion protein of any one of claims 29-35, the composition of claim 36, the complex of any one of claims 37-42, the polynucleotide of any one of claims 43-45, the vector of any one of claims 46-50, the vector of any one of claims 51-55, the vector of any one of claims 56-60, the pharmaceutical composition of any one of claims 61-64, the cell of claim 65, or any combination thereof.
67. A method of mutating one or more nucleotides in a target nucleic acid, the method comprising contacting a cell, or the target nucleic acid, with: (i) the first fusion protein of any one of claims 1-15, the second fusion protein of claims 16-28, and the third fusion protein of claims 29-35, or (ii) the composition of claim 36, or (iii) the complex of any one of claims 37-42, or (iv) the polynucleotide of claim 43, the polynucleotide of claim 44, and the polynucleotide of claim 45, or (v) the vector of any one of claims 46-50, the vector of any one of claims 51-55, the vector of any one of claims 56-60; or (vi) the pharmaceutical composition of any one of claims 61-64, or (vii) the cell of claim 65, or (viii) the kit of claim 66, wherein the contacting results in the binding of the first epitope binding domain of the first fusion protein to the first epitope domain of the third fusion protein and the binding of the second epitope binding domain of the second fusion protein to the second epitope domain of the third fusion protein, wherein upon binding the N-terminal portion of the split deaminase and the C-terminal portion of the split deaminase form a functional deaminase; wherein the DNA binding protein directs the functional deaminase to install one or more mutations at one or more target sites of the target nucleic acid, and wherein the one or more target sites is in a coding region and / or a non-coding region of the target gene. 165 / 169#14463983v168. A method of performing a mutational screen in a target nucleic acid, comprising: (1) contacting the target nucleic acid, or a cell comprising said target nucleic acid, with: (i) the first fusion protein of any one of claims 1-15, the second fusion protein of claims 16-28, and the third fusion protein of claims 29-35, or (ii) the composition of claim 36, or (iii) the complex of any one of claims 37-42, or (iv) the polynucleotide of claim 43, the polynucleotide of claim 44, and the polynucleotide of claim 45, or (v) the vector of any one of claims 46-50, the vector of any one of claims 51-55, the vector of any one of claims 56-60; or (vi) the pharmaceutical composition of any one of claims 61-64, or (vii) the cell of claim 65, or (viii) the kit of claim 66, (2) applying a selection condition to the cells; and (3) determining the mutations caused by step (i) that enable selection.
69. Use of the first fusion protein of any one of claims 1-15, the second fusion protein of any one of claims 16-28, the third fusion protein of any one of claims 29-35 for mutating one or more nucleotides in a target double stranded nucleic acid.
70. Use of the composition of claim 36 or the complex of any one of claims 37-42 for mutating one or more nucleotides in a target double stranded nucleic acid.
71. Use of the polynucleotide of any one of claims 43-45, the vector of any one of claims 46-50, the vector of any one of claims 51-55, the vector of any one of claims 56-60, the pharmaceutical composition of any one of claims 61-64, the cell of claim 65, or the kit of claim 66, for mutating one or more nucleotides in a target double stranded nucleic acid.
72. Use of the first fusion protein of any one of claims 1-15, the second fusion protein of any one of claims 16-28, the third fusion protein of any one of claims 29-35 for performing a mutational screen in a target nucleic acid. 166 / 169#14463983v173. Use of the composition of claim 36 or the complex of any one of claims 37-42 for performing a mutational screen in a target nucleic acid.
74. Use of the polynucleotide of any one of claims 43-45, the vector of any one of claims 46-50, the vector of any one of claims 51-55, the vector of any one of claims 56-60, the pharmaceutical composition of any one of claims 61-64, the cell of claim 65, or the kit of claim 66, for performing a mutational screen in a target nucleic acid.
75. The composition of claims 29, further comprising a gRNA.
76. The complex of any one of claims 37-42, further comprising a gRNA.
77. The method of claim 68 or 69, wherein the target nucleic acid encodes a transcription factor or a cis-regulatory element.
78. The method of claim 77, wherein the cis-regulator element comprises a promoter, an enhancer, and an insulator.
79. A method of performing a mutational screen in a target nucleic acid, comprising contacting the target nucleic acid, or a cell comprising the target nucleic acid, with the first fusion proteins of any one of claims 1-15, the second fusion proteins of any one of claims 16- 28, and the third fusion proteins of any one of claims 29-35.
80. The method of claim 79, further comprising contacting the target nucleic acid, or cell comprising the target nucleic acid with a sgRNA.
81. The method of claim 79 or 80, wherein the target nucleic acid comprises a coding region or a non-coding region of DNA.
82. The method of anyone of claims 79-81, wherein the third fusion protein further comprises a biomarker, optionally, wherein the biomarker comprises a fluorescent maker, optionally, wherein the fluorescent biomarker comprises blue fluorescent protein. 167 / 169#14463983v183. The method of anyone of claims 79-82, wherein the first and / or second proteins further comprises a biomarker, optionally, wherein the biomarker comprises a fluorescent maker, optionally, wherein the fluorescent biomarker comprises mCherry.
84. A method of performing a mutational screen in a target nucleic acid, comprising contacting the target nucleic acid, or a cell comprising the target nucleic acid, with the polynucleotides of claims 43-45.
85. The method of claim 84, wherein at least one polynucleotide further encodes a sgRNA library.
86. The method of claim 84 or 85, wherein the target nucleic acid comprises a coding region or a non-coding region of DNA.
87. The method of any one of claims 84-86, wherein the polynucleotide encoding a third fusion protein further encodes for a fluorescent protein, optionally, wherein the fluorescent protein comprises blue fluorescent protein.
88. The method of anyone of claims 84-87, wherein the polynucleotide encoding a first fusion protein further encodes a fluorescent protein, optionally, wherein the fluorescent protein comprises mCherry.
89. The method of anyone of claims 84-87, wherein the polynucleotide encoding a second fusion protein further encodes a fluorescent protein, optionally, wherein the fluorescent protein comprises mCherry. 168 / 169#14463983v1
Citation Information
Patent Citations
Double-Stranded DNA Deaminases and Uses Thereof
US20230357838A1
Base editors, compositions, and methods for modifying the mitochondrial genome
WO2021155065A1
Evolved double-stranded DNA deaminase base editors and methods of use
WO2022221337A2