Adenosine nucleic acid base editing factor and use of the same
Patent Information
- Application Number
- JP2025077176
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-03-20
- Filing Date
- 2025-05-07
- Publication Date
- 2026-03-05
AI Technical Summary
Existing adenosine deaminases are not capable of efficiently modifying DNA sequences, limiting precise gene editing for therapeutic applications.
Development of engineered adenosine deaminases, such as E. coli TadA variants, fused with Cas9 domains to create nucleobase editors that can catalyze A-to-G mutations in DNA by forming inosine, which pairs with C during replication.
Enables targeted and precise editing of genetic sequences, including single nucleotide polymorphisms in disease-associated genes, by converting A to G, thereby addressing the limitations of existing DNA editing technologies.
Abstract
Description
[Technical Field]
[0001] Background of the Invention Targeted editing of nucleic acid sequences (e.g., targeted cleavage or targeted introduction of specific modifications into genomic DNA) is an extremely promising approach for studying gene function and also holds the promise of providing new therapies for human genetic diseases. While many genetic diseases could in principle be treated by making specific nucleotide changes at specific locations in the genome (e.g., an A→G change or a T→C change at a specific codon for a gene-associated disease), the development of programmable methods for achieving such precise gene editing represents both a powerful new research tool and a promising new approach to gene editing-based therapeutics.
[0002] SUMMARY OF THE INVENTION Provided herein are compositions, kits, and methods for modifying polynucleotides (e.g., DNA) using adenosine deaminase and a nucleic acid-programmable DNA-binding protein (e.g., Cas9). Some aspects of the present disclosure provide nucleic acid base-editing proteins that, in the context of DNA, catalyze the hydrolytic deamination of adenosine (to form inosine, which base-pairs with guanine (G), etc.). No naturally occurring adenosine deaminases are known to act on DNA. Instead, known adenosine deaminases act on RNA (e.g., tRNA or mRNA). To overcome this difficulty, the first deoxyadenosine deaminases were evolved to accept DNA substrates and deaminate deoxyadenosine (dA) to deoxyinosine. An adenosine deaminase (ADC) from Escherichia coli that acts on tRNA (e.g., Cas9) was identified. a denosine d e a Minase acting on t RNA)(ADAT)( t RNA a denosine d e aminase A In honor of TadA, a dCas9 deaminase (TadA) has been covalently fused to a dCas9 domain, and a library of this fusion containing mutations in the deaminase portion of the construct has been constructed. It should be understood that E. coli TadA (ecTadA) deaminase also includes truncations of ecTadA. For example, truncations (e.g., N-terminal truncations) of full-length ecTadA (SEQ ID NO: 84), such as the N-terminally truncated ecTadA set forth in SEQ ID NO: 1, are provided herein for use in the present invention. Furthermore, other adenosine deaminase mutants, such as S. aureus TadA mutants, have also been found to be capable of deaminating adenosine. While not wishing to be bound by any particular theory, truncations of adenosine deaminases (e.g., ecTadA) may have desirable solubility and / or expression properties compared to their full-length counterparts.
[0003] Mutations in the protein-editing nucleobase deaminase domain were made by evolving an adenosine deaminase. Productive variants were identified via chloramphenicol selection for A to G reversion at the active site His codon in the acetyltransferase gene (encoded on a cotransformed selection plasmid). The first round of evolution generated an ecTadA variant, ecTadA D108X (X = G, V, or N), capable of converting A to G in DNA. In some embodiments, the ecTadA variant contains the D108A mutation in SEQ ID NO: 1 or a corresponding mutation in another adenosine deaminase. The first round of evolution also generated an ecTadA variant, ecTadA A106V. Subsequent rounds of evolution yielded another variant, ecTadA D108N_E155X (X = G, V, or D), which allowed E. coli to survive in the presence of high concentrations of chloramphenicol. Additional variants were also identified by evolving ecTadA. For example, an ecTadA variant capable of deaminating adenosine in DNA contains one or more of the following mutations: D108N, A106V, D147, E155V, L84F, H123Y, and I157F of SEQ ID NO: 1. However, it should be understood that homologous mutations can be made in other adenosine deaminases to generate variants capable of deaminating adenosine in DNA. Additional rounds of evolution also provided further ecTadA variants. For example, additional ecTadA variants are shown in Figures 11, 16, 97, 104-106, 125-128, 115 and Table 4.
[0004] In the examples provided herein, an exemplary nucleobase editor with the general structure ecTadA(D108X; X=G, V, or N)-XTEN-nCas9 was evolved to catalyze an A→G transition mutation in cells, such as eukaryotic cells (e.g., Hek293T mammalian cells). In other examples, an exemplary nucleobase editor contains two ecTadA domains and a nucleic acid programmable DNA binding protein (napDNAbp). For example, a nucleobase editor may have the general structure ecTadA(D108N)-ecTadA(D108N)-nCas9. Additional examples of nucleobase editors containing ecTadA variants provided herein demonstrate improved nucleobase editor performance in mammalian cells. For example, one adenosine base editor comprises ecTadA with a D108X (where X = G, V, or N) and / or E155X (where X = B, V, or D) mutation in ecTadA set forth in SEQ ID NO: 1 or another adenine deaminase. In some embodiments, the mutant nucleobase editor is covalently fused to a catalytically dead alkyladenosine glycosylase (AAG), which can protect the edited inosine from base excision repair (or other DNA repair systems) until the T on the opposite strand is changed to a C, e.g., through mismatch repair (or other DNA repair system). Once the base opposite the inosine is changed to a C, the inosine can then be irreversibly and permanently changed to G through the cell's DNA repair processes, resulting in a permanent change from an A:T base pair to a G:C base pair.
[0005] Without wishing to be bound by any particular theory, the adenosine nucleobase editors described herein operate by using ecTadA variants to deaminate A bases in DNA, resulting in A-to-G mutations via inosine formation. Inosine preferentially hydrogen bonds with C, resulting in A-to-G mutations during DNA replication. When covalently tethered to Cas9 (or another nucleic acid-programmable DNA-binding protein), adenosine deaminase (e.g., ecTadA) localizes to a gene of interest and catalyzes A-to-G mutations in ssDNA substrates. This editor can be used to target and revert single nucleotide polymorphisms (SNPs) in disease-associated genes that require A-to-G reversion. This editor can also be used to target and revert single nucleotide polymorphisms (SNPs) in disease-associated genes that require T-to-C reversion by mutating the A opposite the T to G. The T may then be replaced by a C, for example by base excision repair mechanisms, or may be changed in subsequent rounds of DNA replication.
[0006] Some aspects of the present disclosure relate to the discovery that engineered (e.g., evolved) adenosine deaminases are capable of deaminating adenosine in deoxyribonucleic acid (DNA) substrates. In some embodiments, the present disclosure provides such adenosine deaminases. In some embodiments, the adenosine deaminases provided herein are capable of deaminating adenosine in DNA molecules. Another aspect of the present disclosure provides fusion proteins comprising a Cas9 domain and an adenosine deaminase domain (e.g., an engineered deaminase domain capable of deaminating adenosine in DNA). In some embodiments, the fusion protein comprises one or more of a nuclear localization sequence (NLS), an inhibitor of inosine base excision repair (e.g., dISN), and / or a linker.
[0007] In some aspects, the present disclosure provides an adenosine deaminase capable of deaminating adenosine in a deoxyribonucleic acid (DNA) substrate. In some embodiments, the adenosine deaminase is from a bacterium, such as E. coli or S. aureus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA deaminase (ecTadA). In some embodiments, the adenosine deaminase comprises a D108X mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X is any amino acid other than that found in the wild-type protein. In some embodiments, X is G, N, V, A, or Y.
[0008] In some embodiments, the adenosine deaminase comprises an E155X mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X is any amino acid other than that found in the wild-type protein. In some embodiments, X is D, G, or V. It should be understood that the adenosine deaminase provided herein may contain one or more of the mutations provided herein in any combination.
[0009] Some aspects of the present disclosure provide fusion proteins comprising: (i) a Cas9 domain; and (ii) an adenosine deaminase, such as any of the adenosine deaminases provided herein. In some embodiments, the Cas9 domain of the fusion protein is a nuclease-dead Cas9 (dCas9), a Cas9 nickase (nCas9), or a nuclease-active Cas9. In some embodiments, the fusion protein further comprises an inhibitor of inosine base excision repair, such as dISN or a single-stranded DNA binding protein. In some embodiments, the fusion protein comprises one or more linkers used to attach the adenine deaminase (e.g., ecTadA) to the nucleic acid-programmable DNA binding protein (e.g., Cas9). In some embodiments, the fusion protein comprises one or more nuclear localization sequences (NLS).
[0010] The above summary is intended to describe, without limitation, some of the aspects, advantages, features, and uses of the technology disclosed herein. Other aspects, advantages, features, and uses of the technology disclosed herein will be apparent from the detailed description, drawings, examples, and claims. [Brief explanation of the drawings]
[0011] Brief description of the drawings [Figure 1-1] Figure 1 shows the results of high-throughput screening using various deaminases. APOBEC (BE3) is a positive control; ADAR acts on mRNA, ADA acts on deoxyadenosine, and ADAT acts on tRNA. The untreated group is a negative control. The sequence corresponds to SEQ ID NO: 45. [Figure 1-2] Figure 1 shows the results of high-throughput screening using various deaminases. APOBEC (BE3) is a positive control; ADAR acts on mRNA, ADA acts on deoxyadenosine, and ADAT acts on tRNA. The untreated group is a negative control. The sequence corresponds to SEQ ID NO: 45.
[0012] [Figure 2] FIG. 2 is a schematic diagram of the deamination selection plasmid.
[0013] [Figure 3] FIG. 3 shows serial dilutions of the selection plasmid in S1030 cells plated on increasing concentrations of chloramphenicol.
[0014] [Figure 4-1] Figure 4 shows the validation of chloramphenicol selection using the rAPOBEC1-XTEN-dCas9 construct as a positive control. The sequences, from top to bottom, correspond to SEQ ID NOs: 95 (nucleotide sequence), 96 (amino acid sequence), 97 (nucleotide sequence), 98 (amino acid sequence), 95 (nucleotide sequence), and 99 (truncated nucleotide sequence). [Figure 4-2] Figure 4 shows the validation of chloramphenicol selection using the rAPOBEC1-XTEN-dCas9 construct as a positive control. The sequences, from top to bottom, correspond to SEQ ID NOs: 95 (nucleotide sequence), 96 (amino acid sequence), 97 (nucleotide sequence), 98 (amino acid sequence), 95 (nucleotide sequence), and 99 (truncated nucleotide sequence).
[0015] [Figure 5] Figure 5 is a schematic diagram of the deaminase-XTEN-dCas9 construct.
[0016] [Figure 6] Figure 6 shows the sequencing results from the first round of the TadA-XTEN-dCas9 library.
[0017] [Figure 7-1]Figure 7 shows the sequence of the selected plasmid; an A to G reversion was observed. The sequences correspond, from top to bottom, to SEQ ID NOs: 100 (nucleotide sequence), 101 (amino acid sequence), 102 (nucleotide sequence), 103 (amino acid sequence), 104 (nucleotide sequence), and 100 (nucleotide sequence). [Figure 7-2] Figure 7 shows the sequence of the selected plasmid; an A to G reversion was observed. The sequences correspond, from top to bottom, to SEQ ID NOs: 100 (nucleotide sequence), 101 (amino acid sequence), 102 (nucleotide sequence), 103 (amino acid sequence), 104 (nucleotide sequence), and 100 (nucleotide sequence).
[0018] [Figure 8-1] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-2] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-3] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-4] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-5] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-6]Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-7] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-8] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-9] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-10] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-11] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-12] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-13] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-14]Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-15] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607. [Figure 8-16] Figure 8 shows the results of deaminase sequencing, which illustrates the convergence at residue D108. The sequences, from top to bottom, correspond to SEQ ID NOS: 589-607.
[0019] [Figure 9] Figure 9 shows the E. coli TadA crystal structure. Note that residue numbering has been lost in the figure, so D119 corresponds to D108 in the figure.
[0020] [Figure 10-1] 10 shows the crystal structure of TadA (in S. aureus) tRNA and its alignment with TadA from E. coli. The sequences, from top to bottom, correspond to SEQ ID NOs: 105-107. [Figure 10-2] 10 shows the crystal structure of TadA (in S. aureus) tRNA and its alignment with TadA from E. coli. The sequences, from top to bottom, correspond to SEQ ID NOs: 105-107.
[0021] [Figure 11-1] FIG. 11 shows the isolation and challenge results of individual constructs from the ecTadA evolution. [Figure 11-2] FIG. 11 shows the isolation and challenge results of individual constructs from the ecTadA evolution.
[0022] [Figure 12]Figure 12 shows colony forming units (CFU) of various constructs challenged on increasing concentrations of chloramphenicol. Construct numbers correspond to those listed in Figure 11.
[0023] [Figure 13-1] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-2] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-3] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-4] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-5] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-6] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-7] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-8] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-9]Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-10] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-11] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623. [Figure 13-12] Figure 13 shows data from the second round of evolution from a construct containing the D108N mutation. Sequences, from top to bottom, correspond to SEQ ID NOS: 608-623.
[0024] [Figure 14-1] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 14-2] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 14-3] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 14-4] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 14-5] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 14-6] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 14-7] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 14-8] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 14-9] Figure 14 shows A to G editing in mammalian cells. The sequence corresponds to SEQ ID NO: 41.
[0025] [Figure 15] FIG. 15 is a schematic diagram showing the development of ABE.
[0026] [Figure 16] Figure 16 is a table showing the results of clones assayed after the second round of evolution. Columns 1, 8, and 10 represent mutations from the first round of evolution. Columns 11 and 14 represent consensus mutations from the second round of evolution.
[0027] [Figure 17-1] Figure 17 shows the results of antibiotic challenge assays of individual clones. Construct number identification corresponds to the pNMG clone number from Figure 16. [Figure 17-2] Figure 17 shows the results of antibiotic challenge assays of individual clones. Construct number identification corresponds to the pNMG clone number from Figure 16.
[0028] [Figure 18] Figure 18 shows schematic representations of the new constructs developed, which contain the UGI domain, the AAG*E125Q domain, and the EndoV*D35A domain.
[0029] [Figure 19-1] Figure 19 shows the transfection of constructs containing single or double mutations in ecTadA into mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 19-2] Figure 19 shows the transfection of constructs containing single or double mutations in ecTadA into mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 19-3] Figure 19 shows the transfection of constructs containing single or double mutations in ecTadA into mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 19-4] Figure 19 shows the transfection of constructs containing single or double mutations in ecTadA into mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 19-5] Figure 19 shows the transfection of constructs containing single or double mutations in ecTadA into mammalian cells. The sequence corresponds to SEQ ID NO: 41. [Figure 19-6] Figure 19 shows the transfection of constructs containing single or double mutations in ecTadA into mammalian cells. The sequence corresponds to SEQ ID NO: 41.
[0030] [Figure 20-1] Figure 20 shows transfection of a construct in which UGI was added to adenosine nucleobase editor (ABE) (D108N). The sequence corresponds to SEQ ID NO: 41. [Figure 20-2] Figure 20 shows transfection of a construct in which UGI was added to adenosine nucleobase editor (ABE) (D108N). The sequence corresponds to SEQ ID NO: 41. [Figure 20-3] Figure 20 shows transfection of a construct in which UGI was added to adenosine nucleobase editor (ABE) (D108N). The sequence corresponds to SEQ ID NO: 41. [Figure 20-4] Figure 20 shows transfection of a construct in which UGI was added to adenosine nucleobase editor (ABE) (D108N). The sequence corresponds to SEQ ID NO: 41. [Figure 20-5] Figure 20 shows transfection of a construct in which UGI was added to adenosine nucleobase editor (ABE) (D108N). The sequence corresponds to SEQ ID NO: 41. [Figure 20-6] Figure 20 shows transfection of a construct in which UGI was added to adenosine nucleobase editor (ABE) (D108N). The sequence corresponds to SEQ ID NO: 41.
[0031] [Figure 21-1] Figure 21 shows that ABE works best for one of the six genomic sites tested. The sequence corresponds to SEQ ID NO: 46. [Figure 21-2] Figure 21 shows that ABE works best for one of the six genomic sites tested. The sequence corresponds to SEQ ID NO: 46. [Figure 21-3] Figure 21 shows that ABE works best for one of the six genomic sites tested. The sequence corresponds to SEQ ID NO: 46. [Figure 21-4] Figure 21 shows that ABE works best for one of the six genomic sites tested. The sequence corresponds to SEQ ID NO: 46. [Figure 21-5] Figure 21 shows that ABE works best for one of the six genomic sites tested. The sequence corresponds to SEQ ID NO: 46. [Figure 21-6] Figure 21 shows that ABE works best for one of the six genomic sites tested. The sequence corresponds to SEQ ID NO: 46.
[0032] [Figure 22-1] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42. [Figure 22-2] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42. [Figure 22-3] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42. [Figure 22-4] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42. [Figure 22-5] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42. [Figure 22-6] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42. [Figure 22-7] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42. [Figure 22-8] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42. [Figure 22-9] Figure 22 shows that editing at the Hek-3 site is lower compared to editing at the Hek-2 site, even at position 8 of the protospacer. The sequence corresponds to SEQ ID NO: 42.
[0033] [Figure 23-1] Figure 23 shows the inactive C-terminal Cas9 fusion of ecTadA for constructs pNMG-164 to pNMG-173. The sequence corresponds to SEQ ID NO: 41. [Figure 23-2] Figure 23 shows the inactive C-terminal Cas9 fusion of ecTadA for constructs pNMG-164 to pNMG-173. The sequence corresponds to SEQ ID NO: 41. [Figure 23-3] Figure 23 shows the inactive C-terminal Cas9 fusion of ecTadA for constructs pNMG-164 to pNMG-173. The sequence corresponds to SEQ ID NO: 41. [Figure 23-4] Figure 23 shows the inactive C-terminal Cas9 fusion of ecTadA for constructs pNMG-164 to pNMG-173. The sequence corresponds to SEQ ID NO: 41.
[0034] [Figure 24-1] Figure 24 shows the inactive C-terminal Cas9 fusion of ecTadA for constructs pNMG-174 to pNMG-177. The sequence corresponds to SEQ ID NO: 41. [Figure 24-2] Figure 24 shows the inactive C-terminal Cas9 fusion of ecTadA for constructs pNMG-174 to pNMG-177. The sequence corresponds to SEQ ID NO: 41. [Figure 24-3] Figure 24 shows the inactive C-terminal Cas9 fusion of ecTadA for constructs pNMG-174 to pNMG-177. The sequence corresponds to SEQ ID NO: 41. [Figure 24-4] Figure 24 shows the inactive C-terminal Cas9 fusion of ecTadA for constructs pNMG-174 to pNMG-177. The sequence corresponds to SEQ ID NO: 41.
[0035] [Figure 25-1] Figure 25 shows editing results from the ecTadA nucleobase editor (pNMG-143, pNMG-144, pNMG-164, and pNMG-177). The sequence corresponds to SEQ ID NO: 41. [Figure 25-2] Figure 25 shows editing results from the ecTadA nucleobase editor (pNMG-143, pNMG-144, pNMG-164, and pNMG-177). The sequence corresponds to SEQ ID NO: 41. [Figure 25-3] Figure 25 shows editing results from the ecTadA nucleobase editor (pNMG-143, pNMG-144, pNMG-164, and pNMG-177). The sequence corresponds to SEQ ID NO: 41.
[0036] [Figure 26-1] Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41. [Figure 26-2]Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41. [Figure 26-3] Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41. [Figure 26-4] Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41. [Figure 26-5] Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41. [Figure 26-6] Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41. [Figure 26-7] Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41. [Figure 26-8] Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41. [Figure 26-9] Figure 26 shows editing results from the ecTadA nucleobase editor (pNMG-164, pNMG-177, pNMG-178, pNMG-179, and pNMG-180). The sequence corresponds to SEQ ID NO: 41.
[0037] [Figure 27-1] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 27-2] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 27-3] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 27-4] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 27-5] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 27-6] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 27-7] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 27-8] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 27-9] Figure 27 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42.
[0038] [Figure 28-1] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-2] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-3] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-4] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-5] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-6] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-7] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-8] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-9] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-10] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-11] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-12] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-13] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 28-14] Figure 28 shows the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41.
[0039] [Figure 29-1] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-2] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-3] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-4] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-5] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-6] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-7]Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-8] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-9] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-10] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-11] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 29-12] Figure 29 shows the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42.
[0040] [Figure 30-1] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-2] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-3] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-4] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-5] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-6] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-7] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-8] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-9] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-10] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-11] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 30-12] Figure 30 shows the results of editing at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43.
[0041] [Figure 31-1] Figure 31 shows the results of editing at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 31-2] Figure 31 shows the results of editing at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 31-3] Figure 31 shows the results of editing at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 31-4] Figure 31 shows the results of editing at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 31-5] Figure 31 shows the results of editing at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 31-6] Figure 31 shows the results of editing at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 31-7] Figure 31 shows the results of editing at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 31-8] Figure 31 shows the results of editing at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44.
[0042] [Figure 32-1] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-2] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-3] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-4]Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-5] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-6] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-7] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-8] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-9] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-10] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-11] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 32-12] Figure 32 shows the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45.
[0043] [Figure 33-1] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-2] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-3] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-4] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-5] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-6] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-7] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-8] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-9] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-10] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-11] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 33-12] Figure 33 shows the results of editing at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46.
[0044] [Figure 34-1] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-2] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-3] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-4] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-5] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-6] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-7] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-8] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-9]Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-10] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-11] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-12] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-13] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-14] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-15] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 34-16] Figure 34 shows the results of a C-terminal fusion at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41.
[0045] [Figure 35-1] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-2] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-3] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-4] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-5] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-6] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-7]Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-8] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-9] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-10] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-11] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-12] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-13] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 35-14] Figure 35 shows the results of a C-terminal fusion at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42.
[0046] [Figure 36-1] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-2] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-3] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-4] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-5] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-6] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-7]Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-8] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-9] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-10] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-11] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-12] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-13] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-14] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-15] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 36-16] Figure 36 shows the results of a C-terminal fusion at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43.
[0047] [Figure 37-1] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-2] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-3] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-4] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-5]Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-6] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-7] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-8] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-9] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-10] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-11] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-12] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-13] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-14] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-15] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46. [Figure 37-16] Figure 37 shows the results of a C-terminal fusion at the EMX-1 site. The sequence corresponds to SEQ ID NO:46.
[0048] [Figure 38-1] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-2] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-3]Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-4] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-5] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-6] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-7] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-8] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-9] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-10] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-11] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-12] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-13] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 38-14] Figure 38 shows the results of a C-terminal fusion at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44.
[0049] [Figure 39-1] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-2] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-3]Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-4] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-5] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-6] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-7] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-8] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-9] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-10] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-11] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-12] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-13] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-14] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-15] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45. [Figure 39-16] Figure 39 shows the results of a C-terminal fusion at the FANCF site. The sequence corresponds to SEQ ID NO:45.
[0050] [Figure 40-1]Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-2] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-3] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-4] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-5] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-6] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-7] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-8] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-9] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-10] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-11] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-12] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-13] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-14] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-15] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-16] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-17] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-18] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-19] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-20] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-21] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-22] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-23] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-24] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-25] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-26] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-27] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-28]Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-29] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-30] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-31] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-32] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-33] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-34] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-35] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41. [Figure 40-36] Figure 40 shows the results of transfection at the Hek-2 site. The sequence corresponds to SEQ ID NO: 41.
[0051] [Figure 41-1] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-2] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-3] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-4] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-5]Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-6] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-7] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-8] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-9] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-10] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-11] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-12] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-13] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-14] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-15] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-16] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-17] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-18] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-19] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-20] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-21] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-22] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-23] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-24] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-25] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-26] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-27] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-28] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-29] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-30] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-31] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-32]Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-33] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-34] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-35] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-36] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-37] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-38] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-39] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 41-40] Figure 41 shows the results of transfection at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42.
[0052] [Figure 42-1] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-2] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-3] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-4] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-5]Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-6] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-7] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-8] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-9] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-10] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-11] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-12] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-13] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-14] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-15] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-16] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-17] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-18] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-19] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-20] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-21] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-22] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-23] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-24] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-25] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-26] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-27] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-28] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-29] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-30] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-31] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-32]Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-33] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-34] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-35] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-36] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-37] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-38] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-39] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-40] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-41] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-42] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-43] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-44] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-45] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-46] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-47] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-48] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-49] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-50] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-51] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-52] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-53] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-54] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-55] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-56] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-57] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-58] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-59]Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-60] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-61] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-62] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-63] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-64] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-65] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-66] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-67] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-68] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-69] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-70] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-71] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-72] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-73] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-74] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-75] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-76] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-77] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-78] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-79] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-80] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-81] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-82] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-83] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-84] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-85] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-86]Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-87] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-88] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-89] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-90] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-91] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-92] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-93] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-94] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-95] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44. [Figure 42-96] Figure 42 shows the results of transfection at the RNF-2 site. The sequence corresponds to SEQ ID NO: 44.
[0053] [Figure 43-1] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-2] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-3]Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-4] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-5] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-6] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-7] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-8] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-9] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-10] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-11] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-12] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-13] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-14] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-15] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-16] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-17] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-18] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-19] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-20] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-21] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-22] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-23] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-24] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-25] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-26] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-27] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-28] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-29] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-30]Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-31] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-32] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-33] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43. [Figure 43-34] Figure 43 shows the results of transfection at the Hek-4 site. The sequence corresponds to SEQ ID NO: 43.
[0054] [Figure 44-1] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-2] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-3] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-4] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-5] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-6] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-7] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-8] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-9]Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-10] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-11] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-12] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-13] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-14] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-15] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-16] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-17] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-18] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-19] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-20] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-21] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-22] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-23] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-24] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-25] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-26] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-27] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-28] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-29] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46. [Figure 44-30] Figure 44 shows the results of transfection at the EMX-1 site. The sequence corresponds to SEQ ID NO: 46.
[0055] [Figure 45-1] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-2] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-3] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-4] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-5] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-6] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-7] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-8] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-9] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-10] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-11] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-12] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-13] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-14] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-15] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-16] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-17] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-18] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-19]Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-20] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-21] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-22] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-23] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-24] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-25] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-26] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-27] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-28] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 45-29] Figure 45 shows the results of transfection at the FANCF site. The sequence corresponds to SEQ ID NO: 45.
[0056] [Figure 46] Figure 46 shows deaminase editing of sgRNA.
[0057] [Figure 47] Figure 47 shows the constructs developed for fusion at various sites.
[0058] [Figure 48] Figure 48 shows the indel rates for various fusions at various sites.
[0059] [Figure 49] Figure 49 shows the protospacer and PAM sequences of the base editing sites shown in SEQ ID NOs: 46, 45, 6, 42, 43, and 468, from top to bottom, respectively.
[0060] [Figure 50] Figure 50 shows the constructs developed for fusion at various sites, using a further mutated D108 residue.
[0061] [Figure 51] Figure 51 shows, from top to bottom, the protospacer and PAM sequences of the base editing sites shown in SEQ ID NOs: 6, 46, and 42, respectively.
[0062] [Figure 52-1] Figure 52 shows the results of using a mutated D108 residue to prevent the deaminase from accepting RNA as a substrate and alter the editing outcome. [Figure 52-2] Figure 52 shows the results of using a mutated D108 residue to prevent the deaminase from accepting RNA as a substrate and alter the editing outcome.
[0063] [Figure 53-1] Figure 53 shows the results of using a mutated D108 residue to prevent the deaminase from accepting RNA as a substrate and alter the editing outcome. [Figure 53-2] Figure 53 shows the results of using a mutated D108 residue to prevent the deaminase from accepting RNA as a substrate and alter the editing outcome.
[0064] [Figure 54]Figure 54 shows the constructs developed for fusion at various sites.
[0065] [Figure 55] Figure 55 shows, from top to bottom, the protospacer and PAM sequences of the base editing sites shown in SEQ ID NOs: 6, 358, and 359, respectively.
[0066] [Figure 56] Figure 56 shows the results of ABE on HEK site 2.
[0067] [Figure 57-1] Figure 57 shows the results of ABE on HEK site 2. [Figure 57-2] Figure 57 shows the results of ABE on HEK site 2.
[0068] [Figure 58] Figure 58 shows the constructs developed for fusion at various sites using various linker lengths.
[0069] [Figure 59] Figure 59 shows the importance of linker length for base editing function.
[0070] [Figure 60] Figure 60 shows the importance of linker length for base editing function.
[0071] [Figure 61] FIG. 61 is a schematic diagram showing deaminase dimerization.
[0072] [Figure 62] Figure 62 shows the constructs developed for fusion at various sites using various linker lengths.
[0073] [Figure 63]Figure 63 shows the current editing factor architecture (top panel), in trans dimerization (bottom panel, left), and in cis dimerization (bottom panel, right).
[0074] [Figure 64-1] Figure 64 shows dimerization results from base editing. [Figure 64-2] Figure 64 shows dimerization results from base editing.
[0075] [Figure 65-1] Figure 65 shows dimerization results from base editing. [Figure 65-2] Figure 65 shows dimerization results from base editing.
[0076] [Figure 66-1] Figure 66 shows dimerization results from base editing. [Figure 66-2] Figure 66 shows dimerization results from base editing.
[0077] [Figure 67] Figure 67 shows the constructs developed for fusion at various sgRNA sites.
[0078] [Figure 68-1] Figure 68 shows the evolution of ABE editors against new selected sequences. The sequences, from top to bottom and left to right, correspond to SEQ ID NOs: 707-719, respectively. [Figure 68-2] Figure 68 shows the evolution of ABE editors against new selected sequences. The sequences, from top to bottom and left to right, correspond to SEQ ID NOs: 707-719, respectively.
[0079] [Figure 69-1] Figure 69 shows current editing elements that target the Q4 arrest site. The sequences, from top to bottom, correspond to SEQ ID NOs: 624-628. [Figure 69-2]Figure 69 shows current editing elements that target the Q4 arrest site. The sequences, from top to bottom, correspond to SEQ ID NOs: 624-628.
[0080] [Figure 70-1] Figure 70 shows current editing elements that target the W15 stop site. The sequences, from top to bottom, correspond to SEQ ID NOs: 629-633, respectively. [Figure 70-2] Figure 70 shows current editing elements that target the W15 stop site. The sequences, from top to bottom, correspond to SEQ ID NOs: 629-633, respectively.
[0081] [Figure 71] Figure 71 shows the HEK293 site 2 sequence. The sequence corresponds to SEQ ID NO:360.
[0082] [Figure 72-1] FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG. [Figure 72-2] FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG. [Figure 72-3] FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG. [Figure 72-4] FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG. [Figure 72-5] FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG. [Figure 72-6] FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG. [Figure 72-7] FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG. [Figure 72-8] FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG. [Figure 72-9]FIG. 72 shows the results of the first run with various edTadA mutations using the sequence of FIG.
[0083] [Figure 73-1] FIG. 73 shows the results of a second run with various edTadA mutations using the sequences of FIG. [Figure 73-2] FIG. 73 shows the results of a second run with various edTadA mutations using the sequences of FIG. [Figure 73-3] FIG. 73 shows the results of a second run with various edTadA mutations using the sequences of FIG. [Figure 73-4] FIG. 73 shows the results of a second run with various edTadA mutations using the sequences of FIG. [Figure 73-5] FIG. 73 shows the results of a second run with various edTadA mutations using the sequences of FIG. [Figure 73-6] FIG. 73 shows the results of a second run with various edTadA mutations using the sequences of FIG.
[0084] [Figure 74] Figure 74 shows the FANCF sequence. The sequence corresponds to SEQ ID NO:45.
[0085] [Figure 75-1] FIG. 75 shows the results of a second run using various edTadA mutations and the sequence of FIG. [Figure 75-2] FIG. 75 shows the results of a second run using various edTadA mutations and the sequence of FIG.
[0086] [Figure 76-1] Figure 76 shows the results for D108 mutated on all sites. [Figure 76-2] Figure 76 shows the results for D108 mutated on all sites. [Figure 76-3] Figure 76 shows the results for D108 mutated on all sites. [Figure 76-4]Figure 76 shows the results for D108 mutated on all sites.
[0087] [Figure 77-1] Figure 77 shows the in trans data from the previous run (left panel) and the mut-mut fusion interrupted by the ultralong linker. [Figure 77-2] Figure 77 shows the in trans data from the previous run (left panel) and the mut-mut fusion interrupted by the ultralong linker.
[0088] [Figure 78] Figure 78 shows the results of linking mutTadA to ABE.
[0089] [Figure 79] Figure 79 shows the constructs of all inhibitors tested.
[0090] [Figure 80] Figure 80 shows the construct used when AAG was ligated to ABE.
[0091] [Figure 81] Figure 81 is a schematic diagram showing the linking of AAG to ABE.
[0092] [Figure 82-1] Figure 82 shows the results of linking AAG to ABE. [Figure 82-2] Figure 82 shows the results of linking AAG to ABE.
[0093] [Figure 83] Figure 83 shows the construct used when AAG was ligated to ABE with the N-terminus of TadA.
[0094] [Figure 84] Figure 84 is a schematic diagram showing that AAG was linked to an ABE with the N-terminus of TadA.
[0095] [Figure 85] Figure 85 shows the results of linking AAG to ABE.
[0096] [Figure 86] Figure 86 shows the construct used when ligating EndoV to ABE.
[0097] [Figure 87] Figure 87 is a schematic diagram showing EndoV linked to ABE.
[0098] [Figure 88-1] Figure 88 shows the results of linking EndoV to ABE. [Figure 88-2] Figure 88 shows the results of linking EndoV to ABE.
[0099] [Figure 89] Figure 89 shows the constructs used when connecting UGI to ABE.
[0100] [Figure 90-1] Figure 90 shows the results of connecting UGI to the end of ABE. [Figure 90-2] Figure 90 shows the results of connecting UGI to the end of ABE.
[0101] [Figure 91-1] Figure 91 shows the results of various inhibitors that increased A→G editing. [Figure 91-2] Figure 91 shows the results of various inhibitors that increased A→G editing. [Figure 91-3] Figure 91 shows the results of various inhibitors that increased A→G editing. [Figure 91-4] Figure 91 shows the results of various inhibitors that increased A→G editing. [Figure 91-5] Figure 91 shows the results of various inhibitors that increased A→G editing. [Figure 91-6]Figure 91 shows the results of various inhibitors that increased A→G editing.
[0102] [Figure 92-1] Figure 92 shows a sequence alignment of prokaryotic TadA amino acid sequences, which, from top to bottom, correspond to SEQ ID NOs: 634 to 657, respectively. [Figure 92-2] Figure 92 shows a sequence alignment of prokaryotic TadA amino acid sequences, which, from top to bottom, correspond to SEQ ID NOs: 634 to 657, respectively.
[0103] [Figure 93] Figure 93 shows a schematic representation of the relative sequence identity analysis of TadA amino acid sequences.
[0104] [Figure 94] Figure 94 shows a schematic representation of an exemplary adenosine base editing process.
[0105] [Figure 95] Figure 95 shows a schematic of an exemplary adenosine base editor that deaminates adenosine to inosine.
[0106] [Figure 96] Figure 96 shows a schematic diagram of an exemplary base editing selection plasmid.
[0107] [Figure 97] Figure 97 shows a list of clones containing the identified mutations in ecTadA.
[0108] [Figure 98-1] Figure 98 shows an exemplary sequencing analysis of selected plasmids from surviving clones. The sequences, from top to bottom, correspond to SEQ ID NOs: 658-661, respectively. [Figure 98-2] Figure 98 shows an exemplary sequencing analysis of selected plasmids from surviving clones. The sequences, from top to bottom, correspond to SEQ ID NOs: 658-661, respectively.
[0109] [Figure 99] Figure 99 shows a schematic of exemplary adenosine base editors from the third round of evolution.
[0110] [Figure 100] Figure 100 shows the percentage of A→G conversion in Hek293T cells.
[0111] [Figure 101] Figure 101 shows a schematic diagram of an exemplary base editing selection plasmid.
[0112] [Figure 102] Figure 102 shows a schematic diagram of the verdine crystal structure of S. aureus TadA. S. aureus TadA (a homologue of ecTadA) is shown with its co-crystallized tRNA substrate. The red arrows indicate H-bond contacts with various nucleic acids in the tRNA substrate. See Losey, H.C., et al., "Crystal structure of Staphylococcus sureus tRNA adenosine deaminase tadA in complex with RNA," Nature Struct. Mol. Biol. 2, 153-159 (2006).
[0113] [Figure 103] Figure 103 shows a schematic diagram of a construct containing ecTadA_2.2 and dCas9, identifying mutated ecTadA.
[0114] [Figure 104] Figure 104 shows the results of ecTadA evolution (Evolution #4) at sites E25 and R26.
[0115] [Figure 105] Figure 105 shows the results of ecTadA evolution at site R107 (Evolution #4).
[0116] [Figure 106] Figure 106 shows the results of ecTadA evolution (Evolution #4) at sites A142 and A143.
[0117] [Figure 107-1] Figure 107 shows an exemplary sequencing analysis of selected plasmids from surviving colonies. The sequences, from top to bottom, correspond to SEQ ID NOs: 662 to 671, respectively. [Figure 107-2] Figure 107 shows an exemplary sequencing analysis of selected plasmids from surviving colonies. The sequences, from top to bottom, correspond to SEQ ID NOs: 662 to 671, respectively.
[0118] [Figure 108-1] Figure 108 shows an overview of the results of editing at the Hek-2 site. The Hek-2 sequence provided in the figure represents the reverse complement of SEQ ID NO: 41, the DNA strand in which A to G editing occurs. The sequence corresponds to SEQ ID NO: 6. [Figure 108-2] Figure 108 shows an overview of the results of editing at the Hek-2 site. The Hek-2 sequence provided in the figure represents the reverse complement of SEQ ID NO: 41, the DNA strand in which A to G editing occurs. The sequence corresponds to SEQ ID NO: 6. [Figure 108-3] Figure 108 shows an overview of the results of editing at the Hek-2 site. The Hek-2 sequence provided in the figure represents the reverse complement of SEQ ID NO: 41, the DNA strand in which A to G editing occurs. The sequence corresponds to SEQ ID NO: 6. [Figure 108-4] Figure 108 shows an overview of the results of editing at the Hek-2 site. The Hek-2 sequence provided in the figure represents the reverse complement of SEQ ID NO: 41, the DNA strand in which A to G editing occurs. The sequence corresponds to SEQ ID NO: 6.
[0119] [Figure 109-1] Figure 109 shows a summary of the results of editing at the Hek2-3 site. The sequence corresponds to SEQ ID NO: 363. [Figure 109-2] Figure 109 shows a summary of the results of editing at the Hek2-3 site. The sequence corresponds to SEQ ID NO: 363. [Figure 109-3] Figure 109 shows a summary of the results of editing at the Hek2-3 site. The sequence corresponds to SEQ ID NO: 363. [Figure 109-4] Figure 109 shows a summary of the results of editing at the Hek2-3 site. The sequence corresponds to SEQ ID NO: 363.
[0120] [Figure 110-1] Figure 110 shows an overview of the results of editing at the Hek2-6 site. The sequence corresponds to SEQ ID NO: 364. [Figure 110-2] Figure 110 shows an overview of the results of editing at the Hek2-6 site. The sequence corresponds to SEQ ID NO: 364. [Figure 110-3] Figure 110 shows an overview of the results of editing at the Hek2-6 site. The sequence corresponds to SEQ ID NO: 364. [Figure 110-4] Figure 110 shows an overview of the results of editing at the Hek2-6 site. The sequence corresponds to SEQ ID NO: 364.
[0121] [Figure 111-1] Figure 111 shows an overview of the results of editing at the Hek2-7 site. The Hek2-7 sequence provided in the figure represents the reverse complement of the DNA strand in which the A to G edit occurs. The sequence corresponds to SEQ ID NO: 365. [Figure 111-2] Figure 111 shows an overview of the results of editing at the Hek2-7 site. The Hek2-7 sequence provided in the figure represents the reverse complement of the DNA strand in which the A to G edit occurs. The sequence corresponds to SEQ ID NO: 365. [Figure 111-3] Figure 111 shows an overview of the results of editing at the Hek2-7 site. The Hek2-7 sequence provided in the figure represents the reverse complement of the DNA strand in which the A to G edit occurs. The sequence corresponds to SEQ ID NO: 365. [Figure 111-4] Figure 111 shows an overview of the results of editing at the Hek2-7 site. The Hek2-7 sequence provided in the figure represents the reverse complement of the DNA strand in which the A to G edit occurs. The sequence corresponds to SEQ ID NO: 365.
[0122] [Figure 112-1]Figure 112 shows a summary of the results of editing at the Hek2-10 site. The sequence corresponds to SEQ ID NO: 366. [Figure 112-2] Figure 112 shows a summary of the results of editing at the Hek2-10 site. The sequence corresponds to SEQ ID NO: 366. [Figure 112-3] Figure 112 shows a summary of the results of editing at the Hek2-10 site. The sequence corresponds to SEQ ID NO: 366. [Figure 112-4] Figure 112 shows a summary of the results of editing at the Hek2-10 site. The sequence corresponds to SEQ ID NO: 366.
[0123] [Figure 113-1] Figure 113 shows a summary of the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 113-2] Figure 113 shows a summary of the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 113-3] Figure 113 shows a summary of the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42.
[0124] [Figure 114-1] Figure 114 shows a summary of the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 114-2] Figure 114 shows a summary of the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 114-3] Figure 114 shows a summary of the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45.
[0125] [Figure 115-1] Figure 115 shows a summary of the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 367. [Figure 115-2] Figure 115 shows a summary of the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 367. [Figure 115-3]Figure 115 shows a summary of the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 367. [Figure 115-4] Figure 115 shows a summary of the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 367. [Figure 115-5] Figure 115 shows a summary of the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 367. [Figure 115-6] Figure 115 shows a summary of the results of editing at the Hek-2 site. The sequence corresponds to SEQ ID NO: 367.
[0126] [Figure 116-1] Figure 116 shows a summary of the results of editing at the Hek2-2 site. The sequence corresponds to SEQ ID NO: 368. [Figure 116-2] Figure 116 shows a summary of the results of editing at the Hek2-2 site. The sequence corresponds to SEQ ID NO: 368. [Figure 116-3] Figure 116 shows a summary of the results of editing at the Hek2-2 site. The sequence corresponds to SEQ ID NO: 368. [Figure 116-4] Figure 116 shows a summary of the results of editing at the Hek2-2 site. The sequence corresponds to SEQ ID NO: 368.
[0127] [Figure 117-1] Figure 117 shows an overview of the results of editing at the Hek2-3 site. The sequence corresponds to SEQ ID NO: 363. [Figure 117-2] Figure 117 shows an overview of the results of editing at the Hek2-3 site. The sequence corresponds to SEQ ID NO: 363. [Figure 117-3] Figure 117 shows an overview of the results of editing at the Hek2-3 site. The sequence corresponds to SEQ ID NO: 363.
[0128] [Figure 118-1] Figure 118 shows a summary of the results of editing at the Hek2-6 site. The sequence corresponds to SEQ ID NO: 364. [Figure 118-2]Figure 118 shows a summary of the results of editing at the Hek2-6 site. The sequence corresponds to SEQ ID NO: 364. [Figure 118-3] Figure 118 shows a summary of the results of editing at the Hek2-6 site. The sequence corresponds to SEQ ID NO: 364. [Figure 118-4] Figure 118 shows a summary of the results of editing at the Hek2-6 site. The sequence corresponds to SEQ ID NO: 364.
[0129] [Figure 119-1] Figure 119 shows a summary of the results of editing at the Hek2-7 site. The sequence corresponds to SEQ ID NO: 365. [Figure 119-2] Figure 119 shows a summary of the results of editing at the Hek2-7 site. The sequence corresponds to SEQ ID NO: 365. [Figure 119-3] Figure 119 shows a summary of the results of editing at the Hek2-7 site. The sequence corresponds to SEQ ID NO: 365. [Figure 119-4] Figure 119 shows a summary of the results of editing at the Hek2-7 site. The sequence corresponds to SEQ ID NO: 365.
[0130] [Figure 120-1] Figure 120 shows a summary of the results of editing at the Hek2-10 site. The sequence corresponds to SEQ ID NO: 366. [Figure 120-2] Figure 120 shows a summary of the results of editing at the Hek2-10 site. The sequence corresponds to SEQ ID NO: 366. [Figure 120-3] Figure 120 shows a summary of the results of editing at the Hek2-10 site. The sequence corresponds to SEQ ID NO: 366. [Figure 120-4] Figure 120 shows a summary of the results of editing at the Hek2-10 site. The sequence corresponds to SEQ ID NO: 366.
[0131] [Figure 121-1] Figure 121 shows a summary of the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42. [Figure 121-2]Figure 121 shows a summary of the results of editing at the Hek-3 site. The sequence corresponds to SEQ ID NO: 42.
[0132] [Figure 122-1] Figure 122 shows a summary of the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45. [Figure 122-2] Figure 122 shows a summary of the results of editing at the FANCF site. The sequence corresponds to SEQ ID NO: 45.
[0133] [Figure 123] Figure 123 shows the results of ecTadA evolution (Evolution #4) at the HEK2, HEK2-2, HEK2-3, HEK2-6, HEK2-7, and HEK2-10 sites. The constructs used were pNMG-370 (Evolution #2), pNMG-371 (Evolution #3), and pNMG382-389 (Evolution #4). From top to bottom, the sequences correspond to SEQ ID NOs: 7, 368, 363, 364, 369, and 370, respectively.
[0134] [Figure 124] Figure 124 shows a schematic of the construct containing ecTadA and dCas9 used in ecTadA evolution (Evolution #5).
[0135] [Figure 125-1] Figure 125 is a table showing the results of clones assayed after the fifth round of evolution (128ug / mL chlor, 7h). [Figure 125-2] Figure 125 is a table showing the results of clones assayed after the fifth round of evolution (128ug / mL chlor, 7h).
[0136] [Figure 126A] Figures 126A-126E are tables showing the results of subcloned and retransformed clones assayed after round 5 under varying conditions. [Figure 126B]Figures 126A-126E are tables showing the results of subcloned and retransformed clones assayed after round 5 under varying conditions. [Figure 126C] Figures 126A-126E are tables showing the results of subcloned and retransformed clones assayed after round 5 under varying conditions. [Figure 126D] Figures 126A-126E are tables showing the results of subcloned and retransformed clones assayed after round 5 under varying conditions. [Figure 126E] Figures 126A-126E are tables showing the results of subcloned and retransformed clones assayed after round 5 under varying conditions.
[0137] [Figure 127] Figure 127 is a table showing the results of amplification products from spectinomycin-selected clones assayed after the fifth round of evolution.
[0138] [Figure 128-1] Figure 128 is a table showing the results of the clones assayed after the fifth round of evolution. [Figure 128-2] Figure 128 is a table showing the results of the clones assayed after the fifth round of evolution.
[0139] [Figure 129-1] Figure 129 shows a summary of the results of editing at Hek-2 sites using base editors containing engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for editors containing engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 6. [Figure 129-2]Figure 129 shows a summary of the results of editing at Hek-2 sites using base editors containing engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for editors containing engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 6.
[0140] [Figure 130-1] Figure 130 shows an overview of the results of editing at the Hek2-1 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-1 sequence provided in the figure represents the DNA strand on which the A to G edit occurs. The sequence corresponds to SEQ ID NO: 465. [Figure 130-2] Figure 130 shows an overview of the results of editing at the Hek2-1 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-1 sequence provided in the figure represents the DNA strand on which the A to G edit occurs. The sequence corresponds to SEQ ID NO: 465.
[0141] [Figure 131-1] Figure 131 shows a summary of the results of editing at the Hek2-2 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 368. [Figure 131-2]Figure 131 shows a summary of the results of editing at the Hek2-2 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 368.
[0142] [Figure 132] Figure 132 shows a summary of the results of editing at the Hek2-3 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 363.
[0143] [Figure 133] Figure 133 shows an overview of the results of editing at the Hek2-4 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-4 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO: 466.
[0144] [Figure 134-1] Figure 134 shows a summary of the results of editing at the Hek2-6 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 364. [Figure 134-2]Figure 134 shows a summary of the results of editing at the Hek2-6 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 364.
[0145] [Figure 135-1] Figure 135 shows an overview of the results of editing at the Hek2-9 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-9 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO: 467. [Figure 135-2] Figure 135 shows an overview of the results of editing at the Hek2-9 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-9 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO: 467.
[0146] [Figure 136-1] Figure 136 shows an overview of the results of editing at the Hek2-10 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-10 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO: 370. [Figure 136-2]Figure 136 shows an overview of the results of editing at the Hek2-10 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-10 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO: 370.
[0147] [Figure 137-1] Figure 137 shows a summary of the results of editing at the Hek3 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 42. [Figure 137-2] Figure 137 shows a summary of the results of editing at the Hek3 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The sequence corresponds to SEQ ID NO: 42.
[0148] [Figure 138-1] Figure 138 shows an overview of the results of editing at the RNF2 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-10 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO: 468. [Figure 138-2]Figure 138 shows an overview of the results of editing at the RNF2 site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-10 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO: 468.
[0149] [Figure 139-1] Figure 139 shows an overview of the results of editing at the FANCF site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-10 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO:45. [Figure 139-2] Figure 139 shows an overview of the results of editing at the FANCF site using a base editor containing an engineered S. aureus TadA (saTadA) comprising pNMG-346-349. As a comparison, results are shown for an editor containing an engineered E. coli TadA (ecTadA) comprising pNMG-339-341. The Hek2-10 sequence provided in the figure represents the DNA strand in which A-to-G editing occurs. The sequence corresponds to SEQ ID NO:45.
[0150] [Figure 140-1]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-2]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-3]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-4]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-5]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-6]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-7]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-8]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-9]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-10]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140. [Figure 140-11]Figure 140 shows various schematic diagrams of adenosine base editor (ABE) constructs. The identity of the editor, for example "pNMG-367," is shown in Table 4. The following mutations are abbreviated as follows: ecTadA1 (A106V D108N), ecTadA2 (A106V D108N D147Y E155V), ecTadA3 (ecTadA2 + L84F H123Y I156F), ecTadA3 + (ecTadA3 + A142N), ecTadA5a1 (ecTadA3 + H36L R51L S146C K157N), ecTadA5a3 (ecTadA3 + N37S K161T), ecTadA5a11 (ecTadA3 + R51L S146C K157N K161T), ecTadA5a12 (ecTadA3 + S146C K161T), ecTadA5a14 (ecTadA3 + RS146C K161T), ecTadA5a1+ (ecTadA5a1+A142N), ecTadA5a9 (ecTadA3+S146R K161T). Heterodimers of the above three ABE 5a constructs were made and then tested against homodimers. Heterodimeric versions of ABE editing factors typically perform better than the corresponding homodimeric constructs. Both homodimeric and heterodimeric constructs are shown in Figure 140.
[0151] [Figure 141-1] Figure 141 shows the compilation results for various ABE constructs. ABE Plasmid # refers to the pNMG number as shown in Table 4. For example, 367 refers to construct pNMG-367 in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 469 (pNMG-466), 470 (pNMG-467), 471 (pNMG-469), 472 (pNMG-470), 473 (pNMG-501), 474 (pNMG-509), and 475 (pNMG-502), respectively. [Figure 141-2]Figure 141 shows the compilation results for various ABE constructs. ABE Plasmid # refers to the pNMG number as shown in Table 4. For example, 367 refers to construct pNMG-367 in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 469 (pNMG-466), 470 (pNMG-467), 471 (pNMG-469), 472 (pNMG-470), 473 (pNMG-501), 474 (pNMG-509), and 475 (pNMG-502), respectively. [Figure 141-3] Figure 141 shows the compilation results for various ABE constructs. ABE Plasmid # refers to the pNMG number as shown in Table 4. For example, 367 refers to construct pNMG-367 in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 469 (pNMG-466), 470 (pNMG-467), 471 (pNMG-469), 472 (pNMG-470), 473 (pNMG-501), 474 (pNMG-509), and 475 (pNMG-502), respectively.
[0152] [Figure 142-1] Figure 142 shows the compilation results for various ABE constructs at specific sites. The numbers in the top row indicate the pNMG number as shown in Table 4. For example, 107 refers to construct pNMG-107 in Table 4. In some contexts, homodimeric constructs have been shown to perform better than heterodimeric constructs, and vice versa (see, for example, homodimeric construct 371 vs. heterodimeric construct 476). Schematics of these ABE constructs are shown in Figure 140, and the construct architectures are shown in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 478, 478, 514, 516, 516, 520, 520, 521, 521, and 509, respectively. [Figure 142-2]Figure 142 shows the compilation results for various ABE constructs at specific sites. The numbers in the top row indicate the pNMG number as shown in Table 4. For example, 107 refers to construct pNMG-107 in Table 4. In some contexts, homodimeric constructs have been shown to perform better than heterodimeric constructs, and vice versa (see, for example, homodimeric construct 371 vs. heterodimeric construct 476). Schematics of these ABE constructs are shown in Figure 140, and the construct architectures are shown in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 478, 478, 514, 516, 516, 520, 520, 521, 521, and 509, respectively.
[0153] [Figure 143] FIG. 143 shows the percentage of ides formed for the ABE constructs from FIG.
[0154] [Figure 144] Figure 144 shows editing results for various ABE constructs at specific sites. Construct identities are shown in the top row and refer to the pNMG reference numbers in Table 4. The results in Figure 144 show that adding ecTadA monomers to ABE constructs may not improve editing. However, adding longer linkers between monomers may aid editing at some sites (see, for example, the editing results for sgRNA constructs 285b vs. 277 at sites 502, 505, and 507). The identities of the sgRNA constructs are shown in Table 8, and a schematic representation of these is shown in Figure 140. From top to bottom, the sequences correspond to SEQ ID NOs: 478, 480, 480, 514, 517, 517, 517, 517, 519, and 521, respectively.
[0155] [Figure 145]Figure 145 shows the results of ABE constructs at all NAN sites, where target A is at position 5 of the protospacer and PAM sequences. The identity of the ABE constructs shown in the top row refers to the pNMG reference number in Table 4. The numbers represent the % of target A residues edited (e.g., % editing efficiency). From top to bottom, the sequences correspond to SEQ ID NOs: 537-552, respectively.
[0156] [Figure 146] Figure 146 shows the percent A to G editing at the Hek2 site for various ABE constructs (as referenced by their reference pNMG numbers in Table 4).
[0157] [Figure 147] Figure 147 shows the evolution results of evolution round #5b. Numbers represent the % of A to G editing for the indicated sites. From top to bottom, the sequences correspond to SEQ ID NOs: 7, 465, 368, 363, 364, and 370, respectively.
[0158] [Figure 148] Figure 148 shows the editing results for various ABE constructs obtained from different rounds of evolution (evo3 as an example). A general schematic for the ABE constructs is also shown. The identities of the sgRNAs shown in Table 8 and the identities of the base editors shown in Table 4 (see pNMG) are shown. Numbers represent the % of A to G editing for the indicated sites. From top to bottom, the sequences correspond to SEQ ID NOs: 478, 503, 506, 521, 513, 505, 507, and 509, respectively.
[0159] [Figure 149] Figure 149 shows examination of ABE constructs at genomic sites other than the Hek-2 sequence. The Hek-2 site (sgRNA299) is represented by an asterisk. The identity of the sgRNA is shown in Table 8. From top to bottom, the sequences correspond to SEQ ID NOs: 478, 514, 516, 517, 517, 517, 519, 520, 529, and 521, respectively.
[0160] [Figure 150] Figure 150 shows a schematic of a DNA shuffling experiment using nucleotide exchange and excision technology (NExT), referred to as ABE evolution #6. The goal of this approach was to recruit more efficient editors and potentially remove epistatic mutations. DNA shuffling of constructs from various evolutions was used to optimize for desired mutations and to eliminate mutations that negatively affect editing efficiency and / or protein stability.
[0161] [Figure 151-1] Figure 151 shows a schematic diagram of DNA shuffling (NeXT). The spect target sequence is [ka] and the chlor target sequence is [ka] is. [Figure 151-2] Figure 151 shows a schematic diagram of DNA shuffling (NeXT). The spect target sequence is as above, and the chlor target sequence is as above.
[0162] [Figure 152-1] Figure 152 shows the sequence identity of clones from evolution #6 surviving spect only (non-YAC target). The mutations shown are relative to ecTadA (SEQ ID NO: 1). [Figure 152-2] Figure 152 shows the sequence identity of clones from evolution #6 surviving spect only (non-YAC target). The mutations shown are relative to ecTadA (SEQ ID NO: 1). [Figure 152-3]Figure 152 shows the sequence identity of clones from evolution #6 surviving spect only (non-YAC target). The mutations shown are relative to ecTadA (SEQ ID NO: 1). [Figure 152-4] Figure 152 shows the sequence identity of clones from evolution #6 surviving spect only (non-YAC target). The mutations shown are relative to ecTadA (SEQ ID NO: 1).
[0163] [Figure 153-1] Figure 153 shows evolution #6.2, which refers to the enrichment of clones from evolution #6. The mutations shown are relative to ecTadA (SEQ ID NO: 1). A142N is present in nearly all clones sequenced, and the Pro48 mutation is also enriched. Clones were selected against "GAT" in the spectinomycin site. The selection target sequence was [ka] It was. [Figure 153-2] Figure 153 shows evolution #6.2, which indicates the enrichment of clones from evolution #6. The mutations shown are relative to ecTadA (SEQ ID NO: 1). A142N is present in nearly all clones sequenced, and the Pro48 mutation is also enriched. Clones were selected against "GAT" in the spectinomycin site. The selection target sequence was as described above.
[0164] [Fig. 154] Figure 154 shows a schematic diagram of the ABE6 construct. A total of eight new constructs were developed. Mutations from the top two most frequent amplicons in Evo#6 were used in each of the four architectures.
[0165] [Figure 155-1]Figure 155 shows data collected for ABE: Step 1 - Transfection + HTS of key intermediates at six genomic sites, n=3. Transfections were performed with 750 ng ABE + 250 ng gRNA and incubated for 5 days before genomic DNA was extracted and HTS was performed. The identity of each ABE construct is indicated by the pNMG reference number as shown in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 509, 510, 512, 520, 530, and 478, respectively. [Figure 155-2] Figure 155 shows data collected for ABE: Step 1 - Transfection + HTS of key intermediates at six genomic sites, n=3. Transfections were performed with 750 ng ABE + 250 ng gRNA and incubated for 5 days before genomic DNA was extracted and HTS was performed. The identity of each ABE construct is indicated by the pNMG reference number as shown in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 509, 510, 512, 520, 530, and 478, respectively. [Figure 155-3] Figure 155 shows data collected for ABE: Step 1 - Transfection + HTS of key intermediates at six genomic sites, n=3. Transfections were performed with 750 ng ABE + 250 ng gRNA and incubated for 5 days before genomic DNA was extracted and HTS was performed. The identity of each ABE construct is indicated by the pNMG reference number as shown in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 509, 510, 512, 520, 530, and 478, respectively. [Figure 155-4] Figure 155 shows data collected for ABE: Step 1 - Transfection + HTS of key intermediates at six genomic sites, n=3. Transfections were performed with 750 ng ABE + 250 ng gRNA and incubated for 5 days before genomic DNA was extracted and HTS was performed. The identity of each ABE construct is indicated by the pNMG reference number as shown in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 509, 510, 512, 520, 530, and 478, respectively. [Figure 155-5] Figure 155 shows data collected for ABE: Step 1 - Transfection + HTS of key intermediates at six genomic sites, n=3. Transfections were performed with 750 ng ABE + 250 ng gRNA and incubated for 5 days before genomic DNA was extracted and HTS was performed. The identity of each ABE construct is indicated by the pNMG reference number as shown in Table 4. From top to bottom, the sequences correspond to SEQ ID NOs: 509, 510, 512, 520, 530, and 478, respectively.
[0166] [Figure 156-1] Figure 156 shows that ABE editing efficiency improves with iterative rounds of evolution. The top panel shows representative A to G % editing at targeted loci in Hek293T cells using evolved / engineered ABE constructs. The sequence corresponds to SEQ ID NO: 561. The bottom panel shows that iterative rounds of evolution and engineering improve ABE. ABE constructs are indicated by their pNMG reference number as shown in Table 4. The "508" target sequence corresponds to SEQ ID NO: 520. [Figure 156-2] Figure 156 shows that ABE editing efficiency improves with iterative rounds of evolution. The top panel shows representative A to G % editing at targeted loci in Hek293T cells using evolved / engineered ABE constructs. The sequence corresponds to SEQ ID NO: 561. The bottom panel shows that iterative rounds of evolution and engineering improve ABE. ABE constructs are indicated by their pNMG reference number as shown in Table 4. The "508" target sequence corresponds to SEQ ID NO: 520. [Figure 156-3]Figure 156 shows that ABE editing efficiency improves with iterative rounds of evolution. The top panel shows representative A to G % editing at targeted loci in Hek293T cells using evolved / engineered ABE constructs. The sequence corresponds to SEQ ID NO: 561. The bottom panel shows that iterative rounds of evolution and engineering improve ABE. ABE constructs are indicated by their pNMG reference number as shown in Table 4. The "508" target sequence corresponds to SEQ ID NO: 520.
[0167] [Figure 157] Figure 157 shows HTS results of the core six genomic sites from the 10 "best" ABEs. The results show that different editing elements have different local sequence preferences (bottom panel). The graph shows the percent A to G editing at six different loci. ABE constructs are indicated by their pNMG reference numbers as shown in Table 4. The sequences, from top to bottom, correspond to SEQ ID NOs: 509, 510, 512, 520, 530, and 478, respectively.
[0168] [Figure 158-1] Figure 158 shows the transfection of the functional "top 10" ABEs at all genomic sites covering every combination of NAN sequences. Data represents n=1. From top to bottom, the sequences correspond to SEQ ID NOs: 489, 490, 493, 497, 503, 504, 507, 508, 511, and 513, respectively. [Figure 158-2] Figure 158 shows the transfection of the functional "top 10" ABEs at all genomic sites covering every combination of NAN sequences. Data represents n=1. From top to bottom, the sequences correspond to SEQ ID NOs: 489, 490, 493, 497, 503, 504, 507, 508, 511, and 513, respectively.
[0169] [Figure 159-1]Figure 159 shows an ABE window experiment (A at odd positions) to identify which As were edited. ABE pNMG-477, pNMG-586, pNMG-588, BE3 and an untreated control are shown. The sequences for the edits are shown above. The sequence corresponds to SEQ ID NO: 562. [Figure 159-2] Figure 159 shows an ABE window experiment (A at odd positions) to identify which As were edited. ABE pNMG-477, pNMG-586, pNMG-588, BE3 and an untreated control are shown. The sequences for the edits are shown above. The sequence corresponds to SEQ ID NO: 562. [Figure 159-3] Figure 159 shows an ABE window experiment (A at odd positions) to identify which As were edited. ABE pNMG-477, pNMG-586, pNMG-588, BE3 and an untreated control are shown. The sequences for the edits are shown above. The sequence corresponds to SEQ ID NO: 562.
[0170] [Figure 160-1] Figure 160 shows an ABE window experiment (A at even-numbered positions) to identify which As were edited. ABE pNMG-477, pNMG-586, pNMG-588, BE3 and an untreated control are shown. The sequences for the edits are shown above. The sequence corresponds to SEQ ID NO: 563. [Figure 160-2] Figure 160 shows an ABE window experiment (A at even-numbered positions) to identify which As were edited. ABE pNMG-477, pNMG-586, pNMG-588, BE3 and an untreated control are shown. The sequences for the edits are shown above. The sequence corresponds to SEQ ID NO: 563. [Figure 160-3] Figure 160 shows an ABE window experiment (A at even-numbered positions) to identify which As were edited. ABE pNMG-477, pNMG-586, pNMG-588, BE3 and an untreated control are shown. The sequences for the edits are shown above. The sequence corresponds to SEQ ID NO: 563. [Figure 160-4]Figure 160 shows an ABE window experiment (A at even-numbered positions) to identify which As were edited. ABE pNMG-477, pNMG-586, pNMG-588, BE3 and an untreated control are shown. The sequences for the edits are shown above. The sequence corresponds to SEQ ID NO: 563.
[0171] [Figure 161-1] Figure 161 shows an ABE window experiment to identify which A's were edited. ABE pNMG-586, pNMG-560, and an untreated control are shown. The sequences for the edits are shown at the top. From top to bottom, the sequences correspond to SEQ ID NOs: 544 and 541, respectively. [Figure 161-2] Figure 161 shows an ABE window experiment to identify which A's were edited. ABE pNMG-586, pNMG-560, and an untreated control are shown. The sequences for the edits are shown at the top. From top to bottom, the sequences correspond to SEQ ID NOs: 544 and 541, respectively. [Figure 161-3] Figure 161 shows an ABE window experiment to identify which A's were edited. ABE pNMG-586, pNMG-560, and an untreated control are shown. The sequences for the edits are shown at the top. From top to bottom, the sequences correspond to SEQ ID NOs: 544 and 541, respectively. [Figure 161-4] Figure 161 shows an ABE window experiment to identify which A's were edited. ABE pNMG-586, pNMG-560, and an untreated control are shown. The sequences for the edits are shown at the top. From top to bottom, the sequences correspond to SEQ ID NOs: 544 and 541, respectively. [Figure 161-5] Figure 161 shows an ABE window experiment to identify which A's were edited. ABE pNMG-586, pNMG-560, and an untreated control are shown. The sequences for the edits are shown at the top. From top to bottom, the sequences correspond to SEQ ID NOs: 544 and 541, respectively. [Figure 161-6]Figure 161 shows an ABE window experiment to identify which A's were edited. ABE pNMG-586, pNMG-560, and an untreated control are shown. The sequences for the edits are shown at the top. From top to bottom, the sequences correspond to SEQ ID NOs: 544 and 541, respectively.
[0172] [Figure 162-1] Figure 162 shows an ABE window experiment to identify which A's were edited. ABE pNMG-576, pNMG-586, and an untreated control are shown. The sequence for the edit is shown above. The sequence corresponds to SEQ ID NO: 564. [Figure 162-2] Figure 162 shows an ABE window experiment to identify which A's were edited. ABE pNMG-576, pNMG-586, and an untreated control are shown. The sequence for the edit is shown above. The sequence corresponds to SEQ ID NO: 564. [Figure 162-3] Figure 162 shows an ABE window experiment to identify which A's were edited. ABE pNMG-576, pNMG-586, and an untreated control are shown. The sequence for the edit is shown above. The sequence corresponds to SEQ ID NO: 564.
[0173] [Figure 163-1] Figure 163 shows Evolution #7, an attempt to edit multiple A sites. The evolutionary selection design uses two separate gRNAs: one to create a D208N reversion mutation in Kan and one to revert the stop codon to Q. [ka] and [ka] The goal was to target two point mutations in the same gene using . [Figure 163-2]Figure 163 shows Evolution #7, an attempt to edit multiple A sites. The evolutionary selection design was to target two point mutations in the same gene using two separate gRNAs as described above to create a D208N reversion mutation in Kan and to revert the stop codon to Q. [Figure 163-3] Figure 163 shows Evolution #7, an attempt to edit multiple A sites. The evolutionary selection design was to target two point mutations in the same gene using two separate gRNAs as described above to create a D208N reversion mutation in Kan and to revert the stop codon to Q. [Figure 163-4] Figure 163 shows Evolution #7, an attempt to edit multiple A sites. The evolutionary selection design was to target two point mutations in the same gene using two separate gRNAs as described above to create a D208N reversion mutation in Kan and to revert the stop codon to Q.
[0174] [Figure 164-1] Figure 164 shows evolution #7 mutations that evolved to target A within a multi-A site, meaning they are flanked on one or both sides by A. The identity of the mutations relative to SEQ ID NO: 1 is shown. [Figure 164-2] Figure 164 shows evolution #7 mutations that evolved to target A within a multi-A site, meaning they are flanked on one or both sides by A. The identity of the mutations relative to SEQ ID NO: 1 is shown.
[0175] [Figure 165] Figure 165 shows a schematic diagram of ecTadA identifying residues R152 and P48.
[0176] [Figure 166]Figure 166 shows MiSeq results for ABE editing on disease-associated mutations in surrogate cell lines. Nucleofection using the Lonza kit was used with 3 different nucleofection solutions x 16 different electroporation conditions (48 conditions / cell line total). Sequences, from top to bottom, correspond to SEQ ID NOs: 522-524, respectively.
[0177] [Figure 167-1] Figure 167 shows the results of A to G editing at multiple positions for various constructs. ABE constructs are indicated by their pNMG reference number as shown in Table 4. In the top panel, the sequences correspond, from top to bottom, to SEQ ID NOs: 469-471, 567, 475, and 474, respectively. In the bottom panel, the sequences correspond, from top to bottom, to SEQ ID NOs: 469 (pNMG-466), 470 (pNMG-467), 471 (pNMG-469), 567 (pNMG-472), and 474 (pNMG-509), respectively. [Figure 167-2] Figure 167 shows the results of A to G editing at multiple positions for various constructs. ABE constructs are indicated by their pNMG reference number as shown in Table 4. In the top panel, the sequences correspond, from top to bottom, to SEQ ID NOs: 469-471, 567, 475, and 474, respectively. In the bottom panel, the sequences correspond, from top to bottom, to SEQ ID NOs: 469 (pNMG-466), 470 (pNMG-467), 471 (pNMG-469), 567 (pNMG-472), and 474 (pNMG-509), respectively. [Figure 167-3]Figure 167 shows the results of A to G editing at multiple positions for various constructs. ABE constructs are indicated by their pNMG reference number as shown in Table 4. In the top panel, the sequences correspond, from top to bottom, to SEQ ID NOs: 469-471, 567, 475, and 474, respectively. In the bottom panel, the sequences correspond, from top to bottom, to SEQ ID NOs: 469 (pNMG-466), 470 (pNMG-467), 471 (pNMG-469), 567 (pNMG-472), and 474 (pNMG-509), respectively. [Figure 167-4] Figure 167 shows the results of A to G editing at multiple positions for various constructs. ABE constructs are indicated by their pNMG reference number as shown in Table 4. In the top panel, the sequences correspond, from top to bottom, to SEQ ID NOs: 469-471, 567, 475, and 474, respectively. In the bottom panel, the sequences correspond, from top to bottom, to SEQ ID NOs: 469 (pNMG-466), 470 (pNMG-467), 471 (pNMG-469), 567 (pNMG-472), and 474 (pNMG-509), respectively.
[0178] [Figure 168-1] Figure 168 shows the compilation results of various constructs using ABE with various linkers. The ABE constructs are indicated by their pNMG reference number as shown in Table 4. A schematic diagram of the new linker ABE is also shown. From top to bottom, the sequences correspond to SEQ ID NOs: 469 (pNMG-466), 568 (pNMG-468), 471 (pNMG-469), 567 (pNMG-472), 574 (pNGM-509), and 569 (pNMG-539), respectively. [Figure 168-2]Figure 168 shows the compilation results of various constructs using ABE with various linkers. The ABE constructs are indicated by their pNMG reference number as shown in Table 4. A schematic diagram of the new linker ABE is also shown. From top to bottom, the sequences correspond to SEQ ID NOs: 469 (pNMG-466), 568 (pNMG-468), 471 (pNMG-469), 567 (pNMG-472), 574 (pNGM-509), and 569 (pNMG-539), respectively. [Figure 168-3] Figure 168 shows the compilation results of various constructs using ABE with various linkers. The ABE constructs are indicated by their pNMG reference number as shown in Table 4. A schematic diagram of the new linker ABE is also shown. From top to bottom, the sequences correspond to SEQ ID NOs: 469 (pNMG-466), 568 (pNMG-468), 471 (pNMG-469), 567 (pNMG-472), 574 (pNGM-509), and 569 (pNMG-539), respectively.
[0179] [Figure 169] Figure 169 shows the fourth round of evolution. Evolution was done using a monomer construct. Endogenous TadA complements the TadA-dCas9 fusion.
[0180] [Figure 170] Figure 170 shows the results of the fourth round of evolution. The sequences, from top to bottom, correspond to SEQ ID NOs: 7, 368, 363, 364, 369, and 370, respectively.
[0181] [Figure 171-1] Figure 171 shows evolution round #5. Plasmids and experimental schematics are shown (top panel). Graphs illustrate survival on chlor vs. spectinomycin, "TAG" vs. "GAT". The chlor target sequence is [ka] and the spect target sequence is [ka] is. [Figure 171-2] Figure 171 shows evolution round #5. Plasmids and experimental schematics are shown (top panel). Graphs illustrate survival on chlor vs. spectinomycin, "TAG" vs. "GAT." The chlor target sequence is as above, and the spect target sequence is as above.
[0182] [Fig. 172] Figure 172 shows the editing results at chlor and spect sites. Constructs identified from evolution #4 (saturated site / NNK library) appear to edit more efficiently at spect sites rather than at chor sites. ABE constructs are indicated by their pNMG reference numbers as shown in Table 4.
[0183] [Figure 173] Figure 173 shows the fifth round of evolution (part a). The sequence corresponds to SEQ ID NO: 570.
[0184] [Fig. 174] Figure 174 shows the fifth round heterodimer (in trans) results. Round #5a identified mutations that improved both editing efficiency and broadened substrate specificity. The sequences, from top to bottom, correspond to SEQ ID NOs: 7, 368, and 363, 364, 369, and 370, respectively.
[0185] [Figure 175] Figure 175 shows the fifth round heterodimer (in cis) results. Round #5a identified mutations that improved both editing efficiency and broadened substrate specificity, while the cis results conferred higher editing efficiency. ABE constructs are indicated by their pNMG reference numbers as shown in Table 4. The sequences, from top to bottom, correspond to SEQ ID NOs: 7, 571, 465, 368, 363, 466, 364, 369, 572, and 370, respectively.
[0186] [Figure 176] Figure 176 shows the compilation results of various constructs for evolution 5.
[0187] [Figure 177] Figure 177 shows the compilation results of various constructs for evolution 5.
[0188] [Figure 178] Figure 178 shows the gRNA for ABE. The 5a construct is characterized by an A at position 5 in the protospacer on all 16 NAN sequences (left panel). The sequences, from top to bottom, correspond to SEQ ID NOs: 573-578, respectively. Additional sequences beginning with "G" are proposed to minimize variation in the synthesis of the resulting gRNAs (right panel). The sequences, from top to bottom, correspond to SEQ ID NOs: 579-588, respectively.
[0189] [Figure 179] Figure 179 shows %A→G editing of A5 using sgRNA299 and ABE constructs shown in Table 8 (which are designated by their pNMG reference numbers as shown in Table 4). The sequence corresponds to SEQ ID NO:478.
[0190] [Figure 180] Figure 180 shows %A→G editing of A5 using sgRNA469 and ABE constructs shown in Table 8 (which are designated by their pNMG reference numbers as shown in Table 4). The sequence corresponds to SEQ ID NO: 509.
[0191] [Figure 181] Figure 181 shows %A→G editing of A5 using sgRNA470 and ABE constructs shown in Table 8 (which are designated by their pNMG reference numbers as shown in Table 4). The sequence corresponds to SEQ ID NO: 510.
[0192] [Figure 182] Figure 182 shows %A→G editing of A5 using sgRNA472 and ABE constructs shown in Table 8 (which are designated by their pNMG reference numbers as shown in Table 4). The sequence corresponds to SEQ ID NO: 512.
[0193] [Figure 183] Figure 183 shows %A→G editing of A5 using sgRNA508 and ABE constructs shown in Table 8 (which are designated by their pNMG reference numbers as shown in Table 4). The sequence corresponds to SEQ ID NO: 520.
[0194] [Figure 184] Figure 184 shows %A→G editing of A5 using sgRNA536 and ABE constructs shown in Table 8 (which are designated by their pNMG reference numbers as shown in Table 4). The sequence corresponds to SEQ ID NO:530.
[0195] [Figure 185] Figure 185 shows the %A→G editing of the highlighted A (A5) using sgRNA:310, sgRNA:311, sgRNA:314, sgRNA:318, sgRNA:463, and sgRNA:464 for each of the indicated editing elements (which are designated by their pNMG reference numbers as shown in Table 4). Sequences from left to right and top to bottom correspond to SEQ ID NOs: 489, 490, 493, 497, 503, and 504, respectively.
[0196] [Figure 186]Figure 186 shows the %A→G editing of the highlighted A (A5) using sgRNA:466, sgRNA:467, sgRNA:468, sgRNA:471, sgRNA:501, and sgRNA:601 for each of the indicated editing elements (which are designated by their pNMG reference numbers as shown in Table 4). Sequences from left to right and top to bottom correspond to SEQ ID NOs: 506, 507, 508, 511, 513, and 535, respectively.
[0197] definition As used herein and in the claims, the singular forms "a," "an," and "the" include either the singular or the plural unless the context clearly dictates otherwise. Thus, for example, reference to "an agent" includes a singular agent as well as a plurality of such agents.
[0198] The term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as bacteria. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase. In some embodiments, the adenosine deaminase is from a cell such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated E. coli TadA deaminase, e.g., the truncated ecTadA may be missing one or more N-terminal amino acids compared to full-length ecTadA.In some embodiments, the truncated ecTadA may be missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acids compared to full-length ecTadA. In some embodiments, the truncated ecTadA may be missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acids compared to full-length ecTadA. In some embodiments, the ecTadA deaminase does not include an N-terminal methionine.
[0199] In some embodiments, the TadA deaminase is an N-terminally truncated TadA. In certain embodiments, the adenosine deaminase has the amino acid sequence: [ka] Includes:
[0200] In some embodiments, the TadA deaminase is a full-length E. coli TadA deaminase. For example, in some embodiments, the adenosine deaminase has the amino acid sequence: [ka] Includes:
[0201] However, it should be understood that additional adenosine deaminases useful in this application will be apparent to one of skill in the art and are within the scope of this disclosure. For example, the adenosine deaminase may be a homolog of ADAT. Exemplary ADAT homologs include, but are not limited to: [ka] [ka]
[0202] The term "base editor (BE)," or "nucleobase editor (NBE)," refers to an agent comprising a polypeptide capable of modifying a base (e.g., A, T, C, G, or U) in a nucleic acid sequence (e.g., DNA or RNA). In some embodiments, a base editor is capable of deaminating a base in a nucleic acid. In some embodiments, a base editor is capable of deaminating a base in a DNA molecule. In some embodiments, a base editor is capable of deaminating an adenine (A) in DNA. In some embodiments, a base editor is a fusion protein comprising a nucleic acid-programmable DNA-binding protein (napDNAbp) fused to an adenosine deaminase. In some embodiments, a base editor is a Cas9 protein fused to an adenosine deaminase. In some embodiments, a base editor is a Cas9 nickase (nCas9) fused to an adenosine deaminase. In some embodiments, the base editor is fused to a nuclease-inactive Cas9 (dCas9) adenosine deaminase. In some embodiments, the base editor is fused to an inhibitor of base excision repair, such as a UGI domain or a dISN domain. In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and an inhibitor of base excision repair, such as a UGI or dISN domain. In some embodiments, the dCas9 domain of the fusion protein comprises D10A and H840A mutations of SEQ ID NO: 52, or a corresponding mutation in any of SEQ ID NOs: 108-357, which inactivate the nuclease activity of the Cas9 protein. In some embodiments, the fusion protein comprises a D10A mutation and a histidine at residue 840 of SEQ ID NO: 52, or a corresponding mutation in any of SEQ ID NOs: 108-357, which enable Cas9 to cleave only one strand of a nucleic acid duplex. An example of a Cas9 nickase is set forth in SEQ ID NO: 35.
[0203] The term "linker," as used herein, refers to a bond (e.g., a covalent bond), chemical group, or molecule that links two molecules or moieties, e.g., two domains of a fusion protein, such as a nuclease-inactive Cas9 domain and a nucleic acid editing domain (e.g., adenosine deaminase). In some embodiments, the linker connects the gRNA binding domain (including the Cas9 nuclease domain) of an RNA-programmable nuclease with the catalytic domain of a nucleic acid editing protein. In some embodiments, the linker connects dCas9 with a nucleic acid editing protein. Typically, the linker is located between or flanked by two groups, molecules, or other moieties and is connected to each other via a covalent bond, thereby connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the linker is [ka] In some embodiments, the linker comprises the amino acid sequence [ka] In some embodiments, the linker comprises: [ka] or (XP) n motif, or any combination thereof, where n is independently an integer between 1 and 30, and where X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.
[0204] The term "mutation," as used herein, refers to the substitution of a residue in a sequence (e.g., a nucleic acid sequence or an amino acid sequence) with another residue, or the deletion or insertion of one or more residues in a sequence. Mutations are typically described herein by identifying the original residue in the sequence, followed by the identification of the position of said residue and the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012).
[0205] The term "inhibitor of base repair" or "IBR" refers to a protein capable of inhibiting the activity of a nucleic acid repair enzyme, such as a base excision repair enzyme. In some embodiments, the IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7 EndoI, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is catalytically inactive EndoV or catalytically inactive hAAG.
[0206] The term "uracil glycosylase inhibitor" or "UGI," as used herein, refers to a protein capable of inhibiting the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or the UGI set forth in SEQ ID NO: 3. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. For example, in some embodiments, the UGI domain comprises a fragment of the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the UGI fragment comprises an amino acid sequence comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the UGI comprises an amino acid sequence homologous to the amino acid sequence set forth in SEQ ID NO: 3 or a fragment of the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, proteins comprising UGI or a fragment of UGI, or a homolog of UGI or a UGI fragment, are also referred to as "UGI variants." UGI variants share homology with UGI or a fragment thereof. For example, a UGI variant may be at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type UGI or the UGI set forth in SEQ ID NO:3. In some embodiments, the UGI variant comprises a fragment of UGI such that the fragment is at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type UGI or the corresponding fragment of UGI set forth in SEQ ID NO: 3. In some embodiments, UGI comprises the following amino acid sequence: [ka]
[0207] The term "catalytically inactive inosine-specific nuclease" or "inactive inosine-specific nuclease (dISN)" as used herein refers to a protein capable of inhibiting inosine-specific nucleases. Without wishing to be bound by any particular theory, catalytically inactive inosine glycosylases (e.g., alkyladenosine glycosylase [AAG]) will bind to inosine but will not create an abasic site or remove the inosine, thereby sterically blocking the newly formed inosine moiety from DNA damage / repair mechanisms. In some embodiments, catalytically inactive inosine-specific nucleases may be able to bind to inosine in nucleic acids but will not cleave the nucleic acid. Exemplary catalytically inactive inosine-specific nucleases include, but are not limited to, catalytically inactive alkyladenosine glycosylase (AAG nuclease) (e.g., from humans) and catalytically inactive endonuclease V (EndoV nuclease) (e.g., from E. coli). In some embodiments, the catalytically inactive AAG nuclease comprises an E125Q mutation set forth in SEQ ID NO: 32, or a corresponding mutation in another AAG nuclease. In some embodiments, the catalytically inactive AAG nuclease comprises the amino acid sequence set forth in SEQ ID NO: 32. In some embodiments, the catalytically inactive EndoV nuclease comprises a D35A mutation set forth in SEQ ID NO: 32, or a corresponding mutation in another EndoV nuclease. In some embodiments, the catalytically inactive EndoV nuclease comprises the amino acid sequence set forth in SEQ ID NO: 33. It should be understood that other catalytically inactive inosine-specific nucleases (dISNs) will be apparent to those of skill in the art and are within the scope of the present disclosure.
[0208] [ka]
[0209] [ka]
[0210] The term "nuclear localization sequence" or "NLS" refers to an amino acid sequence that facilitates the import of a protein into a cell nucleus, e.g., by nuclear transport. Nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, NLS sequences are described in International PCT Application PCT / EP2000 / 011690 to Plank et al., filed November 23, 2000, and published May 31, 2001 as WO / 2001 / 038547, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In some embodiments, an NLS is a sequence of the amino acid sequence [ka] Includes:
[0211] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to a protein that binds to a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid that guides the napDNAbp to a specific nucleic acid sequence. For example, a Cas9 protein can bind to a guide RNA that guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the napDNAbp is a class 2 microbial CRISPR-Cas effector. In some embodiments, the napDNAbp is a Cas9 domain, e.g., nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2C3, and Argonaute. However, it should be understood that nucleic acid programmable DNA binding proteins also include nucleic acid programmable proteins that bind to RNA. For example, a napDNAbp can be associated with a nucleic acid that directs the napDNAbp to RNA. Other nucleic acid programmable DNA binding proteins are also within the scope of this disclosure, even if not specifically listed herein.
[0212] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease comprising the Cas9 protein or a fragment thereof (e.g., a protein comprising an active, inactive, or partially active DNA-cleavage domain of Cas9 and / or a gRNA that binds to the Cas9 domain). Cas9 nuclease is also sometimes referred to as casn1 nuclease or CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to the antecedent mobile element, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA acts as a guide for ribonuclease 3-aided processing of the pre-crRNA. The Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to the crRNA is first endonucleolytically cleaved and then exonucleolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically require both proteins and RNAs. However, a single guide RNA ("sgRNA" or simply "gRNA") can be engineered to incorporate both the crRNA and tracrRNA into a single RNA species.See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence, which helps distinguish self from non-self.Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);"CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I, Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., the Cas9 is a nickase.
[0213] Nuclease-inactivated Cas9 proteins can be interchangeably referred to as "dCas9" proteins (nuclease-inactive Cas9). Methods for generating Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337:816-821(2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152(5): 1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can suppress the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28; 152(5): 1173-83 (2013)). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA-cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or fragments thereof are referred to as "Cas9 variants." Cas9 variants share homology with Cas9 or fragments thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9.In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), wherein the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9.
[0214] In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or 1300 amino acids in length. In some embodiments, the wild-type Cas9 is [ka] [ka] [ka] [ka] Corresponds to.
[0215] In some embodiments, the wild-type Cas9 corresponds to or comprises SEQ ID NO:49 (nucleotide) and / or SEQ ID NO:50 (amino acid): [ka] [ka] [ka] [ka]
[0216] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_002737.2, SEQ ID NO: 51 (nucleotide); and Uniport Reference Sequence: Q99ZW2, SEQ ID NO: 52 (amino acid)). [ka] [ka] [ka] [ka]
[0217] In some embodiments, Cas9 is: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1);Prevotella intermedia(NCBI Ref: NC_017861.1);Spiroplasma taiwanense(NCBI Ref: NC_021846.1);Streptococcus iniae(NCBI Ref: NC_021314.1);Belliella baltica(NCBI Ref: NC_018010.1) torquisl(NCBI Ref: NC_018721.1); Cas9 from Streptococcus thermophilus (NCBI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: YP_002344900.1), or Neisseria meningitidis (NCBI Ref: YP_002342100.1), or Cas9 from any other organism.
[0218] In some embodiments, the dCas9 corresponds to, or comprises part or all of, a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain comprises the D10A and H840A mutations of SEQ ID NO: 52, or corresponding mutations in another Cas9. In some embodiments, the dCas9 comprises the amino acid sequence of SEQ ID NO: 53. dCas9 (D10A and H840A): [ka]
[0219] In some embodiments, the Cas9 domain contains a D10A mutation, while the residue at position 840 in the amino acid sequence provided by SEQ ID NO:52 or the corresponding position in any of the amino acid sequences provided by SEQ ID NOs:108-357 remains a histidine. Without wishing to be bound by any particular theory, the presence of the catalytic residue H840 maintains the activity of Cas9 to cleave the non-edited (e.g., non-deaminated) strand containing a T opposite the target A. Restoration of H840 (e.g., from A840 in dCas9) does not result in cleavage of the target strand containing an A. Such Cas9 variants are capable of generating a single-strand DNA break (nick) at a specific location based on the target sequence defined by the gRNA, leading to repair of the non-edited strand and ultimately to a T-to-C change in the non-edited strand. A schematic of this process is shown in Figure 94. Briefly, without wishing to be bound by any particular theory, the A of an AT base pair can be deaminated to inosine (I) by an adenosine deaminase, such as an engineered adenosine deaminase that deaminates adenosine in DNA. Nicking the unedited strand bearing the T prompts removal of the T by the mismatch repair mechanism. A UGI domain or catalytically inactive inosine-specific nuclease (dISN) inhibits (e.g., sterically) the inosine-specific nuclease, which prevents removal of the inosine (I).
[0220] In other embodiments, dCas9 variants are provided that have mutations other than D10A and H840A, resulting in, for example, nuclease-inactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, dCas9 variants or homologs (e.g., variants of SEQ ID NO:53) are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to SEQ ID NO:10. In some embodiments, variants of dCas9 (e.g., variants of SEQ ID NO: 53) are provided and have amino acid sequences that are shorter or longer than SEQ ID NO: 10 by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more.
[0221] In some embodiments, the Cas9 fusion proteins provided herein comprise the full-length amino acid sequence of a Cas9 protein, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein comprise only a fragment thereof, rather than the full-length Cas9 sequence. For example, in some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 fragment, which binds to crRNA and tracrRNA or sgRNA, but does not comprise a functional nuclease domain, for example, in that it comprises only a truncated version of the nuclease domain or no nuclease domain at all.
[0222] Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those of skill in the art.
[0223] These include Case9: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); NC_016786.1);Spiroplasma syrphidicola(NCBI Ref: NC_021284.1);Prevotella intermedia(NCBI Ref: NC_017861.1);Spiroplasma taiwan(NCBI Ref: NC_021846.1); NC_021314.1);Belliella baltica(NCBI Ref: NC_018010.1);Psychroflexus torquisI(NCBI Ref: NC_018721.1);Streptococcus thermophilus(NCBI Ref: YP_820832.1); NP_472073.1);Campylobacter jejuni(NCBI Ref: YP_002344900.1);
[0224] It should be understood that additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including variants and homologs thereof, are within the scope of this disclosure. Exemplary Cas9 proteins include, without limitation, those provided below. In some embodiments, the Cas9 protein is a nuclease-dead Cas9 (dCas9). In some embodiments, the dCas9 comprises the amino acid sequence (SEQ ID NO: 34). In some embodiments, the Cas9 protein is a Cas9 nickase (nCas9). In some embodiments, the nCas9 comprises the amino acid sequence (SEQ ID NO: 35). In some embodiments, the Cas9 protein is a nuclease-active Cas9. In some embodiments, the nuclease-active Cas9 comprises the amino acid sequence (SEQ ID NO: 36). Exemplary catalytically inactive Cas9 (dCas9): [ka] [ka] [ka] [ka]
[0225] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., nanoarchaea), which constitute the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, Cas9 refers to CasX or CasY, as described, for example, in Burstein et al., "New CRISPR Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 (the entire contents of which are incorporated herein by reference). Using metagenomics, a number of CRISPR-Cas systems have been identified, including the first reported Cas9 in the archaea domain of life. This diverse Cas9 protein was found in the little-studied nanoarchaea as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered, which are the most compact systems discovered to date. In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins may be used as nucleic acid programmable DNA-binding proteins (napDNAbp) and are within the scope of the present disclosure.
[0226] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any of the fusion proteins provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 417-419. In some embodiments, the napDNAbp comprises the amino acid sequence of any one of SEQ ID NOs: 417-419. It should be understood that CasX and CasY from other bacterial species may be used in accordance with the present disclosure. [ka] [ka]
[0227] As used herein, the term "effective amount" refers to an amount of a biologically active substance that is sufficient to elicit a desired biological response. For example, in some embodiments, an effective amount of a nucleobase editor can refer to the amount of the nucleobase editor that is sufficient to induce mutations at a target site specifically bound and mutated by the nucleobase editor. In some embodiments, an effective amount of a fusion protein provided herein, such as a fusion protein comprising a nucleic acid-programmable DNA-binding protein and a deaminase domain (e.g., an adenosine deaminase domain), can refer to the amount of the fusion protein that is sufficient to induce editing at a target site specifically bound and edited by the fusion protein. As will be understood by those skilled in the art, the effective amount of a substance, such as a fusion protein, nucleobase editor, deaminase, hybrid protein, protein dimer, complex of a protein (or protein dimer) and a polynucleotide, or polynucleotide, can vary depending on various factors, such as the desired biological response, e.g., the particular allele, genome, or target site to be edited, the cell or tissue to be targeted, and the substance to be used.
[0228] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound comprising a nucleobase and an acid moiety, e.g., a nucleoside, a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides, are linear molecules in which adjacent nucleotides are linked by phosphodiester bonds. In some embodiments, "nucleic acid" refers to an individual nucleic acid residue (e.g., a nucleotide and / or a nucleoside). In some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids can occur naturally, for example, in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule can be a non-naturally occurring molecule, e.g., recombinant DNA or RNA, an artificial chromosome, an engineered genome, or a fragment thereof, or synthetic DNA, RNA, or DNA / RNA hybrid, or can include non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms encompass nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced and optionally purified using recombinant expression systems, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, e.g., analogs with chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise indicated.In some embodiments, nucleic acids are selected from natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).
[0229] As used herein, the term "proliferative disorder" refers to any disorder in which cell or tissue homeostasis is disturbed in that a cell or cell population exhibits an abnormally elevated rate of proliferation. Proliferative disorders include hyperproliferative disorders, such as preneoplastic hyperplasia and neoplastic disorders. Neoplastic disorders are characterized by abnormal cell proliferation and include both benign and malignant neoplasms. Malignant neoplasms are also referred to as cancer.
[0230] The terms "protein," "peptide," and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids in length. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of a chemical entity, such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, or a linker for conjugation, functionalization, or other modification. A protein, peptide, or polypeptide may be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide may be simply a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, synthetic, or a combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide containing protein domains from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) or carboxy-terminal (C-terminal) portion of the fusion protein, thus forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. The protein can include different domains, such as a nucleic acid-binding domain of a nucleic acid-editing protein (e.g., the gRNA-binding domain of Cas9, which guides the protein to bind to a target site) and a nucleic acid cleavage domain or catalytic domain. In some embodiments, the protein includes a proteinaceous portion, such as an amino acid sequence constituting a nucleic acid-binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent. In some embodiments, the protein is complexed with or associated with a nucleic acid, such as RNA. Any of the proteins provided herein can be produced by any method known in the art.For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and are described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4). th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference. The terms "RNA-programmable nuclease" and "RNA-guided nuclease" are used interchangeably herein to refer to a nuclease that forms a complex with (e.g., binds to or associates with) one or more RNAs that are not targets for cleavage. In some embodiments, when an RNA-programmable nuclease is complexed with an RNA, it can be referred to as a nuclease:RNA complex. Typically, the bound RNA(s) is referred to as a guide RNA (gRNA). A gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), although "gRNA" is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and directs binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical to or homologous to the tracrRNA provided by Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Provisional Patent Application USSN 61 / 874,682, entitled "Switchable Cas9 Nucleases And Uses Thereof," filed September 6, 2013, and U.S. Provisional Patent Application USSN 61 / 874,746, entitled "Delivery System For Functional Nucleases," filed September 6, 2013, the entire contents of each of which are incorporated herein by reference. In some embodiments, a gRNA includes two or more of domains (1) and (2) and may be referred to as an "extended gRNA."For example, an extended gRNA will bind, e.g., two or more Cas9 proteins and bind to a target nucleic acid at two or more separate regions as described herein. The gRNA comprises a target-complementary nucleotide sequence that mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex.Additionally, RNA-induced RNA-induced polymerization (CRIS PR vaccine) Cas9 strain strain Streptococcus pyogenes and Cas9(Csn1) "Complete genome sequence of an M1 strain of Streptococcus pyogenes ." Ferretti JJ , McShan WM , Ajdic DJ , Savic DJ , Savic G , Lyon K , Primeaux C , Sezate S , Suvorov AN . Kenton S, Lai HS, Lin SP, Qian Y, Jia HG, Ren Q, Zhu H, Song L, White J, Roe BA, McLaughlin RE, Proc RNA and host factor RNase III." Deltcheva E, Chylinski K, Sharma CM, Gonzales K, Chao Y, Pirzada ZA, Eckert MR, Vogel J, Charpentier E, Nature 471 :602–607(2011);および"A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M, Chylinski K, Fonfara I, Hauer M, Doudna JA, Charpentier E. Science 337:816-821(2012) of which the chemical components of this material are prepared in a light-weight manner.
[0231] Because RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, in principle, these proteins are capable of being targeted to any sequence specified by a guide RNA. Methods for using RNA-programmable nucleases, such as Cas9, for site-specific cleavage (e.g., to modify genomes) are known in the art (e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).
[0232] As used herein, the term "subject" refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cattle, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject can be of either sex and at any stage of development.
[0233] The term "target site" refers to a sequence in a nucleic acid molecule that is deaminated by a deaminase or a fusion protein that includes a deaminase (e.g., a dCas9-adenosine deaminase fusion protein provided herein).
[0234] The terms "treatment," "treat," and "treating" refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder described herein, or one or more symptoms thereof. As used herein, the terms "treatment," "treat," and "treating" refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder described herein, or one or more symptoms thereof. In some embodiments, treatment may be administered after one or more symptoms have occurred and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms, or to inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, e.g., to prevent or delay their recurrence.
[0235] The term "recombinant," as used herein in the context of a protein or nucleic acid, refers to a protein or nucleic acid that is not naturally occurring and is the product of human manipulation. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0236] Detailed Description of the Invention Some aspects of this disclosure relate to proteins that deaminate the nucleic acid base adenine. This disclosure provides adenosine deaminase proteins capable of deaminating (i.e., removing the amine group from) adenosine residues in deoxyadenosine (DNA). For example, the adenosine deaminases provided herein are capable of deaminating adenine residues in DNA. It should be understood that prior to the present invention, no adenosine deaminases capable of deaminating deoxyadenosine in DNA were known. Another aspect of this disclosure provides fusion proteins comprising an adenosine deaminase (e.g., an adenosine deaminase that deaminates deoxyadenosine in DNA described herein) and a domain (e.g., a Cas9 or Cpf1 protein) capable of binding to a specific nucleotide sequence. Deamination of adenosine by adenosine deaminase can lead to point mutations, a process referred to herein as nucleic acid editing. For example, adenosine can be converted to an inosine residue, which typically forms base pairs with cytosine residues. Such fusion proteins can be used for targeted editing of DNA in vitro, e.g., to generate mutant cells or animals; for introducing targeted mutations, e.g., to correct genetic defects in cells ex vivo, e.g., in cells obtained from a subject that are then reintroduced into the same or another subject; and for introducing targeted mutations, e.g., to correct genetic defects or introduce inactivating mutations in disease-associated genes in a subject. By way of example, diseases that can be treated by creating A to G or T to C mutations can be treated by using the nucleobase editors provided herein. The present invention provides deaminases, fusion proteins, nucleic acids, vectors, cells, compositions, methods, kits, systems, and the like, that utilize deaminases and nucleobase editors.
[0237] In some embodiments, nucleobase editors provided herein can be generated by fusing one or more protein domains together, thereby generating a fusion protein. In certain embodiments, the fusion proteins provided herein include one or more features that improve the base editing activity (e.g., efficiency, selectivity, and specificity) of the fusion protein. For example, the fusion proteins provided herein can include a Cas9 domain with reduced nuclease activity. In some embodiments, the fusion proteins provided herein can include a Cas9 domain that cleaves one strand of a double-stranded DNA molecule, also referred to as a Cas9 domain without nuclease activity (dCas9) or a Cas9 nickase (nCas9). Without wishing to be bound by any particular theory, the presence of a catalytic residue (e.g., H840) maintains the activity of Cas9 to cleave the unedited (e.g., non-deaminated) strand containing the T opposite the target A. Mutation of the catalytic residue of Cas9 (e.g., D10 to A10) prevents cleavage of the edited strand containing the target A residue. Such Cas9 variants can generate a single-strand DNA break (nick) at a specific location based on the target sequence defined by the gRNA, leading to repair of the unedited strand and ultimately resulting in a T to C change in the unedited strand. In some embodiments, any of the fusion proteins provided herein further comprise an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease (dISN). Without wishing to be bound by any particular theory, the UGI domain or dISN inhibits or prevents base excision repair of deaminated adenosine residues (e.g., inosine), which may improve the activity or efficiency of the base editor.
[0238] Adenosine deaminase Some aspects of the disclosure provide adenosine deaminases. In some embodiments, the adenosine deaminases provided herein are capable of deaminating adenine. In some embodiments, the adenosine deaminases provided herein are capable of deaminating adenine in deoxyadenosine residues in DNA. The adenosine deaminases may be derived from any suitable organism (e.g., E. coli). In some embodiments, the adenosine deaminases are naturally occurring adenosine deaminases containing one or more mutations corresponding to any of the mutations provided herein (e.g., mutations in ecTadA). One skilled in the art would be able to identify corresponding residues in homologous proteins and their respective encoding nucleic acids by methods known in the art, for example, by sequence alignment and determination of homologous residues. Thus, one skilled in the art would be able to generate mutations in any naturally occurring adenosine deaminases (e.g., having homology to ecTadA) corresponding to any of the mutations described herein, for example, any of the mutations identified in ecTadA. In some embodiments, the adenosine deaminase is derived from a prokaryote. In some embodiments, the adenosine deaminase is derived from a bacterium. In some embodiments, the adenosine deaminase is derived from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is derived from E. coli.
[0239] An exemplary alignment of prokaryotic TadA proteins is shown in Figure 92. Residues highlighted in blue are those potentially important for catalyzing the deamination of A to I on ssDNA. Therefore, it should be understood that any mutation identified in ecTadA provided herein can be made at any homologous residue in another adenine deaminase, for example, a TadA deaminase from another bacterium. Figure 93 shows a relative sequence identity analysis (heat map of sequence identity):
[0240] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein. It should be understood that the adenosine deaminase provided herein may contain one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain having a certain percentage of identity in addition to any mutation or combination thereof described herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations compared to any one of the amino acid sequences set forth in SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical consecutive amino acid residues compared to any one of the amino acid sequences set forth in SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein.
[0241] Evolution #1 and #2 Mutations In some embodiments, the adenosine deaminase comprises a D108X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108G, D108N, D108V, D108A, or D108Y mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase. An exemplary alignment of deaminases is shown in Figure 92. However, it should be understood that additional deaminases may be similarly aligned to identify homologous amino acid residues that can be mutated as provided herein.
[0242] In some embodiments, the adenosine deaminase comprises an A106X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0243] In some embodiments, the adenosine deaminase comprises an E155X mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E155D, E155G, or E155V mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0244] In some embodiments, the adenosine deaminase comprises a D147X mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where the presence of X denotes any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D147Y mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0245] It should be understood that any of the mutations provided herein (e.g., based on the ecTadA amino acid sequence of SEQ ID NO: 1) may be introduced into other adenosine deaminases, such as S. aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). It would be apparent to one of skill in the art how to identify amino acid residues from other adenosine deaminases that are homologous to the mutated residues in ecTadA. Thus, any of the mutations identified in ecTadA can be made in other adenosine deaminases that have homologous amino acid residues. It should also be understood that any of the mutations provided herein can be made individually or in any combination in ecTadA or another adenosine deaminase. For example, the adenosine deaminase may contain a D108N, A106V, E155V, and / or D147Y mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the following group of mutations in ecTadA SEQ ID NO: 1 (groups of mutations are separated by ";"), or a corresponding mutation in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V, and E55V; D108N, A106V, and D147Y; D108N, E55V, and D147Y; A106V, E55V, and D147Y; and D108N, A106V, E55V, and D147Y. However, it should be understood that any combination of the corresponding mutations provided herein can be made in adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one or more mutations shown in Table 4, which identify individual mutations and combinations of mutations made in ecTadA and saTadA. In some embodiments, the adenosine deaminase comprises a mutation or combination of mutations shown in Table 4.
[0246] In some embodiments, the adenosine deaminase includes one or more H8X, T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X, and / or K157X mutations in SEQ ID NO: 1, or a corresponding one or more mutations in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E or A56S, E59G, E85K or E85G, M94L, I95I, V102A, F104L, A106V, R107C or R107H or R107P, D108G or D108N, or D108V, or D108A, or D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D, and / or K157R mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more mutations provided in Figure 11 corresponding to SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the mutation(s) or mutations of any one of Constructs 1-16 shown in Figure 11 or the constructs shown in Table 4 corresponding to SEQ ID NO: 1, or one or more corresponding mutation(s) or mutations in another adenosine deaminase.
[0247] In some embodiments, the adenosine deaminase comprises one or more H8X, D108X, and / or N127X mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more H8Y, D108N, and / or N127S mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase.
[0248] In some embodiments, the adenosine deaminase comprises one or more H8X, R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X, E155X, K161X, Q163X, and / or T166X mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the following mutations in SEQ ID NO:1: H8Y, R26W, M61I, L68Q, M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H or Q154R, E155G or E155V or E155D, K161Q, Q163H, and / or T166P, or one or more corresponding mutations in another adenosine deaminase.
[0249] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, D108X, N127X, D147X, R152X, and Q154X in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E155X, and Q163X in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, N127X, E155X, and T166X in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, A106X, D108X, N127X, E155X, and K161X in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, R126X, L68X, D108X, N127X, D147X, and E155X in SEQ ID NO: 1, or the corresponding mutation(s) or mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, A109X, N127X, and E155X in SEQ ID NO: 1, or the corresponding mutation(s) or mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0250] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, D108N, N127S, D147Y, R152C, and Q154H in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, N127S, E155V, and T166P in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, A106T, D108N, N127S, E155D, and K161Q in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, R126W, L68Q, D108N, N127S, D147Y, and E155V in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase.
[0251] In some embodiments, the adenosine deaminase comprises one or more mutations provided in Figure 16 corresponding to SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the mutation of any one of Figure 16 constructs pNMG-149 to pNMG-154 corresponding to SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108N, D108G, or D108V mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V and D108N mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R107C and D108N mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, N127S, D147Y, and Q154H mutations in SEQ ID NO: 1, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, R24W, D108N, N127S, D147Y, and E155V mutations in SEQ ID NO: 1, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises D108N, D147Y, and E155V mutations in SEQ ID NO: 1, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, and S127S mutations in SEQ ID NO: 1, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the A106V, D108N, D147Y and E155V mutations in SEQ ID NO: 1, or corresponding mutations in another adenosine deaminase.
[0252] In some embodiments, the adenosine deaminase comprises one or more S2X, H8X, I49X, L84X, H123X, N127X, I156X, and / or K160X mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase, where the presence of X denotes any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more S2A, H8Y, I49F, L84F, H123Y, N127S, I156F, and / or K160S mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more mutations provided in Figure 97 corresponding to SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the mutation(s) or mutations of any one of clones 1-3 shown in Figure 97, which correspond to SEQ ID NO: 1, or the corresponding mutation(s) or mutations in another adenosine deaminase.
[0253] In some embodiments, the adenosine deaminase comprises an L84X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an L84F mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0254] In some embodiments, the adenosine deaminase comprises an H123X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H123Y mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0255] In some embodiments, the adenosine deaminase comprises an I157X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an I157F mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0256] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84X, A106X, D108X, H123X, D147X, E155X, and I156X in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2X, I49X, A106X, D108X, D147X, and E155X in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, A106X, D108X, N127X, and K160X in SEQ ID NO: 1, or the corresponding mutation(s) or mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0257] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84F, A106V, D108N, H123Y, D147Y, E155V, and I156F in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V in SEQ ID NO: 1, or a corresponding mutation or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, A106T, D108N, N127S, and K160S in SEQ ID NO: 1, or the corresponding mutation(s) in another adenosine deaminase.
[0258] In some embodiments, the adenosine deaminase includes one or more E25X, R26X, R107X, A142X, and / or A143X mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more E25M, E25D, E25A, E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R07K, R107A, R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q and / or A143R mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more mutations provided in Table 7 corresponding to SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the mutation(s) or mutations of any one of clones 1-22 set forth in Table 7 corresponding to SEQ ID NO: 1, or the corresponding mutation(s) or mutations in another adenosine deaminase.
[0259] In some embodiments, the adenosine deaminase comprises an E25X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E25M, E25D, E25A, E25R, E25V, E25S, or E25Y mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0260] In some embodiments, the adenosine deaminase comprises an R26X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R26G, R26N, R26Q, R26C, R26L, or R26K mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0261] In some embodiments, the adenosine deaminase comprises an R107X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R107P, R07K, R107A, R107N, R107W, R107H, or R107S mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0262] In some embodiments, the adenosine deaminase comprises an A142X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N, A142D, A142G mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0263] In some embodiments, the adenosine deaminase comprises an A143X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, or A143R mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0264] In some embodiments, the adenosine deaminase includes one or more H36X, N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X, and / or K161X mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more H36L, N37T, N37S, P48T, P48L, I49V, R51H, R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N and / or K161T mutations in SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more mutations provided in any one of Figures 125-128 corresponding to SEQ ID NO: 1, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the mutation(s) or mutations of any one of clones 1-11 shown in any one of Figures 125-128, which corresponds to SEQ ID NO: 1, or one or more corresponding mutation(s) or mutations in another adenosine deaminase.
[0265] In some embodiments, the adenosine deaminase comprises an H36X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H36L mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0266] In some embodiments, the adenosine deaminase comprises an N37X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an N37T or N37S mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0267] In some embodiments, the adenosine deaminase comprises a P48X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48T or P48L mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0268] In some embodiments, the adenosine deaminase comprises an R51X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R51H or R51L mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0269] In some embodiments, the adenosine deaminase comprises a S146X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a S146R or S146C mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0270] In some embodiments, the adenosine deaminase comprises a K157X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a K157N mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0271] In some embodiments, the adenosine deaminase comprises a P48X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48S, P48T, or P48A mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0272] In some embodiments, the adenosine deaminase comprises an A142X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0273] In some embodiments, the adenosine deaminase comprises a W23X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a W23R or W23L mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0274] In some embodiments, the adenosine deaminase comprises an R152X mutation in ecTadA SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R152P or R52H mutation in SEQ ID NO: 1, or a corresponding mutation in another adenosine deaminase.
[0275] It should be understood that an adenosine deaminase (e.g., a first or second adenosine deaminase) may include one or more mutations provided in any of the adenosine deaminases (e.g., ecTadA adenosine deaminase) shown in Table 4. In some embodiments, an adenosine deaminase includes a combination of mutations in any of the adenosine deaminases (e.g., ecTadA adenosine deaminase) shown in Table 4. For example, an adenosine deaminase may include mutations H36L, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F, and K157N, as shown in the second ecTadA of clone pNMG-477 (relative to SEQ ID NO: 1). In some embodiments, the adenosine deaminase comprises the following combinations of mutations compared to SEQ ID NO:1, where each mutation in the combination is separated by a "_" and each combination of mutations is between parentheses: [ka]
change
change
change
[0276] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations compared to any one of the amino acid sequences set forth in SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 166 identical stretches of amino acid residues compared to any one of the amino acid sequences set forth in SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein. In some embodiments, the adenosine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein. In some embodiments, the adenosine deaminase consists of the amino acid sequence of any one of SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein.The ecTadA sequence provided below is derived from ecTadA (SEQ ID NO: 1) and lacks the N-terminal methionine (M). The saTadA sequence provided below is derived from saTadA (SEQ ID NO: 8) and lacks the N-terminal methionine (M). For clarity, the amino acid numbering scheme used to identify the various amino acid mutations is derived from ecTadA (SEQ ID NO: 1) of E. coli TadA and saTadA (SEQ ID NO: 8) of S. aureus TadA. Amino acid mutations compared to SEQ ID NO: 1 (ecTadA) or SEQ ID NO: 8 (saTadA) are indicated by underlining. [ka] [ka] [ka] [ka] [ka] [ka] [ka] [ka] [ka] [ka] [ka]
[0277] Nucleobase editor Cas9 domain In some aspects, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain. Without limitation, exemplary Cas9 domains are provided herein. The Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase. In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any one of the amino acid sequences set forth in SEQ ID NOs: 108-357. In some embodiments, the Cas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences set forth in SEQ ID NOs: 108-357. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences set forth in SEQ ID NOs: 108-357.In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical stretches of amino acid residues compared to any one of the amino acid sequences set forth in SEQ ID NOs: 108-357.
[0278] In some embodiments, the Cas9 domain is a nuclease-inactivated Cas9 domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactivated dCas9 domain comprises a D10X mutation and a H840X mutation in the amino acid sequence set forth in SEQ ID NO: 52, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-357, where X is any amino acid change. In some embodiments, the nuclease-inactivated dCas9 domain comprises a D10A mutation and a H840A mutation in the amino acid sequence set forth in SEQ ID NO: 52, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-357. In one example, the nuclease-inactivated Cas9 domain comprises the amino acid sequence set forth in SEQ ID NO: 54 (cloning vector pPlatTET-gRNA2, Accession No. BAV54124). [ka] [ka] See, e.g., Qi et al., Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression. Cell. 2013; 152(5): 1173-83, the entire contents of which are incorporated herein by reference.
[0279] Additional suitable nuclease-inactive dCas9 domains are within the scope of this disclosure and will be apparent to those of skill in the art based on this disclosure and knowledge in the art. Exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference). In some embodiments, the dCas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the dCas9 domains provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences set forth in SEQ ID NOs: 108-357.In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 consecutive amino acid residues compared to any one of the amino acid sequences set forth in SEQ ID NOs: 108-357.
[0280] In some embodiments, the Cas9 domain is a Cas9 nickase. The Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of a double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is base-paired with (complementary to) the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase comprises a D10A mutation of SEQ ID NO: 52, a histidine at position 840, or a mutation in any of SEQ ID NOs: 108-357. As an example, the Cas9 nickase can comprise the amino acid sequence set forth in SEQ ID NO: 35. In some embodiments, the Cas9 nickase cleaves the non-target, non-base-edited strand of a double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is not base-paired with the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase comprises an H840A mutation in SEQ ID NO: 52, with an aspartic acid residue at position 10, or a corresponding mutation in any of SEQ ID NOs: 108-357. In some embodiments, the Cas9 nickase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Additional suitable Cas9 nickases are within the scope of this disclosure and will be apparent to those of skill in the art based on this disclosure and knowledge in the art.
[0281] Cas9 domains with reduced PAM exclusivity Some aspects of the present disclosure provide Cas9 domains with different PAM specificities. Typically, Cas9 proteins, such as Cas9 from S. pyogenes (spCas9), require the canonical NGG PAM sequence to bind to specific nucleic acid regions, where the "N" in "NGG" is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. This can limit their ability to edit desired bases in the genome. In some embodiments, the base editing fusion proteins provided herein may require precise placement, for example, the target base is within a four-base region (e.g., a "deamination window") approximately 15 bases upstream of the PAM. See Komor, AC, et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference. In some embodiments, the deamination window is within a 2, 3, 4, 5, 6, 7, 8, 9, or 10 base region. In some embodiments, the deamination window is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 bases upstream of the PAM. Thus, in some embodiments, any of the fusion proteins provided herein can contain a Cas9 domain capable of binding to a nucleotide sequence that does not contain a canonical (e.g., NGG) PAM sequence. Cas9 domains that bind to non-canonical PAM sequences have been described in the art and will be apparent to those of skill in the art.For example, Cas9 domains that bind to non-canonical PAM sequences are described in Kleinstiver, BP, et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities" Nature 523, 481-485 (2015); and Kleinstiver, BP, et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition" Nature Biotechnology 33, 1293-1298 (2015); the entire contents of each are incorporated herein by reference.
[0282] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is nuclease-active SaCas9, nuclease-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 comprises the amino acid sequence SEQ ID NO:55. In some embodiments, the SaCas9 comprises an N579X mutation of SEQ ID NO:55, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs:108-357, where X is any amino acid other than N. In some embodiments, the SaCas9 comprises an N579A mutation of SEQ ID NO:55, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs:108-357.
[0283] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence with a non-canonical PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence with an NNGRRT PAM sequence, where N=A, T, C, or G, and R=A or G. In some embodiments, the SaCas9 domain comprises one or more of the E781X, N967X, and R1014X mutations of SEQ ID NO: 55, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-357, where X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of the E781K, N967K, and R1014H mutations of SEQ ID NO: 55, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-357. In some embodiments, the SaCas9 domain comprises an E781K, N967K, or R1014H mutation of SEQ ID NO: 55, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-357.
[0284] In some embodiments, the Cas9 domain of any of the fusion proteins provided herein is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of any one of SEQ ID NOs: 55-57. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises the amino acid sequence of any one of SEQ ID NOs: 55-57. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein consists of the amino acid sequence of any one of SEQ ID NOs: 55-57.
[0285] [ka]
[0286] The underlined, bolded residue N579 of SEQ ID NO: 55 can be mutated (e.g., to A579) to obtain a SaCas9 nickase.
[0287] [ka]
[0288] The A579 residue of SEQ ID NO: 56, which can be mutated from N579 of SEQ ID NO: 55 to obtain SaCas9 nickase, is underlined and bolded.
[0289] [ka] [ka]
[0290] Residue A579 of SEQ ID NO:57, which can be mutated from N579 of SEQ ID NO:55 to yield SaCas9 nickase, is underlined and bolded. Residues K781, K967, and H1014 of SEQ ID NO:57, which can be mutated from E781, N967, and R1014 of SEQ ID NO:55 to yield SaKKH Cas9, are underlined and italicized.
[0291] In some embodiments, the Cas9 domain is a Cas9 domain from Streptococcus pyogenes (SpCas9). In some embodiments, the SpCas9 domain is nuclease-active SpCas9, nuclease-inactive SpCas9 (SpCas9d), or SpCas9 nickase (SpCas9n). In some embodiments, the SpCas9 comprises the amino acid sequence SEQ ID NO:58. In some embodiments, the SpCas9 comprises a D9X mutation of SEQ ID NO:58, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs:108-357, where X is any amino acid other than D. In some embodiments, the SpCas9 comprises a D9A mutation of SEQ ID NO:58, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs:108-357. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain can bind to a nucleic acid sequence with a non-canonical PAM. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain can bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations of SEQ ID NO: 58, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-35, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134E, R1334Q, and T1336R mutations of SEQ ID NO: 58, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-35. In some embodiments, the SpCas9 domain comprises the D1134E, R1334Q, and T1336R mutations of SEQ ID NO: 58, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-35. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations of SEQ ID NO: 58, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-35, where X is any amino acid.In some embodiments, the SpCas9 domain comprises one or more of the D1134V, R1334Q, and T1336R mutations of SEQ ID NO: 58, or corresponding mutations in any of the amino acid sequences provided by SEQ ID NOs: 108-35. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, R1334Q, and T1336R mutations of SEQ ID NO: 58, or corresponding mutations in any of the amino acid sequences provided by SEQ ID NOs: 108-35. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, G1217X, R1334X, and T1336X mutations of SEQ ID NO: 58, or corresponding mutations in any of the amino acid sequences provided by SEQ ID NOs: 108-35, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, G1217R, R1334Q, and T1336R mutations of SEQ ID NO: 58, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-35. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, G1217R, R1334Q, and T1336R mutations of SEQ ID NO: 58, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-35.
[0292] In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 58-62. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises the amino acid sequence of any one of SEQ ID NOs: 58-62. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein consists of the amino acid sequence of any one of SEQ ID NOs: 58-62. [ka] [ka] [ka] [ka]
[0293] The E1134, Q1334, and R1336 residues of SEQ ID NO:60, which can be mutated from D1134, R1334, and T1336 of SEQ ID NO:58 to obtain SpEQR Cas9, are underlined and bolded. [ka]
[0294] Residues V1134, Q1334, and R1336 of SEQ ID NO:61, which can be mutated from D1134, R1334, and T1336 of SEQ ID NO:58 to obtain SpVQR Cas9, are underlined and bolded. [ka]
[0295] Residues V1134, R1217, Q1334, and R1336 of SEQ ID NO:62, which can be mutated from D1134, G1217, R1334, and T1336 of SEQ ID NO:58 to yield SpVRER Cas9, are underlined and bolded.
[0296] High-fidelity Cas9 domain Some aspects of the present disclosure provide high-fidelity Cas9 domain nucleobase editors, as provided herein. In some embodiments, the high-fidelity Cas9 domain is an edited Cas9 domain that includes one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA compared to the corresponding wild-type Cas9 domain. Without being bound by any particular theory, a high-fidelity Cas9 domain with reduced electrostatic interactions with the sugar-phosphate backbone of DNA may have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain) includes one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, the Cas9 domain comprises one or more mutations that reduce association between the Cas9 domain and the sugar-phosphate backbone of the DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70% or more.
[0297] In some embodiments, any of the Cas9 fusion proteins provided herein comprises one or more N497X, R661X, Q695X, and / or Q926X mutations in the amino acid sequence provided by SEQ ID NO: 52, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-357, where X is any amino acid. In some embodiments, any of the Cas9 fusion proteins provided herein comprises one or more N497A, R661A, Q695A, and / or Q926A mutations in the amino acid sequence provided by SEQ ID NO: 52, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-357. In some embodiments, the Cas9 domain comprises a D10A mutation in the amino acid sequence provided by SEQ ID NO: 52, or a corresponding mutation in any of the amino acid sequences provided by SEQ ID NOs: 108-357. In some embodiments, the Cas9 domain (e.g., of any of the fusion proteins provided herein) comprises the amino acid sequence set forth in SEQ ID NO: 62. High-fidelity Cas9s are known in the art and will be apparent to those of skill in the art. For example, high-fidelity Cas9 domains are described in Kleinstiver, BP, et al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016); and Slaymaker, IM, et al. "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015), the entire contents of each of which are incorporated herein by reference.
[0298] It should be understood that any of the base editors provided herein, e.g., any of the adenosine deaminase base editors provided herein, can be converted to a high-fidelity base editor by modifying the Cas9 domain described herein to generate a high-fidelity base editor, e.g., a high-fidelity adenosine base editor. In some embodiments, the high-fidelity Cas9 domain is a dCas9 domain. In some embodiments, the high-fidelity Cas9 domain is an nCas9 domain.
[0299] [ka]
[0300] Nucleic acid programmable DNA binding proteins Some aspects of the present disclosure provide nucleic acid-programmable DNA-binding proteins that can be used to guide proteins such as base editors to specific nucleic acid (e.g., DNA or RNA) sequences. Nucleic acid-programmable DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2C3, and Argonaute. An example of a nucleic acid-programmable DNA-binding protein with PAM specificity different from Cas9 is clustered regularly interspaced short tandem repeats (Cpf1) from Prevotella and Francisella. Like Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 has been shown to mediate robust DNA interference with characteristics distinct from Cas9. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA and utilizing a T-rich protospacer adjacent motif (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA via double-strand breaks in staggered DNA. Among the 16 Cpf1-family proteins, two enzymes from Acidaminococcus and Lachnospiraceae have been shown to have efficient genome editing activity in human cells. Cpf1 proteins are known in the art and have been previously described, for example, in Yamano et al., "Crystal structure of Cpf1 in complex with guide RNA and target DNA," Cell (165) 2016, pp. 949-962; the entire contents of which are incorporated herein by reference.
[0301] Nuclease-inactive Cpf1 (dCpf1) variants, which can be used as guide nucleotide sequence-programmable DNA-binding protein domains, are also useful in the present compositions and methods. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9, but does not have the HNH endonuclease domain, and the N-terminus of Cpf1 does not have the alpha-helix recognition lobe of Cas9. Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference) showed that the RuvC-like domain of Cpf1 is responsible for cleaving both DNA strands, and that inactivation of the RuvC-like domain inactivates the nuclease activity of Cpf1. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 (SEQ ID NO: 382) inactivate the nuclease activity of Cpf1. In some embodiments, a dCpf1 of the present disclosure contains mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A in SEQ ID NO: 376. Any mutation that inactivates the RuvC domain of Cpf1, e.g., a substitution mutation, deletion, or insertion, can be used in accordance with the present disclosure.
[0302] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any of the fusion proteins provided herein can be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactive Cpf1 (dCpf1). In some embodiments, the Cpf1, nCpf1, or dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 376-382. In some embodiments, dCpf1 comprises an amino acid sequence at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 376-382, and includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A in SEQ ID NO: 376. In some embodiments, dCpf1 comprises the amino acid sequence of any one of SEQ ID NOs: 376-382. It should be understood that Cpf1 from other bacterial species may also be used in accordance with the present disclosure. [ka] [ka] [ka] [ka] [ka] [ka] [ka]
[0303] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) is a nucleic acid programmable DNA-binding protein that does not require a canonical (NGG) PAM sequence. In some embodiments, the napDNAbp is an Argonaute protein. One example of such a nucleic acid programmable DNA-binding protein is the Argonaute protein (NgAgo) from Natronobacterium gregoryi. NgAgo is a ssDNA-guided endonuclease. NgAgo binds to approximately 24 nucleotides of 5'-phosphorylated ssDNA (gDNA), guides it to its target site, and generates a DNA double-strand break at the gDNA site. In contrast to Cas9, the NgAgo-gDNA system does not require a protospacer adjacent motif (PAM). Using a nuclease-inactive NgAgo (dNgAgo) can greatly expand the bases that can be targeted. The characterization and use of NgAgo are described in Gao et al., Nat Biotechnol., 2016 Jul;34(7):768-73. PubMed PMID: 27136078; Swarts et al., Nature. 507(7491) (2014):258-61; and Swarts et al., Nucleic Acids Res. 43(10) (2015):5120-9, each of which is incorporated herein by reference. The Natronobacterium gregoryi argonaute is provided by SEQ ID NO: 416. [ka]
[0304] In some embodiments, the napDNAbp is a prokaryotic homolog of an Argonaute protein. Prokaryotic homologs of Argonaute proteins are known and are described, for example, in Makarova K., et al., "Prokaryotic homologs of Argonaute proteins are predicted to function as key components of a novel system of defense against mobile genetic elements," Biol Direct. 2009 Aug 25;4:29. doi: 10.1186 / 1745-6150-4-29, the entire contents of which are hereby incorporated by reference. In some embodiments, the napDNAbp is a Marinotga piezophila Argonaute (MpAgo) protein. The CRISPR-associated Marinotga piezophila Argonaute (MpAgo) protein cleaves single-stranded target sequences using a 5'-phosphorylated guide. The 5' guide is used by all known Argonautes. The crystal structure of the MpAgo-RNA complex shows a guide strand binding site containing residues that interfere with 5' phosphate interactions. This data suggests the evolution of the Argonaute subclass with noncanonical specificity for 5'-hydroxylated guide RNAs. See, e.g., Kaya et al., "A bacterial Argonaute with noncanonical guide RNA specificity," Proc Natl Acad Sci U S A. 2016 Apr 12;113(15):4057-62, the entire contents of which are incorporated herein by reference. It should be understood that other Argonaute proteins may be used and are within the scope of the present disclosure.
[0305] In some embodiments, a nucleic acid programmable DNA-binding protein (napDNAbp) is a single effector of a microbial CRISPR-Cas system. Single effectors of microbial CRISPR-Cas systems include, but are not limited to, Cas9, Cpf1, C2c1, C2c2, and C2c3. Typically, microbial CRISPR-Cas systems are divided into class 1 and class 2 systems. Class 1 systems have multi-subunit effector complexes, while class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three distinct Class 2 CRISPR-Cas systems (C2c1, C2c2, and C2c3) are described in Shmakov et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems," Mol. Cell, 2015 Nov 5; 60(3): 385-397, the entire contents of which are incorporated herein by reference. The effectors of two systems, C2c1 and C2c3, contain a RuvC-like endonuclease domain related to Cpf1. The third system, C2c2, contains an effector based on two HEPN RNase domains. Unlike CRISPR RNA production by C2c1, mature CRISPR RNA production is independent of tracrRNA. C2c1 relies on both CRISPR RNA and tracrRNA for DNA cleavage. Bacterial C2c2 has been shown to possess a unique RNase activity for CRISPR RNA maturation that is distinct from its active single-stranded RNA degradation activity. These RNase functions are distinct from each other and from the CRISPR RNA processing behavior of CPf1.See, e.g., East-Seletsky, et al., "Two distinct RNase activities of CRISPR-C2c2 enable guide RNA processing and RNA detection," Nature, 2016 Oct 13;538(7624):270-273, the entire contents of which are incorporated herein by reference. In vitro biochemical analysis of C2c2 in Leptotrichia shahii indicates that C2c2 is inducible by a single CRISPR RNA and can be programmed to cleave ssRNA targets carrying a complementary protospacer. Catalytic residues in two conserved HEPN domains mediate cleavage. Mutations in the catalytic residues generate catalytically inactive RNA-binding proteins. See, for example, Abudayyeh et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector", Science, 2016 Aug 5; 353(6299), the entire contents of which are incorporated herein by reference.
[0306] The crystal structure of Alicyclobacillus acidoterrastris C2c1 (AacC2c1) has been reported in complex with a chimeric single-molecule guide RNA (sgRNA). See, e.g., Liu et al., "C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism," Mol. Cell, 2017 Jan 19;65(2):310-322, the entire contents of which are incorporated herein by reference. Crystal structures have also been reported for Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a ternary complex. See, e.g., Yang et al., "PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease," Cell, 2016 Dec 15;167(7):1814-1828, the entire contents of which are incorporated herein by reference. On both the target and non-target DNA strands, the catalytically competent conformation of AacC2c1 is trapped and independently located within a single RuvC catalytic pocket, with C2c1-mediated cleavage resulting in a staggered 7-nucleotide cleavage of the target DNA. Structural comparison between the C2c1 ternary complex and its previously identified Cas9 and Cpf1 counterparts illustrates the diversity of mechanisms employed by CRISPR-Cas9 systems.
[0307] In some embodiments, any fusion protein of a nucleic acid programmable DNA binding protein (napDNAbp) provided herein can be a C2c1, C2c2, or C2c3 protein. In some embodiments, the napDNAbp is a C2c1 protein. In some embodiments, the napDNAbp is a C2c2 protein. In some embodiments, the napDNAbp is a C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring C2c1, C2c2, or C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring C2c1, C2c2, or C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 438 or 439. In some embodiments, the napDNAbp comprises the amino acid sequence of any one of SEQ ID NOs: 438 or 439. It should be understood that C2c1, C2c2, or C2c3 from other bacterial species may also be used in accordance with the present disclosure. [ka] [ka] [ka]
[0308] Fusion proteins containing nuclease-programmable DNA-binding proteins and adenosine deaminase Some aspects of the present disclosure provide fusion proteins comprising a nucleic acid programmable DNA binding protein (napDNAbp) and an adenosine deaminase. In some embodiments, any of the fusion proteins provided herein is a base editor. In some embodiments, the napDNAbp is a Cas9 domain, a Cpf1 domain, a CasX domain, a CasY domain, a C2c1 domain, a C2c2 domain, a C2c3 domain, or an Argonaute domain. In some embodiments, the napDNAbp is any of the napDNAbps provided herein. Some aspects of the present disclosure provide fusion proteins comprising a Cas9 domain and an adenosine deaminase. The Cas9 domain can be a Cas9 domain or a Cas9 protein (e.g., dCas9 or nCas9) provided herein. In some embodiments, any of the Cas9 domains or Cas9 proteins (e.g., dCas9 or nCas9) provided herein can be fused to any of the adenosine deaminase provided herein. In some embodiments, the fusion protein has the following structure: NH2-[adenosine deaminase]-[napDNAbp]-COOH; or NH2-[napDNAbp]-[adenosine deaminase]-COOH Includes:
[0309] In some embodiments, a fusion protein comprising adenosine deaminase and napDNAbp (e.g., a Cas9 domain) does not include a linker sequence. In some embodiments, a linker is present between the adenosine deaminase domain and napDNAbp. In some embodiments, the "-" used in the general configuration above indicates the presence of an optional linker. In some embodiments, adenosine deaminase and napDNAbp are fused via any of the linkers provided herein. For example, in some embodiments, adenosine deaminase and napDNAbp are fused via any of the linkers provided in the section below entitled "Linkers." In some embodiments, adenosine deaminase and napDNAbp are fused via a linker comprising 1 to 200 amino acids. In some embodiments, adenosine deaminase and napDNAbp are 1 to 5, 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 80, 1 to 100, 1 to 150, 1 to 200, 5 to 10, 5 to 20, 5 to 30, 5 to 40, 5 to 60, 5 to 80, 5 to 100, 5 to 150, 5 to 200, 10 to 20, 10 to 30, 10 to 40, 10 to 50, 10 to 60, 10 to 80, 10 to 100, 10 to 150, 10 to 200, 20 to 30, 20 to 40, 20 to 50, 20 to 60, 20 to 80, 20 to 1 and fused via a linker comprising an amino acid length of 00, 20-150, 20-200, 30-40, 30-50, 30-60, 30-80, 30-100, 30-150, 30-200, 40-50, 40-60, 40-80, 40-100, 40-150, 40-200, 50-60, 50-80, 50-100, 50-150, 50-200, 60-80, 60-100, 60-150, 60-200, 80-100, 80-150, 80-200, 100-150, 100-200, or 150-200. In some embodiments, the adenosine deaminase and napDNAbp are fused via a linker comprising 4, 16, 32, or 104 amino acids in length. [ka] In some embodiments, the adenosine deaminase and napDNAbp are fused via a linker comprising the amino acid sequence: [ka] In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker is fused via a linker comprising the amino acid sequence [ka] In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence [ka] In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence [ka] In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence [ka] Includes:
[0310] Fusion proteins containing inhibitors of base repair. Some aspects of the present disclosure provide fusion proteins comprising an inhibitor of base repair (IBR). For example, a fusion protein comprising an adenosine deaminase and a nucleic acid-programmable DNA-binding protein can further comprise an inhibitor of base repair. In some embodiments, the IBR comprises an inhibitor of inosine base repair. In some embodiments, the IBR is an inhibitor of inosine base excision repair. In some embodiments, the inhibitor of inosine base excision repair is a catalytically inactive inosine-specific nuclease (dISN).
[0311] In some embodiments, the fusion proteins provided herein further comprise a catalytically inactive inosine-specific nuclease (dISN). In some embodiments, any of the fusion proteins provided herein comprising a napDNAbp (e.g., a nuclease-active Cas9 domain, a nuclease-inactive dCas9 domain, or a Cas9 nickase) and an adenosine deaminase can be further fused to a catalytically inactive inosine-specific nuclease (dISN) either directly or via a linker. Some aspects of the present disclosure provide fusion proteins comprising an adenosine deaminase (e.g., an adenosine deaminase designed to deaminate adenosine in DNA), a napDNAbp (e.g., dCas9 or nCas9), and a dISN. Without being bound by any particular theory, the cellular DNA-repair response to the presence of I:T in heteroduplex DNA may contribute to the reduced nucleobase editing efficiency in cells. For example, AAG can initiate base excision repair, catalyzing the removal of inosine (I) from DNA in cells, most often resulting in the conversion of an I:T pair back to an A:T pair. In some embodiments, catalytically inactive inosine-specific nucleases can bind to inosine in nucleic acids without cleaving the nucleic acid, preventing the removal of inosine residues in DNA (e.g., by cellular DNA repair mechanisms).
[0312] In some embodiments, dISNs inhibit inosine-removing enzymes from excising inosine residues from DNA (e.g., by steric hindrance). For example, catalytically deficient inosine glycosylases (e.g., alkyladenine glycosylase [AAG]) bind to inosine but do not create an abasic site or remove the inosine, thereby sterically preventing potential DNA damage / repair mechanisms from de novo forming the inosine moiety. Thus, the present disclosure contemplates fusion proteins comprising napDNAbp and adenosine deaminase, further fused to a dISN. The present disclosure also contemplates fusion proteins comprising any Cas9 domain, such as a Cas9 nickase (nCas9) domain, a catalytically inactive Cas9 (dCas9) domain, a high-fidelity Cas9 domain, or a Cas9 domain with reduced PAM exclusivity. It should be understood that the use of a dISN can increase the editing efficiency of adenosine deaminases capable of catalyzing the A to I change. For example, a fusion protein containing a dISN domain may be more efficient at deaminating A residues. In some embodiments, the fusion protein has the following structure: NH2-[adenosine deaminase]-[napDNAbp]-[dISN]-COOH; NH2-[adenosine deaminase]-[dISN]-[napDNAbp]-COOH; NH2-[dISN]-[adenosine deaminase]-[napDNAbp]-COOH; NH2-[napDNAbp]-[adenosine deaminase]-[dISN]-COOH; NH2-[napDNAbp]-[dISN]-[adenosine deaminase]-COOH; or NH2-[dISN]-[napDNAbp]-[adenosine deaminase]-COOH Includes:
[0313] In some embodiments, the fusion proteins provided herein do not contain a linker. In some embodiments, a linker is present between two domains or proteins (e.g., adenosine deaminase, napDNAbp, or dISN). In some embodiments, the "-" used in the general configuration above indicates the presence of an optional linker sequence. In some embodiments, the dISN comprises an inosine-specific nuclease with reduced or no nuclease activity. In some embodiments, the dISN has up to 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the nuclease activity of a corresponding (e.g., wild-type) inosine-specific nuclease. In some embodiments, the dISN is a wild-type inosine-specific nuclease that contains one or more mutations that reduce or eliminate the nuclease activity of the wild-type inosine-specific nuclease. Exemplary catalytically inactive inosine-specific nucleases include, but are not limited to, catalytically inactive AAG nucleases and catalytically inactive EndoV nucleases. In some embodiments, the catalytically inactive AAG nuclease contains an E125Q mutation compared to SEQ ID NO: 32, or a corresponding mutation in another AAG nuclease. In some embodiments, the catalytically inactive AAG nuclease contains the amino acid sequence set forth in SEQ ID NO: 32. In some embodiments, the catalytically inactive EndoV nuclease contains a D35A mutation compared to SEQ ID NO: 32, or a corresponding mutation in another EndoV nuclease. In some embodiments, the catalytically inactive EndoV nuclease contains the amino acid sequence set forth in SEQ ID NO: 33. It should be understood that other catalytically inactive inosine-specific nucleases (dISNs) will be apparent to those skilled in the art and are within the scope of the present disclosure.
[0314] In some embodiments, the dISN proteins provided herein include fragments of dISN proteins and protein homologs of dISN or dISN fragments. For example, in some embodiments, dISN comprises a fragment of the amino acid sequence set forth in SEQ ID NO: 32 or 33. In some embodiments, dISN fragments comprise an amino acid sequence comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid sequence set forth in SEQ ID NO: 32 or 33. In some embodiments, dISN comprises an amino acid sequence homologous to the amino acid sequence set forth in SEQ ID NO: 32 or 33, or an amino acid sequence homologous to a fragment of the amino acid sequence set forth in SEQ ID NO: 32 or 33. In some embodiments, proteins comprising dISN or a fragment of dISN or a homolog of dISN or dISN fragments are referred to as "dISN variants." dISN variants share homology with dISN or fragments thereof. For example, dISN variants are at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type dISN or the dISN set forth in SEQ ID NO: 32 or 33. In some embodiments, dISN variants comprise dISN fragments such that the fragments are at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to the corresponding fragment of wild-type dISN or the dISN set forth in SEQ ID NO: 32 or 33. In some embodiments, dISN has the following amino acid sequence: [ka] Includes:
[0315] Suitable dISN proteins are provided herein, and additional suitable dISN proteins are known to those skilled in the art, including, for example, AAG, EndoV, and variants thereof. It should be understood that additional proteins that prevent or inhibit base excision repair, such as inosine excision, are also within the scope of the present disclosure. In some embodiments, proteins that bind to inosine in DNA are used.
[0316] Some aspects of the present disclosure relate to fusion proteins containing MBD4 or TDG that can be used as inhibitors of base repair. Accordingly, the present disclosure contemplates fusion proteins containing napDNAbp and adenosine deaminase further fused to MBD4 or TDG. The present disclosure also contemplates fusion proteins containing any Cas9 domain, such as a Cas9 nickase (nCas9) domain, a catalytically inactive Cas9 (dCas9) domain, a high-fidelity Cas9 domain, or a Cas9 domain with reduced PAM exclusivity. It should be understood that the use of MBD4 or TDG can increase the editing efficiency of adenosine deaminase capable of catalyzing the A to I change. For example, fusion proteins containing MBD4 or TDG may be more efficient at deaminating A residues. In some embodiments, the fusion protein has the following structure: NH2-[adenosine deaminase]-[napDNAbp]-[MBD4 or TDG]-COOH; NH2-[adenosine deaminase]-[MBD4 or TDG]-[napDNAbp]-COOH; NH2-[MBD4 or TDG]-[adenosine deaminase]-[napDNAbp]-COOH; NH2-[napDNAbp]-[adenosine deaminase]-[MBD4 or TDG]-COOH; NH2-[napDNAbp]-[MBD4 or TDG]-[adenosine deaminase]-COOH; or NH2-[MBD4 or TDG]-[napDNAbp]-[adenosine deaminase]-COOH Includes:
[0317] In some embodiments, the fusion proteins provided herein do not include a linker. In some embodiments, a linker is present between two domains or proteins (e.g., adenosine deaminase, napDNAbp, MBD4, or TDG). In some embodiments, the "-" used in the general configuration above indicates the presence of an optional linker sequence. In some embodiments, the MBD4 or TDG is wild-type MBD4 or TDG. Illustratively, MBD4 and TDG amino acid sequences will be apparent to those skilled in the art and include, but are not limited to, the amino acid sequences of MBD4 and TDG provided below. [ka]
[0318] In some embodiments, the MBD4 or TDG proteins provided herein include fragments of the MBD4 or TDG proteins and protein homologs of the MBD4 or TDG fragments. For example, in some embodiments, the MBD4 or TDG protein comprises a fragment of the amino acid sequence set forth in SEQ ID NO: 689 or 690. In some embodiments, the MBD4 or TDG fragment comprises an amino acid sequence comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid sequence set forth in SEQ ID NO: 689 or 690. In some embodiments, the MBD4 or TDG protein comprises an amino acid sequence homologous to the amino acid sequence set forth in SEQ ID NO: 689 or 690, or an amino acid sequence homologous to a fragment of the amino acid sequence set forth in SEQ ID NO: 689 or 690. In some embodiments, proteins comprising MBD4 or TDG, or fragments of MBD4 or TDG, or homologs of MBD4 or TDG fragments are referred to as "MBD4 variants" or "TDG variants." MBD4 or TDG variants share homology with MBD4 or TDG, or fragments thereof. For example, an MBD4 or TDG variant is at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type MBD4 or TDG or the MBD4 or TDG set forth in SEQ ID NO: 689 or 690.In some embodiments, the MBD4 or TDG variant comprises a fragment of MBD4 or TDG such that the fragment is at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type MBD4 or TDG or the corresponding fragment of MBD4 or TDG set forth in SEQ ID NO: 689 or 690. In some embodiments, the dISN comprises the following amino acid sequence:
[0319] Some aspects of the present disclosure relate to fusion proteins comprising a uracil glycosylase inhibitor (UGI) domain. In some embodiments, any of the fusion proteins provided herein comprising a napDNAbp (e.g., a nuclease-active Cas9 domain, a nuclease-inactive dCas9 domain, or a Cas9 nickase) and an adenosine deaminase may be further fused to a UGI domain, either directly or via a linker. Some aspects of the present disclosure provide fusion proteins comprising an adenosine deaminase (e.g., an edited adenosine deaminase that deaminates deoxyadenosine in DNA), a napDNAbp (e.g., dCas9 or nCas9), and a UGI domain. Without being bound by any particular theory, the cellular DNA-repair response to the presence of I:T in heteroduplex DNA may be a factor in the reduced nucleobase editing efficiency in cells. For example, alkyladenosine glycosylase (AAG) is involved in inosine (I)-linked DNA repair and catalyzes the removal of I from DNA in cells. This can initiate base excision repair, most often resulting in the conversion of an I:T pair back to an A:T pair. A UGI domain can inhibit inosine-removing enzymes from excising inosine residues from DNA (e.g., by steric hindrance). Thus, the present disclosure contemplates fusion proteins comprising a Cas9 domain and an adenosine deaminase domain fused to a UGI domain. The present disclosure also contemplates fusion proteins comprising any nucleic acid-programmable DNA-binding protein, such as a Cas9 nickase (nCas9) domain, a catalytically inactive Cas9 (dCas9) domain, a high-fidelity Cas9 domain, or a Cas9 domain with reduced PAM exclusivity. It should be understood that the use of a UGI domain can increase the editing efficiency of adenosine deaminases capable of catalyzing the A to I change. For example, a fusion protein containing a UGI domain may be more efficient at deaminating adenosine residues. In some embodiments, the fusion protein has the following structure: NH2-[adenosine deaminase]-[napDNAbp]-[UGI]-COOH; NH2-[adenosine deaminase]-[UGI]-[napDNAbp]-COOH; NH2-[UGI]-[adenosine deaminase]-[napDNAbp]-COOH; NH2-[napDNAbp]-[adenosine deaminase]-[UGI]-COOH; NH2-[napDNAbp]-[UGI]-[adenosine deaminase]-COOH; or NH2-[UGI]-[napDNAbp]-[adenosine deaminase]-COOH Includes:
[0320] In some embodiments, the fusion proteins provided herein do not include a linker. In some embodiments, a linker is present between any of the domains or proteins (e.g., adenosine deaminase, napDNAbp, and / or UGI domains). In some embodiments, the "-" used in the general configuration above indicates the presence of an optional linker.
[0321] In some embodiments, the UGI domain comprises wild-type UGI or the UGI set forth in SEQ ID NO: 3. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. For example, in some embodiments, the UGI domain comprises a fragment of the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the UGI fragment comprises an amino acid sequence comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, UGI comprises an amino acid sequence homologous to the amino acid sequence set forth in SEQ ID NO: 3 or an amino acid sequence homologous to a fragment of the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, proteins comprising UGI or a fragment of UGI or a homolog of UGI or a UGI fragment are referred to as "UGI variants." UGI variants share homology with UGI or a fragment thereof. For example, a UGI variant is at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type UGI or the UGI set forth in SEQ ID NO: 3. In some embodiments, the UGI variant comprises a fragment of UGI such that the fragment is at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to the corresponding fragment of wild-type UGI or the UGI set forth in SEQ ID NO: 3. In some embodiments, UGI has the following amino acid sequence: [ka] Includes:
[0322] Suitable UGI protein and nucleotide sequences are provided herein; additional suitable UGI sequences will be known to those of skill in the art, see, e.g., Wang et al., Uracil-DNA glycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNA glycosylase. J. Biol. Chem. 264:1163-1171(1989); Lundquist et al., Site-directed mutagenesis and characterization of uracil-DNA glycosylase inhibitor protein. Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J. Biol. Chem. 272:21408-21419(1997); Ravishankar et al., X-ray analysis of a complex of Escherichia coli uracil DNA glycosylase (EcUDG) with a proteinaceous inhibitor. The structure elucidation of a prokaryotic UDG. Nucleic Acids Res. 26:4880-4887 (1998); and Putnam et al., Protein mimicry of DNA from crystal structures of the uracil-DNA glycosylase inhibitor protein and its complex with Escherichia coli uracil-DNA glycosylase. J. Mol. Biol. 287:331-346 (1999), the entire contents of each of which are incorporated herein by reference.
[0323] It should be understood that additional proteins that prevent or inhibit base excision repair, such as inosine base excision, are also within the scope of the present disclosure. In some embodiments, DNA-binding proteins are used. In other embodiments, substitutions for UGI are used. In some embodiments, the uracil glycosylase inhibitor is a protein that binds to single-stranded DNA. For example, the uracil glycosylase inhibitor can be Erwinia tasmaniensis single-stranded binding protein. In some embodiments, the single-stranded binding protein comprises the amino acid sequence (SEQ ID NO: 29). In some embodiments, the uracil glycosylase inhibitor is a protein that binds to uracil. In some embodiments, the uracil glycosylase inhibitor is a protein that binds to uracil in DNA. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive uracil DNA-glycosylase protein. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive uracil DNA-glycosylase protein that does not excise uracil from DNA. For example, the uracil glycosylase inhibitor is UdgX. In some embodiments, UdgX comprises the amino acid sequence (SEQ ID NO: 30). As another example, the uracil glycosylase inhibitor is catalytically inactive UDG. In some embodiments, the catalytically inactive UDG comprises the amino acid sequence (SEQ ID NO: 31). It should be understood that other uracil glycosylase inhibitors will be apparent to those of skill in the art and are within the scope of the present disclosure. In some embodiments, the uracil glycosylase inhibitor is a protein homologous to any one of SEQ ID NOs: 29-31. In some embodiments, the uracil glycosylase inhibitor is a protein that is at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical to any one of SEQ ID NOs: 29-31. [ka]
[0324] Fusion protein containing nuclear localization sequence (NLS) In some embodiments, the fusion proteins provided herein further comprise one or more nuclear targeting sequences, e.g., a nuclear localization sequence (NLS). In some embodiments, the NLS comprises an amino acid sequence that facilitates import of the protein, including the NLS, into the cell nucleus (e.g., by nuclear transport). In some embodiments, any of the fusion proteins provided herein further comprise a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of an IBR (e.g., dISN). In some embodiments, the NLS is fused to the C-terminus of an IBR (e.g., dISN). In some embodiments, the NLS is fused to the N-terminus of a napDNAbp. In some embodiments, the NLS is fused to the C-terminus of adenosine deaminase. In some embodiments, the NLS is fused to the C-terminus of adenosine deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, the NLS comprises the amino acid sequence of any one of the NLS sequences provided or referenced herein. In some embodiments, the NLS comprises the amino acid sequence set forth in SEQ ID NO: 4 or SEQ ID NO: 5. Additional nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690, the contents of which are incorporated herein by reference for disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS is [ka] It contains the amino acid sequence of
[0325] In some embodiments, the general configuration of an exemplary fusion protein involving adenosine deaminase and napDNAbp comprises any one of the following structures, where the NLS is a nuclear localization sequence (e.g., any NLS provided herein), NH2 is the N-terminus of the fusion protein, and COOH is the C-terminus of the fusion protein:
[0326] A fusion protein containing adenosine deaminase, napDNAbp, and an NLS NH2-[NLS]-[adenosine deaminase]-[napDNAbp]-COOH; NH2-[adenosine deaminase]-[NLS]-[napDNAbp]-COOH; NH2-[adenosine deaminase]-[napDNAbp]-[NLS]-COOH; NH2-[NLS]-[napDNAbp]-[adenosine deaminase]-COOH; NH2-[napDNAbp]-[NLS]-[adenosine deaminase]-COOH; NH2-[napDNAbp]-[adenosine deaminase]-[NLS]-COOH;
[0327] In some embodiments, the fusion proteins provided herein do not include a linker. In some embodiments, a linker is present between one or more of the domains or proteins (e.g., adenosine deaminase, napDNAbp, and / or NLS). In some embodiments, the "-" used in the general configuration above indicates the presence of an optional linker.
[0328] A fusion protein containing adenosine deaminase, napDNAbp, and inhibitor of base repair (IBR) NH2-[IBR]-[adenosine deaminase]-[napDNAbp]-COOH; NH2-[adenosine deaminase]-[IBR]-[napDNAbp]-COOH; NH2-[adenosine deaminase]-[napDNAbp]-[IBR]-COOH; NH2-[IBR]-[napDNAbp]-[adenosine deaminase]-COOH; NH2-[napDNAbp]-[IBR]-[adenosine deaminase]-COOH; NH2-[napDNAbp]-[adenosine deaminase]-[IBR]-COOH;
[0329] In some embodiments, the fusion proteins provided herein do not include a linker. In some embodiments, a linker is present between one or more of the domains or proteins (e.g., adenosine deaminase, napDNAbp, and / or IBR). In some embodiments, the "-" used in the general configuration above indicates the presence of an optional linker.
[0330] A fusion protein containing adenosine deaminase, napDNAbp, inhibitor of base repair (IBR), and NLS NH2-[IBR]-[NLS]-[adenosine deaminase]-[napDNAbp]-COOH; NH2-[NLS]-[IBR]-[adenosine deaminase]-[napDNAbp]-COOH; NH2-[NLS]-[adenosine deaminase]-[IBR]-[napDNAbp]-COOH; NH2-[NLS]-[adenosine deaminase]-[napDNAbp]-[IBR]-COOH; NH2-[IBR]-[adenosine deaminase]-[NLS]-[napDNAbp]-COOH; NH2-[adenosine deaminase]-[IBR]-[NLS]-[napDNAbp]-COOH; NH2-[adenosine deaminase]-[NLS]-[IBR]-[napDNAbp]-COOH; NH2-[adenosine deaminase]-[NLS]-[napDNAbp]-[IBR]-COOH; NH2-[IBR]-[adenosine deaminase]-[napDNAbp]-[NLS]-COOH; NH2-[adenosine deaminase]-[IBR]-[napDNAbp]-[NLS]-COOH; NH2-[adenosine deaminase]-[napDNAbp]-[IBR]-[NLS]-COOH; NH2-[adenosine deaminase]-[napDNAbp]-[NLS]-[IBR]-COOH; NH2-[IBR]-[NLS]-[napDNAbp]-[adenosine deaminase]-COOH; NH2-[NLS]-[IBR]-[napDNAbp]-[adenosine deaminase]-COOH; NH2-[NLS]-[napDNAbp]-[IBR]-[adenosine deaminase]-COOH; NH2-[NLS]-[napDNAbp]-[adenosine deaminase]-[IBR]-COOH; NH2-[IBR]-[napDNAbp]-[NLS]-[adenosine deaminase]-COOH; NH2-[napDNAbp]-[IBR]-[NLS]-[adenosine deaminase]-COOH; NH2-[napDNAbp]-[NLS]-[IBR]-[adenosine deaminase]-COOH; NH2-[napDNAbp]-[NLS]-[adenosine deaminase]-[IBR]-COOH; NH2-[IBR]-[napDNAbp]-[adenosine deaminase]-[NLS]-COOH; NH2-[napDNAbp]-[IBR]-[adenosine deaminase]-[NLS]-COOH; NH2-[napDNAbp]-[adenosine deaminase]-[IBR]-[NLS]-COOH; NH2-[napDNAbp]-[adenosine deaminase]-[NLS]-[IBR]-COOH;
[0331] In some embodiments, the fusion proteins provided herein do not include a linker. In some embodiments, a linker is present between one or more of the domains or proteins (e.g., adenosine deaminase, napDNAbp, NLS, and / or IBR). In some embodiments, the "-" used in the general configuration above indicates the presence of an optional linker.
[0332] Some aspects of the present disclosure provide fusion proteins comprising a nucleic acid programmable DNA binding protein (napDNAbp) and at least two adenosine deaminase domains. Without being bound by any particular theory, dimerization of adenosine deaminases (e.g., in cis or trans) may improve the ability (e.g., efficiency) of the fusion protein to modify nucleic acid bases, such as deaminating adenine. In some embodiments, any of the fusion proteins may contain two, three, four, or five adenosine deaminase domains. In some embodiments, any of the fusion proteins provided herein contains two adenosine deaminases. In some embodiments, any of the fusion proteins provided herein contains only two adenosine deaminases. In some embodiments, the adenosine deaminases are the same. In some embodiments, the adenosine deaminases are any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminases are different. In some embodiments, the first adenosine deaminase is any of the adenosine deaminases provided herein, and the second adenosine deaminase is any of the adenosine deaminases provided herein, but is not identical to the first adenosine deaminase. As one example, a fusion protein can include a first adenosine deaminase and a second adenosine deaminase, both of which comprise the amino acid sequence of SEQ ID NO: 72, which contains the A106V, D108N, D147Y, and E155V mutations from ecTadA (SEQ ID NO: 1). As another example, the fusion protein may comprise a first adenosine deaminase domain comprising the amino acid sequence of SEQ ID NO: 72, which contains the A106V, D108N, D147Y, and E155V mutations from ecTadA (SEQ ID NO: 1), and a second adenosine deaminase comprising the amino acid sequence of SEQ ID NO: 421, which contains the L84F, A106V, D108N, H123Y, D147Y, E155V, and I156F mutations from ecTadA (SEQ ID NO: 1).
[0333] In some embodiments, the fusion protein comprises two adenosine deaminase (e.g., a first adenosine deaminase and a second adenosine deaminase). In some embodiments, the fusion protein comprises a first adenosine deaminase and a second adenosine deaminase. In some embodiments, in the fusion protein, the first adenosine deaminase is N-terminal to the second adenosine deaminase. In some embodiments, in the fusion protein, the first adenosine deaminase is C-terminal to the second adenosine deaminase. In some embodiments, the first adenosine deaminase and the second deaminase are fused directly or via a linker. In some embodiments, the linker is any of the linkers provided herein, for example, any of the linkers described in the "Linkers" section. In some embodiments, the linker comprises any one of the amino acid sequences of SEQ ID NOs: 10, 37-40, 384-386, or 685-688. In some embodiments, the first adenosine deaminase is the same as the second adenosine deaminase. In some embodiments, the first adenosine deaminase and the second adenosine deaminase are any of the adenosine deaminases described herein. In some embodiments, the first adenosine deaminase and the second adenosine deaminase are different. In some embodiments, the first adenosine deaminase is any of the adenosine deaminases provided herein. In some embodiments, the second adenosine deaminase is any of the adenosine deaminases provided herein, but is not identical to the first adenosine deaminase. In some embodiments, the first adenosine deaminase is ecTadA adenosine deaminase.In some embodiments, the first adenosine deaminase is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein. In some embodiments, the first adenosine deaminase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the second adenosine deaminase is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences set forth in SEQ ID NOs: 1, 64-84, 420-437, 672-684, or any adenosine deaminase provided herein. In some embodiments, the second adenosine deaminase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the first adenosine deaminase and the second adenosine deaminase of the fusion protein comprise a mutation in ecTadA (SEQ ID NO: 1) and a corresponding mutation in another adenosine deaminase as shown in any one of the constructs provided in Table 4 (e.g., pNMG-371, pNMG-477, pNMG-576, pNMG-586, and pNMG-616). In some embodiments, the fusion protein comprises two adenosine deaminase (e.g., a first adenosine deaminase and a second adenosine deaminase) in any one of the constructs in Table 4 (e.g., pNMG-371, pNMG-477, pNMG-576, pNMG-586, and pNMG-616).
[0334] In some embodiments, the general configuration of an exemplary fusion protein involving a first adenosine deaminase, a second adenosine deaminase, and napDNAbp comprises any one of the following structures, wherein the NLS is a nuclear localization sequence (e.g., any NLS provided herein):
[0335] A fusion protein comprising a first adenosine deaminase, a second adenosine deaminase, and napDNAbp. NH2-[first adenosine deaminase]-[second adenosine deaminase]-[napDNAbp]-COOH; NH2-[first adenosine deaminase]-[napDNAbp]-[second adenosine deaminase]-COOH; NH2-[napDNAbp]-[first adenosine deaminase]-[second adenosine deaminase]-COOH; NH2-[second adenosine deaminase]-[first adenosine deaminase]-[napDNAbp]-COOH; NH2-[second adenosine deaminase]-[napDNAbp]-[first adenosine deaminase]-COOH; NH2-[napDNAbp]-[second adenosine deaminase]-[first adenosine deaminase]-COOH;
[0336] In some embodiments, the fusion proteins provided herein do not include a linker. In some embodiments, a linker is present between one or more domains or proteins (e.g., a first adenosine deaminase, a second adenosine deaminase, and / or napDNAbp). In some embodiments, as used in the general structure above, "-" indicates the presence of an optional linker.
[0337] The fusion protein comprises a first adenosine deaminase, a second adenosine deaminase, napDNAbp, and an NLS. NH2-[NLS]-[first adenosine deaminase]-[second adenosine deaminase]-[napDNAbp]-COOH; NH2-[first adenosine deaminase]-[NLS]-[second adenosine deaminase]-[napDNAbp]-COOH; NH2-[first adenosine deaminase]-[second adenosine deaminase]-[NLS]-[napDNAbp]-COOH; NH2-[first adenosine deaminase]-[second adenosine deaminase]-[napDNAbp]-[NLS]-COOH; NH2-[NLS]-[first adenosine deaminase]-[napDNAbp][second adenosine deaminase]-COOH; NH2-[first adenosine deaminase]-[NLS]-[napDNAbp]-[second adenosine deaminase]-COOH; NH2-[first adenosine deaminase]-[napDNAbp]-[NLS]-[second adenosine deaminase]-COOH; NH2-[first adenosine deaminase]-[napDNAbp]-[second adenosine deaminase]-[NLS]-COOH; NH2-[NLS]-[napDNAbp]-[first adenosine deaminase]-[second adenosine deaminase]-COOH; NH2-[napDNAbp]-[NLS]-[first adenosine deaminase]-[second adenosine deaminase]-COOH; NH2-[napDNAbp]-[first adenosine deaminase]-[NLS]-[second adenosine deaminase]-COOH; NH2-[napDNAbp]-[first adenosine deaminase]-[second adenosine deaminase]-[NLS]-COOH; NH2-[NLS]-[second adenosine deaminase]-[first adenosine deaminase]-[napDNAbp]-COOH; NH2-[second adenosine deaminase]-[NLS]-[first adenosine deaminase]-[napDNAbp]-COOH; NH2-[second adenosine deaminase]-[first adenosine deaminase]-[NLS]-[napDNAbp]-COOH; NH2-[second adenosine deaminase]-[first adenosine deaminase]-[napDNAbp]-[NLS]-COOH; NH2-[NLS]-[second adenosine deaminase]-[napDNAbp]-[first adenosine deaminase]-COOH; NH2-[second adenosine deaminase]-[NLS]-[napDNAbp]-[first adenosine deaminase]-COOH; NH2-[second adenosine deaminase]-[napDNAbp]-[NLS]-[first adenosine deaminase]-COOH; NH2-[second adenosine deaminase]-[napDNAbp]-[first adenosine deaminase]-[NLS]-COOH; NH2-[NLS]-[napDNAbp]-[second adenosine deaminase]-[first adenosine deaminase]-COOH; NH2-[napDNAbp]-[NLS]-[second adenosine deaminase]-[first adenosine deaminase]-COOH; NH2-[napDNAbp]-[second adenosine deaminase]-[NLS]-[first adenosine deaminase]-COOH; NH2-[napDNAbp]-[second adenosine deaminase]-[first adenosine deaminase]-[NLS]-COOH;
[0338] In some embodiments, the fusion proteins provided herein do not include a linker. In some embodiments, a linker is present between one or more domains or proteins (e.g., a first adenosine deaminase, a second adenosine deaminase, a napDNAbp, and / or an NLS). In some embodiments, as used in the general structure above, "-" indicates the presence of an optional linker.
[0339] It should be appreciated that the fusion proteins of the present disclosure may include one or more additional features. For example, in some embodiments, the fusion protein may include, for example, a cytoplasmic localization sequence, an export sequence, a nuclear export sequence, or other localization sequence, and a sequence tag useful for solubilizing, purifying, or detecting the fusion protein. Suitable protein tags provided herein include, but are not limited to, a biotin carboxylase carrier protein (BCCP) tag, a myc tag, a calmodulin tag, a FLAG tag, a hemagglutinin (HA) tag, a polyhistidine tag, also referred to as a histidine tag or His tag, a maltose-binding protein (MBP) tag, a nus tag, a glutathione-S-transferase (GST) tag, a green fluore...
Claims
1. 1. An adenosine deaminase capable of deaminating adenine of deoxyadenosine in deoxyribonucleic acid (DNA), adenosine deaminase comprising one or more mutations with reference to SEQ ID NO: 1, wherein the one or more mutations are S2A, H8Y, T17S, L18E, W23R, W23L, W23G, D24G, E25M, E25D, E25A, E25R, E25V, E25S, E25Y, E25G, R26W, R26G, R26N, R26Q, R26C, R26L, R26K, L34S, H36L, N37T, N37S, W45L, P48A, P48S, P48L, P48T, I49F , I49V, R51L, R51H, R52H, A56E, A56S, E59A, E59G, M61I, G67V, L68Q, M70V, M70L, Q71R, Q71L, N72S, N72D, R74A, R74Q, D77G, L8 4F, E85K, E85G, A91T, M94L, I95L, H96L, S97C, R98Q, V102A, F104I, F104L, A106V, A106T, R107C, R107H, R107N, R107K, R107P, R 107A, R107W, R107S, D108Y, D108N, D108G, D108R, D108Q, D108M, D108L, D108K, D108I, D108F, D108A, D108V, A109T, K110I, M1 18K, H123Y, G125A, N127S, R129Q, E134G, L137M, A138V, A142N, A142G, A142D, A143D, A143G, A143E, A143L, A143W, A143M, A14 3S, A143Q, A143R, S146C, S146T, S146R, D147Y, F149Y, M151V, R152P, R152H, R152C, R153C, Q154H, Q154L, Q154R, E155V, E155G, E155D, I156F, I156Y, I156D, K157N, K157R, L157N, Q159L, K160S, K160E, K161Q, K161T, Q163H, and T166P; and The adenosine deaminase is TadA deaminase. The adenosine deaminase.
2. 2. The adenosine deaminase of claim 1, wherein the one or more mutations are selected from the group consisting of W23R, R26W, R26G, R26N, H36L, P48T, P48A, R51L, M70L, L84F, A106V, D108Y, D108N, D108G, D108R, D108Q, D108L, D108A, H123Y, S146T, S146R, S146C, D147Y, F149Y, R152P, E155D, E155V, I156F, K157N, and K157R.
3. 3. The adenosine deaminase of claim 1 or 2, wherein the one or more mutations are selected from the group consisting of W23R, R26G, H36L, P48A, R51L, M70L, L84F, A106V, D108N, H123Y, R129Q, E134G, S146C, F149Y, R152P, E155V, I156F, and K157N.
4. The adenosine deaminase according to any one of claims 1 to 3, wherein the one or more mutations are selected from the group consisting of W23R, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, F149Y, R152P, E155V, I156F, and K157N.
5. 1. An adenosine deaminase capable of deaminating an adenine of a deoxyadenosine in a deoxyribonucleic acid (DNA), wherein the adenosine deaminase comprises two or more mutations with reference to SEQ ID NO:1, wherein the two or more mutations are selected from the group consisting of W23R, R26G, H36L, P48A, R51L, M70L, L84F, A106V, D108N, H123Y, R129Q, E134G, S146C, F149Y, R152P, E155V, I156F, and K157N.
6. 1. An adenosine deaminase capable of deaminating adenine of deoxyadenosine in deoxyribonucleic acid (DNA), wherein the adenosine deaminase comprises two or more mutations with reference to SEQ ID NO:1, wherein the two or more mutations comprise two or more mutations selected from the group consisting of W23R, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, F149Y, R152P, E155V, I156F, and K157N.
7. The adenosine deaminase according to any one of claims 1 to 6, wherein the adenosine deaminase is derived from a bacterium.
8. A fusion protein comprising a first adenosine deaminase and a nucleic acid programmable DNA binding protein (napDNAbp), wherein the first adenosine deaminase is the adenosine deaminase described in any one of claims 1 to 7.
9. The fusion protein of claim 8, wherein the napDNAbp is Cas9, Cpf1, Casx, Casy, C2c1, C2c2, or C2c3.
10. The fusion protein of claim 8 or 9, wherein the napDNAbp is nuclease-inactive or a nickase.
11. 11. The fusion protein of claim 10, wherein the napDNAbp domain comprises an inactive Cas9 domain, a Cas9 nickase, or a nuclease-active Cas9.
12. 12. The fusion protein of claim 11, wherein the inactive Cas9 domain, Cas9 nickase, or nuclease-active Cas9 is capable of binding to a nucleotide sequence that does not contain a canonical PAM sequence.
13. The fusion protein according to any one of claims 8 to 12, further comprising a second adenosine deaminase.
14. 14. The fusion protein of claim 13, wherein the first adenosine deaminase and the second adenosine deaminase are different.
15. structure: NH2-[first adenosine deaminase]-[second adenosine deaminase]-[napDNAbp]-COOH; NH2-[first adenosine deaminase]-[napDNAbp]-[second adenosine deaminase]-COOH; or NH2-[napDNAbp]-[first adenosine deaminase]-[second adenosine deaminase]-COOH, The fusion protein of claim 13 or 14, comprising:
16. The fusion protein according to any one of claims 13 to 15, wherein the second adenosine deaminase is TadA adenosine deaminase.
17. The fusion protein of any one of claims 13 to 16, wherein the second adenosine deaminase comprises an amino acid sequence that is at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence represented by SEQ ID NO: 1, 8, 9, 371, 372, 373, 374, or 375.
18. A complex comprising the fusion protein of any one of claims 8 to 17 and a guide RNA bound to the nucleic acid programmable DNA binding protein (napDNAbp) of the fusion protein.
19. 19. The complex of claim 18, wherein the target nucleic acid sequence is in the genome of a prokaryote or a eukaryote.
20. A polynucleotide encoding the adenosine deaminase according to any one of claims 1 to 7 or the fusion protein according to any one of claims 8 to 17.
21. 21. A vector comprising the polynucleotide of claim 20, wherein the vector comprises a heterologous promoter driving expression of the polynucleotide.
22. A cell comprising the adenosine deaminase according to any one of claims 1 to 7, the fusion protein according to any one of claims 8 to 17, the complex according to claim 18 or 19, the polynucleotide according to claim 20, or the vector according to claim 21.
23. A pharmaceutical composition comprising the adenosine deaminase according to any one of claims 1 to 7, the fusion protein according to any one of claims 8 to 17, the complex according to claim 18 or 19, the polynucleotide according to claim 20, the vector according to claim 21, or the cell according to claim 22.
24. 24. The adenosine deaminase according to any one of claims 1 to 7, the fusion protein according to any one of claims 8 to 17, the complex according to claim 18 or 19, the polypeptide according to claim 20, the vector according to claim 21, the cell according to claim 22, or the pharmaceutical composition according to claim 23, for use as a pharmaceutical.
25. 1. An in vitro or ex vivo method for editing nucleobases of a DNA sequence, comprising: (a) contacting the DNA sequence with the adenosine deaminase of any one of claims 1 to 7, the fusion protein of any one of claims 8 to 17, or the complex of claim 18 or 19, thereby converting a first nucleobase of the DNA sequence to a second nucleobase. The method.
26. 26. The method of claim 25, wherein the activity of the complex results in the correction of a point mutation in a nucleic acid molecule associated with a disease or disorder.
27. 27. The method of claim 25 or 26, wherein the first nucleobase is adenine.
28. 28. The method of any one of claims 25 to 27, wherein the second nucleobase is inosine.
29. 29. The method of any one of claims 25 to 28, wherein a third nucleobase complementary to the first nucleobase is replaced by a fourth nucleobase complementary to the second nucleobase.
30. 29. The method of any one of claims 25 to 28, wherein the second nucleobase is replaced with a fifth nucleobase that is complementary to the fourth nucleobase.
31. 31. The method of claim 30, wherein the fifth nucleobase is guanine.
32. 32. The method of any one of claims 25 to 31, wherein the DNA sequence comprises a point mutation associated with a disease or disorder.
33. 33. The method of any one of claims 25-32, wherein the DNA sequence comprises a G to A point mutation associated with a disease or disorder; and deamination of the mutated A base results in a sequence that is not associated with the disease or disorder.
34. 34. The method of any one of claims 25-33, wherein the target nucleic acid sequence comprises a C to T point mutation associated with a disease or disorder; and deamination of the A base that is complementary to the T base of the C to T point mutation results in a sequence that is not associated with the disease or disorder.
35. 35. The method of any one of claims 25-34, wherein the efficiency of deaminating A bases is at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 98%.
36. 36. The method of any one of claims 32-35, wherein the point mutation is in the codon and the contacting step results in a change in the amino acid encoded by the mutant codon compared to the wild-type codon, or results in the removal of a stop codon.
37. 37. The method of any one of claims 25 to 36, wherein the contacting step results in the introduction of a splice site, results in the removal of a splice site, results in the introduction of a mutation in a gene promoter, or results in the introduction of a mutation in a gene repressor.
38. 38. The method of any one of claims 25-37, which results in less than 20%, 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1%, 0.5%, 0.2%, or 0.1% indel formation.