Evolved cytidine deaminase and method for editing DNA using same
Through the development of directional evolution, TadA-CD deaminase and TadCBE, the balance between existing CBEs in high school target activity and low off-target editing is solved, achieving efficient, low off-target cytosine base editing and is sized suitable for delivery through a single AAV vector.
Patent Information
- Application Number
- CN202380071979.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-08-15
- Publication Date
- 2025-05-16
AI Technical Summary
Existing cytosine base editors (CBEs) have high Cas-independent off-target editing problems while maintaining high school target activity, and are large in size, making it difficult to reach cells through certain delivery methods, such as a single AAV vector.
A new class of highly selective cytidine deaminases and CBEs, called TadA-CD deaminases and TadCBEs, has the ability to preferentially deaminases in DNA, and combines the napDNAbp domain to reduce off-target editing and achieve minimization of size.
A significant reduction in Cas-independent off-target editing while maintaining high editing efficiency is achieved and the size is less than 4.7 kb, allowing it to be delivered to cells via a single AAV vector.
Smart Images

Figure CN120019143A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority under 35 USC § 119 (e) to a U.S. Provisional Application (USSN 63 / 398,483) filed by David R. Liu et al. on August 16, 2022, entitled “Evolved Cytosine Deaminases and Methods of Editing DNA Using the Same”; and to a U.S. Provisional Application (USSN 63 / 380,523) filed by David R. Liu et al. on October 21, 2022, entitled “Evolved Cytosine Deaminases and Methods of Editing DNA Using the Same”. Both applications are incorporated herein by reference in their entirety.
[0003] Government support
[0004] This invention was made with support from the National Institutes of Health under Grant Nos. RM1HG009490, R01EB027793, R01EB031172, R35GM118062, and U01AI142756. The government has certain rights in this invention.
[0005] References to Electronic Sequence Listings
[0006] The contents of the electronic Sequence Listing (B119570170WO00-SEQ-JQM.xml; size: 457,025 bytes; and creation date: August 15, 2023) are incorporated herein by reference in its entirety. Background Art
[0007] Base editors (BEs) are useful tools for in vivo forward genetic mutagenesis screening and have the potential to correct pathogenic point mutations by precisely installing targeted point mutations. BEs contain a fusion of a Cas protein and a base modification enzyme (e.g., a deaminase). Cytosine base editors (CBEs) convert C·G base pairs to T·A base pairs, while adenine base editors (ABEs) convert A·T base pairs to G·C base pairs. Together, CBEs and ABEs can mediate all four possible conversion mutations (e.g., C to T, A to G, T to C, and G to A). See International Patent Application No. PCT / US2017 / 045381 (published on February 8, 2018), International Patent Application No. PCT / US2018 / 056146 (published as WO 2019 / 079347 on April 25, 2019), Koblan et al., Nat Biotechnol (2018), and Gaudelli et al., Nature 551, 464-471 (2017).
[0008] Highly active naturally occurring DNA-modifying cytidine deaminases (such as APOBEC family enzymes) can deaminize transiently exposed single-stranded DNA fragments outside of the R-loops generated by the Cas protein domains of CBEs, leading to low-level but widespread Cas-independent genome modification. 13-15,19 Similarly, highly active cytidine deaminases that can robustly bind RNA can also mediate unintended RNA deamination independent of guide RNA hybridization. 21 The significant Cas-independent off-target DNA and RNA editing observed in editing with existing CBEs may limit the use of these CBEs in applications where minimization of off-target editing is desired. 15 Existing CBEs include BE3 (which contains the structure NH 2 -[NLS]-[rAPOBEC1 deaminase]-[Cas9 nickase (D10A)]-[UGI domain]-[NLS]-COOH), BE4 (which contains the structure NH 2 -[NLS]-[rAPOBEC1 deaminase]-[Cas9 nickase (D10A)]-[UGI domain]-[UGI domain]-[NLS]-COOH) and BE4max (a version of BE4 in which the base editor encoding construct has been codon-optimized for expression in human cells). Cas-independent off-target effects arise from the random association of base editors with DNA sites due to the inherent affinity of overexpressed base editors for DNA. Cas-independent off-target DNA editing of several TadA*-based ABEs has been found to be undetectable or much less frequent 13 Although low levels of RNA deamination can still be detected from overexpression of some ABEs8,9,34 .
[0009] There is a need in the art for novel cytidine deaminases and cytosine base editors that exhibit low Cas-independent off-target editing while maintaining high on-target activity. There is also a need in the art for CBEs of smaller size, e.g., small enough to be encoded by a single adeno-associated virus (AAV) vector (e.g., with a packaging capacity of approximately 4.7 kb). Summary of the invention
[0010] The present disclosure provides the first directed evolution of deaminase to selectively deaminate different bases. The present disclosure provides variants of adenosine deaminase engineered to preferentially deaminate cytidine in DNA. Therefore, the present disclosure provides a cytidine deaminase as a variant of adenosine deaminase (e.g., wild-type or engineered tRNA adenosine deaminase (TadA)). The present disclosure provides a cytidine base editor comprising a deaminase variant domain and a nucleic acid programmable binding protein (napDNAbp) domain comprising preferential deaminase cytidine in DNA, wherein the adenosine deaminase variant can deaminate cytidine in nucleic acid molecules to a similar or identical degree to existing cytidine deaminase. In some aspects, the present disclosure provides deaminase variants with minimized size, which provide base editors with reduced off-target effects while maintaining the high editing efficiency of existing CBEs compared to existing cytosine base editors (CBEs). In some aspects, the present disclosure provides base editors, complexes, nucleic acids, vectors, cells, compositions, methods, kits, and uses utilizing deaminases and base editors provided herein.
[0011] The present disclosure is based at least in part on the hypothesis that adenosine deaminases may further evolve to recognize cytosine as a substrate, and that such evolution may lead to a new class of highly selective cytidine deaminases and CBEs with high editing efficiency and lower off-target Cas-independent DNA and RNA editing (compared to naturally occurring cytidine deaminases). Wild-type TadA is evolutionarily related to cytidine deaminases. In addition, low levels of cytidine deamination have been reported in evolved ABE variants. 11,31,32 In addition, mutagenesis of TadA7.10 (TadA-7.10P48R) was shown to disrupt adenosine selectivity and increase the 5′-T at position 6 of the SpCas9 protospacer in the editing window (counting the SpCas9 protospacer adjacent motif PAM as positions 21-23). C Cytidine deamination in context 32 , although adenosine deamination is still preferred in other contexts and positions. Finally, adenosine deaminases acting on RNA (ADARs) have evolved to perform both cytidine and adenosine deamination in RNA. 33 .
[0012] The present disclosure generally relates to base editors (BEs) for gene editing. Base editors reported to date comprise a programmable DNA-binding protein domain (e.g., Cas9) fused to a deaminase (e.g., a "base" modification domain). In some cases, BEs may also include additional domains that alter cellular DNA repair processes to improve the efficiency, incorporation, and / or stability of the resulting single nucleotide changes. The programmable DNA-binding domain directs the deaminase to directly convert one base to another at a target site programmed by the guide RNA. Two major classes of BEs have been developed to date: cytidine BEs (CBEs), which convert C·G to T·A; and adenine BEs (ABEs), which convert A·T to G·C. Collectively, CBEs and ABEs are able to correct all four types of conversion mutations (C to T, G to A, A to G, and T to C). Since half of the known disease-associated gene variants are point mutations, and conversion mutations account for approximately 60% of known pathogenic point mutations, BEs are widely used to study and treat genetic diseases in various cell types and organisms, including animal models of human genetic diseases.
[0013] CBEs and ABEs may include any programmable DNA binding domain known to those skilled in the art. CBEs further comprise a deaminase configured to deaminate cytidine; whereas ABEs comprise a deaminase configured to deaminate adenosine. Without wishing to be bound by any particular theory, it is generally believed that current CBEs comprise naturally occurring deaminases or variants thereof configured to deaminate cytidine to uracil. In contrast, ABEs comprise tRNA-specific adenosine deaminases that have evolved (using laboratory techniques such as PACE and PANCE mutations) to accept DNA substrates such as those described in International Patent Application No. PCT / US2021 / 016827 (filed on February 5, 2021, incorporated herein by reference) to achieve A·T to G·C editing. All ABEs reported to date 4,8–10 , including those that have entered clinical trials 2 Or has been approved for clinical trials 1 Those that do use TadA7.10 or an evolved or engineered variant of this deaminase. TadA7.10 is the adenosine deaminase of the most advanced ABE (ABE7.10), which is disclosed in International Publication No. WO 2018 / 027078 (published on August 2, 2018). TadA7.10 is also the deaminase domain of ABEmax, a variant of ABE7.10 that has been codon-optimized for expression in human cells. For example, the current generation of ABE variant ABE8e (which contains the TadA-8e mutant adenosine deaminase) generally achieves higher editing efficiencies than existing CBEs, despite the strong preference of wild-type TadA for tRNA substrates. 9,11,12. TadA-8e and ABE8e are described in International Publication No. WO 2021 / 158921 (published on August 12, 2021).
[0014] ABEs have several advantages over their CBE counterparts. For example, TadA enzymes are less processive than most CBE deaminases and thus are generally able to achieve higher single-nucleotide editing precision. 3,7,8,11 ABEs also provided lower levels of Cas-independent off-target editing compared to CBEs 8,9,13–15 This advantage may be due to the fact that the K m = 830 nm), the tighter non-assisted binding of cytidine deaminase to nucleic acid substrates (the Michaelis constant K for APOBEC1 binding to mRNA m This may also stem from the fact that wild-type TadA cannot process DNA and that TadA-8e evolved to use TadA7.10 only in a Cas-dependent manner. 19 Protein engineering and proteomic analysis have provided alternative cytidine deaminases with reduced Cas-independent DNA and RNA editing, but to date these variants have suffered from reduced on-target editing activity and / or large size. 15,20-24 .
[0015] The evolved TadA adenosine deaminase is 166 amino acids long, which is longer than commonly used cytidine deaminases such as APOBEC1 (227 amino acids), AID (182 amino acids) and 25 、CDA (207 amino acids) 7 or APOBEC3A (198 amino acids) 26 Significantly smaller, which makes TadA-derived base editors more easily delivered into cells using size-restricted methods and systems such as AAV. In fact, the small size of TadA enables ABEs, but not CBEs, to be delivered into animal tissues in vivo using a single AAV. 27,28 .
[0016] The inventors of the present disclosure hypothesize that directing an adenosine deaminase to deaminate cytidine by directed evolution may generate CBEs that maintain high on-target activity but inherit the lower Cas-independent off-target editing and smaller size of current ABEs (e.g., making them easier to deliver to cells by size-restricted methods such as AAV). Therefore, in some embodiments, the present disclosure provides CBEs comprising a mutant adenosine deaminase (preferentially deaminates cytidine in DNA) and a napDNAbp domain (e.g., a Cas9 nickase). The cytidine deaminase evolved from the TadA deaminase described herein is referred to as "TadA-CD", and the CBE containing TadA-CD disclosed herein is referred to herein as "TadCBE".
[0017] Thus, aspects of the present disclosure relate to CBEs comprising a programmable DNA binding protein (e.g., Cas9) and an evolved deaminase that preferentially deaminates pyrimidines in DNA, particularly cytidine. For example, the disclosed TadA-CD deaminase variants exhibit a ratio of cytidine deamination to adenine deamination of about 10: 1, 15: 1, 20: 1, or more than 20: 1. In a specific embodiment, the disclosed deaminase variants exhibit a ratio of cytidine deamination to adenine deamination of about 20: 1. One or more TadA-CD deaminases described herein comprise multiple mutations that are located on a loop near the active site and are critical for the selectivity of the conversion of adenosine to cytidine. These mutations give TadA-CD the unique advantage of having a low off-target editing frequency exhibited by adenosine deaminases used in existing ABEs (e.g., TadA-8e) while being active against cytidines in the DNA target region. They also have the advantage of being minimized in size (e.g., <4.7 kb), which gives the ability to encode TadCBEs containing these deaminase variants in a single AAV vector, rather than by two intein-mediated split AAV vectors, or using engineered virus-lipid particles (e.g., those described herein). In some embodiments, the TadCBE further comprises any napDNAbp domain useful for cytidine base editing activity, and a uracil glycosylase inhibitor (UGI) domain. These TadA-CD variants were generated by continuous and / or non-continuous evolution methods, including PACE experiments performed on TadA-8e substrates (or starting points).
[0018] Other aspects of the present disclosure relate to phage-assisted evolution selection systems (e.g., PACE and / or PANCE) to enhance the substrate specificity of the adenosine deaminase domain of an ABE for cytosine (wherein the ABE contains Cas9 or a Cas9 ortholog). In some embodiments, the selection technique comprises a vector system for PACE evolution, which comprises a low stringency vector and a high stringency vector. Additional aspects relate to cells containing these vectors or disclosed vector systems. For example, in some embodiments, the highly active adenosine deaminase TadA-8e is evolved (e.g., mutated) by PACE to deaminize cytidine. The evolved TadA-CD contains mutations located on the d-loop near the active site of the deaminase, which is critical for the selectivity of the conversion of adenosine to cytidine.
[0019] Compared to the most commonly used naturally occurring CBEs (such as BE4max and its variants), the disclosed TadCBEs provide comparable or higher on-target activity, smaller size and / or significantly lower Cas-independent DNA and RNA off-target editing activity, both of which can be further inhibited by introducing the V106W mutation without reducing on-target editing. These TadCBEs can be used for single or multiplexed base editing at therapeutically relevant genomic loci in mammalian cells (such as primary human T cells and hematopoietic stem and progenitor cells), as shown herein. Other cells are also possible and disclosed elsewhere herein. The creation of TadCBEs expands the practicality of cytosine base editors in gene editing.
[0020] In some embodiments, the evolved TadA-CD may comprise mutations at residues E27, V28, and H96 in the amino acid sequence of SEQ ID NO: 41 (i.e., TadA-8e deaminase), and may further comprise at least one mutation at a residue selected from R26, M61, Y73, I75, M151, Q154, and A158, or a corresponding mutation in a homologous adenosine deaminase. Exemplary homologous deaminases include TadA deaminases derived from any one of Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, and Bacillus subtilis. Thus, in some embodiments, the evolved TadA-CD may comprise one or more mutations in any one of SEQ ID NOs: 317-323, 354, and 355, which confer cytidine activity. In some embodiments, the evolved TadA-CD may comprise one or more mutations in any of SEQ ID NOs: 34-40, 42-54, 33, 315, and 326 that confer cytidine activity. The deaminases of the present disclosure may be evolved from any adenosine deaminase reported to date that has adenosine deaminase activity.
[0021] In some embodiments, the disclosed TadA-CD variants comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of TadA-8e (SEQ ID NO:41), wherein the amino acid corresponding to residue 27 of SEQ ID NO:41 is any amino acid except E.
[0022] In some embodiments, the TadA-CD variant comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:41, wherein the amino acid corresponding to residue 28 of SEQ ID NO:41 is any amino acid except V.
[0023] In other embodiments, the TadA-CD variant comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:41, wherein the amino acid corresponding to residue 96 of SEQ ID NO:41 is any amino acid except H.
[0024] In some embodiments, the disclosed TadA-CD variants comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of any one of SEQ ID NOs: 34-40. In other embodiments, the TadA-CD variants comprise the amino acids of any one of SEQ ID NOs: 34-40.
[0025] The disclosed TadA-CD variants may further comprise a V106W mutation. In some embodiments, the V106W mutation results in less than or equal to 1.5%, less than or equal to 1%, less than or equal to 0.75%, less than or equal to 0.5%, less than or equal to 0.25%, less than or equal to 0.1%, less than or equal to 0.05%, or less than or equal to 0.01% adenine base editing on the assessed target (the above editing frequencies may represent average or maximum values).
[0026] Other aspects of the present disclosure relate to base editors comprising a programmable DNA binding domain (e.g., napDNAbp) and a disclosed evolutionary TadA-CD domain. In some embodiments, the napDNAbp of a base editor is a Cas9 protein, such as a Cas9 nickase. In some embodiments, the napDNAbp of a base editor is an Nme2Cas9 protein (such as an eNme2Cas9 nickase) or an Nme2Cas9 variant. In some embodiments, the napDNAbp of a base editor is any protein listed in Table 6. In some embodiments, the base editor further comprises a UGI domain. In some embodiments, the base editor further comprises a nuclear localization domain. Therefore, TadCBE is provided herein. In another aspect, the present disclosure describes a complex comprising any disclosed base editor and a guide RNA bound to the napDNAbp domain of a base editor.
[0027] In some aspects, the present disclosure relates to TadA-derived cytidine deaminases that provide efficient conversion of target cytosine to thymine and target adenine to guanine (referred to herein as "TadA-dual" deaminases and base editors). TadA-dual deaminases are capable of editing C and A bases within the original spacer, and particularly within the editing window of the original spacer. These editors install A to G and C to T editing with approximately equal efficiency (e.g., base editors comprising TadA-dual, SEQ ID NO: 39).
[0028] In some embodiments, the TadA-dual deaminase is mutated relative to TadA-8e (SEQ ID NO. 41). In some embodiments, the TadA-dual deaminase comprises a cytidine deaminase comprising one, two, three, four or five mutations selected from R26G, V28A, A48R, Y73S and H96N (e.g., TadA-CDf, SEQ ID NO: 39).
[0029] In some embodiments, the TadA-dual deaminase is mutated relative to TadA-CDf (SEQ ID NO: 39). In some embodiments, the TadA-dual deaminase comprises a mutation at position N46 in the amino acid sequence of SEQ ID NO: 39. In some embodiments, the TadA-dual deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, at least 99.5%, or at least 99.8% identical to the sequence identity of SEQ ID NOs: 39-54.
[0030] In some embodiments, TadA-dual deaminase has an increased affinity for cytosine relative to adenosine. For example, in some embodiments, dual editors provide editing of A to G and C to T at a ratio of 0.7: 1, 0.8: 1, 0.9: 1, 1: 1, 1.1: 1, 1.2: 1, 1.3: 1, 1.4: 1, or 1.5: 1. However, in some embodiments, TadA-dual deaminase has a higher specificity for cytosine than for adenosine.
[0031] In other embodiments, the TadA-dual (e.g., SEQ ID NO: 39) deaminase can be further mutated (e.g., using PANCE and / or PACE) to generate a cytidine deaminase with increased affinity for cytosine relative to adenosine. For example, in some embodiments, the ratio of adenosine deamination activity to cytidine deamination activity of the deaminase is at least about 0.001:1, 0.005:1, 0.007:1, 0.01:1, 0.05:1, 0.07:1, or 0.1:1.
[0032] Other aspects of the present disclosure relate to polynucleotides, vectors and cells encoding napDNAbp, cytidine deaminase and fusion proteins thereof. In some embodiments, the base editor of the present disclosure may be encoded in a polynucleotide as disclosed herein. In some embodiments, the deaminase variants of the present disclosure may be encoded in a polynucleotide as disclosed herein. In certain embodiments, the disclosed vector comprises a polynucleotide encoding any one of the base editors disclosed herein. In other embodiments, the present disclosure provides cells and compositions comprising any one of the deaminase variants, base editors, complexes, nucleic acids or vectors described herein. In addition, an AAV vector encoding any disclosed base editor and optionally a guide RNA is also provided herein.
[0033] Other aspects of the present disclosure provide pharmaceutical compositions comprising any of the cytidine deaminase or variants thereof, base editors, complexes, viruses, nucleic acids and / or vectors described herein.
[0034] In some aspects, the present disclosure encompasses methods comprising contacting a nucleic acid molecule (e.g., DNA) with any of the base editors or complexes described herein. For example, in some embodiments, the method comprises contacting any of the BEs described herein with sgRNA and DNA. The contact in these methods can be performed in vivo, in vitro, or ex vivo.
[0035] Other embodiments describe methods of using the base editors described herein. In some embodiments, the method includes using (a) any base editor of the present invention and (b) targeting the base editor of (a) to a C:G nucleotide base pair in a double-stranded DNA molecule in DNA editing. Guide RNA. In other embodiments, the method includes using a base editor, complex, or pharmaceutical composition of the present invention as a drug. In certain embodiments, the method includes using a base editor, complex, or pharmaceutical composition of the present invention as a drug to treat a disease, disorder, or condition (e.g., sickle cell disease or HIV / AIDS).
[0036] In some embodiments, the present disclosure provides methods for selecting (e.g., evolving, engineering, etc.) cytosine base editors. These methods may include evolving adenosine base editors through several rounds of continuous PACE and / or PANCE evolution. In certain embodiments, the method includes a selection phage encoding a mutant TadA-8e protein fused to an NpuN intein, a first plasmid encoding an NpuC intein fused to dCas9-UGI, a second plasmid encoding gIII driven by a T7 or proT7 promoter and encoding an sgRNA, and a third plasmid encoding a T7 RNA polymerase-degradation determinant fusion.
[0037] In another aspect, the disclosure encompasses methods of generating one or more of the base editors described herein using any of the vectors described herein.
[0038] A further aspect of the present disclosure is also directed to a kit comprising a nucleic acid construct comprising (a) a nucleic acid sequence encoding any one of the base editors described herein, and (b) a nucleic acid sequence encoding a guide RNA. In some embodiments, the nucleic acid construct further comprises one or more heterologous promoters driving the expression of the sequence of (a) and / or the sequence of (b).
[0039] In certain aspects, the base editors described herein can be applied to a subject to treat a disease or condition. Therefore, a method is provided in which the TadCBE is applied to a subject and a target sequence in the subject's genome is edited. The target sequence may include a mutant C: G base pair, such as a mutant C: G base pair associated with a disease or condition. In various embodiments of these methods, the degree of cytidine deamination of the base editor exceeds the degree of adenosine deamination by 10 times, 15 times, 20 times, or more than 20 times (ratio of 10: 1, 15: 1, 20: 1, or more than 20: 1).
[0040] The present disclosure also provides the use of any of the base editors described herein and a guide RNA that targets the base editor to a target C:G base pair in a nucleic acid molecule in the preparation of a kit or composition for nucleic acid editing, wherein nucleic acid editing includes contacting the nucleic acid molecule with the base editor and the guide RNA under conditions suitable for deamination of cytosine (C) of the C:G nucleobase pair. The present disclosure also provides the use of any of the base editors described herein and a guide RNA that targets the base editor to a target C:G base pair in a nucleic acid molecule in the preparation of a kit for evaluating the off-target effects of a base editor.
[0041] Other advantages and novel features of the present disclosure will become apparent upon consideration of the detailed description of various non-limiting embodiments of the disclosure in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Non-limiting embodiments of the present disclosure will be described by way of example with reference to the accompanying drawings, which are schematic and are not intended to be drawn to scale. In the accompanying drawings, each identical or nearly identical component is generally represented by a single numeral. For clarity, not every component is labeled in every figure, nor is every component of every embodiment shown without the need for illustration to allow a person of ordinary skill in the art to understand the present disclosure. In the accompanying drawings:
[0043] Figures 1A-1E .Cytidine deaminase from phage-assisted evolution of TadA-8e. Figure 1A ) The evolutionary trajectory of TadA-based cytidine deaminases from tRNA deaminase TadA. Figure 1B ) Overview of PACE. The phage of choice (purple) encodes the protein being evolved. The E. coli host (grey) contains 1) a mutagenic plasmid (red) for diversifying the phage and 2) a plasmid system to regulate the expression of pill (blue, encoded by gIII). Only variants with the desired activity trigger the production of pill and propagate. Phage without the desired activity cannot propagate and are diluted out of the lagoon. ( Figure 1C ) The selection circuit for cytidine deamination. The TadA-8e variant is encoded on the selection phage (SP, purple). E. coli carries four helper plasmids in addition to the mutagenesis plasmid to establish the selection circuit: P1 contains the Cas9-UGI components of the base editor. After phage infection, the complete base editor is reconstructed by splitting the Npu intein system (yellow). P2 encodes the guide RNA and gIII, which are under the transcriptional control of the T7 promoter. P3 contains the T7 RNA polymerase inactivated by fusion with a degradation determinant tag. The C·G to T·A editing activity inserts a stop codon between T7 RNAP and the degradation determinant, thereby generating active T7 RNAP, which leads to transcription of gIII and phage propagation. ( Figure 1D )Two versions of the CBE circuit described here. In both cases, C·G to T·A editing inserts a stop codon before the degron tag, resulting in active T7 RNAP. The less stringent circuit requires a C·G to T·A edit on the noncoding strand (above) and can tolerate one undesired A to G edit. The more stringent circuit requires a C·G to T·A edit on the coding strand and cannot tolerate any undesired A·T to G·C edits. ( Figure 1E ) phage-assisted non-continuous evolution of cytidine deaminase from TadA-8e. The ProD (stronger, less stringent) or ProA (weaker, more stringent) promoter used in each PANCE passage is as indicated. In each passage, unless otherwise stated, phage was diluted 1:50. After several rounds of evolution, despite the increase in dilution rate between passages, phage titers stabilized, indicating evolution of cytidine deamination activity.
[0044] Figures 2A-2D .Evolved TadA* variants catalyze cytidine deamination. ( Figure 2A ) A summary of TadA-8e variants evolved and characterized herein. These variants represent conservative mutations after nine PANCE passages or after 159 hours of PACE. For a complete list of mutations, see Figure 13-14 . ( Figure 2B ) Method for evaluating base editing of a target plasmid in Escherichia coli. Cells were co-transformed with a target plasmid (blue) and a base editor plasmid (purple). Base editor expression was induced with arabinose. After 16 hours, cells were harvested and the target plasmid was analyzed by high-throughput sequencing. ( Figure 2C ) Base editing of the protospacer matching the selection loop target site in Escherichia coli. C·G to T·A editing is shown in blue. A·T to G·C editing is shown in magenta. Points represent a single biological replicate, and bars represent the mean ± SD from four independent biological replicates. ( Figure 2D ) Cryo-EM structure of ABE8e (PDB: 6VPC) 18 The location of the evolved mutations.
[0045] Figure 3.Characterization of evolved TadCBEs with SpCas9 domains in mammalian cells. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected with each of the nine guide RNAs targeting the protospacer region shown in each figure. Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent mean ± SD of three independent biological replicates. HEK293T site 3 is abbreviated as HEK3, and HEK293T site 4 is abbreviated as HEK4.
[0046] Figure 4 Characterization of evolved deaminases with evolved eNme2-C Cas9 domains in mammalian cells. The eNme2-C Cas9 nickase domain (PAM=N) was used in the BE4max architecture or ABE8e with 2xUGI. 4 CN) were transfected with each of the six guide RNAs targeting the protospacer region shown in each figure. Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.
[0047] Figures 5A-5D . Characterization of the base editing window of TadCBE and Cas-independent off-target DNA and RNA editing. Figure 5A ) Base editing activity windows of ABE8e, TadCBEa, and TadCBEa-V106W with 2xUGI across nine different target genomic sites in HEK293T. Points represent average editing across all sites containing a specified base at a specified position within the protospacer. Individual data points used for this analysis are Figures 2A-2D , Fig.14 and Figures 16A-16B middle.( Figure 5B Methods for measuring Cas-independent off-target DNA editing using an orthogonal R-loop assay 15 . ( Figure 5C ) Average Cas-independent off-target editing between all cytosines within the six orthogonal R-loops (SaR1-SaR6) generated by dead S. aureus Cas9. ( Figure 5D ) Off-target RNA editing. RNA was harvested from HEK293T cells 48 h after transfection with the indicated base editors. After cDNA synthesis, CTNNB1, IP90, and RSL1D1 were amplified and analyzed by high-throughput sequencing. Figure 5C-5D , Dots represent single biological replicates, and bars represent the mean ± SD of three independent biological replicates.
[0048] Figure 6 .Base editing at therapeutically relevant loci by TadCBE in primary human T cells and hematopoietic stem and progenitor cells. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into human T cells (n=4 donors) together with synthetic guide RNAs targeting (top) CXCR4 or (middle) CCR5 at the indicated protospacer. Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. After 3 days, genomic DNA was harvested from T cell lysates and analyzed by high-throughput sequencing. The gray boxes indicate the expected positions for stop codon installation in CXCR4 and CCR5. Produced after cytosine base editing T AG(CXCR4) and T The targeted cytidine of the AA (CCR5) stop codon is underlined. The figure below shows that mRNA encoding the indicated base editor or GFP as a negative control was electroporated into hematopoietic stem and progenitor cells together with a synthetic guide RNA targeting the BCL11A enhancer. After 3 days, genomic DNA was harvested from the cell lysates and analyzed by high-throughput sequencing. C·G to T·A base editing is shown in blue shading, and A·T to G·C base editing is shown in magenta shading. Points represent individual biological replicates, and bars represent the mean ± SD of n = 4 donors (top and middle) or n = 3 donors (bottom).
[0049] Figure 7 The basis for the selective selection of deamination in the PACE and PANCE circuits. In circuit 1, only when the base editor is on A 7 and A 8 Only when both are deaminated, the formation of the stop codon is hindered. Therefore, loop 1 tolerates a moderate level of A deamination. In loop 2, the single adenine A 6 Deamination of adenosine would prevent stop codon formation and hamper loop activation and phage propagation. Therefore, loop 2 is more selective against adenosine deamination.
[0050] Fig. 8A and 8B .PANCE titers and evolved TadA-CD genotypes. ( Fig. 8A) Phage titers of pools 1-7 during PANCE. Stringency was modulated by increasing promoter strength from ProD (strongest, least stringent) to ProA (weakest, most stringent), increasing dilution factor, and by switching from loop 1 to loop 2. Pools 1-6 were inoculated with phage encoding TadA8e-NpuN, while pool 7 was inoculated with phage encoding TadA8e A48R-NpuN. Figure 8B ) Genotypes of various PANCE culture pools (L1-L7) after PANCE.
[0051] Figures 9A-9C .PACE titers and evolved TadA-CD genotypes. ( Fig.9A ) Phage titers and pool flow rates during PACE. Pool 1 showed activity-independent reproduction in S2060 cells after t=43 hours, indicating that the phage evolved selection-independent replication and did not proceed. ( Fig. 9B ) Genotypes of evolved TadA* variants from pool 1 at t = 43 hours, before selection-independent reproduction occurred. ( Fig. 9C ) Genotypes of evolved TadA* variants from pool 2 at different time points.
[0052] Figures 10A-10C .AlphaFold model of TadA-CDa. ( Fig. 10A ) shows the cryo-EM structure of ABE8e (PDB ID 6VPC) 1 Binding of DNA to the 8-azancbularine (8Az) substrate mimetic containing adenosine. Val 28 (magenta) supports the correct positioning of the adenine substrate relative to the catalytic zinc. Fig. 10B ) Using Chimera software 2 The "Swapna" function in replaces 8Az with cytidine. In the resulting model, C4 of cytosine (which is targeted for nucleophilic attack during deamination) is approximately 1.3 Å away from the target carbon of 8Az. Therefore, movement of the DNA substrate may be required for productive catalysis. Val 28 may hinder this movement of the DNA substrate deep into the TadA-8e pocket. Fig. 10C ) Using AlphaFold 3To generate a model of the evolved TadA-CDa. The ABE8e structure was superimposed to generate a model with the DNA substrate R loop from 6VPC. The evolved enzyme is not expected to adopt any significant differences in secondary structure compared to TadA8e. The evolutionary substitution of Val 28 in TadA-8e to the smaller Ala or Gly residues found in TadA-CD may alleviate the steric constraints that are expected to hinder the productive positioning of the target C4 in the cytosine relative to the catalytic zinc ion.
[0053] Fig.11 .Indels and C·G to G·C editing by SpCas9 variants at nine genomic target sites. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the nine guide RNAs targeting the protospacer region shown in each figure. C·G to G·C base edits are shown in blue shading. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. The corresponding on-target editing data can be found in Figure 3 .
[0054] Fig.12 .Indels and C·G to G·C editing by eNme2-C Cas9 variants at six genomic target sites. The indicated base editors using the eNme2-C Cas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the six guide RNAs targeting the protospacer region shown in each figure. C·G to G·C base edits are shown in blue shading. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. The corresponding on-target data are in Figure 4 middle.
[0055] Fig.13 .Proximity of V106W to TadA-CD mutations. Mutations generated during the evolution of TadA-CD are shown in blue. Residue V106 is shown in red. Addition of V106W to TadA-7.10, TadA-8e, and TadA-8.17 reduces off-target editing activity 4-6 Addition of V106W to TadCBEa-e increased selectivity for cytidine deamination relative to adenosine and reduced off-target editing activity.
[0056] Fig.14.Base editing by the V106W variant at six genomic target sites. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the six guide RNAs targeting the protospacer region shown in each figure. Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.
[0057] Fig.15 Indels and C·G to G·C editing by the V106W variant at six genomic target sites. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the six guide RNAs targeting the protospacer region indicated in each figure. C·G to G·C base edits are shown in blue shading. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.
[0058] Figures 16A-16B .Base editing, indel formation, and C·G to G·C editing of the TadA-CD(V106W) variant at three additional genomic target sites. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the three guide RNAs targeting the protospacer region indicated in each figure. ( Fig.16A ) Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. C·G to T·A base editing is shaded in blue. A·T to G·C base editing is shaded in magenta. ( Fig. 16B ) C·G to G·C base editing is shown in blue shades. Indels are shown in gray. Points represent individual values, and bars represent mean ± SD of three independent biological replicates.
[0059] Fig.17 .CBE base editing activity window across nine genomic target sites. Points represent average editing across all sites containing a given base at a given position within the protospacer. Individual data points used for this analysis are Figures 2A-2D , Fig.14 and Figures 16A-16B middle.
[0060] Fig.18On-target editing of EMX1 in Cas-independent R-loop editing experiments. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells together with SpCas9 guide RNA targeting EMX1 and the indicated SaCas9 sgRNA. 5 and C 6 The average on-target C·G to T·A base editing of . Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. The corresponding Cas-dependent off-target data are Figure 5C , Fig.19 and Fig. 20 Displayed in.
[0061] Fig.19 Cas-independent off-target C·G to T·A editing at a single site within six orthogonal R-loops generated by SaCas9. Orthogonal R-loop assay performed on CBE variants in the BE4max architecture 7 Cells were transfected with base editors and one SpCas9 sgRNA targeting the EMX1 locus, as well as the orthogonal dead SaCas9 and one SaCas9 sgRNA corresponding to Sa sites 1-6. Points represent a single biological replicate, and bars represent the mean ± SD of three independent biological replicates. The corresponding on-target data are in Fig.18 middle.
[0062] Fig. 20 Cas-independent off-target C·G to T·A editing by TadCBEe V106W at a single site within six orthogonal R-loops generated by SaCas9. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture. Cells were transfected with base editors and one SpCas9 sgRNA targeting the EMX1 locus along with orthogonal dead SaCas9 and one SaCas9sgRNA corresponding to Sa sites 1-6. Points represent a single biological replicate, and bars represent the mean ± SD of three independent biological replicates.
[0063] Figures 21A-21C Cas-independent off-target DNA editing by TadCBEe V106W at six genomic SaCas9 R-loops. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture. Cells were transfected with base editors and one SpCas9 sgRNA targeting the EMX1 locus (on-target) as well as an orthogonal dead SaCas9 and one SaCas9 sgRNA corresponding to Sa sites 1-6 (SaR1-SaR6). Fig.21A On-target editing at the EMX1 locus is shown. Fig.21B Shown is the average C·G to T·A base editing across all adenines within the indicated protospacer regions depicted in the figure. ( Fig. 21C ) The average A·T to G·C base editing between all adenines in the indicated protospacer is depicted. Points represent individual biological replicates, and bars represent the mean ± SD of three independent biological replicates.
[0064] Fig. 22 .Cas-independent off-target DNA editing at six genomic SaCas9 R-loops. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture. Cells were transfected with base editors and one SpCas9 sgRNA targeting the EMX1 locus (on-target) and an orthogonal dead SaCas9 and one SaCas9 sgRNA corresponding to Sa sites 1-6 (SaR1-SaR6). The average A·T to G·C base editing between all adenines within the indicated protospacer is depicted. Points represent individual biological replicates, and bars represent the mean ± SD of three independent biological replicates.
[0065] Fig.23A and 23B .Cas-independent off-target RNA editing of all cytosines and adenines tested across three transcripts for TadCBEe V106W. Total RNA was harvested from HEK293T cells 48 hours after transfection with the indicated base editors. After cDNA synthesis, CTNNB1, IP90, and RSL1D1 were amplified and analyzed by high-throughput sequencing. At the same time, genomic DNA was harvested from another plate transfected in parallel. Analyze on-target editing of EMX1 in genomic DNA as a control for base editor activity. Fig.23A On-target editing of EMX1 in samples corresponding to RNA editing analysis is shown. Fig. 23B Mean C to U (blue shading) or A to I (magenta shading) are shown. Points represent a single biological replicate, and bars represent the mean ± SD of three independent biological replicates.
[0066] Fig.24 On-target editing of EMX1 in an RNA off-target editing experiment. The indicated base editors were transfected into HEK293T cells in two parallel plates. In one plate, RNA was harvested from HEK293T cells 48 hours after transfection with the indicated base editors and expressed as Figures 23A-23B The analysis was performed as described. At the same time, genomic DNA was harvested from another plate transfected in parallel. On-target editing of EMX1 in genomic DNA was analyzed as a control for base editor activity. Points represent a single biological replicate, and bars represent the mean ± standard deviation of three independent biological replicates.
[0067] Fig.25 .Cas-dependent editing of known off-target sites of HEK3. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells together with a guide RNA targeting HEK293T site 3 (HEK3). 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Tables 2A-2E. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.
[0068] Fig.26 .Cas-dependent editing of known off-target sites of HEK4. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells together with a guide RNA targeting HEK293T site 4 (HEK4). 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Tables 2A-2E. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.
[0069] Figures 27A-27B .Cas-dependent editing of known off-target sites of EMX1. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells together with a guide RNA targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Tables 2A-2E. C·G to T·A base edits are shown in blue shading. A·T to G·C base edits are shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. The corresponding on-target data are in Fig.35 Displayed in.
[0070] Fig.28.Cas-dependent editing of known off-target sites of BCL11A. The specified base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into primary human CD34-positive hematopoietic stem and progenitor cells (n=3 donors) together with a guide RNA targeting BCL11A. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Tables 2A-2E. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± standard deviation of three independent biological replicates.
[0071] Figures 29A-29B C·G to G·C editing and indels in T cell experiments targeting CXCR4 and CCR5. mRNA encoding the indicated base editors or GFP as a negative control was co-expressed with targeting ( Fig.29A )CXCR4 or ( Fig.29B ) Two synthetic guide RNAs for CCR5 were electroporated together into primary human T cells (n=4 donors). After 3 days, genomic DNA was harvested from T cell lysates and analyzed by high-throughput sequencing. C·G to G·C base editing is shown in blue shading. Indels are shown in gray. Points represent individual values, and bars represent the mean ± standard deviation of three independent biological replicates.
[0072] Fig.30 .Cas-dependent off-target editing of T cell experiments targeting CXCR4 and CCR5. mRNA encoding the indicated base editors or GFP as a negative control was electroporated into primary human T cells (n=4 donors) together with two synthetic guide RNAs targeting CXCR4 or CCR5 at the specified protospacer. After 3 days, genomic DNA was harvested from T cell lysates and known off-target sites were amplified using the primers in Tables 2A-2E. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± standard deviation of three independent biological replicates.
[0073] Figures 31A-31B .C·G to G·C editing, indels, and Cas-dependent off-target editing of BCL11A in hematopoietic stem and progenitor cells. mRNA encoding the indicated base editors or GFP as a negative control was electroporated into CD34-positive human hematopoietic stem and progenitor cells (n=3 donors) together with a synthetic guide RNA targeting BCL11A at the specified protospacer. After 3 days, genomic DNA was harvested from cell lysates and analyzed by high-throughput sequencing. ( Fig.31A) C·G to G·C base edits are shaded in blue. Indels are shown in grey. Fig.31B ) Amplification of known Cas-dependent off-target sites using primers listed in Tables 2A-2E. C·G to G·C base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent mean ± SD of three independent biological replicates.
[0074] Figures 32A-32C . Characterization of TadCBE using a genomically integrated mESC target sequence library. Fig.32A The overall efficiency and selectivity of base editors analyzed by editing libraries are shown. The data show the average fraction of edited sequencing reads across all library members between protospacer positions -9 to 20, where positions 21-23 are PAMs. Fig.32B Editing spectra of BE4max, TadCBEa-d, TadCBEd V106W, and the dual base editor TadDE across 10,683 genomic integration target sites are shown. The editing window is defined as the protospacer position where the average editing efficiency is ≥ 30% of the peak editing efficiency. Window plots for all variants tested in the library experiment can be found in Fig.39 . Fig.32C The sequence motifs of the TadCBEd and TadCBEd V106W base editing outcomes for cytosine and adenine, determined by performing editing efficiency regression, are shown. The opacity of the sequence motifs is proportional to the test R on the retained sequence set. The complete sequence motif map for all variants is in Fig.41A and 41B Displayed in.
[0075] Fig.33 . Testing of single mutations in TadCBE. Base editing of protospacers matching the target site of the selection loop in Escherichia coli. Cells were co-transformed with target and base editor plasmids. Expression of the base editors was induced with arabinose. After 16 h, cells were harvested and the target plasmids were analyzed by high-throughput sequencing. (Top) Single mutations identified by evolution were not sufficient for CBE generation when added to ABE8e. (Middle) Analysis of mutations in TadCBEa-c and TadCBEe. Mutations in the ABE8e loop region confer selectivity for cytidine deamination, while auxiliary mutations enhance activity. (Bottom) Analysis of mutations in TadCBEd. C·G to T·A editing is shown in blue. A·T to G·C editing is shown in magenta. Points represent individual biological replicates, and bars represent the mean ± SD of four independent biological replicates.
[0076] Fig.34.Regression analysis of TadCBE. Base editing of the protospacer matching the selection loop target site in Escherichia coli. Cells were co-transformed with the target plasmid and the base editor plasmid. Expression of the base editor was induced with arabinose. After 16 hours, the cells were harvested and the target plasmid was analyzed by high-throughput sequencing. (Top) Mutations shown are relative to TadCBEa (grey box). (Bottom) Mutations shown are relative to TadCBEe (grey box). C·G to T·A editing is shown in blue. A·T to G·C editing is shown in magenta. Points represent individual biological replicates, and bars represent the mean ± SD of four independent biological replicates.
[0077] Fig.35 .On-target editing of EMX1. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells together with a guide RNA targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Table 1. C·G to T·A base edits are shown in blue shading. A·T to G·C base edits are shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. The corresponding off-target data are in Fig.27A and 27B Displayed in.
[0078] Fig.36 .On-target and off-target editing of EMX1 by TadCBE V106W. The indicated base editors using the SpCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells together with a guide RNA targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Table 4. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.
[0079] Fig.37 Schematic diagram of the mESC library experiment. Thousands of pairs of sgRNAs and corresponding target sites are integrated into mESCs and treated with base editors. Cells containing base editors are enriched by antibiotic selection, and the library cassettes are amplified for high-throughput sequencing.
[0080] Fig.38 Correlation between replicates in the mESC library experiment. Uncorrected C·G to T·A editing efficiency at each target site for each replicate. The red dashed line is the overall least squares regression line.
[0081] Fig.39.Editing window of the TadCBE V106W variant in the mESC library editing experiment. The editing window is defined as the position within the protospacer where the average fraction of converted bases at this position is at least 30% of the average edit at the position of maximal editing. C·G to T·A base edits are shown in blue. A·T to G·C base edits are shown in red.
[0082] Fig.40A and 40B .Effect of V106W on peak editing in mESC library experiments. Fig.40A Shown is the C·G to T·A editing efficiency using TadCBEd (with and without the V106W substitution) for each library member containing a cytosine at protospacer position 6. The red dashed line is the overall least squares regression line. Fig.40B Shown is the A·T to G·C editing efficiency using TadCBEd (with and without V106W) for each library member containing an adenine at protospacer position 6. The red dashed line is the overall least squares regression line.
[0083] Fig.41A and 41B .Base editing activity of TadCBE is determined by performing editing efficiency regression. The opacity of the flag is proportional to R on the retained test set. C·G to T·A base editing is provided ( Fig.41A ) and A·T to G·C base editing ( Fig.41B )’s picture.
[0084] Fig.42 Characterization of evolved deaminases with evolved eNme2-C Cas9 domains. The eNme2-C Cas9 nickase domain (PAM = N 4 CN) were transfected into HEK293T cells with each of the six guide RNAs targeting the protospacer region shown in each figure. Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.
[0085] Fig.43.Indels and C·G to G·C editing by eNme2-C Cas9 variants at six genomic target sites. The indicated base editors using the eNme2-C Cas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the six guide RNAs targeting the protospacer region shown in each figure. C·G to G·C base edits are shown in blue shading. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. The corresponding on-target data are in Fig.42 Displayed in.
[0086] Fig.44 .Characterization of evolved deaminases with SaCas9 domains. The indicated base editors using the SaCas9 nickase domain (PAM=NNGRRT) in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the nine guide RNAs targeting the protospacer region shown in each figure. Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± standard deviation of three independent biological replicates.
[0087] Fig.45 .Indels and C·G to G·C editing by SaCas9 variants at nine genomic target sites. The indicated base editors using the SaCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the six guide RNAs targeting the protospacer region shown in each figure. C·G to G·C base edits are shown in blue shading. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. The corresponding on-target data are shown in Figure 2. Fig.44 shown.
[0088] Fig.46 .Characterization of TadDE with SpCas9 in mammalian cells. The indicated base editors using the SaCas9 nickase domain (PAM=NGG) in the BE4max architecture or ABE8e with 2xUGI were transfected with each of the nine guide RNAs targeting the protospacer region shown in each figure. Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.
[0089] Fig.47 .Indels and C·G to G·C editing by SpCas9 variants at nine genomic target sites. The indicated base editors using the SaCas9 nickase domain in the BE4max architecture or ABE8e with 2xUGI were transfected into HEK293T cells with each of the six guide RNAs targeting the protospacer region shown in each figure. C·G to G·C base edits are shown in blue shading. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. The corresponding on-target data are in Fig.46 middle.
[0090] Fig.48 .On-target editing of the V106W variant in T cell experiments targeting CXCR4 and CCR5. mRNA encoding the indicated base editors or GFP as a negative control was electroporated into human T cells (n=4 donors) together with two synthetic guide RNAs targeting (top) CXCR4 or (bottom) CCR5 at the indicated protospacer. Target cytosine is blue, target adenine is magenta, and the PAM sequence is underlined. After 3 days, genomic DNA was extracted from T cell lysates and analyzed by high-throughput sequencing. The gray boxes indicate the expected positions for stop codon installation in CXCR4 and CCR5. Generation by cytosine base editing T AG(CXCR4) and T The targeted cytidine of the AA (CCR5) stop codon is underlined. Points represent individual values and bars represent the mean ± SD of three independent biological replicates.
[0091] Fig.49 .C·G to G·C editing and indels in T cell experiments targeting CXCR4 and CCR5 using the TadCBEe V106W variant. mRNA encoding the indicated base editors or GFP as a negative control was electroporated into primary human T cells (n=4 donors) together with two synthetic guide RNAs targeting (above) CXCR4 or (below) CCR5 at the specified protospacer. After 3 days, genomic DNA was extracted from T cell lysates and analyzed by high-throughput sequencing. C·G to G·C base editing is shown in blue shading. Indels are shown in gray. Points represent individual values, and bars represent the mean ± standard deviation of three independent biological replicates.
[0092] Fig.50.Cas-dependent off-target editing in T cell experiments targeting CXCR4 and CCR5 using the TadCBEe V106W variant. mRNA encoding the indicated base editors or GFP as a negative control was electroporated into primary human T cells (n=4 donors) together with two synthetic guide RNAs targeting (above) CXCR4 or (below) CCR5 at the specified protospacer. After 3 days, genomic DNA was harvested from T cell lysates and known off-target sites were amplified using the primers in Table 4. C·G to T·A base editing is shown in blue shading. A·T to G·C base editing is shown in magenta shading. Points represent individual values, and bars represent the mean ± standard deviation of three independent biological replicates.
[0093] Fig.51A -51F. Prophetic use of active and selective cytosine base editors for installation of stop codons at disease-associated sites. Residual A to G editing prevents correct stop codon installation ( Fig.51A ). Schematic diagram of the evolution of cytosine base editors from TadA dual base editor (TadA-DE) ( Fig.51B ). Schematic diagram depicting phage-assisted continuous evolution or PACE (left) and a selection circuit used according to some embodiments (right). In some embodiments, a continuous flow of E. coli host cells using a selection circuit and a mutagenic plasmid (red) are infected with a selection phage (SP) encoding a partial deaminase. In a specific embodiment of this selection circuit, phage propagation is associated with the expression of gIII (P2), which can only be transcribed in the presence of an active T7 RNA polymerase. In some embodiments, the T7 RNA polymerase (P3) is fused to a C-terminal degron and the deaminase must undergo C to U editing to install a stop codon before the degron, thereby producing an active T7 RNA polymerase. In the case of phage infection, a full deaminase is completed using a split intein system (P1), and mutations can occur on the deaminase. Beneficial mutations lead to the propagation and enrichment of phages in the culture tank, while less adaptable phages are unable to propagate and are subsequently washed away by the constant outflow ( Fig.51C ). Evolutionary trajectories of active and selective cytosine base editors from TadA-DE. TadA-DE was subjected to phage-assisted non-continuous evolution (PANCE) until phage titers increased despite increasing stringency of dilution factor and promoter strength, indicating that beneficial mutations had occurred. The resulting variants recognized a conservative mutation at position N46 of the deaminase, so an NNK library was constructed at position N46 and PANCE was performed on these variants. To increase stringency even further, the resulting variants from both PANCEs were subjected to PACE for >100 hours. The dilution factor is indicated on the right y-axis ( Fig.51DThe cryo-EM structure of ABE8e (PDB: 6VPC) has the marked novel conservative mutations ( Fig.51E ).
[0094] Figures 52A-52E . Genotypes from PANCE culture pools (L1–L2) after PANCE ( Fig.52B ). Genotypes from PANCE pools (L1–L3) after PANCE using the NNK library at N46 ( Fig.52B ). Genotypes from the PANCE culture pool (L1) after PANCE using the NNK library at N46 ( Fig.52C Genotypes at different time points from the PANCE culture pool (L1) after PANCE using the NNK library at N46 ( Fig.52D Genotypes at different time points from the PANCE culture pool (L2) after PANCE using the NNK library at N46 ( Fig.52E ). Fig.52A Selected sequences are shown in .
[0095] Fig.53 .Profiling of the activity and sequence context specificity of TadCBE in E. coli. The bars indicate the average activity of the CBE variants when tested on a substrate library designed to contain a target base (A or C) at position 6 of the protospacer with 5' and 3' base changes of A, T, C, or G. Each dot represents the percentage of sequencing reads containing the specified edit in a given sequence context. The dots are colored according to the 5' context of the base (A, red; C, green; G, blue; T, yellow). Mutations in the newly evolved mutations are listed relative to TadDE. TadDE = TadA8e R26G V28A A48 Y73S H96N. While eTdCBEmax, CBET-1.52, TadCBEd, TadDE N46I Y73P, and TadDE N46C Y73P showed lower activities depending on the sequence context, the evolved variants TadDE N46V Y73P and TadDE N46L Y73P showed more than 80% editing regardless of the sequence context.
[0096] Figure 54. Comparison of evolved active and selective cytosine base editors with existing cytosine base editors in mammalian cells. The TadDE N46 variant was transfected into HEK293T cells with existing cytosine base editors with SpCas9 nickase in the BE4max architecture and guide RNAs targeting three protospacers. The TadDE N46 variant showed comparable on-target activity with no residual A to G editing. The dots represent individual values of independent biological replicates. The PAM sequence is underlined. HEK293T site 2 is abbreviated as HEK2, and HEK293T site 4 is abbreviated as HEK4. The TadDE N46 variant was transfected into HEK293T cells with existing cytosine base editors with eNme-Cas9 nickase in the BE4max architecture and guide RNAs targeting two protospacers. The TadDE N46 variant showed higher or comparable on-target activity with no residual A to G editing. The dots represent individual values of independent biological replicates. The PAM sequence is underlined.
[0097] Figure 55. Cas9-independent and RNA off-target editing by TadCBE. Average Cas9-independent off-target editing across all cytosines in four orthogonal R-loops (SaR1–SaR4) generated by dead S. aureus Cas9. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant showed similar off-target editing compared to TadCBEd. Points represent individual values of independent biological replicates ( Fig.55A ). Off-target RNA editing. The TadDE N46 variant showed similar off-target editing compared to TadCBEd. Points represent individual values from independent biological replicates ( Fig.55B ).
[0098] Fig.56 .Stop codon installation at therapeutically relevant loci by TadCBE in HEK293T cells. TadCBE was used to install a stop codon in PCSK9, a therapeutic strategy being explored to lower blood cholesterol. Grey boxes indicate the desired position for stop codon installation. Mutations in the newly evolved mutations are listed relative to TadDE. Residual A to G editing from TadCBEd resulted in stop codon erasure, indicating that the lack of residual A to G editing in the TadDE N46 variant is critical for stop codon installation. Dots represent individual values from independent biological replicates. The PAM sequence is underlined.
[0099] Fig.57On-target and Cas-dependent editing of known off-target sites of HEK3. The TadDE N46 variant was transfected into HEK293T cells along with existing cytosine base editors with SpCas9 nickase in the BE4max architecture and guide RNA targeting HEK3. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant showed similar off-target editing compared to TadCBEd. Points represent individual values of independent biological replicates.
[0100] Fig.58 On-target and Cas-dependent editing of known off-target sites of HEK4. The TadDE N46 variant was transfected into HEK293T cells along with existing cytosine base editors with SpCas9 nickase in the BE4max architecture and guide RNA targeting HEK4. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant showed similar off-target editing compared to TadCBEd. Points represent individual values of independent biological replicates.
[0101] Fig.59 Targeted and Cas-dependent editing of known off-target sites of EMX1. The TadDE N46 variant was transfected into HEK293T cells along with existing cytosine base editors with SpCas9 nickase in the BE4max architecture and guide RNA targeting EMX1. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant showed similar off-target editing compared to TadCBEd. Points represent individual values of independent biological replicates.
[0102] Fig.60 .On-target and Cas-dependent editing of known off-target sites of BCL11a. The TadDE N46 variant was transfected into HEK293T cells along with existing cytosine base editors with SpCas9 nickase in the BE4max architecture and guide RNA targeting BCL11a. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant showed similar off-target editing compared to TadCBEd. Points represent individual values of independent biological replicates.
[0103] Fig.61 On-target editing at EMX1 associated with Cas-independent off-target editing. The TadDE N46 variant was transfected into HEK293T cells along with existing cytosine base editors with SpCas9 nickase in the BE4max architecture and SpCas9 guide RNA targeting EMX1 along with SaCas9 guide RNA. Mutations in the newly evolved mutations are listed relative to TadDE. Points represent individual values from independent biological replicates.
[0104] Fig.62 .On-target editing at EMX1 associated with off-target editing of RNA. The TadDE N46 variant was transfected into HEK293T cells in two plates along with an existing cytosine base editor with the SpCas9 nickase in the BE4max architecture. In one plate, RNA was harvested 48 hours after transfection and in the other plate, genomic DNA was harvested. Genomic DNA was analyzed for on-target editing of EMX1. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant showed similar off-target editing compared to TadCBEd. Points represent individual values of independent biological replicates.
[0105] Fig.63 . Fig.53 The continuation of. Fig.53 Each figure in Fig.63 However, Fig.53 Each data point (denoted as a point) in Fig.63 Displayed as bars.
[0106] definition
[0107] As used herein and in the claims, the singular forms "a," "an," and "the" include singular and plural references unless the context clearly dictates otherwise. Thus, for example, reference to "an agent" includes a single agent and a plurality of such agents.
[0108] AAV
[0109] "Adeno-associated virus" or "AAV" is a virus that infects humans and some other primate species. The wild-type AAV genome is a single-stranded deoxyribonucleic acid (ssDNA) that can be either positive or negative sense. The genome contains two inverted terminal repeats (ITRs), one at each end of the DNA strand, and two open reading frames (ORFs): rep and cap between the ITRs. The rep ORF contains four overlapping genes encoding the Rep proteins required for the AAV life cycle. The cap ORF contains overlapping genes encoding the capsid proteins: VP1, VP2, and VP3, which interact together to form the viral capsid. VP1, VP2, and VP3 are translated from one mRNA transcript, which can be spliced in two different ways: either a longer or shorter intron can be excised, resulting in two isoforms of mRNA: one about 2.3 kb and one about 2.6 kb long. The capsid forms a supramolecular assembly composed of approximately 60 individual capsid protein subunits, forming a non-enveloped T-1 icosahedral lattice capable of protecting the AAV genome. The mature capsid contains VP1, VP2, and VP3 (molecular weights of approximately 87 kDa, 73 kDa, and 62 kDa, respectively) in a ratio of approximately 1:1:10.
[0110] The rAAV particle may comprise a nucleic acid vector (e.g., a recombinant genome) that may comprise at least: (a) one or more heterologous nucleic acid regions comprising a sequence encoding a protein or polypeptide of interest (e.g., split Cas9 or split nucleobase) or an RNA of interest (e.g., gRNA), or one or more nucleic acid regions comprising a sequence encoding a Rep protein; and (b) one or more regions flanking the one or more nucleic acid regions (e.g., heterologous nucleic acid regions) comprising an inverted terminal repeat (ITR) sequence (e.g., a wild-type ITR sequence or an engineered ITR sequence). In some embodiments, the size of the nucleic acid vector is 4 kb to 5 kb (e.g., a size of 4.2 to 4.7 kb). In some embodiments, the nucleic acid vector further comprises a region encoding a Rep protein. In some embodiments, the nucleic acid vector is circular. In some embodiments, the nucleic acid vector is single-stranded. In some embodiments, the nucleic acid vector is double-stranded. In some embodiments, the double-stranded nucleic acid vector may be, for example, a self-complementary vector containing a region of the nucleic acid vector that is complementary to another region of the nucleic acid vector, thereby initiating the formation of double-stranded nucleic acid vectors.
[0111] Deaminase
[0112] The term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine (or adenine) deaminase, which catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA) to inosine. In other embodiments, the deaminase is a cytidine (or cytosine) deaminase, which catalyzes the hydrolytic deamination of cytidine or cytosine.
[0113] The deaminases provided herein can be from any organism, such as bacteria. In some embodiments, the deaminases or deaminases domains are variants of naturally occurring deaminases from organisms. In some embodiments, the deaminases or deaminases domains do not exist in nature. For example, in some embodiments, the deaminases or deaminases domains are at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to naturally occurring deaminases.
[0114] As used herein, the term "adenosine deaminase" or "adenosine deaminase domain" refers to a protein or enzyme that catalyzes adenosine (or adenine) deamination reaction. For the purposes of this disclosure, the terms "adenosine" and "adenine" can be used interchangeably. For example, for the purposes of this disclosure, reference to "adenine base editor" (ABE) refers to the same entity as "adenosine base editor" (ABE). Similarly, for the purposes of this disclosure, reference to "adenine deaminase" refers to the same entity as "adenosine deaminase". However, it will be understood by those of ordinary skill in the art that "adenine" refers to a purine base, while "adenosine" refers to a larger nucleoside molecule including a purine base (adenine) and a sugar moiety (e.g., ribose or deoxyribose). In certain embodiments, the present disclosure provides a base editor fusion protein comprising one or more adenosine deaminase domains. For example, an adenosine deaminase domain may include a heterodimer of a first adenosine deaminase and a second deaminase domain connected by a linker. Adenosine deaminase (e.g., engineered adenosine deaminase or evolved adenosine deaminase) provided herein is an enzyme that converts adenine (A) in DNA or RNA to inosine (I). Such adenosine deaminase can result in A: T to G: C base pair conversion. In some embodiments, the deaminase is a variant of a naturally occurring deaminase from an organism. In some embodiments, the deaminase does not exist in nature. For example, in some embodiments, the deaminase is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a naturally occurring deaminase.
[0115] In some embodiments, the adenosine deaminase is derived from a bacterium, such as Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, or Caulobacter crescentus. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is an Escherichia coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated Escherichia coli TadA deaminase. For example, the truncated ecTadA may lack one or more N-terminal amino acids relative to the full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19 or 20 N-terminal amino acid residues relative to the full-length ecTadA. In some embodiments, the truncated ecTadA can lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19 or 20 C-terminal amino acid residues relative to the full-length ecTadA. In some embodiments, the ecTadA deaminase does not contain an N-terminal methionine. Reference is made to U.S. Patent Publication No. 2018 / 0073012 published on March 15, 2018, which is incorporated herein by reference.
[0116] As used herein, the term "cytidine deaminase" or "cytidine deaminase domain" refers to a protein or enzyme that catalyzes the deamination reaction of cytidine or cytosine. The terms "cytidine" and "cytosine" are used interchangeably in the present disclosure. For example, for the purposes of this disclosure, reference to a "cytosine base editor" (CBE) refers to the same entity as a "cytosine base editor" (CBE). Similarly, for the purposes of this disclosure, reference to a "cytidine deaminase" refers to the same entity as a "cytosine deaminase". However, one of ordinary skill in the art will understand that "cytosine" refers to a pyrimidine base, while "cytidine" refers to a larger nucleoside molecule comprising a pyrimidine base (cytosine) and a sugar moiety (e.g., ribose or deoxyribose). Cytidine deaminase is encoded by the CDA gene and is an enzyme that catalyzes the removal of amine groups from cytidine (i.e., the base cytosine when attached to the ribose ring, i.e., a nucleoside called cytidine) to uridine (C to U) and from cytidine to deoxyuridine (C to U). A non-limiting example of a cytidine deaminase is APOBEC1 ("apolipoprotein B mRNA editing enzyme, catalytic polypeptide 1"). Another example is AID ("activation-induced cytidine deaminase"). Under standard Watson-Crick hydrogen bond pairing, cytosine bases form hydrogen bonds with guanine bases. When cytidine is converted to uridine (or cytidine is converted to deoxyuridine), uridine (or the uracil base of uridine) undergoes hydrogen bond pairing with adenine bases. Therefore, cytidine deaminase converts "C" to uridine ("U"), which will result in the insertion of "A" instead of "G" during cell repair and / or replication. As adenine "A" pairs with thymine "T", cytidine deaminase works in conjunction with DNA replication, resulting in the conversion of C·G pairing in double-stranded DNA molecules into T·A pairing.
[0117] Antisense strand
[0118] In genetics, the "antisense" strand of a fragment in a double-stranded DNA is a template strand, and it is considered to run in a 3' to 5' direction. In contrast, a "sense" strand is a fragment running from 5' to 3' in a double-stranded DNA, and it is complementary to the DNA antisense strand or template strand running from 3' to 5'. In the case of a DNA fragment encoding a protein, a sense strand is a DNA strand with the same sequence as mRNA, which uses an antisense strand as its template during transcription, and eventually undergoes (usually, not always) translation into protein. Therefore, the antisense strand is responsible for the RNA that is subsequently translated into protein, and the sense strand has a composition almost identical to that of mRNA. Note that for each fragment of dsDNA, there may be two groups of sense strands and antisense strands, depending on the direction of reading (because sense strands and antisense strands are related to the viewing angle). Ultimately, it is the gene product or mRNA that determines which strand of a fragment of dsDNA is called a sense strand or an antisense strand.
[0119] Base editing
[0120] "Base editing" refers to a genome editing technique that involves converting a specific nucleic acid base into another at a targeted genomic locus. In certain embodiments, this can be achieved without the need for double-stranded DNA breaks (DSBs) or single-strand breaks (i.e., nicks). To date, other genome editing techniques, including CRISPR-based systems, have all started with the introduction of DSBs at the locus of interest. Subsequently, cellular DNA repair enzymes repair the breaks, typically resulting in random insertions or deletions (indels) of bases at the DSB site. However, when it is desired to introduce or correct point mutations at the target locus rather than randomly destroy the entire gene, these genome editing techniques are inappropriate because the correction rate is low (e.g., typically 0.1% to 5%), and the main genome editing product is an indel. In order to improve the efficiency of gene correction without simultaneously introducing random indels, the inventors previously modified the CRISPR / Cas9 system to directly convert one DNA base into another without forming a DSB. See, Komor, AC, et al., Programmable editing of atarget base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference.
[0121] Base editors
[0122] As used herein, the term "base editor (BE)" refers to a polypeptide that includes a base (e.g., A, T, C, G, or U) that can be modified in a nucleic acid sequence (e.g., DNA or RNA) to convert one base into another base (e.g., A to G, A to C, A to T, C to T, C to G, C to A, G to A, G to C, G to T, T to A, T to C, T to G). In some embodiments, the base editor is capable of deaminating a base in a nucleic acid (such as a base in a DNA molecule). In the case of an adenine base editor, the base editor is capable of deaminating adenine (A) in DNA. Such a base editor may include a nucleic acid programmable DNA binding protein (napDNAbp) fused to an adenosine deaminase. Some base editors include CRISPR-mediated fusion proteins used in the base editing methods described herein. In some embodiments, the base editor comprises a nuclease-inactivated Cas9 (dCas9) fused to a deaminase that binds to nucleic acids in a guide RNA-programmed manner via the formation of an R-loop, but does not cut nucleic acids. For example, the dCas9 domain of the fusion protein may include D10A and H840A mutations (which enable Cas9 to cut only one strand of a nucleic acid duplex), as described in PCT / US2016 / 058344, which was published as WO2017 / 070632 on April 27, 2017, and is incorporated herein by reference in its entirety. The DNA cleavage domain of Streptococcus pyogenes Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cuts the strand complementary to the gRNA (the "targeting strand", or the strand in which editing or deamination occurs), while the RuvC1 subdomain cuts the non-complementary strand containing the PAM sequence (the "non-editing strand"). The RuvC1 mutant D10A generates a nick in the targeted strand, whereas the HNH mutant H840A generates a nick on the non-edited strand (see Jinek et al., Science, 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013).
[0123] As used herein, the terms cytidine, cytosine, and deoxycytidine are all synonyms and refer to cytidine that can be edited using CBE. Similarly, the terms adenosine, adenine, and deoxyadenine all refer to adenine that can be edited using ABE. In addition, the terms cytidine base editor, cytosine base editor, etc. are synonyms. Similarly, the terms adenosine base editor, adenine base editor, etc. are synonyms.
[0124] In some embodiments, a nucleobase editor is a macromolecule or macromolecular complex that primarily (e.g., greater than 80%, greater than 85%, greater than 90%, greater than 95%, greater than 99%, greater than 99.9%, or 100%) results in the conversion of a nucleobase to another nucleobase (i.e., a transition or transversion) in a polynucleic acid sequence using a combination of 1) an enzyme that modifies nucleotides, nucleosides, or nucleobases and 2) a nucleic acid binding protein that can be programmed to bind to a specific nucleic acid sequence.
[0125] In some embodiments, the nucleobase editor comprises a DNA binding domain (e.g., a programmable DNA binding domain, such as dCas9 or nCas9), which guides the base editor to the target sequence. In some embodiments, the nucleobase editor comprises a nucleobase modifying enzyme fused to a programmable DNA binding domain (e.g., dCas9 or nCas9). "Nucleobase modifying enzyme" is an enzyme that can modify a nucleobase and convert one nucleobase into another (e.g., a deaminase such as cytidine deaminase or adenosine deaminase). In some embodiments, the nucleobase editor can target cytosine (C) bases in a nucleic acid sequence and convert C into thymine (T) bases. In some embodiments, the editing of C to T is performed by a deaminase (e.g., cytidine deaminase). Base editors that can perform other types of base conversions (e.g., adenosine (A) to guanine (G), C to G) are also contemplated.
[0126] In some embodiments, the nucleobase editor that converts C to T comprises a cytidine deaminase. "Cytidine deaminase" refers to an enzyme that catalyzes the chemical reaction "cytosine + H 2 O→uracil+NH 3 ” or “5-methyl-cytosine+H 2 O→thymine + NH 3" enzyme. As is apparent from the reaction formula, such chemical reactions result in a C to U / T nucleobase change. In the case of a gene, such a nucleotide change or mutation can in turn result in an amino acid change in the protein, which can affect the function of the protein, e.g., loss of function or gain of function. In some embodiments, the C to T nucleobase editor comprises a dCas9 or nCas9 fused to a cytidine deaminase. In some embodiments, the cytidine deaminase domain is fused to the N-terminus of the dCas9 or nCas9. In some embodiments, the base editor further comprises a domain that inhibits uracil glycosylase and / or a nuclear localization signal. Such nucleobase editors have been described in the art, e.g., Rees & Liu, Nat Rev Genet. 2018; 19(12):770-788, Rees, et al. Sci. Advances 5, eaax5717 (2019), and Koblan et al. al., Nat Biotechnol. 2018; 36(9): 843-846; and U.S. Patent Publication No. 2018 / 0073012, published on March 15, 2018, which was issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693, published on May 4, 2017, which was issued as U.S. Patent No. 10,167,457 on January 1, 2019; International Publication No. WO 2017 / 0121693, published on April 27, 2017 2017 / 070633; U.S. Patent Publication No. 2015 / 0166980, published on June 18, 2015; U.S. Patent No. 9,840,699, issued on December 12, 2017; U.S. Patent No. 10,077,453, issued on September 18, 2018; PCT Publication No. WO 2019 / 023680, published on January 31, 2019; International Publication No. WO 2019 / 03680, published on September 27, 2018 2018 / 0176009; International Application No. PCT / US2019 / 033848 filed on May 23, 2019; International Application No. PCT / US2019 / 47996 filed on August 23, 2019; U.S. Provisional Application No. 62 / 835,490 filed on April 17, 2020; International Application No. PCT / US2019 / 61685 filed on November 15, 2019; International Application No. PCT / US2019 / 57956 filed on October 24, 2019; U.S. Provisional Application No. 62 / 858,958 filed on June 7, 2019; International Publication No. PCT / US2019 / 58678 filed on October 29, 2019, the contents of each of which are incorporated herein by reference in their entirety.
[0127] In some embodiments, the nucleobase editor converts A to G. In some embodiments, the nucleobase editor comprises an adenosine deaminase. "Adenosine deaminase" is an enzyme involved in purine metabolism. It is required for the breakdown of adenosine from food and the turnover of nucleic acids in tissues. Its main function in the human body is the development and maintenance of the immune system. Adenosine deaminase catalyzes the hydrolytic deamination of adenosine (forming inosine, which base pairs with G) in the context of DNA. No known adenosine deaminase acts on DNA. In contrast, known adenosine deaminases only act on RNA (tRNA or mRNA). Evolved adenosine deaminases that accept a DNA substrate and deaminate dA to deoxyinosine have been described below: for example, PCT application PCT / US2017 / 045381 filed on August 3, 2017 (published as WO 2018 / 027078) and PCT application number PCT / US2019 / 033848 filed on May 23, 2019 (published as WO 2019 / 226953), the contents of each of which are incorporated herein by reference.
[0128] Exemplary adenine base editors (ABEs) (or "adenosine base editors") and cytosine base editors (CBEs) (or "cytosine base editors") are also described below: Rees & Liu, Base editing: precision chemistry on the genome and transcriptome of living cells, Nat. Rev. Genet. 2018; 19(12): 770-788; and U.S. Patent Publication No. 2018 / 0073012, published on March 15, 2018, which was issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693, published on May 4, 2017, which was issued as U.S. Patent No. 10,167,457 on January 1, 2019; International Publication No. WO 2017 / 0121693, published on April 27, 2017 2017 / 070633; U.S. Patent Publication No. 2015 / 0166980, published on June 18, 2015; U.S. Patent No. 9,840,699, issued on December 12, 2017; and U.S. Patent No. 10,077,453, issued on September 18, 2018, the contents of each of which are incorporated herein by reference in their entirety.
[0129] Cas9
[0130] The term "Cas9" or "Cas9 nuclease" refers to an RNA-guided nuclease comprising a Cas9 domain or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9 and / or a gRNA binding domain of Cas9). As used herein, a "Cas9 domain" is a protein fragment comprising an active or inactive cleavage domain of Cas9 and / or a gRNA binding domain of Cas9. A "Cas9 protein" is a full-length Cas9 protein. Cas9 nucleases are sometimes also referred to as casn1 nucleases or CRISPR ( C lustered R egularly I Interspaced S hort P alindromic Repeat, clustered regularly interspaced short palindromic repeats) related nuclease. CRISPR is an adaptive immune system that can provide protection for mobile genetic elements (viruses, transposable elements and conjugative plasmids). The CRISPR cluster contains a spacer, a sequence complementary to the previous mobile element, and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and Cas9 domains. TracrRNA is used as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleases cut linear or circular dsDNA targets complementary to the spacer. First, the target chain that is not complementary to crRNA is cut endonucleases, and then 3'-5' is trimmed exonucleases. In nature, DNA binding and cutting usually require proteins and two RNAs. However, a single guide RNA ("sgRNA", or simply "gNRAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self.Cas9 nuclease sequence and structure are well known to those skilled in the art (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma C. M., Gonzales K., Chao Y., Pirzada Z. A., Eckert M. R., Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, Streptococcus pyogenes and Streptococcus thermophilus.Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.
[0131] Nuclease-inactivated Cas9 domains are interchangeably referred to as "dCas9" proteins (for nuclease-"dead" Cas9). Methods for generating Cas9 domains (or fragments thereof) with inactivated DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337: 816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152 (5): 1173-83, the entire contents of each of which are incorporated herein by reference). For example, it is known that the DNA cleavage domain of Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cuts the strand complementary to the gRNA, while the RuvC1 subdomain cuts the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337: 816-821 (2012); Qi et al., Cell. 28; 152 (5): 1173-83 (2013)). In some embodiments, a protein comprising a Cas9 fragment is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) a gRNA binding domain of Cas9; or (2) a DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". Cas9 variants share homology with Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 200).In some embodiments, the Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 200). In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 200). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 200).
[0132] As used herein, the term "nCas9" or "Cas9 nickase" refers to Cas9 or a variant thereof that cuts only one strand of the target cleavage site or nicks one strand of the target cleavage site, thereby introducing a nick in the double-stranded DNA molecule instead of generating a double-strand break. This can be achieved by introducing an appropriate mutation into the wild-type Cas9 that inactivates one of the two endonuclease activities of Cas9. It is contemplated that any suitable mutation that inactivates one Cas9 endonuclease activity but leaves the other intact, such as one of the D10A or H840A mutations in the wild-type Streptococcus pyogenes Cas9 amino acid sequence, or the D10A mutation in the wild-type Staphylococcus aureus Cas9 amino acid sequence, can be used to form nCas9.
[0133] cDNA
[0134] The term "cDNA" refers to a DNA strand copied from an RNA template. The cDNA is complementary to the RNA template.
[0135] Circular replacement
[0136] As used herein, the term "cyclic substitution" refers to a protein or polypeptide (e.g., Cas9) comprising a cyclic substitution, which is a change in the structural configuration of a protein, involving a change in the order of amino acids that appear in the amino acid sequence of the protein. In other words, a cyclic substitution is a protein with an altered N-terminus and C-terminus compared to a wild-type counterpart, for example, the wild-type C-terminal half of a protein becomes a new N-terminal half. Circular substitution (or CP) is essentially a topological rearrangement of the primary sequence of a protein, typically using a peptide linker to connect its N-terminus and C-terminus, while splitting its sequence at different positions to produce new adjacent N-termini and C-termini. The result is that the protein structure has different connectivity, but generally has the same overall similar three-dimensional (3D) shape, and may include improved or altered properties, including reduced proteolytic sensitivity, improved catalytic activity, altered substrate or ligand binding and / or improved thermal stability. Circular substitution proteins can exist in nature (e.g., concanavalin A and lectins). In addition, circular permutations can occur as a result of post-translational modifications or can be engineered using recombinant techniques (e.g., see Oakes et al., "Protein Engineering of Cas9 for enhanced function," Methods Enzymol, 2014, 546: 491-511 and Oakes et al., "CRISPR-Cas9 Circular Permutants as Programmable Scaffolds for Genome Modification," Cell, January 10, 2019, 176: 254-267, each of which is incorporated herein by reference).
[0137] CRISPR
[0138] CRISPR is a family of DNA sequences in bacteria and archaea (i.e., CRISPR clusters), which represent fragments of previous infections by viruses that have invaded prokaryotes. DNA fragments are used by prokaryotic cells to detect and destroy DNA from subsequent attacks of similar viruses, and together with arrays of CRISPR-related proteins (including Cas9 and homologues thereof) and CRISPR-related RNAs, effectively form a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), the correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and Cas9 domains. TracrRNA is used as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleases cut linear or circular dsDNA targets complementary to RNA. Specifically, endonucleases first cut target chains that are not complementary to crRNA, and then exonucleases trim 3'-5'. In nature, DNA binding and cleavage usually require proteins and two RNAs. However, a single guide RNA ("sgRNA", or simply "gRNA") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species-guide RNA. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes short motifs (PAM or protospacer adjacent motifs) in CRISPR repeat sequences to help distinguish self from non-self.CRISPR biology and Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski E., et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” K., Sharma C. M., Gonzales K., Chao Y., Pirzada Z. A., Eckert M. R., Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, Streptococcus pyogenes and Streptococcus thermophilus.Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.
[0139] In certain types of CRISPR systems (e.g., type II CRISPR systems), trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and Cas9 proteins are required for the correct processing of pre-crRNA. TracrRNA acts as a guide for ribonuclease 3 to assist in the processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleases cut linear or circular nucleic acid targets complementary to RNA. Specifically, the target strand that is not complementary to crRNA is first endonucleases, and then 3'-5' exonucleases are trimmed. In nature, DNA binding and cutting generally require proteins and two kinds of RNA. However, a single guide RNA ("sgRNA" or "gRNA" for short) can be engineered to incorporate the embodiments of both crRNA and tracrRNA into a single RNA species-guide RNA.
[0140] In general, "CRISPR system" refers to the collective term for transcripts and other elements that participate in CRISPR-associated ("Cas") gene expression or guide Cas gene activity, including sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr pairing sequences (covering "direct repeats" and partial direct repeats processed by tracrRNA in the context of endogenous CRISPR systems), guide sequences (also referred to as "spacers" in the context of endogenous CRISPR systems) or other sequences and transcripts from CRISPR loci. The tracrRNA of the system is complementary (completely or partially) to the tracr pairing sequence present on the guide RNA.
[0141] Degron
[0142] The term "degrader" or "degrader domain" refers to a portion of a polypeptide that affects, controls, directs or otherwise regulates the rate of polypeptide degradation. Degraders can be highly variable and can include short amino acid sequences, structural motifs and / or exposed amino acids. In addition, degraders can be located at any position within a polypeptide (e.g., at the N-terminus, C-terminus, or at an internal position within the primary structure). The specific mechanism of polypeptide degradation regulated by degraders is not limited and can include ubiquitin-dependent degradation (i.e., degradation involving proteasome-based degradation) or ubiquitin-independent degradation. For example, the 4-amino acid sequence tail of NH3-EMLA-COOH (SEQ ID NO: 384) encoded by exon 8 of the SMN2 gene acts as a degrader, thereby initiating the degradation of SMN2.
[0143] Effective amount
[0144] As used herein, the term "effective amount" refers to the amount of a bioactive agent sufficient to cause a desired biological response. For example, in some embodiments, the effective amount of a base editor may refer to the amount of a base editor sufficient to edit a target site nucleotide sequence (e.g., a genome). In some embodiments, the effective amount of a base editor provided herein (e.g., a base editor comprising a Cas9 nickase domain and a nucleobase modification domain (e.g., a deaminase domain)) may refer to the amount of a base editor sufficient to induce editing of a target site specifically bound and edited by a base editor. In some embodiments, the effective amount of a base editor provided herein may refer to the amount of a base editor sufficient to induce editing with the following characteristics: >50% product purity, <5% indels and / or an editing window of 2-8 nucleotides in the region surrounding the target sequence. In other embodiments, the effective amount of a base editor may refer to the amount of a base editor sufficient to induce the following editing: >45% product purity, <10% indels, the ratio of expected point mutations to indels is at least 5:1 and / or the editing window is 2-10 nucleotides. As will be appreciated by those skilled in the art, the effective amount of an agent (e.g., a base editor, a nuclease, a deaminase, a hybrid protein, a complex of a protein and a polynucleotide, or a polynucleotide (e.g., a gRNA)) can vary depending on various factors, such as, for example, the desired biological response, for example, according to the specific allele, genome, or target site to be edited, according to the target cell or tissue (i.e., the cell or tissue to be edited), and according to the agent used.
[0145] Off-target and on-target editing
[0146] As used herein, the term "off-target editing" refers to the introduction of unintended modifications (e.g., deamination) to nucleotides (e.g., cytosine) in sequences outside the typical base editor binding window (i.e., from one protospacer position to another protospacer position, typically 2 to 8 nucleotides long). Off-target editing may be due to weak or nonspecific binding of the gRNA sequence to the target sequence. Off-target editing may also be due to the intrinsic association of the nucleotide modification domain (e.g., deaminase domain) of the base editor with nucleobases in a locus unrelated to the target sequence.
[0147] The term "Cas9-dependent off-target editing" refers to the introduction of unintended modifications caused by weak or non-specific binding of a Cas9-gRNA complex (e.g., a complex between a gRNA and the Cas9 domain of a base editor) to a nucleic acid site that has a fairly high (e.g., greater than 60%, or less than 6 mismatches relative to the target sequence) sequence identity with the target sequence. In contrast, the term "Cas9-independent off-target editing" refers to the introduction of unintended modifications caused by weak association of a base editor (e.g., a nucleotide modification domain) with a nucleic acid site that does not have a high sequence identity (about 60% or less, or 6-8 or more mismatches relative to the target sequence) with the target sequence. Because these associations occur independently of any hybridization between the Cas9-gRNA complex and the associated nucleic acid site, they are referred to as "Cas9-independent."
[0148] As used herein, the term "on-target editing" refers to the introduction of a desired modification (e.g., deamination) to a nucleotide (e.g., cytosine) in a target sequence, for example using a base editor described herein.
[0149] As used herein, the terms "on-target editing frequency" and "on-target editing efficiency" refer to the number or proportion of expected base pairs edited. For example, if a base editor edits 10% of its expected targeted base pairs (e.g., within a cell or within a cell population), the efficiency of the base editor can be described as 10%. Certain aspects of editing efficiency include modifying (e.g., deaminating) specific nucleotides within DNA without generating a large number or a high percentage of insertions or deletions (i.e., indels). It is generally accepted that editing constitutes a high editing efficiency when less than 5% indels are generated in the region immediately surrounding the target sequence (as measured on the total target nucleotide substrate). Generating more than 20% indels is generally considered to be poor or low editing efficiency.
[0150] As used herein, the term "off-target editing efficiency" refers to the number or ratio of unexpected base pairs edited. The on-target and off-target editing frequencies can be measured by the methods and assays described herein, further taking into account the techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves hybridization of nucleic acid primers (e.g., DNA primers) complementary to the nucleic acid (e.g., DNA) region just upstream or downstream of the target sequence or off-target sequence of interest. Because the DNA target sequence and the Cas9-independent off-target sequence are innately known in the methods disclosed herein, techniques known in the art (e.g., PhusionU PCR kit (Life Technologies), Phusion HS II kit (Life Technologies) and Illumina MiSeq kit) can be used to design nucleic acid primers with sufficient complementarity to the target sequence of interest and the region upstream or downstream of the Cas9-independent off-target sequence. Since many Cas9-dependent off-target sites have high sequence identity with the target site of interest, it is also possible to use techniques known in the art and kits to design nucleic acid primers with sufficient complementarity to the region upstream or downstream of the Cas9-dependent off-target site. These kits utilize polymerase chain reaction (PCR) amplification, which produces an amplicon as an intermediate product. The target sequence and off-target sequence may include a genomic locus further containing a protospacer and a PAM. Therefore, as used herein, the term "amplicon" may refer to a nucleic acid molecule constituting an aggregate of a genomic locus, a protospacer, and a PAM. The high-throughput sequencing technology used herein may further include Sanger sequencing and / or whole genome sequencing (WGS). The off-target effects of the disclosed base editors can be measured using the assays and methods disclosed in International Application No. PCT / US2020 / 624628, filed on November 25, 2020, which are incorporated herein by reference.
[0151] Functional equivalents
[0152] The term "functional equivalent" refers to a second biomolecule that is functionally equivalent to a first biomolecule but not necessarily structurally equivalent. For example, a "Cas9 equivalent" refers to a protein that has the same or substantially the same function as Cas9 but does not necessarily have the same amino acid sequence. In the context of the present disclosure, this specification refers to "protein X or its functional equivalent" throughout. In this context, a "functional equivalent" of protein X includes any homolog, paralog, fragment, naturally occurring, engineered, ring-substituted, mutated or synthetic form of protein X with equivalent function.
[0153] Fusion Protein
[0154] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) portion of the fusion protein or the carboxyl-terminal (C-terminal) protein, thereby forming an "amino-terminal fusion protein" or a "carboxyl-terminal fusion protein", respectively. The protein can contain different domains, for example, a nucleic acid binding domain (e.g., a gRNA binding domain of Cas9 that guides the protein to bind to the target site) and a nucleic acid cleavage domain or catalytic domain of a nucleic acid editing protein. Another example includes Cas9 fused to adenosine deaminase or its equivalent. Any protein provided herein can be produced by any method known in the art. For example, the protein provided herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins comprising peptide linkers. Methods for recombinant protein expression and purification are well known and include Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.
[0155] Guide nucleic acid
[0156] The term "guide nucleic acid" or "napDNAbp-programmed nucleic acid molecule" or equivalently "guide sequence" refers to one or more nucleic acid molecules that associate with the napDNAbp protein and guide or otherwise program the napDNAbp protein to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to the one or more nucleic acid molecules (or portions or regions thereof) associated with the protein, thereby binding the napDNAbp protein to the nucleotide sequence at the specific target site. A non-limiting example is the guide RNA of the Cas protein of the CRISPR-Cas genome editing system.
[0157] Guide RNA is a special type of guide nucleic acid, which is most commonly associated with the Cas protein of CRISPR-Cas9 and associated with Cas9, guiding the Cas9 protein to a specific sequence in a DNA molecule, which is complementary to the original spacer sequence of the guide RNA. As used herein, "guide RNA" refers to a synthetic fusion of endogenous bacterial crRNA and tracrRNA, which provides the Cas9 nuclease with targeting specificity and scaffold and / or binding ability to the target DNA. This synthetic fusion does not exist in nature and is also commonly referred to as sgRNA. However, the term also includes an equivalent guide nucleic acid molecule associated with a Cas9 equivalent, homologue, orthologue or paralogue, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and it is otherwise programmed to locate to a specific target nucleotide sequence to the Cas9 equivalent. Cas9 equivalents can include other napDNAbps from any type of CRISPR system (e.g., type II, type V, type VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system) and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353 (6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein. In addition, methods for designing suitable guide RNA sequences are provided herein.
[0158] Guide RNA ("gRNA")
[0159] As used herein, "guide RNA" is a specific type of guide nucleic acid, which is generally associated with the Cas protein of CRISPR-Cas9, and is combined with Cas9 to guide the Cas9 protein to a specific sequence in a DNA molecule, which contains complementarity with the original spacer sequence of the guide RNA. However, the term also covers equivalent guide nucleic acid molecules associated with Cas9 equivalents, homologs, orthologs, or paralogs (whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant)), and it otherwise programs Cas9 equivalents to locate to a specific target nucleotide sequence. Cas9 equivalents can include other napDNAbps from any type of CRISPR system (e.g., II, V, VI types), including Cpf1 (V-type CRISPR-Cas system), C2c1 (V-type CRISPR-Cas system), C2c2 (VI-type CRISPR-Cas system) and C2c3 (V-type CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein.
[0160] The guide RNA can contain various structural elements, including but not limited to (a) a spacer sequence—a sequence in the guide RNA (having a length of approximately 20 nt) that binds to the complementary strand of the target DNA (and having the same sequence as the protospacer of the DNA) and (b) a gRNA core (or gRNA scaffold or backbone sequence)—refers to the sequence within the gRNA responsible for Cas9 binding, which does not include the approximately 20 bp spacer sequence used to guide Cas9 to the target DNA.
[0161] As used herein, "guide RNA target sequence" refers to about 20 nucleotides that are complementary to the protospacer sequence in the PAM strand. The target sequence is the sequence that anneals to or is targeted by the spacer sequence of the guide RNA. The spacer sequence and the protospacer of the guide RNA have the same sequence (except that the spacer sequence is RNA, while the protospacer is DNA).
[0162] As used herein, the term "guide RNA scaffold sequence" refers to the sequence within the gRNA responsible for Cas9 binding, which does not include the 20 bp spacer / targeting sequence used to guide Cas9 to the target DNA.
[0163] Host cells
[0164] As used herein, the term "host cell" refers to cells that can carry, replicate and transfer phage vectors that can be used for the continuous evolution process provided herein. In the embodiment where the vector is a viral vector, a suitable host cell is a cell that can be infected by a viral vector, can replicate the viral vector, and can package the viral vector into a viral particle that can infect a fresh host cell. If the cell supports the expression of viral vector genes, the replication of the viral genome, and / or the generation of viral particles, the cell can carry the viral vector. One criterion for determining whether a cell is a suitable host cell for a given viral vector is to determine whether the cell can support the viral life cycle of the wild-type viral genome derived from the viral vector. For example, if the viral vector is a modified M13 phage genome, as provided in some embodiments described herein, a suitable host cell will be any cell that can support the wild-type M13 phage life cycle. Suitable host cells for viral vectors that can be used for the continuous evolution process are well known to those skilled in the art, and the present disclosure is not limited in this regard. In some embodiments, the viral vector is a phage and the host cell is a bacterial cell. In some embodiments, the host cell is an Escherichia coli cell. Suitable E. coli host strains are apparent to those skilled in the art, and include, but are not limited to, New England Biolabs (NEB) Turbo, Top10F', DH12S, ER2738, ER2267, and XL1-Blue MRF'. These strain names are well-recognized in the art, and the genotypes of these strains have been well characterized. It should be understood that the above strains are exemplary only, and the present invention is not limited in this regard. The term "fresh" is used interchangeably herein with the term "non-infected" or "uninfected" in the context of host cells, and refers to a host cell that is not infected by a viral vector containing a gene of interest used in the continuous evolution process provided herein. However, fresh host cells may have been infected by a viral vector unrelated to the vector to be evolved, or may have been infected by a vector of the same or similar type but not carrying a gene of interest.
[0165] In some embodiments, the host cell is a prokaryotic cell, such as a bacterial cell. In some embodiments, the host cell is an Escherichia coli cell. In some embodiments, the host cell is a eukaryotic cell, such as a yeast cell, an insect cell, or a mammalian cell. Of course, the type of host cell will depend on the viral vector used, and suitable host cell / viral vector combinations will be apparent to those skilled in the art.
[0166] Inteins and split inteins
[0167] As used herein, the term "intein" refers to a self-processing polypeptide domain present in organisms across all domains of life. An intein (intervening protein) undergoes a unique self-processing event called protein splicing, in which it excises itself from a larger precursor polypeptide by cleaving two peptide bonds and in the process joins the flanking extein (exoprotein) sequences by forming new peptide bonds. This rearrangement occurs after translation (or possibly concurrently with translation) as intein genes are found embedded within the framework of other protein-coding genes. Furthermore, intein-mediated protein splicing is spontaneous; it does not require external factors or energy sources, only the intein domain needs to be folded. This process is also known as cis-protein splicing, as opposed to the natural process of trans-protein splicing using "split inteins."
[0168] Split inteins are a subclass of inteins. Unlike the more common continuous inteins, split inteins are transcribed and translated into two separate polypeptides, the N-intein and the C-intein, each fused to an extein. After translation, the intein fragments spontaneously and non-covalently assemble into a typical intein structure to undergo trans-protein splicing.
[0169] Inteins and split inteins are the protein equivalents of self-splicing RNA introns (see Perler et al., Nucleic Acids Res. 22: 1125-1127 (1994)), which catalyze their own excision from a precursor protein and are accompanied by fusion of flanking protein sequences called exteins (reviewed in Perler et al., Curr. Opin. Chem. Biol. 1: 292-299 (1997); Perler, FB Cell 92(1): 1-4 (1998); Xu et al., EMBO J. 15(19): 5146-5153 (1996)).
[0170] As used herein, the term "protein splicing" refers to the process of removing the internal region of the precursor protein (intein) and connecting the flanking regions of the protein (extein) to form a mature protein. This natural process has been observed in many proteins of prokaryotes and eukaryotes (Perler, FB, Xu, MQ, Paulus, H. Current Opinion in Chemical Biology 1997, 1, 292-299; Perler, FB Nucleic Acids Research 1999, 27, 346-347). The intein unit contains the necessary components required to catalyze protein splicing and usually contains an endonuclease domain involved in the movement of the intein (Perler, FB, Davis, EO, Dean, GE, Gimble, FS, Jack, WE, Neff, N., Noren, CJ, Thomer, J., Belfort, M. Nucleic Acids Research 1994, 22, 1127-1127). However, the proteins produced are linked rather than expressed as separate proteins. Protein splicing can also occur in trans, where split inteins expressed on separate polypeptides spontaneously join to form a single intein, which then undergoes the protein splicing process to be linked to separate proteins.
[0171] The elucidation of the protein splicing mechanism has led to a large number of intein-based applications (Comb, et al., U.S. Pat. No. 5,496,714; Comb, et al., U.S. Pat. No. 5,834,247; Camarero and Muir, J. Amer. Chem. Soc., 121:5597-5598 (1999); Chong, et al., Gene, 192:271-281 (1997), Chong, et al., Nucleic Acids Res., 26:5109-5115 (1998); Chong, et al., J. Biol. Chem., 273:10567-10577 (1998); Cotton, et al. J. Am. Chem. Soc., 121:1100-1101 (1999); Evans, et al., Gene, 192:271-281 (1997), Chong, et al., Nucleic Acids Res., 26:5109-5115 (1998); Chong, et al., J. Biol. Chem., 273:10567-10577 (1998); Cotton, et al. J. Am. Chem. Soc., 121:1100-1101 (1999); Evans, et al., Gene, 192:271-281 (1997), Chong, et al., Nucleic Acids Res., 26:5109-5115 (1998); al., J. Biol. Chem., 274: 18359-18363 (1999); Evans, et al., J. Biol. Chem., 274: 3923-3926 (1999); Evans, et al., Protein Sci., 7: 2256-2264 (1998); Evans, et al. al., J. Biol. Chem., 275: 9091-9094 (2000); Iwai and Pluckthun, FEBS Lett. 459: 166-172 (1999); Mathys, et al., Gene, 231: 1-13 (1999); Mills, et al., Proc. Natl. Acad. Sci. USA 95:3543-3548(1998); Muir, et al. al., Proc. Natl. Acad. Sci. USA 95: 6705-6710 (1998); Otomo, et al., Biochemistry 38: 16040-16044 (1999); Otomo, et al., J. Biolmol. NMR 14: 105-114 (1999); Scott, et al. al., Proc. Natl. Acad. Sci. USA 96: 13638-13643 (1999); Severinov and Muir, J. Biol. Chem., 273: 16205-16209 (1998); Shingledecker, et al., Gene, 207: 187-195 (1998); Southworth, et al., EMBO J.17:918-926(1998); Southworth, et al., Biotechniques, 27: 110-120 (1999); Wood, et al., Nat. Biotechnol., 17: 889-892 (1999); Wu, et al., Proc. Natl. Acad. Sci. USA 95: 9226-9231 (1998a); Wu, et al., Biochim Biophys Acta 1387: 422-432 (1998b); Xu, et al., Proc. Natl. Acad. Sci. USA 96: 388-393 (1999); Yamazaki, et al., J. Am. Chem. Soc., 120: 5591-5592 (1998)). Each reference is incorporated herein by reference.
[0172] Connectors
[0173] As used herein, the term "linker" refers to a chemical group or molecule connecting two molecules or domains, for example, dCas9 and deaminase. Typically, the linker is located between or flanking two groups, molecules, or other domains, and is connected to each group, molecule, or other domain via a covalent bond, thereby connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical domain. Chemical groups include, but are not limited to, disulfide, hydrazone, and azide domains. In some embodiments, the length of the joint is 5-100 amino acid, for example, a length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150 or 150-200 amino acid. Longer or shorter joints are also contemplated. In some embodiments, the joint is an XTEN joint. In some embodiments, the joint is a 32 amino acid joint. In other embodiments, the joint is a 30, 31, 33 or 34 amino acid joint.
[0174] mutation
[0175] As used herein, the term "mutation" refers to the replacement of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) by another residue; the deletion or insertion of one or more residues within a sequence; or the replacement of a residue within the genomic sequence of a subject to be corrected. Mutations are generally described herein by identifying the original residue, followed by the position of the residue in the sequence, and the identity of the newly replaced residue. Various methods for preparing amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). Mutations can include a variety of categories, such as single base polymorphisms, microrepetitive regions, indels, and inversions, and are not meant to be limiting in any way. Mutations can include "loss of function" mutations, which are mutations that reduce or eliminate protein activity. Most loss of function mutations are recessive because in heterozygotes, the second chromosome copy carries an unmutated version of the gene encoding a fully functional protein, the presence of which compensates for the effect of the mutation. There are some exceptions where loss of function mutations are dominant, an example being haploinsufficiency, where the organism cannot tolerate the approximately 50% reduction in protein activity experienced by heterozygotes. This is an explanation for some genetic diseases in humans, including Marfan syndrome, which is caused by mutations in the gene for a connective tissue protein called fibrillin. Mutations also include "gain of function" mutations, which are mutations that confer abnormal activity to a protein or cell that is not present under normal conditions. Many gain of function mutations are located in regulatory sequences rather than in coding regions and can therefore have many consequences. For example, a mutation may cause one or more genes to be expressed in the wrong tissues, which gain functions that they normally lack. Alternatively, a mutation may cause one or more genes involved in controlling the cell cycle to be overexpressed, leading to uncontrolled cell division and, in turn, cancer. Due to their nature, gain-of-function mutations are often dominant.
[0176] napDNAbp
[0177] The term "napDNAbp," which stands for "nucleic acid programmable DNA binding protein," refers to any protein that can associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., which can be broadly referred to as "napDNAbp programming nucleic acid molecules," and include, for example, guide RNAs in the case of Cas systems) that direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to the one or more nucleic acid molecules (or portions or regions thereof) associated with the protein, thereby causing the protein to bind to the nucleotide sequence at the specific target site. The term napDNAbp includes CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and can include Cas9 equivalents from any type of CRISPR system (e.g., type II, V, VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, Cas13d, Cas14, Argonaute, and nCas9. Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated herein by reference. However, nucleic acid programmable DNA binding proteins (napDNAbp) that can be used in conjunction with the present invention are not limited to CRISPR-Cas systems. The present invention includes any such programmable proteins, such as Argonaute proteins from Natronobacterium gregoryi (NgAgo), which can also be used for DNA-guided genome editing. The NgAgo-guided DNA system does not require a PAM sequence or a guide RNA molecule, which means that genome editing can be simply performed by expressing a universal NgAgo protein and introducing a synthetic oligonucleotide on any genomic sequence. See Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute. Nature Biotechnology 2016; 34(7): 768-73, which is incorporated herein by reference.
[0178] In some embodiments, napDNAbp is an RNA programmable nuclease, which, when in a complex with RNA, can be referred to as a nuclease: RNA complex. Typically, the bound RNA is referred to as a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), although "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and guides the binding of the Cas9 (or equivalent) complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and comprises a stem-loop structure. For example, in some embodiments, domain (2) is similar to that of Jinek et al., Science 337:816-821 (2012) Figure 1E, the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Patent No. 9,340,799, entitled "mRNA-Sensing Switchable gRNAs", and International Patent Application No. PCT / US2014 / 054247, filed on September 6, 2013, published as WO 2015 / 035136 and entitled "Delivery System For Functional Nucleases", the entire contents of each of which are incorporated herein by reference. In some embodiments, the gRNA comprises two or more of domains (1) and (2), and can be referred to as an "extended gRNA". For example, an extended gRNA will, for example, bind to two or more Cas9 proteins and bind to a target nucleic acid at two or more different regions, as described herein. The gRNA comprises a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site, providing sequence specificity of the nuclease: RNA complex. In some embodiments, the RNA programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 from Streptococcus pyogenes (Csn1) (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti JJ et al.., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and hostfactor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference.
[0179] napDNAbp nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, and these proteins can in principle target any sequence specified by a guide RNA. Methods for site-specific cleavage (e.g., to modify a genome) using napDNAbp nucleases such as Cas9 are known in the art (see, e.g., Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al. Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature Biotechnology 31, 227-229 (2013); Jinek, M. et al. RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acid Res. (2013); Jiang, W. et al. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).
[0180] Nicking enzyme
[0181] The term "nickase" refers to a napDNAbp that has only a single nuclease activity that cuts only one strand of the target DNA instead of both strands (e.g., one of the two nuclease domains is inactivated). Therefore, the nickase-type napDNAbp does not leave double-strand breaks. In some embodiments, any disclosed base editor or vector may include a S.pyogenes Cas9 nickase (SpCas9n or nCas9) containing a D10A mutation. In some embodiments, any disclosed base editor may include an Nme2Cas9 nickase (Nme2Cas9n) containing a D16A mutation.
[0182] nuclear localization signal
[0183] Nuclear localization signal or sequence (NLS) is an amino acid sequence that marks, specifies or otherwise identifies a protein to be imported into the nucleus by nuclear transport. Typically, the signal consists of one or more short sequences of positively charged lysine or arginine exposed on the surface of the protein. Different nuclear localization proteins can share the same NLS. NLS has a function opposite to the nuclear export signal (NES), which targets proteins outside the nucleus. Therefore, a single nuclear localization signal can guide an entity associated with it to the nucleus. Such a sequence can be of any size and composition, for example, more than 25, 25, 15, 12, 10, 8, 7, 6, 5 or 4 amino acids, but preferably comprises at least 4 to 8 amino acid sequences known to act as nuclear localization signals (NLS).
[0184] Nucleic acid molecules
[0185] As used herein, the term "nucleic acid molecule" refers to RNA and single-stranded and / or double-stranded DNA. Nucleic acid molecules can be naturally occurring, such as in the context of genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cosmids, chromosomes, chromatids or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or include non-naturally occurring nucleotides or nucleosides. In addition, the terms "nucleic acid", "DNA", "RNA" and / or similar terms include nucleic acid analogs, for example, with analogs other than the phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using a recombinant expression system, and optionally purified, chemically synthesized, etc. Where appropriate, such as in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as analogs with chemically modified bases or sugars and backbone modifications. Unless otherwise indicated, nucleic acid sequences are presented in 5' to 3' directions. In some embodiments, the nucleic acid is or comprises natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, adenosine, deoxythymidine, deoxyguanosine, and cytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine); , 7-deazaadenosine, 7-deazaguanosine, inosinedenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite bonds).
[0186] PACE
[0187] As used herein, the term "phage-assisted continuous evolution (PACE)" refers to continuous evolution using bacteriophage as a viral vector. The general concept of PACE technology has been described below: for example, international PCT application PCT / US2009 / 056194 filed on September 8, 2009, published as WO 2010 / 028347 on March 11, 2010; international PCT application PCT / US2011 / 066747 filed on December 22, 2011, published as WO 2012 / 088381 on June 28, 2012; U.S. application, U.S. Patent No. 9,023,594 issued on May 5, 2015, international PCT application PCT / US2015 / 012022 filed on January 20, 2015, published as WO 2010 / 028347 on March 11, 2010; international PCT application PCT / US2011 / 066747 filed on December 22, 2011, published as WO 2012 / 088381 on June 28, 2012; U.S. application, U.S. Patent No. 9,023,594 issued on May 5, 2015, international PCT application PCT / US2015 / 012022 filed on January 20, 2015, published as WO 2015 / 134121, and international PCT application PCT / US2016 / 027795 filed on April 15, 2016, published as WO 2016 / 168631 on October 20, 2016, the entire contents of each of which are incorporated herein by reference.
[0188] Promoter
[0189] The term "promoter" is recognized in the art and refers to a nucleic acid molecule having a sequence that is recognized by the cell transcription machinery and can initiate transcription of downstream genes. A promoter can be constitutively active, meaning that the promoter is always active in a given cell environment, or conditionally active, meaning that the promoter is active only under specific conditions. For example, a conditional promoter can be active only in the presence of a specific protein that connects a protein associated with a regulatory element in the promoter to a basic transcription machinery, or only in the absence of an inhibitory molecule. A subclass of conditionally active promoters is an inducible promoter, which requires the presence of a small molecule "inducer" to obtain activity. Examples of inducible promoters include, but are not limited to, arabinose inducible promoters, Tet-on promoters, and tamoxifen inducible promoters. Various constitutive, conditional, and inducible promoters are well known to those skilled in the art, and those skilled in the art will be able to determine various such promoters that can be used to implement the present invention, and the present invention is not limited to this aspect. In various embodiments, the present disclosure provides a vector having a suitable promoter for driving the expression of a nucleic acid sequence encoding a fusion protein (or one or more of its individual components).
[0190] Product purity
[0191] As used herein, the term "product purity" refers to the percentage of the desired product in the base editing reaction to the total product. For example, the product purity of a CBE can be measured as the percentage of total edited sequencing reads in which the target C is edited to T (reads in which the target C has been converted to a different base) to the portion of interest in the nucleic acid. Product purity includes the absence of indels, as well as the desired product of the base conversion.
[0192] The term "R loop" refers to a triplex structure in which the two strands of double-stranded DNA are separated by a stretch of nucleotides and separated by a single-stranded RNA molecule (e.g., gRNA). R loop formation can be induced by hybridization of a gRNA with complementarity to DNA associated with a napDNAbp protein or domain (e.g., Cas9). When the mechanisms that generate their formation (e.g., napDNAbp-gRNA complexes) act independently of each other, two R loops are referred to as "orthogonal."
[0193] Protospacer
[0194] As used herein, the term "protospacer" refers to a sequence (about 20 bp) in DNA adjacent to a PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals with the complementary sequence of the protospacer sequence on the target DNA (specifically, one of its strands, i.e., the "target strand" and the "non-target strand" of the target DNA sequence). In order for Cas9 to work, it also requires a specific protospacer adjacent motif (PAM), which varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease is derived from Streptococcus pyogenes and recognizes the PAM sequence of NGG present directly downstream of the target sequence in genomic DNA, located on the non-target strand. The technician will understand that the prior art literature sometimes refers to the "protospacer" as the approximately 20nt target-specific guide sequence on the guide RNA itself, rather than referring to it as a "spacer". Therefore, in some cases, the term "protospacer" used herein can be used interchangeably with the term "spacer". The context in which the term “protospacer” or “spacer” appears in the specification will help inform the reader whether the term refers to the gRNA or the DNA target.
[0195] Protospacer adjacent motif (PAM)
[0196] As used herein, the term "protospacer adjacent sequence" or "PAM" refers to a DNA sequence of about 2-6 base pairs, which is an important targeting component of the Cas9 nuclease. Typically, the PAM sequence is on either chain and downstream of the Cas9 cleavage site in the 5' to 3' direction. A typical PAM sequence (i.e., a PAM sequence associated with the Cas9 nuclease or SpCas9 of Streptococcus pyogenes)) is 5'-NGG-3', where "N" is any nucleobase followed by two guanine ("G") nucleobases. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease (e.g., SpCas9) can be modified to change the PAM specificity of the nuclease so that the nuclease recognizes an alternative PAM sequence.
[0197] For example, with reference to the typical SpCas9 amino acid sequence is SEQ ID NO: 200, the PAM sequence can be modified by introducing one or more mutations, including (a) D1135V, R1335Q and T1337R "VQR variant", which changes the specificity of the PAM to NGAN or NGNG, (b) D1135E, R1335Q and T1337R "EQR variant", which changes the specificity of the PAM to NGAG, and (c) D1135V, G1218R, R1335E and T1337R "VRER variant", which changes the specificity of the PAM to NGCG. In addition, the D1135E variant of the typical SpCas9 still recognizes NGG, but has higher selectivity than the wild-type SpCas9 protein.
[0198] It should also be understood that Cas9 enzymes (i.e., Cas9 orthologs) from different bacterial species can have different PAM specificities. For example, Cas9 (SaCas9) from Staphylococcus aureus recognizes NGRRT or NGRRN. In addition, Cas9 (NmCas) from Neisseria meningitidis recognizes NNNNGATT. In another example, Cas9 (StCas9) from Streptococcus thermophilus recognizes NNAGAAW. In yet another example, Cas9 (TdCas) from Treponema denticola recognizes NAAAC. These are examples and are not meant to be limiting. It will be further understood that non-SpCas9 binds to a variety of PAM sequences, which makes them useful when there is no suitable SpCas9 PAM sequence at the desired target cleavage site. In addition, non-SpCas9 may have other features that make them more useful than SpCas9. For example, Cas9 (SaCas9) from Staphylococcus aureus is about 1 kilobase smaller than SpCas9, so it can be packaged into adeno-associated virus (AAV). Further reference may be made to Shah et al., “Protospacer cognitive motifs: mixed identities and functional diversity,” RNA Biology, 10(5): 891-899 (which is incorporated herein by reference).
[0199] Chain of Justice
[0200] In genetics, the "sense" strand is a fragment running from 5' to 3' in a double-stranded DNA, and it is complementary to the antisense strand or template strand of the DNA running from 3' to 5'. In the case of a DNA fragment encoding a protein, the sense strand is a DNA strand with the same sequence as mRNA, which uses the antisense strand as its template during transcription, and eventually undergoes (usually, not always) translation into protein. Therefore, the antisense strand is responsible for the RNA that is subsequently translated into protein, and the sense strand has a composition almost identical to that of mRNA. Note that for each fragment of dsDNA, there may be two groups of sense strands and antisense strands, depending on the direction of reading (because sense strands and antisense strands are related to the viewing angle). Ultimately, it is the gene product or mRNA that determines which strand of a fragment of dsDNA is called the sense strand or antisense strand.
[0201] Spacer sequence
[0202] As used herein, the term "spacer sequence" in relation to guide RNA refers to a portion of about 20 nucleotides in the guide RNA that contains a nucleotide sequence complementary to the original spacer sequence in the target DNA sequence. The spacer sequence anneals with the original spacer sequence to form a ssRNA / ssDNA hybrid structure at the target site, as well as a corresponding R-loop ssDNA structure of the endogenous DNA chain complementary to the original spacer sequence.
[0203] Subjects
[0204] As used herein, the term "subject" refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, an amphibian, a reptile, a fish, an insect, a fly, or a nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject can be of any gender and can also be at any stage of development. In some embodiments, the subject is a plant.
[0205] Target site
[0206] The term "target site" refers to a sequence within a nucleic acid molecule that is edited by a fusion protein (e.g., a dCas9-deaminase fusion protein provided herein). A target site also refers to a sequence within a nucleic acid molecule to which a complex of a fusion protein and a gRNA binds.
[0207] Transcription terminator
[0208] A "transcription terminator" is a nucleic acid sequence that causes transcription to stop. A transcription terminator can be unidirectional or bidirectional. It consists of a DNA sequence that participates in the specific termination of RNA transcripts by RNA polymerase. The transcription terminator sequence prevents the transcriptional activation of the downstream nucleic acid sequence by the upstream promoter. The transcription terminator may be necessary in vivo to achieve a desired expression level or to avoid transcription of certain sequences. When a transcription terminator is able to terminate the transcription of the sequence to which it is connected, it is considered that the transcription terminator is "operably linked" to a nucleotide sequence.
[0209] The most commonly used terminator type is the forward terminator. When the forward transcription terminator is located downstream of the nucleic acid sequence that is usually transcribed, it will cause transcription to stop. In some embodiments, a bidirectional transcription terminator is provided, which usually causes transcription to stop on the forward strand and the reverse strand. In some embodiments, a reverse transcription terminator is provided, which usually only stops transcription on the reverse strand.
[0210] In prokaryotic systems, terminators are generally divided into two categories: (1) rho-independent terminators and (2) rho-dependent terminators. Rho-independent terminators are generally composed of a palindromic sequence that forms a stem-loop rich in GC base pairs followed by several T bases. Without being bound by theory, the conventional model for transcription termination is that the stem-loop causes RNA polymerase to pause, while transcription of the poly-A tail causes the RNA:DNA duplex to unwind and dissociate from the RNA polymerase.
[0211] In eukaryotic systems, terminator regions can include specific DNA sequences that allow site-specific cutting of new transcripts to expose polyadenylation sites. This signals specialized endogenous polymerases to add a fragment (polyA) of approximately 200 A residues to the 3' end of the transcript. RNA molecules modified with this polyA tail appear to be more stable and more efficiently translated. Therefore, in some embodiments relating to eukaryotic organisms, terminator regions can include signals for cutting RNA. In some embodiments, terminator signals promote the polyadenylation of information. Terminator and / or polyadenylation site elements can be used to enhance output nucleic acid levels and / or minimize read-through between nucleic acids.
[0212] The terminator used according to the present disclosure includes any transcription terminator described herein or known to those of ordinary skill in the art. Examples of terminators include, but are not limited to, the terminator sequence of a gene such as, for example, a bovine growth hormone terminator, and viral terminator sequences such as, for example, SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA, and arcA terminators. In some embodiments, the termination signal can be a sequence that cannot be transcribed or translated, such as those produced by sequence truncation.
[0213] Conversion
[0214] As used herein, "transition" refers to the interchange of purine nucleobases. or pyrimidine nucleobase interchange Such swaps involve nucleobases of similar shape. The compositions and methods disclosed herein can induce one or more transitions in a target DNA molecule. The compositions and methods disclosed herein can also induce both transitions and transversions in the same target DNA molecule. These changes involve or In the context of double-stranded DNA with Watson-Crick paired nucleobases, a transversion refers to the following base pair exchange: or The compositions and methods disclosed herein are capable of inducing one or more transitions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing both transitions and transversions, as well as other nucleotide changes, including deletions and insertions, in the same target DNA molecule.
[0215] treat
[0216] The term "treatment" (treatment, treat, treating) refers to a clinical intervention intended to reverse a disease or disorder or one or more symptoms thereof, alleviate a disease or disorder or one or more symptoms thereof, delay the onset of a disease or disorder or one or more symptoms thereof, or inhibit the progression of a disease or disorder or one or more symptoms thereof, as described herein. As used herein, the term "treatment" (treatment, treat, treating) refers to a clinical intervention intended to reverse a disease or disorder or one or more symptoms thereof, alleviate a disease or disorder or one or more symptoms thereof, delay the onset of a disease or disorder or one or more symptoms thereof, or inhibit the progression of a disease or disorder or one or more symptoms thereof, as described herein. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after the disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, for example, to prevent symptoms or delay the onset of symptoms or inhibit the onset or progression of a disease. For example, treatment may be administered to susceptible individuals (for example, according to a history of symptoms and / or according to genetic or other susceptibility factors) before the onset of symptoms. Treatment may also continue after symptoms have subsided, for example, to prevent or delay their recurrence.
[0217] Uracil glycosylase inhibitors
[0218] As used herein,The term "uracil glycosylase inhibitor" or "UGI" refers to a protein that can inhibit the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises a wild-type UGI or a UGI as shown in SEQ ID NO: 272. In some embodiments, the UGI protein provided herein includes a UGI fragment and a protein homologous to UGI or a UGI fragment. For example, in some embodiments, the UGI domain comprises a fragment of the amino acid sequence shown in SEQ ID NO: 272. In some embodiments, the UGI fragment comprises the following amino acid sequence: comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% of the amino acid sequence shown in SEQ ID NO: 272. In some embodiments, UGI comprises an amino acid sequence homologous to the amino acid sequence shown in SEQ ID NO: 272, or an amino acid sequence homologous to a fragment of the amino acid sequence shown in SEQ ID NO: 272. In some embodiments, a protein comprising a UGI or a fragment of UGI or a homolog of a UGI or a fragment of UGI is referred to as a "UGI variant". A UGI variant shares homology with a UGI or a fragment thereof. For example, a UGI variant is at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to a wild-type UGI or a UGI as shown in SEQ ID NO: 272. In some embodiments, a UGI variant comprises a fragment of UGI such that the fragment is at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to a corresponding fragment of a wild-type UGI or a UGI as shown in SEQ ID NO: 272. In some embodiments, UGI comprises the following amino acid sequence:
[0219] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 272) (P14739|UNGI_BPPB2 uracil-DNA glycosylase inhibitor).
[0220] Variants
[0221] As used herein, the term "variant" refers to a protein that has features that deviate from proteins that exist in nature but retains at least one function (i.e., binding, interaction or enzymatic ability and / or therapeutic properties). The variant is at least about 70% identical to the wild-type protein, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. For example, a variant of Cas9 may include a Cas9 having one or more changes in amino acid residues compared to a wild-type Cas9 amino acid sequence. As another example, a variant of a deaminase may include a deaminase having one or more changes in amino acid residues compared to a wild-type deaminase amino acid sequence, for example, after reconstruction of the ancestral sequence of the deaminase. These changes include chemical modifications, including replacement, truncation, covalent addition (e.g., labeling) of different amino acid residues, and any other mutations. The term also encompasses circular substitutions, mutants, truncations or domains of a reference sequence that exhibit the same or substantially the same functional activity or activities as the reference sequence. The term also includes fragments of a wild-type protein.
[0222] The level or degree of retained properties may be reduced relative to the wild-type protein, but is generally identical or similar in kind. Typically, variants are very similar overall and, in many regions, identical to the amino acid sequence of the protein described herein. Those skilled in the art will understand how to prepare and use variants that retain all or at least some functional capabilities or properties.
[0223] The variant protein can comprise, or alternatively consist of, an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical to the amino acid sequence of, for example, a wild-type protein or any protein provided herein (e.g., an SMN protein).
[0224] By a polypeptide having an amino acid sequence that is at least, for example, 95% "identical" to a query amino acid sequence, it is meant that the amino acid sequence of the subject polypeptide is identical to the query sequence except that the subject polypeptide sequence may include up to five amino acid changes for every 100 amino acids of the query amino acid sequence. In other words, to obtain a polypeptide having an amino acid sequence that is at least 95% identical to the query amino acid sequence, up to 5% of the amino acid residues in the subject sequence may be inserted, deleted, or substituted with another amino acid. These changes to the reference sequence may occur at the amino-terminal or carboxyl-terminal positions of the reference amino acid sequence or anywhere between these terminal positions, interspersed individually between residues in the reference sequence or interspersed in one or more contiguous groups within the reference sequence.
[0225] In practice, whether any particular polypeptide is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to, for example, the amino acid sequence of a protein such as the SMN protein can be determined conventionally using known computer programs. A preferred method for determining the best overall match (also known as a global sequence alignment) between a query sequence (a sequence of the invention) and a subject sequence can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6:237-245 (1990)). In a sequence alignment, both the query sequence and the subject sequence are nucleotide sequences or both are amino acid sequences. The results of the global sequence alignment are expressed as percent identity. Preferred parameters for FASTDB amino acid alignment are: matrix=PAM 0, k-tuple=2, mismatch penalty=1, ligation penalty=20, randomization group length=0, cutoff score=1, window size=sequence length, gap penalty=5, gap size penalty=0.05, window size=500 or the length of the subject amino acid sequence, whichever is shorter.
[0226] If the subject sequence is shorter than the query sequence due to N-terminal or C-terminal deletions, rather than due to internal deletions, the result must be manually corrected. This is because the FASTDB program does not consider the N-terminal and C-terminal truncations of the subject sequence when calculating the global identity percentage. For the subject sequence that is truncated at the N-terminal and C-terminal relative to the query sequence, the percentage identity is corrected by calculating the number of residues located at the N-terminal and C-terminal of the subject sequence in the query sequence (these residues do not match / align with the corresponding subject residues) as a percentage of the total bases of the query sequence. Whether the residue matches / aligns is determined by the result of the FASTDB sequence comparison. The percentage is then subtracted from the percentage identity calculated using the specified parameters by the above-mentioned FASTDB program to obtain the final percentage identity score. This final percentage identity score is used for the purpose of the present invention. For the purpose of manually adjusting the identity percentage score, only the residues at the N-terminal and C-terminal of the subject sequence that do not match / align with the query sequence are considered. That is, only the residue positions outside the farthest N-terminal and C-terminal residues of the query sequence are queried.
[0227] Carrier
[0228] As used herein, the term "vector" refers to a nucleic acid that can be modified to encode a gene of interest and is capable of entering a host cell, mutating and replicating within the host cell, and then transferring the replicated form of the vector to another host cell. Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phages, and conjugative plasmids. Based on this disclosure, additional suitable vectors will be apparent to those skilled in the art.
[0229] wild type
[0230] As used herein, the term "wild type" is a term understood by those skilled in the art and means the typical form of an organism, strain, gene or trait occurring in nature, as distinguished from mutant or variant forms. DETAILED DESCRIPTION
[0231] The present disclosure provides a cytosine base editor comprising an evolutionarily directed adenosine deaminase domain (e.g., a variant of adenosine deaminase TadA that preferentially deaminates cytidine in DNA as described herein) and a napDNAbp domain (e.g., Cas9 protein) that can bind to a specific nucleotide sequence, wherein the adenosine deaminase variant provides a smaller size and lower off-target effect for the base editor (TadCBE) while maintaining the high editing efficiency of the existing CBE. TadCBE deamination of cytidine can result in a point mutation from cytosine (C) to (T) (a process referred to herein as nucleic acid editing), thereby converting a C·G base pair to a T·A base pair. Such base editors can be used for targeted editing of nucleic acid sequences, such as DNA molecules. Such base editors are particularly useful for in vitro targeted editing of DNA, for example, for generating mutant cells or animals. Such base editors can be used to introduce targeted mutations in living mammalian cells. Such base editors can also be used to introduce targeted mutations in ex vivo cells to correct genetic defects, such as in cells obtained from a subject and subsequently reintroduced into the same or another subject, or for multiplexed editing of multiple genes in a genome. These base editors can be used to introduce targeted mutations in vivo, such as correcting genetic defects or introducing inactivating mutations of disease-related genes in a subject, or for multiplexed editing of a genome. The cytosine base editors described herein can be used for targeted editing of T to C mutations (e.g., targeted genome editing). The present invention provides deaminases, base editors, nucleic acids, vectors, cells, compositions, methods, kits, and uses utilizing deaminases and base editors provided herein.
[0232] Here, PACE and PANCE were used to alter the substrate specificity of TadA-8e, generating a new class of selective cytidine deaminases (TadA-CDs) and cytosine base editors ( Figure 1A). To achieve cytidine deamination, TadA-CD variants acquired mutations at residues that interact with the DNA backbone near the active site. The disclosed TadA-CD cytosine base editor (TadCBE) is highly active and exhibits C·G to T·A conversion efficiencies comparable to or higher than current BE4max, evoAPOBEC1-BE4max (evoA), and evoFERNY-BE4max (evoFERNY) CBEs at various sites in mammalian cells. TadA-CD also interacts with SpCas9 (PAM = NGG) and evolved eNme2-C Cas9 (PAM = N 4 CN) variants, promoting broad target accessibility. Off-target analysis showed that TadCBE induced lower Cas-independent off-target DNA and RNA editing than the widely used APOBEC-based CBE variants. Addition of V106W mutation 9,34 Further reduce off-target editing of TadCBEs, optimize their editing windows, and improve selectivity from C·G to T·A while maintaining peak on-target editing efficiency. In this paper, the evolved TadCBEs were extensively characterized using a library of 10,638 genomically integrated, highly variable target sites in mouse embryonic stem cells (mESCs) to determine the selectivity and sequence context preferences of TadCBEs. TadA-CD is also compatible with both SpCas9 and evolved eNme2-C Cas9 variants, facilitating broad target accessibility. The disclosed TadCBEs can be used for efficient cytosine base editing at therapeutically relevant loci in human cells, including multiplexed editing, and in particular for cytosine editing at therapeutically relevant sites in primary human hematopoietic stem and progenitor cells (HSPCs). These disclosed TadCBEs exhibit more precise editing windows and less bystander editing than existing CBEs at, for example, the CXCR5 and CCR5 genes in primary human T cells. The present disclosure provides a new class of small CBEs with high on-target activity, a well-defined editing window that facilitates precise base editing, and low off-target activity, and establishes the potential for adenosine deaminases to evolve into selective cytidine deaminases.
[0233] In some aspects, the present disclosure relates to an adenosine deaminase (e.g., TadA-CD) having cytosine targeting activity. In some embodiments, TadA-CD is evolved from an E. coli tRNA adenosine deaminase (e.g., TadA-8e) that was previously engineered to act on single-stranded DNA (rather than RNA) for adenosine base editing applications. Those skilled in the art will appreciate that the PACE and PANCE methods can be used to introduce additional mutations in the TadA-8e domain, thereby changing the substrate specificity of the enzyme to produce TadA-CD. In some embodiments, TadA-CD (e.g., a mutant TadA-8e deaminase) has 80% to 99.5% sequence homology with the parent TadA-8e. In some cases, the TadA-CD deaminase comprises mutations at E27, V28, and H96, and further comprises at least one mutation at a residue selected from R26, M61, Y73, I76, M151, Q154, and A158, relative to parent TadA-8e.
[0234] In some embodiments, the TadA-CD variant has enhanced selectivity and deamination activity for cytosine relative to adenosine relative to the parent TadA-8e variant. For example, in some embodiments, the TadA-CD deaminase is at C 4 and C 5 The base editor containing the TadA-8e deaminase converted 85% to 92% (depending on the variant type) of CT base pairs to TA base pairs, while editing less than 2% of adenine; 6 The position converts approximately 90% of AT base pairs to GC base pairs, with less than 2% editing CG to TA base pairs (see Example 2). This represents a more than 3000-fold change in the ability of TadA-CD to edit cytosine and adenine bases compared to the TadA-8e variant.
[0235] In some aspects, the present disclosure relates to a cytosine base editor (CBE) comprising a nucleic acid programmable DNA binding protein (e.g., Cas9) domain fused to a TadA-CD deaminase (e.g., TadCBE) having cytidine activity. In some embodiments, the napDNAbp domain comprises a Cas homolog, paralog, ortholog, or an analog. The napDNAbp domain can be selected from Cas9, Cas9n (e.g., SpCas9n), dCas9, CasX, CasY, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, Cas9-NG, LbCas12a, enAsCas12a, SaCas9, SaCas9-KKH, a circularly arranged Cas9, an Argonaute (Ago) domain, SmacCas9, Spy-macCas9, SpCas9-VRQR, SpCas9-NRRH, SpaCas9-NRTH, SpCas9-NRCH, eNme2Cas9, eNme2-C Cas9, enCjCas9, SauriCas9, Cas9-NG-VRQR and / or variants thereof. In certain embodiments, the napDNAbp domain includes or is derived from a Cas9 domain or a Cas12a domain of Streptococcus pyogenes or Staphylococcus aureus. In some cases, the napDNAbp domain is derived from the Nme2Cas9 domain of Neisseria meningitidis. In some embodiments, the napDNAbp domain includes a nuclease-dead Cas9 (dCas9) domain, a Cas9 nickase (nCas9) domain, or a nuclease-active Cas9 domain. In some cases, the napDNAbp domain is CjCas9. In various embodiments, the napDNAbp domain is a nickase.
[0236] The disclosed CBEs exhibit low levels of undesired editing, such as low Cas9-independent off-target editing. The disclosed CBEs exhibit fewer insertions and / or deletions (indels) and undesired editing of RNA molecules following methods for editing target sequences in nucleic acids. The disclosed CBEs also exhibit editing efficiencies that exceed those of the most commonly used CBEs for several therapeutically relevant sites and cell types.
[0237] In some aspects, TadA-CD exhibits a narrower editing window than natural cytosine base editors while maintaining comparable or higher maximum editing efficiency. Taken together, the small size of TadCBE, its compatibility with eNme2Cas9 (and eNme2-CCas9), its more focused editing window, and its high editing efficiency and selectivity for cytosine relative to adenine base editors demonstrate their suitability for a variety of precision cytosine base editing applications.
[0238] Other aspects of the present disclosure relate to compositions comprising a TadCBE described herein and one or more guide RNAs (e.g., a single guide RNA "sgRNA"). In addition, the present disclosure provides nucleic acid molecules encoding and / or expressing the TadCBE described herein, and expression vectors or constructs (e.g., AAV vectors) for expressing the TadCBE and / or gRNA described herein, host cells comprising the nucleic acid molecules and expression vectors, and one or more gRNAs, and compositions for delivering and / or administering the nucleic acid-based embodiments described herein. In particular, the present disclosure provides improved methods for delivering the disclosed base editors, such as to a subject. Delivery of the disclosed TadCBE variants as RNPs rather than DNA plasmids generally increases the on-target: off-target DNA editing ratio. Delivery of the disclosed TadCBE variants as mRNA molecules (e.g., using electroporation) may increase editing efficiency. CBEs with an apparent on-target editing efficiency of about 50% in vivo have been described in International Publication No. WO / 2019 / 226953 (published on November 28, 2019) and Komoret al., Sci. Adv. 2017; 3:eaao4774, each of which is incorporated herein by reference. The disclosed CBEs can exhibit higher on-target editing efficiency for target cytosine bases.
[0239] Also provided herein are methods of contacting any disclosed TadCBE with a nucleic acid molecule, e.g., a nucleic acid molecule (e.g., DNA) comprising a target sequence. In some embodiments of the disclosed methods, low off-target DNA and / or RNA editing effects are observed. In some embodiments, the nucleic acid molecule comprises DNA, e.g., single-stranded DNA or double-stranded DNA. The target sequence of the nucleic acid molecule can comprise a target nucleobase pair comprising cytosine (C). The target sequence can be contained within a genome, e.g., within a human genome. The target sequence can comprise a sequence associated with a disease or condition such as sickle cell disease or HIV / AIDS, e.g., a target sequence having a point mutation. In other embodiments, the target nucleotide sequence is in the genome of a rodent, such as a mouse or rat. In other embodiments, the target nucleotide sequence is in the genome of a domestic animal, such as a horse, cat, dog, or rabbit. In some embodiments, the target nucleotide sequence is in the genome of a research animal. In some embodiments, the target nucleotide sequence is in the genome of a genetically engineered non-human subject. In some embodiments, the target nucleotide sequence is in the genome of a plant. In some embodiments, the target nucleotide sequence is in the genome of a microorganism, such as a bacterium.
[0240] In addition, the present disclosure provides methods for generating TadCBE described herein, and methods for using base editors or nucleic acid molecules encoding any of these base editors in applications including editing nucleic acid molecules (e.g., genomes). In certain embodiments, the method for engineering the base editor provided herein relates to a phage-assisted continuous evolution (PACE) system or a non-continuous system (e.g., PANCE), which can be used to evolve one or more components of the base editor (e.g., deaminase domain). In certain embodiments, after successfully evolving one or more components of the base editor (e.g., deaminase domain), the method for manufacturing the base editor includes recombinant protein expression methods and techniques known to those skilled in the art. An exemplary base editor is manufactured by fusing or associating an adenosine deaminase domain with any of the various napDNAbp domains (e.g., Cas9 domains) disclosed herein.
[0241] Without wishing to be bound by any particular theory, the TadCBE described herein induces editing in a nucleic acid substrate by deaminating C bases using TadA variants, causing C to T mutations via the formation of uracil. It is believed that fusion of one or more uracil DNA glycosylase inhibitors to the deaminase and napDNAbp domains of the CBE inhibits innate DNA repair processes, resulting in conversion of the original C·G base pair to a T·A base pair when coupled with a nucleic acid programmable DNA binding protein (e.g., dCas9) engineered to nick the unedited DNA strand (e.g., the strand containing the G of the original CG target base pair). Without wishing to be bound by any particular theory, it is believed that the mutations of residues 26-28 in the disclosed deaminase (relative to the TadA8e deaminase) promote "slippage" of the backbone of the DNA substrate, enabling the binding pocket of the adenosine deaminase to accept cytosine.
[0242] In some embodiments, the TadCBE described herein has been engineered to exhibit highly targeted and efficient editing capabilities. Such TadCBEs can be used, for example, to target and reverse single nucleotide polymorphisms (SNPs) in disease-related genes, such as genes associated with sickle cell disease and HIV / AIDS. However, in some cases, the TadCBE described herein can allow the replacement of target C to a mixture of T, A, and G. For example, TadCBEs lacking a UGI domain can be used as a screening platform for targeted random in vivo mutagenesis. More specifically, they can be used as forward genetic tools to screen for gain-of-function and / or loss-of-function variants at base resolution.
[0243] Deaminase domain
[0244] The present disclosure provides a cytidine base editor (TadCBE) evolved from the adenosine deaminase domain of an existing adenosine base editor (ABE). The adenosine deaminase used herein is evolved using standard methods to convert adenosine (A) to inosine (I) in mammalian DNA. Such adenosine deaminase can cause A: T to G: C base pair conversion. The most advanced ABE is ABE7.10, which is disclosed in International Publication No. WO 2018 / 027078 (published on August 2, 2018). The most recently generated ABE is ABE8e, which contains an adenosine deaminase domain comprising a single deaminase variant called TadA8e, as described in International Publication No. WO2021 / 158921 (published on August 12, 2021). TadA8e contains nine mutations relative to TadA7.10 (adenosine deaminase of ABE7.10). TadA7.10 is also the deaminase domain of ABEmax, a variant of ABE7.10 that has been codon-optimized for expression in human cells.
[0245] In some embodiments, the adenosine deaminase is a variant of the known adenosine deaminase TadA7.10, which comprises the following mutations compared to wild-type ecTadA (SEQ ID NO: 325): W23R, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, R152P, E155V, I156F and K157N. In some embodiments, the disclosed adenosine deaminase is a variant of TadA derived from species other than Escherichia coli, such as Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis.
[0246] The substrate for the evolution experiments disclosed herein was TadA-8e, which contains the following mutations relative to TadA7.10: A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and D167N. For disclosure of phage-assisted evolution experimental methods, please refer to International Publication No. WO 2018 / 027078; International Publication No. WO 2019 / 079347 (published on April 25, 2019); International Publication No. WO 2019 / 226593 (published on November 28, 2019); U.S. Patent Publication No. 2018 / 0073012 (published on March 15, 2018), which was issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693 (published on May 4, 2017), which was issued as U.S. Patent No. 10,167,457 on January 1, 2019; International Publication No. WO 2020 / 214842 (published on October 22, 2020), and International Patent Application No. PCT / US2020 / 033873 (filed on May 20, 2020), International Publication No. WO 2020 / 236982 (published on November 26, 2020), and International Publication No. WO 2021 / 158921, the contents of each of which are incorporated herein by reference in their entirety.
[0247] Provided herein are exemplary, non-limiting embodiments for evolved adenosine deaminases. In some embodiments, the adenosine deaminase domain of any disclosed base editor comprises a single adenosine deaminase or a monomer. In some embodiments, the adenosine deaminase domain comprises 2, 3, 4, or 5 adenosine deaminases. In some embodiments, the adenosine deaminase domain comprises two adenosine deaminases or dimers. In some embodiments, the deaminase domain comprises an engineered (or evolved) deaminase and a dimer of a wild-type deaminase (e.g., a wild-type Escherichia coli derived deaminase). It should be understood that the mutations provided herein (e.g., mutations in ecTadA) can be applied to adenosine deaminases in other adenosine base editors, such as those provided in: International Publication No. WO 2018 / 027078 (published on August 2, 2018); International Publication No. WO 2019 / 079347 (published on April 25, 2019); International Application No. PCT / US2019 / 033848 (filed on May 23, 2019), which was published as International Publication No. WO on November 28, 2019. 2019 / 226593; U.S. Patent Publication No. 2018 / 0073012 (published on March 15, 2018), which issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693 (published on May 4, 2017), which issued as U.S. Patent No. 10,167,457 on January 1, 2019; International Publication No. WO 2017 / 070633 (published on April 27, 2017); U.S. Patent Publication No. 2015 / 0166980 (published on June 18, 2015); U.S. Patent No. 9,840,699 (issued on December 12, 2017); U.S. Patent No. 10,077,453 (issued on September 18, 2018) and International Patent Application No. PCT / US2020 / 28568 (filed on April 16, 2020); all of which are incorporated herein by reference in their entirety.
[0248] Exemplary adenosine deaminase substrates that can be evolved into cytidine deaminases according to the present disclosure are disclosed below. Exemplary TadA deaminases derived from Bacillus subtilis (full sequence as shown in SEQ ID NO:318), Staphylococcus aureus (SEQ ID NO:317) and Streptococcus pyogenes (SEQ ID NO:354) are provided. Amino acid substitutions in Escherichia coli TadA-8e, as well as homologous mutations in Bacillus subtilis, Staphylococcus aureus and Streptococcus pyogenes TadA deaminases are shown. Therefore, one skilled in the art will be able to generate mutations corresponding to any mutation described herein (e.g., any mutation identified in ecTadA) in any naturally occurring adenosine deaminase (e.g., having homology to ecTadA). In some embodiments, the adenosine deaminase is derived from a prokaryotic organism. In some embodiments, the adenosine deaminase is from bacteria. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from Escherichia coli. Those skilled in the art will be able to identify the corresponding residues in any homologous proteins and in the corresponding encoding nucleic acids by methods well known in the art (e.g., by sequence alignment and determination of homologous residues).
[0249] In some embodiments, the adenosine deaminase substrate comprises TadA9 or a variant thereof. TadA9 contains V82S and Q154R substitutions relative to TadA-8e. (In other words, TadA9 contains Y147R, Q154R, and I76Y mutations relative to TadA7.10.) In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of TadA9 (SEQ ID NO: 33). TadA9 may be referred to in the literature as TadA*8.9. ABEs containing TadA9 deaminase are referred to herein as ABE9. TadA9 is described in more detail in Gaudelli et al., Nat Biotechnol. 2020 Jul; 38(7): 892-900 and PCT Publication No. WO 2021 / 050571 (published on March 18, 2021), each of which is incorporated herein by reference.
[0250] In some embodiments, the adenosine deaminase substrate comprises TadA20, TadA-8.17-m (TadA17), or a variant thereof. TadA20 contains I76Y, V82S, Y123H, Y147R, and Q154R substitutions relative to TadA7.10. TadA17 contains V82S and Q154R substitutions relative to TadA7.10. TadA20 and TadA17 are described in more detail in Gaudelli et al., Nat Biotechnol. 2020 Jul; 38(7): 892-900 and WO 2021 / 050571 (published on March 18, 2021). TadA20 may be referred to as TadA*8.20 in the literature. In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of TadA20 (SEQ ID NO: 326). An ABE containing a TadA20 deaminase is referred to herein as ABE20. It may be referred to in the art as ABE8.20, ABE8.20-d, or ABE8.20-m. An ABE containing a TadA17 deaminase is referred to herein as ABE17. It may be referred to in the art as ABE8.17 or ABE8.17-m.
[0251] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences in SEQ ID NOs: 317-323.
[0252] In certain embodiments, the adenosine deaminase domain comprises an adenosine deaminase having a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% sequence identity to one of:
[0253] TadA 7.10 (E. coli):
[0254] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD(SEQ ID NO:315)
[0255] TadA-8e (Escherichia coli):
[0256] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO:350)
[0257] Tad1:
[0258] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDF (SEQ ID NO:2)
[0259] Tad2:
[0260] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO:2)
[0261] Tad3:
[0262] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYGLIDATLYVTFEPCVMCAGAIIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO:3)
[0263] Tad4:
[0264] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO:4)
[0265] Tad6:
[0266] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYGLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO:5)
[0267] Tad6-SR:
[0268] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYGLIDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRRVFNAQKKAQSSIN(SEQ ID NO:6)
[0269] TadA9:
[0270] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILANECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO:33)
[0271] TadA20
[0272] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLYDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQKKAQSSTD(SEQ ID NO:326)
[0273] Bacillus subtilis TadA:
[0274] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN(SEQ ID NO:317)
[0275] Staphylococcus aureus TadA:
[0276] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE(SEQ ID NO:318)
[0277] Salmonella typhimurium TadA:
[0278] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV(SEQ ID NO:319)
[0279] Shewanella putrefaciens TadA:
[0280] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE(SEQ ID NO:320)
[0281] Haemophilus influenzae F3031 (H. influenzae) TadA:
[0282] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK(SEQ ID NO:321)
[0283] Caulobacter crescentus (C. crescentus) TadA:
[0284] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI(SEQ ID NO:322)
[0285] Geobacter sulfurreducens (G. sulfurreducens) TadA:
[0286] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP(SEQ ID NO:323)
[0287] Streptococcus pyogenes TadA:
[0288] MPYSLEEQTYFMQEALKEAEKSLQKAEIPIGCVIVKDGEIIGRGHNAREESNQAIMHAEIMAINEANAHEGNWRLLDTTLFVTIEPCVMCSGAIGLARIPHVIYGASNQKFGGADSLYQILTDERLNHRVQVERGLLAADCANIMQTFFRQGRERKKIAKHLIKEQSDPFD(SEQ ID NO:354)
[0289] A.aeolicus TadA:
[0290] MGKEYFLKVALREAKRAFEKGEVPVGAIIVKEGEIISKAHNSVEELKDPTAHAEMLAIKEACRRLNTKYLEGCELYVTLEPCIMCSYALVLSRIEKVIFSALDKKHGGVVSVFNILDEPTLNHRVKWEYYPLEEASELLSEFFKKLRNNII(SEQ ID NO:355)
[0291] In some embodiments, the TadA deaminase is full-length Escherichia coli TadA deaminase (ecTadA). For example, in certain embodiments, the adenosine deaminase domain comprises a deaminase comprising the following amino acid sequence:
[0292] MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD(SEQ ID NO:325)
[0293] TadA-derived cytidine deaminase (TadA-CD)
[0294] Certain aspects of the present disclosure relate to evolved adenosine deaminases with enhanced cytosine specificity and cytidine deamination activity. According to certain embodiments, the evolved deaminases are capable of deaminating cytidine in DNA. In some embodiments, the deaminases are evolved from parent adenosine deaminases using continuous and / or non-continuous laboratory-directed methods (e.g., PACE and PANCE). In some embodiments, the parent adenosine deaminases evolved using PACE and / or PANCE have cytidine deaminase activity. The deaminases of the present disclosure can be evolved from any adenosine deaminase reported to date to have adenosine deaminase activity, such as, for example, those described in International Patent Application No. PCT / US2017 / 045381 (filed on August 3, 2017), International Patent Application No. PCT / US2020 / 028568 (filed on April 16, 2020), International Patent Application No. PCT / US2021 / 016827 (filed on February 5, 2021), International Patent Application No. PCT / US2022 / 073781 (filed on July 15, 2022), all of which are incorporated herein by reference in their entirety. In some cases, the parent deaminase comprises Escherichia coli tRNA adenosine deaminase (TadA). The deaminases of the present application can be evolved from previously mutated (i.e., evolved) parent TadA variants, such as those described in, for example, International Patent Application No. PCT / US2021 / 016827 (filed on February 5, 2021, published as WO 2021 / 158921 on August 12, 2021). For example, in some embodiments, the parent adenosine deaminase is TadA7.10. In other embodiments, the parent adenosine deaminase is a TadA8e variant, which contains an additional 8 mutations relative to TadA7.10: A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and D167N. Other parent adenosine deaminase substrates are also possible.
[0295] In some embodiments, the TadA-derived cytidine deaminase of the present application is derived from a parent adenosine deaminase (e.g., TadA-8e) using a combination of phage-assisted continuous evolution (PACE) and discontinuous evolution (PANCE). According to certain embodiments, the parent adenosine deaminase comprises an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, at least 95%, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO: 41. In some cases, the parent adenosine deaminase comprises the sequence of SEQ ID NO: 41.
[0296] In some embodiments, the evolved TadA-derived cytidine deaminase is at least partially homologous to the parent TadA-8e variant. For example, according to certain embodiments, the TadA-derived cytidine deaminase (e.g., TadA-CD) comprises an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO: 41, wherein residue 27 of SEQ ID NO: 41 is any amino acid except E (glutamic acid). TadA-CDs with other sequence homology are also possible. For example, in certain embodiments, the TadA-derived cytidine deaminase (e.g., TadA-CD) comprises an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO: 41, wherein residue 28 of SEQ ID NO: 41 is any amino acid except V (valine). In another exemplary embodiment, the TadA-derived cytidine deaminase is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO: 41, wherein residue 96 of SEQ ID NO: 41 is any amino acid except H (histidine).
[0297] Those skilled in the art will appreciate that TadA derived cytidine deaminase (e.g., TadA-CD) can comprise multiple mutations relative to the parent adenosine deaminase (e.g., TadA-8e). In some embodiments, the deaminase (e.g., TadA-CD) of the present application comprises mutations at residues E27, V28, and H96. In some embodiments, the disclosed deaminase further comprises at least one mutation at a residue selected from R26, M61, Y73, I76, M151, Q154, and A158 in the amino acid sequence of SEQ ID NO: 41, or a corresponding mutation in a homologous adenosine deaminase.
[0298] In some embodiments, the deaminase comprises at least one mutation selected from E27A, E27K, V28G, V28A and H96N in the amino acid sequence of SEQ ID NO: 41, and further comprises at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase. Other mutations are also possible. For example, in certain embodiments, the TadA-CD enzyme comprises at least one mutation selected from E27A, V28G and H96N in the amino acid sequence of SEQ ID NO: 41, and further comprises at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase.
[0299] Other exemplary embodiments may include: (1) a deaminase comprising E27K, V28G and H96N mutations in the amino acid sequence of SEQ ID NO:41, and further comprising at least one mutation selected from the group consisting of R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase; (2) a deaminase comprising E27A, V28A and H96N mutations in the amino acid sequence of SEQ ID NO:41, and further comprising at least one mutation selected from the group consisting of R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase; (3) a deaminase comprising E27A, V28A and H96N mutations in the amino acid sequence of SEQ ID NO:41, and further comprising at least one mutation selected from the group consisting of R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase; The amino acid sequence of NO:41 comprises E27K, V28A and H96N mutations, and further comprises at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or corresponding mutations in homologous adenosine deaminase.
[0300] In some embodiments, the TadA-derived cytidine deaminase (TadA-CD) comprises at least two mutations (relative to the parent deaminase) at residues selected from R26, M61, Y73, I76, M151, Q154, and A158. In other embodiments, the TadA-CD comprises at least two mutations at residues selected from R26G, M61I, Y73H, I76F, M151I, Q154H, Q154R, and A158S.
[0301] In certain aspects, a TadA-derived cytidine deaminase is provided that can retain some A to G base editing activity. Without wishing to be bound by any particular theory, it has been determined via a back analysis that residues 26-28 of TadA-8e deaminase (as shown in SEQ ID NO: 41), which are located on a loop near the active site, are crucial for the selectivity of the conversion of adenosine to cytidine. It is further believed that the positioning of the substrate in the active site is a key determinant of the deamination selectivity, and the sequence context may affect the selective deamination of the target base, because the interaction between TadA-CD and 5' and 3' may affect the positioning of the substrate in the active site.
[0302] Again, without wishing to be bound by theory, it is believed that residual A to T editing is highest when the adenine is located in the center of the editing window (e.g., protospacer position 5 or 6 of SpCas9 with the PAM at position 21-23) and is preceded by a T or C. In some embodiments, adding the V106W mutation improves selectivity by suppressing A deamination to a greater extent than C deamination.
[0303] In some aspects, TadA-derived cytidine deaminases are provided that provide efficient conversion of target cytosine to thymine and target adenine to guanine (referred to herein as "TadA-Dual" deaminases and base editors). TadA-Dual deaminases are capable of editing C and A bases within the original spacer, and particularly within the editing window of the original spacer. These editors install both A to G and C to T editing with roughly comparable efficiency.
[0304] For example, the disclosed TadA dual deaminase installs A to G and C to T editing at a ratio of about 1.1: 1. In some embodiments, the dual editor provides A to G and C to T editing at a ratio of 0.7: 1, 0.8: 1, 0.9: 1, 1: 1, 1.1: 1, 1.2: 1, 1.3: 1, 1.4: 1, or 1.5: 1. Other ranges are also possible, including ratios exceeding 1.5: 1. These evolved TadA deaminases and "dual" editors containing these deaminases can edit A·T to G·C with almost the same efficiency as C·G to T·A, and may be suitable for screening applications, such as methods for screening novel Cas homology domains and other napDNAbp domains for editing activity against various target sequences. These deaminases and dual editors can also be used for mutagenesis applications, such as in vivo forward genetic mutagenesis screening or targeted random mutagenesis screening. These dual editors can also be used for multiplexed editing applications.
[0305] Dual Editor
[0306] In some embodiments, the TadA-based dual editor comprises a cytidine deaminase comprising one, two, three, four or five mutations selected from R26G, V28A, A48R, Y73S and H96N. The dual editor is referred to herein as TadDE, and the dual editing deaminase is referred to herein as TadA-CDf (e.g., TadA-Dual), which has an amino acid sequence as shown in SEQ ID NO: 39.
[0307] Thus, in some embodiments, provided herein are deaminases comprising mutations at residues R26, V28, A48, and Y73 in the amino acid sequence of SEQ ID NO: 41, or corresponding mutations in homologous adenosine deaminases. Also provided herein are deaminases comprising mutations at residues R26, E27, V28, A48, and Y73 in the amino acid sequence of SEQ ID NO: 41 (i.e., further comprising E27 mutations). In specific embodiments, these deaminases comprise mutations R26G, V28A, A48R, Y73S, and H96N. In some embodiments, these deaminases comprise mutations R26G, V28G, A48R, and Y73C.
[0308] As described above and herein, preferred TadA-derived cytidine deaminases evolved using the PACE and PANCE methods can comprise one or more mutations. For example, a TadA-CD variant can comprise at least one mutation selected from the group consisting of R26G, E27A, V28G, I76F, H96N, and M151I (e.g., TadA-CDa, SEQ ID NO: 34); R26G, E27A, V28G, I76F, H96N, and A158S (e.g., TadA-CDb, SEQ ID NO: 35); R26G, E27A, V28G, I76F, H96N, Q154R, and A158S (e.g., TadA-CDc, SEQ ID NO: 36); E27A, V28G, Y73H, H96N, Q154H, and A158S (e.g., TadA-CDd, SEQ ID NO: 37); NO:37); R26G, V28A, A48R, Y73S and H96N (e.g., TadA-CDe, SEQ ID NO:38); V28A, A48R and Y73S (e.g., TadA-CDf, SEQ ID NO:39), and R26G, V28G, A48R and Y73C (e.g., TadA-CDg, SEQ ID NO:40).
[0309] In some preferred embodiments, the deaminase comprises the following mutations: R26G, E27A, V28G, I76F, H96N and A158S (e.g., TadA-CDa, SEQ ID NO: 34), R26G, E27A, V28G, I76F, H96N, Q154R and A158S (e.g., TadA-CDb, SEQ ID NO: 35), R26G, E27A, V28G, I76F, H96N and M151I (e.g., TadA-CDc, SEQ ID NO: 36), E27K, V28A, M61I and H96N (e.g., TadA-CDd, SEQ ID NO: 37), E27A, V28G, Y73H, H96N, Q154H and A158S (e.g., TadA-CDe, SEQ ID NO: NO:38), R26G, V28A, A48R, Y73S and H96N (e.g., TadA-CDf, SEQ ID NO:39), and R26G, V28G, A48R and Y73C (e.g., TadA-CDg, SEQ ID NO:40).
[0310] Those of ordinary skill in the art will appreciate that the evolved deaminases described herein may exhibit different specificities and / or deamination activities for cytosine and / or adenosine bases due to different inherited mutation types and combinations. In some embodiments, the cytidine deamination activity of TadA-CD exceeds the cytidine deamination activity of TadA-8e. For example, the cytidine deamination activity of the TadA-CD variant may be greater than or equal to 10 times, greater than or equal to 20 times, greater than or equal to 40 times, greater than or equal to 80 times, greater than or equal to 100 times, greater than or equal to 200 times, greater than or equal to 400 times, greater than or equal to 800 times, greater than or equal to 1000 times, greater than or equal to 2000 times, greater than or equal to 3000 times, greater than or equal to 4000 times the cytidine deamination activity of TadA-8e. In other embodiments, the cytidine deamination activity of the TadA-CD variant is less than or equal to 4000 times, less than or equal to 2000 times, less than or equal to 1000 times, less than or equal to 800 times, less than or equal to 400 times, less than or equal to 200 times, less than or equal to 100 times, less than or equal to 80 times, less than or equal to 40 times, less than or equal to 20 times, or less than or equal to 10 times the cytidine deamination activity of TadA-8e.
[0311] In some embodiments, the adenosine deamination activity of the TadA-CD deaminase is lower than the deamination activity of TadA-8e. For example, in some cases, the adenosine deamination activity of the TadA-CD variant is less than or equal to 4000 times, less than or equal to 2000 times, less than or equal to 1000 times, less than or equal to 800 times, less than or equal to 400 times, less than or equal to 200 times, less than or equal to 100 times, less than or equal to 80 times, less than or equal to 40 times, less than or equal to 20 times, or less than or equal to 10 times the adenosine deamination activity of TadA-8e.
[0312] In some embodiments, the TadA-CD variants described above and herein may also include a V106W mutation. It has recently been discovered that adenosine deaminase TadA variants comprising a V106W mutation (such as those described in International Patent Publication Nos. WO 2021 / 214842 and WO2021 / 158921, each of which is incorporated herein by reference) have reduced Cas-independent off-target editing of DNA and RNA while maintaining high levels of on-target adenosine deamination activity. In some embodiments, the average peak editing efficiency of the TadA-CD variants comprising the V106W mutation is greater than or equal to 50%, greater than or equal to 60%, greater than or equal to 70%, greater than or equal to 80%, and greater than or equal to 90%. In other embodiments, the average peak editing efficiency of the TadA-CD variants comprising the V106W mutation is less than or equal to 90%, less than or equal to 80%, less than or equal to 70%, less than or equal to 60%, or less than or equal to 50%. ABEs containing only a single TadA deaminase domain (rather than a single-chain dimer) allow for a reduction in the size of the editor 30,31 . In addition, although SaCas9 is small enough (1053 amino acids in length, SEQ ID NO: 347) to provide a single AAV-compatible base editor, the rarity of its NNGRRT PAM greatly limits its utility. Since base editing requires the presence of a suitable PAM to place the target nucleotide within the editing window, TadCBE, which provides broad PAM compatibility and simple and efficient in vivo delivery as a whole, will promote the in vivo application of base editing.
[0313] In some embodiments, any one of the deaminases listed in Table 10 may further comprise a V106W mutation. In some embodiments, the TadA-CD variant comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% of any of the amino acid sequences listed in Table 10, wherein any one of the sequences listed in Table 10 further comprises a V106W mutation.
[0314] In some embodiments, the TadA variant comprises at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identity to any of the amino acid sequences listed in Table 10.
[0315] Table 10.
[0316]
[0317]
[0318] In some embodiments, the dual editor deaminase (e.g., TadA-CDf or TadA-Dual, SEQ ID NO: 39) of the TadDE dual editor can be further evolved, for example, using the PACE and / or PANCE assays further described below or elsewhere herein. In some embodiments, the TadA-Dual deaminase (e.g., TadA-CDf, SEQ ID NO: 39) is further evolved to enhance specificity for cytosine bases and reduce specificity for adenosine bases. For example, Fig.51E Shown is a table listing the evolved TadA-Dual deaminases (eg, TadDE-1 to TadDE-5) and their mutations relative to the unmutated TadA-Dual deaminases and their parental TadA-8e deaminase.
[0319] In some embodiments, such as Fig.51C As shown, TadA-Dual deaminase was mutated using PACE. In some embodiments, phage-assisted continuous evolution or PACE ( Fig.51C , left) and the selection loop ( Fig.51C , right) in combination. In some embodiments, the selected phage encoding the partial deaminase (SP) infects the continuous flow of E. coli host cells. Those skilled in the art will understand that the E. coli host cell must also contain a plasmid defining the selection loop and a mutagenic plasmid. In the selection loop, phage propagation is associated with the expression of gIII (P2), and gIII can only be transcribed using an active T7 RNA polymerase. In some embodiments, T7 RNA polymerase (P3) is fused to a C-terminal degradation determinant, and the deaminase must be C to U edited to install a stop codon before the degradation determinant, thereby producing an active T7 RNA polymerase. In the case of phage infection, a complete deaminase is completed using a split intein system (P1), and mutations can occur on the deaminase. Beneficial mutations cause phages to propagate and enrich in the culture tank, while phages with poor adaptability cannot propagate and are subsequently washed out by a constant outflow.
[0320] In some embodiments, such as Fig.51DAs shown, TadA-Dual deaminase is mutated using phage-assisted non-continuous evolution (PANCE). In some embodiments, PANCE is performed on TadA-Dual deaminase (SEQ ID NO: 39) until phage titer increases despite higher stringency of dilution factor and promoter strength, indicating that a beneficial mutation has occurred. In some embodiments, the beneficial mutation comprises a mutation at position N46 in the deaminase.
[0321] In some embodiments, PANCE is performed on the NNK library at position N46 to further identify beneficial mutations. In some embodiments, a combination of mutagenesis assays can be used. For example, in some embodiments, PACE is performed for more than 100 hours on the resulting variants from the PANCE study.
[0322] In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46I, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-1, Fig.51E , PANCE, SEQ ID NO:42). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46T, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-2, Fig.51E , PANCE, SEQ ID NO:43). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46T, A48R, Y73S, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-3, Fig.51E, PANCE, SEQ ID NO:44). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73S, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-4, PANCE of NNK library at N46, SEQ ID NO:45). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-5, PANCE of NNK library at N46, SEQ ID NO:46). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46L, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-6, PANCE of NNK library at N46, SEQ ID NO: 47). In some embodiments, the evolved TadA-Dual deaminase comprises V28A, N46L, A48P, and Y73P mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-7, PANCE of NNK library at N46, SEQ ID NO: 48). In some embodiments, the evolved TadA-Dual deaminase comprises V28A, N46C, A48P, and Y73P mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-8, PANCE of NNK library at N46, SEQ ID NO: 49). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-9, Fig.51E , PACE, SEQ ID NO:50). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Q71H, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-10, Fig.51E , PACE, SEQ ID NO:51). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46L, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-11, Fig.51E, PACE, SEQ ID NO:52). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46C, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-12, Fig.51E , PACE, SEQ ID NO:53). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46C, A48R, Y73P, H96N, and A162V mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-13, Fig.51E , PACE, SEQ ID NO:54).
[0323] In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46I, A48R, Y73S, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-14, Fig.51E , PANCE, SEQ ID NO:359). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, A48R, Q71S, Y73S, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-15, Fig.51E , PANCE, SEQ ID NO:360). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46L, A48R, and Y73P mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-16, Fig.51E , PANCE, SEQ ID NO:361). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46L, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-17, Fig.51E , PANCE, SEQ ID NO:362). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-18, Fig.51E, PANCE, SEQ ID NO:363). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73S, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-19, Fig.51E , PANCE, SEQ ID NO:364). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-20, Fig.51E , PANCE, SEQ ID NO: 365). In some embodiments, the evolved TadA-Dual deaminase comprises R26G and N46L mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-21, Fig.51E , PANCE, SEQ ID NO: 366). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46I, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-22, Fig.51E , PANCE, SEQ ID NO:367). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-23, Fig.51E , PANCE, SEQ ID NO:368). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, A48P, Y73H, T79P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-24, Fig.51E , PANCE, SEQ ID NO:369). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, N46I, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-25, Fig.51E , PANCE, SEQ ID NO:370).
[0324] In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-26, Fig.51E, PANCE, SEQ ID NO:371). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46L, A48R, Y73S, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-27, Fig.51E , PANCE, SEQ ID NO: 372). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46C, A48R, H96N, and A162V mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-28, Fig.51E , PANCE, SEQ ID NO:373). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Q71H, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-29, Fig.51E , PANCE, SEQ ID NO: 374). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46C, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-30, Fig.51E , PANCE, SEQ ID NO:375). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46C, A48R, Y73P, H96N, and A162V mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-31, Fig.51E , PANCE, SEQ ID NO:376).
[0325] In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-32, Fig.51E , PANCE, SEQ ID NO: 377). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48R, Y73S, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-33, Fig.51E, PANCE, SEQ ID NO:378). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46V, A48P, Y73S, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-34, Fig.51E , PANCE, SEQ ID NO:379). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46C, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-35, Fig.51E , PANCE, SEQ ID NO: 380). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, L34M, N46L, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-36, Fig.51E , PANCE, SEQ ID NO:381). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46L, A48R, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO:41 (TadA-CD-37, Fig.51E , PANCE, SEQ ID NO: 382). In some embodiments, the evolved TadA-Dual deaminase comprises R26G, V28A, N46L, A48P, R64K, Y73P, and H96N mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-38, Fig.51E , PANCE, SEQ ID NO:383).
[0326] In some embodiments, the evolved TadA-Dual deaminase comprises N46I, S73P, and H154Q mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-1, Fig.51E , PANCE, SEQ ID NO:42). In some embodiments, the evolved TadA-Dual deaminase comprises an N46T mutation relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-2, Fig.51E , PANCE, SEQ ID NO:43). In some embodiments, the evolved TadA-Dual deaminase comprises N46T and H154Q mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-3, Fig.51E, PANCE, SEQ ID NO:44). In some embodiments, the evolved TadA-Dual deaminase comprises N46V and H154Q mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-4, PANCE performed on the NNK library at N46, SEQ ID NO:45). In some embodiments, the evolved TadA-Dual deaminase comprises N46V, S73P, G105S, and H154Q mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-5, PANCE performed on the NNK library at N46, SEQ ID NO:46). In some embodiments, the evolved TadA-Dual deaminase comprises N46L, S73P, and H154Q mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-6, PANCE performed on the NNK library at N46, SEQ ID NO:47). In some embodiments, the evolved TadA-Dual deaminase comprises G26R, N46L, R48P, S73P, N96H, and H154Q mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-7, PANCE of NNK library at N46, SEQ ID NO: 48). In some embodiments, the evolved TadA-Dual deaminase comprises N46C, N96H, and H154Q mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-8, PANCE of NNK library at N46, SEQ ID NO: 49). In some embodiments, the evolved TadA-Dual deaminase comprises N46V, S73P, and H154Q mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-9, Fig.51E , PACE, SEQ ID NO: 50). In some embodiments, the evolved TadA-Dual deaminase comprises N46V, Q71H, S73P, and H154Q mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-10, Fig.51E , PACE, SEQ ID NO:51). In some embodiments, the evolved TadA-Dual deaminase comprises N46L and H154Q mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-11, Fig.51E , PACE, SEQ ID NO:52). In some embodiments, the evolved TadA-Dual deaminase comprises N46C, S73P, and H154Q mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-12, Fig.51E, PACE, SEQ ID NO:53). In some embodiments, the evolved TadA-Dual deaminase comprises N46C, S73P, H154Q, and A162V mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-13, Fig.51E , PACE, SEQ ID NO:54).
[0327] In some embodiments, the evolved TadA-Dual deaminase comprises N46I and H154Q mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-14, Fig.51E , PACE, SEQ ID NO: 359). In some embodiments, the evolved TadA-Dual deaminase comprises Q71S and H154Q mutations relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-15, Fig.51E , PANCE, SEQ ID NO:360). In some embodiments, the evolved TadA-Dual deaminase comprises N46L, S73P, N79T, and N96H mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-16, Fig.51E , PANCE, SEQ ID NO:361). In some embodiments, the evolved TadA-Dual deaminase comprises N46L, S73P, and N79T mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-17, Fig.51E , PANCE, SEQ ID NO:362). In some embodiments, the evolved TadA-Dual deaminase comprises R48A, S73P, and N79T mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-18, Fig.51E , PANCE, SEQ ID NO: 363). In some embodiments, the evolved TadA-Dual deaminase comprises N46V and N79T mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-19, Fig.51E , PANCE, SEQ ID NO:364). In some embodiments, the evolved TadA-Dual deaminase comprises N46V, S73P, and N79T mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-20, Fig.51E, PANCE, SEQ ID NO:365). In some embodiments, the evolved TadA-Dual deaminase comprises A28V, N46L, R48A, S73Y, N79T, and N96H mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-21, Fig.51E , PANCE, SEQ ID NO:366). In some embodiments, the evolved TadA-Dual deaminase comprises N46I, S73P, and N79T mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-22, Fig.51E , PANCE, SEQ ID NO:367). In some embodiments, the evolved TadA-Dual deaminase comprises N46V, S73P, N79T, and G106S mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-23, Fig.51E , PANCE, SEQ ID NO: 368). In some embodiments, the evolved TadA-Dual deaminase comprises R48P, S73H, and N79P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-24, Fig.51E , PANCE, SEQ ID NO:369). In some embodiments, the evolved TadA-Dual deaminase comprises A28V, N46I, R48A, S73Y, and N79T mutations relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-25, Fig.51E , PANCE, SEQ ID NO:370).
[0328] In some embodiments, the evolved TadA-Dual deaminase comprises N46V and S73P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-26, Fig.51E , PANCE, SEQ ID NO:371). In some embodiments, the evolved TadA-Dual deaminase comprises an N46L mutation relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-27, Fig.51E , PANCE, SEQ ID NO: 372). In some embodiments, the evolved TadA-Dual deaminase comprises N46C, S73Y, and A162V mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-28, Fig.51E, PANCE, SEQ ID NO: 373). In some embodiments, the evolved TadA-Dual deaminase comprises N46V, Q71H, and S73P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-29, Fig.51E , PANCE, SEQ ID NO: 374). In some embodiments, the evolved TadA-Dual deaminase comprises N46C and S73P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-30, Fig.51E , PANCE, SEQ ID NO: 375). In some embodiments, the evolved TadA-Dual deaminase comprises N46C, S73P, and A162V mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-31, Fig.51E , PANCE, SEQ ID NO: 376). In some embodiments, the evolved TadA-Dual deaminase comprises N46V and S73P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-32, Fig.51E , PANCE, SEQ ID NO: 377). In some embodiments, the evolved TadA-Dual deaminase comprises an N46V mutation relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-33, Fig.51E , PANCE, SEQ ID NO: 378). In some embodiments, the evolved TadA-Dual deaminase comprises N46V and R48P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-34, Fig.51E , PANCE, SEQ ID NO: 379). In some embodiments, the evolved TadA-Dual deaminase comprises N46CV and S73P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-35, Fig.51E , PANCE, SEQ ID NO: 380). In some embodiments, the evolved TadA-Dual deaminase comprises L34M, N46L, and S73P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-36, Fig.51E , PANCE, SEQ ID NO: 381). In some embodiments, the evolved TadA-Dual deaminase comprises N46L and S73P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-37, Fig.51E, PANCE, SEQ ID NO: 382). In some embodiments, the evolved TadA-Dual deaminase comprises N46L, r48P, R64K, and S73P mutations relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-38, Fig.51E , PANCE, SEQ ID NO:383).
[0329] In some embodiments, the TadA-CD deaminase evolved from the TadA-Dual deaminase has improved specificity for cytosine bases. In some embodiments, the evolved TadA-CD deaminase exhibits cytosine target activity similar to other evolved deaminases described herein. In some embodiments, the deaminase evolved from the TadA-Dual deaminase has increased specificity for cytosine bases and reduced specificity for adenosine bases. In some embodiments, the deaminase evolved from the TadA-Dual deaminase exhibits no residual A to G base editing (e.g., TadA-CD-1 to TadA-CD-38).
[0330] In some embodiments, TadA-CD-1 exhibits no residual A to G base editing when incorporated into the BE4max architecture. In some embodiments, TadA-CD-2 exhibits no residual A to G base editing when incorporated into the BE4max architecture. In some embodiments, TadA-CD-3 exhibits no residual A to G base editing when incorporated into the BE4max architecture. In some embodiments, TadA-CD-4 exhibits no residual A to G base editing when incorporated into the BE4max architecture. In some embodiments, TadA-CD-5 exhibits no residual A to G base editing when incorporated into the BE4max architecture. The above description is not intended to be limiting in any way, and the evolved TadA-CD deaminases described herein can be used with any suitable architecture known to those of skill in the art.
[0331] In some embodiments, the TadA-CD evolved from TadA-dual comprises at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to any of the amino acid sequences listed in Table 11.
[0332] In some embodiments, any one of the deaminases listed in Table 11 may further comprise a V106W mutation. In some embodiments, the TadA-CD variant is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences listed in Table 10, wherein any one of the sequences listed in Table 11 further comprises a V106W mutation.
[0333] Table 11. List of exemplary mutant TadA-CD relatives derived from TadA-Dual (SEQ ID NO: 39). The sequences of TadA-8e and TadA-Dual are provided as reference.
[0334]
[0335]
[0336]
[0337]
[0338] napDNAbp domain
[0339] The base editor described herein comprises a nucleic acid programmable DNA binding (napDNAbp) domain. napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which locates napDNAbp to a DNA sequence comprising a DNA chain (i.e., a target chain) complementary to a guide nucleic acid or a portion thereof (e.g., a protospacer of a guide RNA). In other words, the guide nucleic acid "programs" the napDNAbp domain to locate and bind to the complementary sequence of the target chain. The binding of the napDNAbp domain to the complementary sequence enables the base editor's nucleobase modification domain (i.e., adenosine deaminase domain) to approach and enzymatically deaminate the target base in the target chain.
[0340] napDNAbp can be a CRISPR (clustered regularly interspaced short palindromic repeats) related nuclease. As outlined above, CRISPR is an adaptive immune system that provides protection for mobile genetic elements (viruses, transposable elements, and conjugative plasmids). The CRISPR cluster contains a spacer (a sequence complementary to the previous mobile element) and targets invading nucleic acids. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required for correct processing of pre-crRNA. TracrRNA serves as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleases cut linear or circular double-stranded DNA targets complementary to the spacer. The target chain that is not complementary to crRNA is first cut endonucleases and then trimmed by 3'-5' exonucleases. In nature, DNA binding and cutting usually require proteins and two RNAs. However, a single guide RNA ("sgRNA" or simply "gRNA") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, e.g., Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference.
[0341] The description of various napDNAbps that can be used in combination with the disclosed adenosine deaminase is not intended to be limiting in any way. The base editor may include a typical SpCas9, or any orthologous Cas9 protein, or any variant Cas9 protein (including any naturally occurring variant, mutant, or other engineered version of Cas9), which is known or can be made or evolved by directed evolution or other mutagenesis processes. In various embodiments, napDNAbp has nickase activity, i.e., only one chain of the target DNA sequence is cut. In other embodiments, napDNAbp has an inactivated nuclease, such as a "dead" protein. Other variant Cas9 proteins that can be used are those with a smaller molecular weight (e.g., for easier delivery) relative to typical SpCas9 or those with a modified or rearranged primary amino acid sequence (e.g., a circular arrangement). The base editor described herein may also include a Cas9 equivalent, including a Cas12a / Cpf1 protein. The napDNAbp used herein (e.g., SpCas9, SaCas9 or SaCas9 variants or SpCas9 variants) can also contain various modifications that change / enhance its PAM specificity. The present disclosure contemplates any Cas9, Cas9 variants or Cas9 equivalents having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.9% sequence identity with any Cas9 protein disclosed herein. In some embodiments, the napDNAbp domain comprises a nickase variant of wild-type Cas9. In some embodiments, the napDNAbp domain comprises any Cas9 nickase disclosed herein.
[0342] In some embodiments, napDNAbp guides the cutting of one or two chains at the target sequence position, such as in the target sequence and / or in the complementary sequence of the target sequence.In some embodiments, napDNAbp guides the cutting of one or two chains in about 1,2,3,4,5,6,7,8,9,10,15,20,25,50,100,200,500 or more base pairs from the first or last nucleotide of the target sequence.For example, aspartic acid to alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from Streptococcus pyogenes converts Cas9 from the nuclease of cutting two chains to a nickase (cutting single strand).Other mutation examples that make Cas9 a nickase include but are not limited to H840A, N854A and N863A relative to typical SpCas9 sequences, and H588A and D16A relative to Nme2Cas9 sequences, and equivalent amino acid positions in other Cas9 variants or Cas9 equivalents.
[0343] As used herein, the term "Cas protein" refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequence different from that of a naturally occurring Cas protein, or any Cas protein fragment that still retains all or most of the necessary basic functions required for the disclosed method, i.e., (i) having programmable nucleic acid binding of the Cas protein to the target DNA, and (ii) the ability to nick the target DNA sequence on one strand. Cas proteins contemplated herein encompass CRISPR Cas9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease-inactivated Cas9 (dCas9)) homologs, orthologs or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and may include Cas9 equivalents from any type of CRISPR system (e.g., type II, type V, type VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). More Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353 (6299), the contents of which are incorporated herein by reference.
[0344] The term "Cas9" or "Cas9 domain" includes any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or functional fragment, any Cas9 homolog, ortholog or paralog, and any mutant or variant of naturally occurring or engineered Cas9. The term Cas9 is not intended to be particularly limited and may be referred to as "Cas9 or equivalent". Exemplary Cas9 proteins are further described herein and / or described in the art and are incorporated herein by reference. The present disclosure is not limited to the specific napDNAbp used in the base editor of the present disclosure.
[0345] As used herein, the terms "compact Cas9 protein," "compact napDNAbp," and "compact variant [of a Cas protein]" refer to a Cas9 protein or variant having an amino acid length of less than about 1250 amino acids. In some embodiments, the compact Cas9 protein or compact napDNAbp contains less than 1250 amino acids, less than 1240 amino acids, less than 1230 amino acids, less than 1220 amino acids, less than 1210 amino acids, less than 1200 amino acids, less than 1190 amino acids, less than 1180 amino acids, less than 1170 amino acids, less than 1160 amino acids, less than 1150 amino acids, less than 1140 amino acids, less than 1130 amino acids, less than 1120 amino acids, less than 1110 amino acids, less than 1100 amino acids, less than 1050 amino acids, less than 1000 amino acids, less than 950 amino acids, less than 900 amino acids, less than 850 amino acids, less than 800 amino acids, less than 750 amino acids, less than 700 amino acids, less than 650 amino acids, less than 600 amino acids, less than 550 amino acids, or less than 500 amino acids in length. These terms also include any Cas9 protein or variant encoded by a nucleic acid sequence having a length of less than about 3750 nucleotides. The base editor of the present disclosure may include a compact napDNAbp and / or a compact Cas9 protein. In some embodiments, the compact Cas9 protein is about 350 amino acids shorter than SpCas9. In some embodiments, the length of the compact Cas9 protein is about 1000 amino acids. In some embodiments, the compact protein is a compact variant of Streptococcus pyogenes Cas9 (SpCas9), Cpf1, CasX, CasY, C2c1, C2c2, C2c3, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Cas3 or CasΦ. "Compact variant" may refer to a Cas9 protein with one or more truncations or one or more deletions relative to a wild-type Cas9 protein (such as wild-type SpCas9 or Cpf1).
[0346] Additional Cas9 sequences and structures are known to those skilled in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” K., Sharma C. M., Gonzales K., Chao Y., Pirzada Z. A, Eckert M. R, Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A, Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference), and are also provided below.
[0347] Examples of Cas9 and Cas9 equivalents are provided below; however, these specific examples are not intended to be limiting. The base editors of the present disclosure can use any suitable napDNA bp, including any suitable Cas9 or Cas9 equivalents.
[0348] In some embodiments, Cas9 comprises or is derived from wild-type SaCas9 (e.g., Staphylococcus aureus, 1053AA, 123 kDa). In some embodiments, wild-type SaCas9 comprises the following amino acid sequence:
[0349]
[0350] In some embodiments, the sequence of SaCas9 comprises at least at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, at least 99.5%, or at least 99.8% sequence identity to SEQ ID NO:347.
[0351] In some embodiments, Cas9 comprises or is derived from wild-type SpCas9 (e.g., SpCas9, Streptococcus pyogenes M1, SwissProt accession number Q99ZW2, wild-type). In some embodiments, wild-type SaCas9 comprises the following amino acid sequence:
[0352]
[0353] In some embodiments, the sequence of SpCas9 comprises at least at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, at least 99.5%, or at least 99.8% sequence identity to SEQ ID NO:200.
[0354] Cas nickase
[0355] In some embodiments, the disclosed base editors may include a napDNAbp domain containing a Cas nickase. In some embodiments, the base editors described herein include a Cas9 nickase. In some embodiments, any disclosed base editor or vector may include a Streptococcus pyogenes Cas9 nickase (SpCas9n or nCas9) containing a D10A mutation. In some embodiments, any disclosed base editor may include an Nme2Cas9 nickase (Nme2Cas9n) or an eNme2-C Cas9 nickase (eNme2-C Cas9n), each of which contains a D16A mutation.
[0356] The term "Cas9 nickase" or "nCas9" refers to a Cas9 variant that can introduce single-strand breaks in a double-stranded DNA molecular target. In some embodiments, the Cas9 nickase comprises only a single functional nuclease domain. Wild-type Cas9 (e.g., typical SpCas9) comprises two independent nuclease domains, namely a RuvC domain (which cuts a non-protospacer DNA strand) and a HNH domain (which cuts a protospacer DNA strand). In one embodiment, the Cas9 nickase comprises a mutation in the RuvC domain that inactivates the RuvC nuclease activity. For example, mutations at aspartic acid (D) 10, histidine (H) 983, aspartic acid (D) 986, or glutamic acid (E) 762 have been reported as loss-of-function mutations in the RuvC nuclease domain and create functional Cas9 nickases (e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, which is incorporated herein by reference). Thus, nickase mutations in the RuvC domain may include D10X, H983X, D986X, or E762X, wherein X is any amino acid other than the wild-type amino acid. In certain embodiments, the nickase may be D10A, H983A, D986A, or E762A, or a combination thereof.
[0357] In some embodiments, the napDNAbp domain of any disclosed base editor comprises Streptococcus pyogenes Cas9 nickase (SpCas9n). In some embodiments, the napDNAbp domain of any disclosed base editor comprises at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 343. In some embodiments, the napDNAbp domain of any disclosed base editor comprises the amino acid sequence of SEQ ID NO: 343.
[0358] In some embodiments, the napDNAbp domain of any disclosed base editor comprises Staphylococcus aureus Cas9 nickase (SaCas9n). In some embodiments, the napDNAbp domain of any disclosed base editor comprises at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 351. In some embodiments, the napDNAbp domain of any disclosed base editor comprises the amino acid sequence of SEQ ID NO: 351.
[0359] In some embodiments, the napDNAbp domain of any disclosed base editor comprises Neisseria meningitidis Cas9 nickase (Nme2Ca9n) or a variant thereof. In some embodiments, the napDNAbp domain of any disclosed base editor comprises at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 352 or 353. In some embodiments, the napDNAbp domain of any disclosed base editor comprises the amino acid sequence of SEQ ID NO: 352 or 353. In some embodiments, the napDNAbp domain comprises the amino acid sequence of SEQ ID NO: 353. eNme2 -C Cas9 (SEQ ID NO: 353) variants showed activity against NNNNCN (N 4 CN) PAM preference of base editors containing this eNme2-C variant in human cells 4 Base editing efficiencies of approximately 60% or higher were generated on CC PAMs, representing a two-fold improvement relative to base editors containing wild-type Nme2Cas9.
[0360] In some embodiments, the napDNAbp domain of any disclosed base editor comprises a wild-type Nme2Cas9 nuclease (SEQ ID NO: 349).
[0361] In various embodiments, the Cas nickase can have a mutation in the RuvC nuclease domain and have one of the following amino acid sequences, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0362]
[0363]
[0364]
[0365]
[0366]
[0367] Cas9 equivalents
[0368] In some embodiments, the base editor described herein may include any Cas9 equivalent. As used herein, the term "Cas9 equivalent" is a broad term that covers any napDNAbp that performs the same function as Cas9 in the base editor of the present invention, although its primary amino acid sequence and / or its three-dimensional structure may be different and / or may be irrelevant from an evolutionary perspective. Therefore, although Cas9 equivalents include any Cas9 orthologs, homologs, mutants or variants described or covered herein that are evolutionarily related, Cas9 equivalents also include proteins that may have evolved through convergent evolution processes to have the same or similar functions as Cas9, but these proteins do not necessarily have any similarities in amino acid sequence and / or three-dimensional structure. The base editor described herein covers any Cas9 equivalent that can provide the same or similar functions as Cas9, even if the Cas9 equivalent may be based on a protein produced by convergent evolution. For example, if Cas9 refers to a type II enzyme of a CRISPR-Cas system, then a Cas9 equivalent may refer to a type V or type VI enzyme of a CRISPR-Cas system.
[0369] For example, Cas12e (CasX) is a Cas9 equivalent that is reported to have the same function as Cas9 but evolved through convergent evolution. Therefore, Liu et al., "CasX enzymes comprise a distinct family of RNA-guided genome editors," Nature, 2019, Vol. 566: 218-223 describes Cas12e (CasX) proteins that are considered for use with base editors described herein. In addition, any variant or modification of Cas12e (CasX) is also conceivable and within the scope of the present disclosure.
[0370] Cas9 is a bacterial enzyme that has evolved in multiple species. However, the Cas9 equivalents considered herein may also be obtained from archaea, which constitute a domain and kingdom of unicellular prokaryotic microorganisms distinct from bacteria.
[0371] In some embodiments, the Cas9 equivalent may refer to Cas12e (CasX) or Cas12d (CasY), which have been described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21, the entire contents of which are incorporated herein by reference. Using genome-analyzed metagenomics, many CRISPR-Cas systems have been identified, including the first reported Cas9 in the field of archaeal life. This different Cas9 protein was found in less studied nanoarchaea as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-Cas12e and CRISPR-Cas12d, were found, which are one of the most compact systems discovered so far. In some embodiments, Cas9 refers to Cas12e or a variant of Cas12e. In some embodiments, Cas9 refers to Cas12d or a variant of Cas12d. It should be understood that other RNA-guided DNA binding proteins can also be used as nucleic acid programmable DNA binding proteins (napDNAbp) and are within the scope of the present disclosure. See also Liu et al., "CasX enzymes comprise a distinct family of RNA-guided genomeeditors," Nature, 2019, Vol. 566: 218-223. All of these Cas9 equivalents are considered.
[0372] In some embodiments, the Cas9 equivalent comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a naturally occurring Cas12e (CasX) or Cas12d (CasY) protein. In some embodiments, the napDNAbp is a naturally occurring Cas12e (CasX) or Cas12d (CasY) protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a wild-type Cas portion or any Cas portion provided herein.
[0373] In various embodiments, nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), C2C3Cas12e (CasX), Cas12d (CasY), Cas12a (Cpf1), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), Argonaute. An example of a nucleic acid programmable DNA binding protein with a different PAM specificity from Cas9 is the clustered regularly interspaced short palindromic repeats 1 (i.e., Cas12a (Cpf1)) from Prevotella and Francisella. Similar to Cas9, Cas12a (Cpf1) is also a Class 2 CRISPR effector, but it is a member of the Type V enzyme subgroup, not the Type II subgroup. It has been demonstrated that Cas12a (Cpf1) mediates powerful DNA interference with different characteristics from Cas9. Cas12a (Cpf1) is a single RNA-guided endonuclease lacking tracrRNA, and it utilizes a T-rich protospacer adjacent motif (TTN, TTTN or YTN). In addition, Cpf1 cuts DNA via staggered DNA double-strand breaks. Among the 16 Cpf1 family proteins, two enzymes from the genus Acidococcus and Lachnospiraceae show efficient genome editing activity in human cells. Cpf1 protein is known in the art and has been described before, such as Yamano et al., "Crystal structure of Cpf1in complex with guide RNA and target DNA." Cell (165) 2016, p.949-962; the entire contents of which are incorporated herein by reference.
[0374] In other embodiments, the Cas protein may include any CRISPR-associated protein, including but not limited to Cas12a, Cas12b1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, C sn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof, and preferably comprises a nickase mutation (e.g., a mutation corresponding to the D10A mutation of the wild-type Cas9 polypeptide of SEQ ID NO: 200).
[0375] In various other embodiments, the napDNAbp can be any of the following proteins: Cas9, C2c3Cas12a (Cpf1), Cas12e (CasX), Cas12d (CasY), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, a circularly arranged Cas9 or Argonaute (Ago) domain, or a variant thereof.
[0376] Cas9 variants with modified PAM specificity
[0377] The base editor of the present disclosure may also include a Cas9 variant with modified PAM specificity. For example, the base editor described herein may utilize any naturally occurring or engineered SpCas9 variants with extended and / or relaxed PAM specificity, which have been described in the literature, including in Nishimasu et al., "Engineered CRISPR-Cas9 nuclease with expanded targeting space," Science, 2018, 361: 1259-1262; Chatterjee et al., "Robust Genome Editing of Single-Base PAM Targets with Engineered ScCas9 Variants," BioRxiv, April 26, 2019. Some aspects of the present disclosure provide a Cas9 protein that exhibits activity on a target sequence that does not include a typical PAM (5'-NGG-3', wherein N is A, C, G, or T) at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NGG-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NNG-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NNA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NNC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NNT-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NGT-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NGA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NGC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NAA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NAC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NAT-3' PAM sequence at its 3' end. In other embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NAG-3' PAM sequence at its 3' end.
[0378] The above description of various napDNAbps that can be used with the currently disclosed base editors is not intended to be limiting in any way. The base editor may include a typical SpCas9, or any orthologous Cas9 protein, or any Cas9 variant (including any naturally occurring variant, mutant, or other engineered version of Cas9), which is known or can be made or evolved by directed evolution or other mutagenesis processes. In various embodiments, Cas9 or Cas9 variants have nickase activity, i.e., only one strand of the target DNA sequence is cut. In other embodiments, Cas9 or Cas9 variants have an inactivated nuclease, i.e., a "dead" Cas9 protein. Other Cas9 variants that can be used include those with a molecular weight less than that of a typical SpCas9 (e.g., for easier delivery) or variants with a modified or rearranged primary amino acid structure (e.g., a cyclic substitution form). The base editor described herein may also include Cas9 equivalents, including Cas12a / Cpf1 and Cas12b proteins, which are the result of convergent evolution. The napDNAbp used herein (e.g., SpCas9, Cas9 variants, or Cas9 equivalents) may also contain various modifications that alter / enhance their PAM specificity. Finally, the present application contemplates any Cas9, Cas9 variants, or Cas9 equivalents having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity with a reference Cas9 sequence (e.g., a reference SpCas9 typical sequence or a reference Cas9 equivalent (e.g., Cas12a / Cpf1)).
[0379] In some embodiments, SpCas9(H840A) comprises a sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or at least 99.5% identical to the amino acid sequence in SEQ ID NO:480.
[0380] SpCas9(H840A)
[0381]
[0382] In a specific embodiment, the Cas9 variant with expanded PAM capacity is SpCas9(H840A)VRQR, having the following amino acid sequence (wherein the V, R, Q, R substitutions relative to SpCas9(H840A) of SEQ ID NO: 480 are shown in bold underline. In addition, for SpCas9(H840A)VRQR ("SpCas9-VRQR"), the methionine residue in SpCas9(H840) is removed. This SpCas9 variant has an altered PAM specificity that recognizes a PAM of 5'-NGA-3' instead of the typical 5'-NGG-3' PAM:
[0383] SpCas9-VRQR
[0384]
[0385] In another specific embodiment, the Cas9 variant with expanded PAM capacity is SpCas9(H840A)VQR, having the following amino acid sequence (wherein the V, Q, R substitutions relative to SpCas9(H840A) of SEQ ID NO: 480 are shown in bold underline. In addition, for SpCas9(H840A)VRQR ("SpCas9-VQR"), the methionine residue in SpCas9(H840) is removed. This SpCas9 variant has an altered PAM specificity that recognizes a PAM of 5'-NGA-3' instead of the typical 5'-NGG-3' PAM:
[0386] SpCas9-VQR
[0387]
[0388] In another specific embodiment, the Cas9 variant with expanded PAM capacity is SpCas9(H840A)VRER, having the following amino acid sequence (wherein the V, R, E, R substitutions relative to SpCas9(H840A) of SEQ ID NO: 480 are shown in bold underline. In addition, for SpCas9(H840A)VRER ("SpCas9-VRER"), the methionine residue in SpCas9(H840) is removed. This SpCas9 variant has an altered PAM specificity that recognizes a PAM of 5'-NGCG-3' instead of the typical 5'-NGG-3' PAM:
[0389] SpCas9-VRER
[0390]
[0391] In another embodiment, the Cas9 variant with expanded PAM capability is SpCas9-NG, as reported by Nishimasu et al., "Engineered CRISPR-Cas9 nuclease with expanded targeting space," Science, 2018, 361: 1259-1262, which is incorporated herein by reference. Relative to the typical SpCas9 sequence (SEQ ID NO: 200), SpCas9-NG (VRVRFRR) has the following amino acid sequence substitutions: R1335V, L1111R, D1135V, G1218R, E1219F, A1322R, and T1337R. The SpCas9 has a relaxed PAM specificity, i.e., it is active against the PAM of NGH (where H = A, T, or C). See Nishimasu et al., “Engineered CRISPR-Cas9 nuclease with expanded targeting space,” Science, 2018, 361: 1259-1262, which is incorporated herein by reference.
[0392] SpCas9-NG
[0393]
[0394] In addition, any available method can be used to obtain or construct variant or mutant Cas9 proteins. As used herein, the term "mutation" refers to a residue in a sequence (such as a nucleic acid sequence or an amino acid sequence) being replaced by another residue, or the insertion or deletion of one or more residues in the sequence. Mutation is generally described herein by identifying the original residue, followed by the position of the residue in the sequence and the identity of the newly replaced residue. The various methods provided herein for performing amino acid replacement (mutation) are well known in the art, and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). Mutation can include a variety of categories, such as single base polymorphisms, micro-repeat regions, insertions and deletions, and inversions, and are not intended to be limited in any way. Mutation can include "loss of function" mutations, which are normal results of mutations that reduce or abolish protein activity. Most loss of function mutations are recessive, because in heterozygotes, the second chromosome copy carries an unmutated gene encoding a fully functional protein, and its presence compensates for the effects of the mutation. Mutations also include "gain of function" mutations, which are mutations that confer abnormal activities on a protein or cell that it does not normally possess. Many gain of function mutations are located in regulatory sequences rather than coding regions, and can therefore have a variety of consequences. For example, a mutation can cause one or more genes to be expressed in the wrong tissues, which gain functions that they normally lack. By their nature, gain of function mutations are often dominant.
[0395] Site-directed mutagenesis can be used to introduce mutations into the reference Cas9 protein. Older site-directed mutagenesis methods known in the art rely on subcloning the sequence to be mutated into a vector, such as an M13 phage vector, which allows the separation of a single-stranded DNA template. In these methods, a mutagenic primer (i.e., a primer that can anneal to a site to be mutated but carries one or more mismatched nucleotides at the site to be mutated) is annealed to a single-stranded template, and then the complementary strand of the template is polymerized from the 3' end of the mutagenic primer. The resulting duplex is then transformed into a host bacterium, and plaques are screened for the desired mutation. Recently, site-directed mutagenesis has adopted a PCR method, which has the advantage of not requiring a single-stranded template. In addition, a method that does not require subcloning has also been developed. When performing PCR-based site-directed mutagenesis, several issues must be considered. First, in these methods, it is desirable to reduce the number of PCR cycles to prevent the amplification of undesired mutations introduced by polymerase. Secondly, selection must be adopted to reduce the number of unmutated parent molecules that persist in the reaction. Third, it is preferred to use a PCR method that extends the length so as to allow the use of a single PCR primer set. And fourth, due to the template-independent end extension activity of some thermostable polymerases, it is often necessary to incorporate an end-polishing step into the procedure prior to blunt-end ligation of PCR-generated mutant products.
[0396] Mutations can also be introduced by directed evolution processes, such as phage-assisted continuous evolution (PACE) or phage-assisted non-continuous evolution (PANCE). As used herein, the term "phage-assisted continuous evolution (PACE)" refers to continuous evolution using bacteriophage as a viral vector. The general concepts of PACE technology have been described, for example, in International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, and published as WO 2010 / 028347 on March 11, 2010; International PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, and published as WO 2012 / 088381 on June 28, 2012; U.S. Application No. 9,023,594, issued May 5, 2015, International PCT Application No. PCT / US2015 / 012022, filed January 20, 2015, and published as WO 2012 / 088381 on September 28, 2012; 2015 / 134121, and international PCT application PCT / US2016 / 027795, filed on April 15, 2016, published as WO 2016 / 168631 on October 20, 2016, the entire contents of each of which are incorporated herein by reference. Variant Cas9 can also be obtained by phage-assisted non-continuous evolution (PANCE), as used herein, the term "phage-assisted non-continuous evolution (PANCE)" refers to non-continuous evolution using phage as a viral vector. PANCE is a simplified technique for rapid in vivo directed evolution, using evolving "selection phage" (SP) containing the gene of interest to be evolved in fresh E. coli host cells for continuous bottle transfer, thereby allowing the genes in the host E. coli to remain unchanged while the genes contained in the SP continue to evolve. Continuous bottle transfer has long been a widely accessible method for laboratory microbial evolution, and recently, similar methods have also been developed for phage evolution. The PANCE system is characterized by lower stringency than the PACE system.
[0397] Compact Cas9 variants with modified PAM specificity
[0398] In some embodiments, napDNAbp comprises a compact Cas protein, such as Cas9 derived from Campylobacter jejuni (C.jejuni), Staphylococcus auricularis (S.auricularis), Neisseria meningitidis (N.meningitidis) or Staphylococcus aureus (S.aureus). In an exemplary embodiment, napDNAbp comprises CjCas9 nickase, SauriCas9 nickase, Nme2Cas9 nickase, SaCas9 nickase or SaKKH-Cas9 nickase. In some embodiments, napDNAbp is not Nme2Cas9 protein or nickase. In some embodiments, napDNAbp is not SaCas9 protein or nickase.
[0399] In some embodiments, the base editor of the present disclosure comprises a napDNAbp domain comprising a Cas9 ortholog derived from Neisseria meningitidis (Nme or Nme2). In some embodiments, the napDNAbp domain comprises Nme2Cas9. In other embodiments, the napDNAbp domain is an Nme2Cas9 domain. In some embodiments, the base editor of the present disclosure comprises an Nme2Cas9 nickase. Nme2Cas9 recognizes a simple dinucleotide PAM, NNNNCC or N 4 CC (wherein N is any nucleotide), as described in Edraki et al., Molecular Cell 73, 714-726, which is incorporated herein by reference. In other embodiments, the napDNAbp domain comprises an Nme2Cas9 variant. The variants of Nme2Cas9 can recognize a wider array of PAMs. In some embodiments, the Nme2Cas9 variants disclosed herein recognize single nucleotide pyrimidine PAMs. In some embodiments, the Nme2Cas9 variant recognition sequence is a PAM of NYN, wherein Y is any pyrimidine (ie, C, T, or U). In other embodiments, the Nme2Cas9 variant recognition sequence is NNNNCN or N 4 CN. In some embodiments, the Nme2Cas9 variant is an eNme2Cas9 nickase (SEQ ID NO: 439). In some embodiments, the Nme2Cas9 variant is an eNme2-C Cas9 nickase (SEQ ID NO: 353).
[0400] The sequence of wild-type Nme2Cas9 is shown in SEQ ID NO: 349. In some embodiments, the base editor of the present disclosure comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 349. In some embodiments, the base editor of the present disclosure comprises a napDNAbp comprising SEQ ID NO: 5. The protein may be referred to herein as engineered Nme2Cas9 or eNme2Cas9. In various embodiments, any disclosed TadCBE comprises Nme2Cas9 or a variant of Nme2Cas9.
[0401] Wild-type Nme2Cas9
[0402]
[0403] The "e" at the beginning of the Nme2Cas9 variants described herein indicates an "evolved" Nme2 variant. Amino acid substitutions relative to wild-type Nme2Cas9 are indicated in bold underline.
[0404] eNme2-C Cas9
[0405]
[0406]
[0407] In some embodiments, the base editor of the present disclosure comprises napDNAbp, which comprises a compact Cas9 ortholog derived from Campylobacter jejuni (CjCas9). In some embodiments, napDNAbp comprises CjCas9. In some embodiments, the base editor of the present disclosure comprises a CjCas9 nickase. CjCas9 recognizes NNNNACA and NNNNACACPAM. See Kim et al., Nature Communications 8(14500):1-12(2017), which is incorporated herein by reference. The sequence of CjCas9 (nickase) is shown in SEQ ID NO:348. In some embodiments, the base editor of the present disclosure comprises a napDNAbp domain having a sequence identical to SEQ ID NO:348 at least 90%, at least 95%, at least 98% or at least 99%. In some embodiments, the disclosed base editor comprises napDNAbp, which comprises SEQ ID NO:348. The length of the protein is 984 amino acids. The protein may be referred to herein as engineered CjCas9 or enCjCas9. Rationally engineered CjCas9 variants (enCjCas9) are described in Nakagawa, et al., Communications Biology, (2022) 5:211, which is incorporated herein by reference. In various embodiments, any disclosed TadCBE comprises a variant of CjCas9 or enCjCas9 (SEQ ID NO: 348).
[0408] (SEQ ID NO:348)
[0409] The base editor of the present disclosure may also include a Cas9 variant with a modified PAM specificity. Some aspects of the present disclosure provide a Cas9 protein that exhibits activity on a target sequence that does not include a typical PAM (5'-NGG-3', wherein N is A, C, G, or T) at its 3' end. In some embodiments, the Cas9 protein exhibits activity to a target sequence that includes a 5'-NGG-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to a target sequence that includes a 5'-NNG-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to a target sequence that includes a 5'-NNA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to a target sequence that includes a 5'-NNC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to a target sequence that includes a 5'-NNC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to a target sequence that includes a 5'-NNT-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NGT-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NGA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NGC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NAA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NAC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NAT-3'PAM sequence at its 3' end. In other embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5'-NAG-3'PAM sequence at its 3' end.
[0410] In some embodiments, the base editor of the present disclosure comprises a napDNAbp domain, which comprises SpCas9-NG, which has a PAM corresponding to NGN. In some embodiments, the disclosed base editor comprises a napDNAbp domain, which has a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SpCas9-NG. The sequence of SpCas9-NG is as follows:
[0411]
[0412] In some embodiments, the base editor disclosed herein comprises a napDNAbp domain comprising a Staphylococcus aureus Cas9 nickase KKH, or SaCas9-KKH or SaKKH-Cas9 having a PAM corresponding to NNNRRT or NNGRRT. The Cas9 variant contains amino acid substitutions D10A, E782K, N968K, and R1015H ("KKH") relative to wild-type SaCas9 (as shown in SEQ ID NO: 347). In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SaCas9-KKH. SaCas9 (and SaKKH-Cas9) has a length of 1053 amino acids. The sequence of SaCas9-KKH (nickase) is as follows:
[0413] Staphylococcus aureus Cas9 nickase KKH (SaCas9-KKH)
[0414]
[0415] In some embodiments, the disclosed base editor comprises a napDNAbp comprising a Cas9 protein derived from Staphylococcus auris (S. auri Cas9 or SauriCas9). In some embodiments, the disclosed base editor comprises a SauriCas9 nickase. SauriCas9 recognizes NNGG and NNNGG PAM. The sequence of SauriCas9 (nickase) is shown in SEQ ID NO: 358. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 358. In some embodiments, the disclosed base editor comprises a napDNAbp comprising SEQ ID NO: 358. The protein is 1061 amino acids long.
[0416]
[0417] In some embodiments, the napDNAbp comprises a SauriCas9-KKH variant or a SauriCas9-KKH nickase variant. SauriCas9-KKH contains the corresponding triple KKH mutations: Q788K, Y973K, and R1020H. See Hu et al. (2020) PLoS Biol. 18(3): e3000686, which is incorporated herein by reference.
[0418] In some embodiments, a base editor of the present disclosure comprises a napDNAbp domain comprising the Streptococcus pyogenes Cas9 nickase KKH, or SpCas9-KKH, having a PAM corresponding to NNNRRT.
[0419] In some embodiments, the Cas variant is a variant of SpRY having a mutation that confers high fidelity. This variant is referred to as SpRY-HF or SpRY-HF1. The high-fidelity variant of SpRY or any Cas variant provided herein may comprise one or more of the N497A, R661A, Q695A and / or Q926A mutations relative to SEQ ID NO: 74, or the corresponding mutations in any Cas9 provided herein. Cas9 variants with high fidelity are known in the art and are apparent to those skilled in the art. For example, high-fidelity Cas9 domains have been described in Kleinstiver, BP, et al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016); and Slaymaker, IM, et al. "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015), each of which is incorporated herein by reference.
[0420] In some embodiments, the disclosed Cas variants include Cas9 derived from Streptococcus macacae (e.g., Streptococcus macacae NCTC 11558), or a variant of SmacCas9. In some embodiments, the Cas variant comprises a hybrid variant of SmacCas9, which incorporates a SpCas9 domain and a SmacCas9 domain, and is referred to as Spy-macCas9 or a variant thereof. In some embodiments, the Cas variant comprises a hybrid variant of SmacCas9, which incorporates an enhanced nucleic acid cleavage variant of the SpCas9 (iSpy Cas9) domain, and is referred to as iSpy-macCas9. Relative to Spymac-Cas9, iSpyMac-Cas9 contains two mutations, R221K and N394K, which are identified by deep mutation scanning of Spy Cas9, which can increase the modification rate of the protein on most targets. See Jakimo et al., bioRxiv, ACas9 with Complete PAM Recognition for Adenine Dinucleotides (Sep 2018), which is incorporated herein by reference. Jakimo et al. showed that the hybrid Spy-macCas9 and iSpy-macCas9 recognized the short 5'-NAA-3' PAM and recognized all evaluated adenine dinucleotide PAM sequences and had strong editing efficiency in human cells. Liu et al. engineered a base editor containing Spy-mac Cas9 and demonstrated that cytidine and adenine base editors containing Spymac domains can induce efficient C to T and A to G conversion in vivo. In addition, Liu et al. proposed that the PAM range of Spy-mac Cas9 can be 5'-TAAA-3' instead of 5'-NAA-3' reported by Jakimo et al. (see Liu et al. Cell Discovery (2019) 5:58, which is incorporated herein by reference).
[0421] Any of the above references related to Cas9 variants or Cas9 equivalents are hereby incorporated by reference in their entirety, if not already stated so.
[0422] The following table provides a comparison of the PAM preferences targeted by the Nme2 Cas variants of the present disclosure and other Cas homologs (R = any purine, Y = any pyrimidine, N = any nucleotide).
[0423] Table 6 - PAM preferences of exemplary Cas homologs
[0424]
[0425]
[0426] PAM / protospacer sequence
[0427] For typical (i.e., Streptococcus pyogenes Cas9-derived) base editors, base editing requires the presence of a protospacer adjacent motif (PAM) located approximately 15 base pairs adjacent to the target nucleotide. Each programmable DNA-binding protein domain recognizes a different PAM sequence. Only about a quarter of pathogenic transition point mutations have a suitably positioned typical PAM "NGG" sequence that is compatible with Streptococcus pyogenes Cas9 (SpCas9)-derived base editors. Naturally occurring cytidine deaminases have shown broad compatibility with many Cas homologs, including Staphylococcus aureus Cas9 (SaCas9) 98 、SaCas9-KKH 8 , Cas12a(Cpf1) 9,10 、SpCas9-NG 11 and circularly arranged CP-Cas9 7 , greatly expanding their targeting range.
[0428] In some embodiments, napDNAbp comprises a protospacer sequence having a sequence and located upstream of a sequence having a sequence. In some embodiments, the protospacer sequence is located upstream of a PAM having a sequence of TGG. In other embodiments, the protospacer sequence is located upstream of a PAM having a sequence of GGG. In other embodiments, the protospacer sequence is located upstream of a PAM having a sequence of AGG. In some embodiments, the protospacer sequence is located upstream of a PAM having a sequence of CGG. In some embodiments, the protospacer sequence is located upstream of a PAM having a sequence of AGACCC. In other embodiments, the protospacer sequence is located upstream of a PAM having a sequence of ACCTCA. In some embodiments, the protospacer sequence is located upstream of a PAM having a sequence of GGGGCG. In other embodiments, the protospacer sequence is located upstream of a PAM having a sequence of CAGCCG. In some embodiments, the protospacer sequence is located upstream of a PAM having a sequence of GCGGCT. In other embodiments, the protospacer sequence is located upstream of a PAM having a sequence of GGGGCA. In some embodiments, the protospacer sequence is located upstream of a PAM having a sequence of AAGGGT. In other embodiments, the protospacer sequence is located upstream of a PAM having the sequence TCGGGT. In some embodiments, the protospacer sequence is located upstream of a PAM having the sequence GAGAGT. In some embodiments, the protospacer sequence is located upstream of a PAM having the sequence CAGAAT. In some embodiments, the protospacer sequence is located upstream of a PAM having the sequence CTGGGT.
[0429] In some embodiments, the expected editing base pair is located 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides upstream of the PAM site. In some embodiments, the expected editing base pair is located downstream of the PAM site. In some embodiments, the expected editing base pair is located 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides downstream of the PAM site. In some embodiments, the method does not require a typical (eg, NGG) PAM site. In some embodiments, the target region comprises a target window, wherein the target window comprises a target nuclear base pair.
[0430] The protospacer sequences of the present disclosure may include but are not limited to the following sequences:
[0431]
[0432]
[0433]
[0434] Edit Window
[0435] The base editor of the present disclosure may have a variable target region including a target window (e.g., an editing window or a deamination window) of a target nuclear base pair, wherein a nuclear base editor is installed in the target region. In some embodiments, TadA-CD has a C to T base editing window corresponding to the original spacer position 2-12 of the original spacer. In some embodiments, the TadA-CD base editor has a C to T base editing window corresponding to the original spacer position 2-12 of the original spacer. In a specific embodiment, the TadA-CD base editor has a C to T base editing window corresponding to the original spacer position 3 to 8. The base editor of the present disclosure may have particularly high editing activity on cytosine between original spacer positions 5 to 7.
[0436] In some embodiments, the target window (e.g., editing window) comprises 1-10 nucleotides. In some embodiments, the length of the editing window is 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, or 1 nucleotide. In some embodiments, the length of the target window (e.g., editing window) is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides. In some embodiments, the expected editing base pair is located within the editing window. In some embodiments, the editing window comprises an expected editing base pair.
[0437] In some cases, the TadA-CD base editing window starts after position 2, after position 3, after position 4, after position 5, after position 6, after position 7, after position 8, after position 9, after position 10, and after position 11 of the protospacer. In some embodiments, the editing window ends before position 12, before position 11, before position 10, before position 9, before position 8, before position 7, before position 6, before position 5, before position 4, and before position 3 of the protospacer.
[0438] In some embodiments, a TadA-CD base editor comprising a V106W mutation has a narrower editing window relative to a TadA-CD base editor lacking the mutation. For example, the base editing window of TadA-CDa (SEQ ID NO: 34) is between about position 4 and about position 9 of the original spacer. In certain embodiments, a TadA-CD base editor comprising a V106W mutation (e.g., TadA-CDa V106W and TadA-CDd V106W) has a C to T base editing window between position 3 and position 9 of the original spacer, or any combination thereof. For example, the editor can install a C to T substitution at position 3, position 4, position 5, position 6, position 7, position 8, or position 9 of the original spacer, or any combination thereof.
[0439] In some cases, the TadA-CD V106W base editing window starts after position 2, after position 4, after position 5, after position 6, after position 7, after position 8, or after position 9 of the protospacer. In some embodiments, the editing window ends before position 10, before position 9, before position 8, before position 7, before position 6, before position 5, and before position 4 of the protospacer.
[0440] In some embodiments, the TadA-CD base editor has an A to G base editing window between position 4 and position 7 of the protospacer. In some cases, the TadA-CD base editor installs an A to G edit at position 4, position 5, position 6, or position 7 of the protospacer, or any combination thereof. According to some embodiments, by including the V106W mutation, the A to G base editing property of TadA-CD can be narrowed to between position 5 and position 7 of the protospacer.
[0441] Those skilled in the art will appreciate that the TadA-CD base editors described above and herein have narrower C to T base editing windows than several existing cytidine deaminases such as rAPOBEC1, evoAPOBEC1 (evoA), evoFERNY, and YE1. For example, BE4max and evoABE4max exhibit C to T editing windows ranging from position 1 to position 14 of the original spacer; evoFERNY-BE4 exhibits a C to T editing window from position 1 to approximately position 10; and YE1-BE4 exhibits a C to T editing window from position 3 to position 9 of the original spacer (see Figure 3 ).
[0442] Those skilled in the art will also appreciate that the TadA-CD base editors described above and herein have a narrower A to G base editing window and a wider C to T base editing window than the parent adenosine deaminase from which it evolved (e.g., TadA-8e). For example, TadA-8e exhibits an A to G base editing window between positions 1 and 15 of the original spacer, and a C to T base editing window between positions 4 to 7 of the original spacer.
[0443] The TadA-CD base editor of the present disclosure can convert one or more target cytosines to thymines within the original spacer sequence. For example, in some embodiments, TadA-CB can convert 2 cytosines, 3 cytosines, 4 cytosines, or 5 cytosines within the original spacer sequence.
[0444] Editing efficiency
[0445] Aspects of the present disclosure relate to the efficiency of editing a DNA target sequence in a target region of a target window containing a target base pair by a cytosine base editor described herein. In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the expected base pairs are edited. In some embodiments, the efficiency of C to T conversion of any disclosed base editor or method using these base editors is at least 80% in all sequencing reads. In a specific embodiment, TadCBEa achieves an average conversion efficiency of 51%-60% of the target cytosine.
[0446] In some embodiments, any disclosed base editor or method using the base editor provides an average 70% cytosine conversion efficiency in clinically relevant genes (such as the CXCR5 and CCR5 genes, which are associated with HIV / AIDS).
[0447] In some embodiments, the cytidine deamination activity of the disclosed deaminase (and therefore the cytosine editing activity of the disclosed base editor) exceeds the adenosine deamination activity of the deaminase in a significant proportion. For example, the ratio of the cytidine deamination activity to the adenosine deamination activity of the disclosed Tad-CD deaminase is at least about 5: 1, 6: 1, 7: 1, 8: 1, 9: 1, 10: 1, 11: 1, 12: 1, 13: 1, 14: 1, 15: 1, 17: 1, 19: 1, 20: 1, 21: 1, 23: 1, 25: 1, 30: 1 or more than 30: 1. In some embodiments, the ratio of the cytidine deamination activity to the adenosine deamination activity of the deaminase is at least about 10: 1. In some embodiments, the ratio is at least about 20: 1. In some embodiments, the ratio is about 5:1-7.5:1, 7.5:1-9.5:1, 5:1-10:1, 10:1-15:1, 15:1-20:1, 10-17:1, 12:1-17:1, 20:1-21:1, 21:1-25:1, 20:1-30:1, 25:1-35:1, 30:1-35:1, 30:1-40:1, 40:1-42:1, 21 :1-42:1, 25:1-40:1, 10:1-40:1, 25-45:1, 30:1-50:1, 45:1-50:1, 50:1-60:1, 55:1-65:1, 60:1-70:1, 70:1-80:1, 80:1-85:1, 10:1-80:1, 40:1-80:1, 20:1-60:1, 20:1-80:1 or 75:1-85:1.
[0448] In some embodiments, the peak editing efficiency of TadA-CD is comparable to that of a natural cytosine base editor (e.g., a BE4max editor containing APOBEC1, evoFERNY, or evoA deaminase). In some embodiments, the editing efficiency of the TadA-CD base editor is higher than that of a natural cytosine base editor. For example, in some embodiments, TadA-CDa, TadA-CDb, and TadA-CDc edit the Nme50 gene at positions 3-8 of the original spacer with an efficiency between 5% and 48%. In some cases, TadA-CD comprises a V106W substitution that maintains editing efficiency while narrowing the editing window of the base editor.
[0449] In some embodiments, the disclosed TadCBE and editing methods include a step of contacting DNA with any disclosed TadCBE, resulting in at least about 20%, 21%, 25%, 30%, 35%, 40%, 50%, 60%, 70%, 80%, 85% or more than 85% on-target (C to T) base editing efficiency at the target nucleobase pair in all sequencing reads. The contacting step can result in at least about 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 52%, 55%, 60%, 62%, 65%, 70%, 72%, 75%, 80%, 82%, 85% or more than 85% C to T base editing efficiency. In particular, the contacting step results in an on-target base editing efficiency greater than 75%. In certain embodiments, a base editing efficiency of 99% can be achieved.
[0450] In some cases, the TadA-CD base editors described herein have a C to T editing efficiency between 20% and 80%. In some embodiments, the C to T editing efficiency is greater than or equal to 10%, 20%, 25%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, or 80%. In other embodiments, the C to T editing efficiency is less than or equal to 95%, less than or equal to 90%, less than or equal to 80%, less than or equal to 70%, less than or equal to 60%, less than or equal to 50%, less than or equal to 40%, less than or equal to 30%, less than or equal to 20%, less than or equal to 10%, less than or equal to 5%, or less than or equal to 1%.
[0451] In some cases, the base editors of the present disclosure may have different base editing efficiencies (e.g., converting C to T) for targeted nucleotides within a given protospacer sequence. In other words, the TadA-CD base editors of the present invention may preferentially edit a certain position (or multiple positions) in the protospacer sequence. For example, in some embodiments, the TadA-CDa variant preferentially edits the protospacer sequence GC 2 A 3 A 4 GA 6 GC 8 A 9 C 10 A 11 A 12 GAGGAAGAGAG AGACCC (SEQ ID NO:385) The C8 position of , where the PAM is underlined; whereas the TadA-CDc variant edits both the C8 and C10 positions with similar efficiency.
[0452] Thus, in some embodiments, the editing efficiency of TadCBE at each position of the protospacer within the editing window (e.g., position 1 to position 15) is between 20% and 80%. In some embodiments, the editing efficiency at each position of the protospacer within the editing window (e.g., position 1 to position 15) is greater than or equal to 10%, greater than or equal to 20%, greater than or equal to 30%, greater than or equal to 40%, greater than or equal to 50%, greater than or equal to 60%, greater than or equal to 70%, greater than or equal to 80%, or greater than or equal to 85%. In other embodiments, the editing efficiency at each position of the protospacer within the editing window (e.g., position 1 to position 15) is less than or equal to 85%, less than or equal to 80%, less than or equal to 70%, less than or equal to 60%, less than or equal to 50%, less than or equal to 40%, less than or equal to 30%, less than or equal to 20%, less than or equal to 10%, less than or equal to 5%, or less than or equal to 1%.
[0453] Thus, in some embodiments, the TadCBE of the present application provides at least 20%, 21%, 25%, 30%, 35%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 52%, 55%, 60%, 62%, 65%, 70%, 72%, 75%, 80%, 82%, 85%, or more than 85% C to T base conversion when contacted with DNA comprising a target sequence. The target sequence is selected from the group consisting of: CTT, CTC, CTA, CTG, CCT, CCC, CCA, CCG, CAT, CAC, CAA, CAG, CGT, CGC, CGA, CGG, TCT, TCC, TCG, ACT, ACC, ACA, ACG, GCT, GCC, GCA, GCG, TTC, TAC, TGC, ATC, AAC, AGC, GTC, GAC and GGC.
[0454] The disclosed TadCBE has a higher affinity and specificity for cytosine bases and is therefore less prone to deaminating adenosine residues. In some cases, the TadA-CD base editor described herein has a lower residual A to G editing efficiency, for example, between 0.1% and 20%. In some embodiments, the A to G editing efficiency is greater than or equal to 0.1%, greater than or equal to 5%, greater than or equal to 10%, or greater than or equal to 20%. In other embodiments, the A to G editing efficiency is less than or equal to 95%, less than or equal to 90%, less than or equal to 80%, less than or equal to 70%, less than or equal to 60%, less than or equal to 50%, less than or equal to 40%, less than or equal to 30%, less than or equal to 20%, less than or equal to 10%, less than or equal to 5%, or less than or equal to 0.1%.
[0455] In some cases, the TadA-CD base editors described herein have residual (off-target) C to G editing ability. In some embodiments, the V106W variant has reduced C to G editing compared to the natural TadA-CD base editor. In some embodiments, the TadA-CD V016W mutant reduces C to G editing in HEK293T cells, T cells, and HSPCs.
[0456] Wizard Sequence
[0457] The present disclosure further provides guide RNAs for use according to the disclosed editing methods. The present disclosure provides guide RNAs (gRNAs) designed to recognize target sequences. Such gRNAs can be designed to have guide sequences (or "spacers") that are complementary to the original spacers within the target sequence.
[0458] Also provided are guide RNAs for use with one or more disclosed adenine base editors, such as in the disclosed methods of editing nucleic acid molecules. Such gRNAs can be designed to have a guide sequence complementary to the original spacer within the target sequence to be edited, and a backbone sequence that specifically interacts with the napDNAbp domain of any disclosed base editor (e.g., the Cas9 nickase domain of the disclosed base editor).
[0459] In various embodiments, the base editor can be compounded, bound or otherwise associated with one or more guide sequences (e.g., via any type of covalent or non-covalent bond). The guide sequence becomes associated or bound to the base editor and guides it to locate a specific target sequence that is complementary to the guide sequence or a portion thereof. The specific design implementation of the guide sequence will depend on the nucleotide sequence of the genomic target sequence (i.e., the desired site to be edited) and the type of napDNAbp (e.g., Cas9 protein type) present in the base editor, as well as other factors such as PAM sequence position, percentage G / C content in the target sequence, degree of microhomology region, secondary structure, etc.
[0460] Typically, a guide sequence is any polynucleotide sequence that has sufficient complementarity to a target polynucleotide sequence to hybridize with the target sequence and guide sequence-specific binding of a napDNAbp (e.g., Cas9 or a Cas9 variant) to the target sequence. In some embodiments, when optimally aligned using a suitable alignment algorithm, the degree of complementarity between the guide sequence and its corresponding target sequence is about or greater than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more. Optimal alignment can be determined by using any suitable sequence alignment algorithm, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transform (e.g., Burrows WheelerAligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies), ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).
[0461] In some embodiments, the guide sequence is about or more than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, 80, 85, 90, 95, 100, or more nucleotides in length. In other embodiments, the guide sequence is about or more than about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, or 200 nucleotides in length. In some embodiments, each gRNA comprises a guide sequence of at least 10 consecutive nucleotides (e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 consecutive nucleotides) complementary to the target sequence (or off-target site). In some embodiments, each gRNA comprises a guide sequence of at least 15 consecutive nucleotides complementary to the target sequence (or off-target site). In other embodiments, each gRNA comprises a guide sequence of at least 20 consecutive nucleotides complementary to the target sequence (or off-target site).
[0462] In some embodiments, the length of the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 or less nucleotides. The ability of the guide sequence to guide the base editor to specifically bind to the target sequence can be evaluated by any suitable assay. For example, the components of the base editor (including the guide sequence to be tested) can be provided to a host cell with a corresponding target sequence, such as by transfection with a vector encoding the base editor components disclosed herein, and then evaluating the preferential cutting within the target sequence. Similarly, the cutting of the target polynucleotide sequence can be evaluated in situ by providing a target sequence, a component of a base editor, including a guide sequence to be tested and a control guide sequence different from the test guide sequence, and comparing the binding or cutting rate on the target sequence between the test guide sequence reaction and the control guide sequence reaction. Other assays are also possible, and those skilled in the art can think of it.
[0463] The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of the cell. Exemplary target sequences include those that are unique in the target genome.
[0464] In some embodiments, the guide sequence is selected to reduce the degree of secondary structure within the guide sequence. Secondary structure can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. An example of such an algorithm is mFold, as described by Zuker & Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example folding algorithm is the online web server RNAfold, developed by the Institute of Theoretical Chemistry of the University of Vienna, using the central structure prediction algorithm (see, for example, ARGruber et al., 2008, Cell 106 (1): 23-24; and PA Carr & GM Church, 2009, Nature Biotechnology 27 (12): 1151-62). Other algorithms can be found in Chuai, G. et al., Deep CRISPR: optimized CRISPR guide RNA design by deep learning, Genome Biol. 19:80 (2018), and U.S. Application Serial No. 61 / 836,080 and U.S. Patent No. 8,871,445 (issued on October 28, 2014), the entire contents of each of which are incorporated herein by reference.
[0465] The guide sequence of the gRNA is linked to a tracr mate (also referred to as a "backbone") sequence, which in turn hybridizes to the tracr sequence. The tracr mate sequence includes any sequence that has sufficient complementarity to the tracr sequence to promote one or more of the following: (1) excision of the guide sequence flanked by the tracr mate sequence in a cell containing the corresponding tracr sequence; and (2) formation of a complex at the target sequence, wherein the complex comprises the tracr mate sequence hybridized to the tracr sequence. In general, the degree of complementarity refers to the optimal alignment of the tracr mate sequence and the tracr sequence along the shorter sequence length of the two sequences. The optimal alignment can be determined by any suitable alignment algorithm, and can further take into account secondary structures, such as self-complementarity within the tracr sequence or the tracr mate sequence. In some embodiments, when optimally aligned, the degree of complementarity of the tracr sequence and the tracr mate sequence along the shorter sequence length of the two sequences is about or greater than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more. In some embodiments, the length of the tracr sequence is about or more than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 or more nucleotides. In some embodiments, the tracr sequence and the tracr mate sequence are contained in a single transcript so that hybridization between the two produces a transcript with a secondary structure (such as a hairpin). The preferred loop sequence for the hairpin structure is four nucleotides in length and most preferably has the sequence GAAA. However, longer or shorter loop sequences can be used, and alternative sequences can also be used. These sequences preferably include nucleotide triplets (e.g., AAA) and additional nucleotides (e.g., C or G). Examples of loop sequences include CAAA and AAAG. In embodiments of the present invention, the transcript or transcribed polynucleotide sequence has at least two or more hairpins. In certain embodiments, the transcript has two, three, four or five hairpins. In further embodiments of the present invention, the transcript has up to five hairpins. In some embodiments, the single transcript also includes a transcription termination sequence; preferably this is a polyT sequence, eg, six T nucleotides.
[0466] A non-limiting example of a single (DNA) polynucleotide comprising a guide sequence, a tracr partner sequence, and a tracr sequence is as follows (listed from 5' to 3'), where "N" represents a base of the guide sequence, the first set of lowercase letters represents the tracr partner sequence, and the second set of lowercase letters represents the tracr sequence, and the last poly-T sequence represents a transcription terminator:
[0467] (1) NNNNNNNNNNgtttttgtactctcaagatttaGAAAtaaatcttgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT (SEQ ID NO:264);
[0468] (2) NNNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT (SEQ ID NO:265);
[0469] (3) NNNNNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcattttatggcagggtgtTTTTT (SEQ ID NO:266);
[0470] (4) NNNNNNNNNNNNNNNNNNNgttttagagctaGAAAtagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcTTTTTT (SEQ ID NO:267);
[0471] (5) NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAATAGcaagttaaaataaggctagtccgttatcaacttgaaaaagtgTTTTTTT (SEQ ID NO:268); and
[0472] (6) NNNNNNNNNNNNNNNNNNNNNgttttagagctagAAATAGcaagttaaaataaggctagtccgttatcaTTTTTTTT (SEQ ID NO:269).
[0473] In some embodiments, sequences (1) to (3) are used in combination with Cas9 from S. thermophilus CRISPR1. In some embodiments, sequences (4) to (6) are used in combination with Cas9 from S. pyogenes. In some embodiments, the tracr sequence is a separate transcript from a transcript comprising a tracr partner sequence.
[0474] In some embodiments, the guide RNA used according to the disclosed editing method comprises a synthetic single guide RNA (sgRNA) containing modified ribonucleotides. In some embodiments, the guide RNA contains modifications such as 2'-O-methylated nucleotides and phosphorothioate bonds. In some embodiments, the guide RNA contains 2'-O-methyl modifications in the first three and last three nucleotides and contains phosphorothioate bonds between the first three and last three nucleotides. Exemplary modified synthetic sgRNAs are disclosed in Hendel A. et al., Nat. Biotechnol. 33, 985-989 (2015), which is incorporated herein by reference.
[0475] In some embodiments, the guide RNA used according to the disclosed editing method comprises a main chain structure recognized by a Streptococcus pyogenes Cas9 protein or domain (such as the SpCas9 domain of the disclosed base editor). The main chain structure recognized by the SpCas9 protein may include the sequence 5'-[guide sequence]-guuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuu-3' (SEQ ID NO: 339), wherein the guide sequence comprises a sequence complementary to the original spacer sequence of the target sequence. See U.S. Publication No. 2015 / 0166981 (published on June 18, 2015), the disclosure of which is incorporated herein by reference. The guide sequence is typically 20 nucleotides long.
[0476] In other embodiments, the guide RNA used according to the disclosed editing method comprises a backbone structure recognized by the Staphylococcus aureus Cas9 protein. The backbone structure recognized by the SaCas9 protein can comprise the sequence 5'-[guide sequence]-guuuuaguacucuguaaugaaaauuacagaaucuacuaaaacaaggcaaaaugccguguuuaucucgucaacuuguuggcgagauuuuuuuu-3' (SEQ ID NO: 78).
[0477] In some embodiments, the guide RNA used according to the disclosed editing method comprises a main chain structure recognized by a Neisseria meningitidis Cas9 protein or domain (such as an Nme2Cas9 domain). The main chain structure (or scaffold) recognized by the Nme2Cas9 protein can include the following provided sequence: 5'-[guide sequence]-gttgtagctccctttctcatttcggaaacgaaatgagaaccgttgctacaataaggccgtctgaaaagatgtgccgcaacgctctgccccttaaagcttctgctttaaggggcatcgttta-3' (SEQ ID NO: 274). The scaffold sequence is recognized by Nme1Cas9, Nme2Cas9 and Nme3Cas9 proteins.
[0478] Based on the present disclosure, suitable guide RNA sequences for targeting the disclosed TadCBE to a specific genomic site will be apparent to those skilled in the art. Such suitable guide RNA sequences typically comprise a guide sequence complementary to a nucleic acid sequence within 50 nucleotides upstream or downstream of the tar...
Claims
1. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:41, wherein the amino acid corresponding to residue 26 of SEQ ID NO:41 is any amino acid except R.
2. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:41, wherein the amino acid corresponding to residue 27 of SEQ ID NO:41 is any amino acid except E.
3. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:41, wherein the amino acid corresponding to residue 28 of SEQ ID NO:41 is any amino acid except V.
4. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:8, wherein the amino acid corresponding to residue 48 of SEQ ID NO:41 is any amino acid except R.
5. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:41, wherein the amino acid corresponding to residue 73 of SEQ ID NO:41 is any amino acid except Y.
6. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:41, wherein the amino acid corresponding to residue 96 of SEQ ID NO:41 is any amino acid except H.
7. A deaminase comprising mutations at residues E27, V28 and H96 in the amino acid sequence of SEQ ID NO: 41, and further comprising at least one mutation at a residue selected from the group consisting of R26, M61, Y73, I76, M151, Q154 and A158, or a corresponding mutation in a homologous adenosine deaminase.
8. The deaminase according to any one of claims 1 to 7, wherein the deaminase is capable of deaminating cytidine in DNA.
9. The deaminase according to any one of claims 7-8, wherein the deaminase comprises at least one mutation selected from E27A, E27K, V28G, V28A and H96N in the amino acid sequence of SEQ ID NO: 41, and further comprises at least one mutation at a residue selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase.
10. The deaminase according to any one of claims 7-8, wherein the deaminase comprises the mutations E27A, V28G and H96N in the amino acid sequence of SEQ ID NO: 41, and further comprises at least one mutation selected from the group consisting of R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase.
11. The deaminase according to any one of claims 7-8, wherein the deaminase comprises the mutations E27K, V28G and H96N in the amino acid sequence of SEQ ID NO: 41, and further comprises at least one mutation selected from the group consisting of R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase.
12. The deaminase according to any one of claims 7-8, wherein the deaminase comprises the mutations E27A, V28A and H96N in the amino acid sequence of SEQ ID NO: 41, and further comprises at least one mutation selected from the group consisting of R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase.
13. The deaminase according to any one of claims 7-8, wherein the deaminase comprises the mutations E27K, V28A and H96N in the amino acid sequence of SEQ ID NO: 41, and further comprises at least one mutation selected from the group consisting of R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H and A158S, or a corresponding mutation in a homologous adenosine deaminase.
14. The deaminase according to any one of claims 1 to 13, further comprising a mutation at position V106.
15. The deaminase according to any one of claims 1 to 13, further comprising the mutation V106W.
16. The deaminase according to any one of claims 1 to 15, comprising at least two mutations at residues selected from the group consisting of R26, M61, Y73, I76, M151, Q154 and A158.
17. The deaminase according to any one of claims 1 to 16, comprising at least two mutations at residues selected from the group consisting of R26G, M61I, Y73H, I76F, M151I, Q154H, Q154R and A158S.
18. The deaminase according to any one of claims 1 to 17, comprising at least one mutation selected from E27A, V28G, I76F and M151I.
19. The deaminase according to any one of claims 1 to 17, comprising at least one mutation selected from E27A, V28G, I76F and A158S.
20. The deaminase according to any one of claims 1-17, comprising at least one mutation selected from E27A, V28G, I76F, Q154R and A158S.
21. The deaminase according to any one of claims 1-17, comprising at least one mutation selected from E27K, V28A and M61I.
22. The deaminase according to any one of claims 1-17, comprising at least one mutation selected from E27A, V28G, Y73H, Q154H and A158S.
23. The deaminase according to any one of claims 1-18, comprising the mutations R26G, E27A, V28G, I76F, H96N and M151I.
24. The deaminase according to any one of claims 1-17, comprising the mutations R26G, E27A, V28G, I76F, H96N and A158S.
25. The deaminase according to any one of claims 1-17, comprising the mutations R26G, E27A, V28G, I76F, H96N, Q154R and A158S.
26. The deaminase according to any one of claims 1-17, comprising the mutations E27K, V28A, M61I and H96N.
27. The deaminase according to any one of claims 1-17, comprising the mutations E27A, V28G, Y73H, H96N, Q154H and A158S.
28. The deaminase of any one of claims 7-27, wherein the cytidine deamination activity of the deaminase exceeds the cytidine deamination activity of TadA-8e.
29. The deaminase of any one of claims 7-28, wherein the cytidine deamination activity of the deaminase exceeds the adenosine deamination activity of the deaminase.
30. The deaminase of any one of claims 7-29, wherein the deaminase has a ratio of cytidine deamination activity to adenosine deamination activity of at least about 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 17:1, 19:1, 20:1, 21:1, 23:1, 25:1, 30:1, or greater than 30:
1.
31. The deaminase of any one of claims 7-30, wherein the deaminase has a ratio of cytidine deamination activity to adenosine deamination activity of at least about 10:
1.
32. The deaminase of any one of claims 7-31, wherein the deaminase has a ratio of cytidine deamination activity to adenosine deamination activity of at least about 20:
1.
33. The deaminase of any one of claims 2-32, wherein the deaminase achieves an efficiency of at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, or 80% conversion of cytidine to thymine.
34. The deaminase of any one of claims 2-33, wherein the deaminase achieves an efficiency of at least about 75% conversion of cytidine to thymine.
35. A deaminase comprising mutations at residues R26, V28, A48 and Y73 in the amino acid sequence of SEQ ID NO: 41, or corresponding mutations in a homologous adenosine deaminase (eg TadA-dual, SEQ ID NO: 39).
36. The deaminase of claim 35, wherein the deaminase is capable of deaminating cytidine in DNA.
37. The deaminase of claim 35 or 36, wherein the deaminase further comprises a mutation at residue H96.
38. The deaminase according to any one of claims 35-37, comprising the mutations R26G, V28A, A48R, Y73S and H96N.
39. The deaminase according to any one of claims 35-38, comprising the mutations R26G, V28G, A48R and Y73C.
40. The deaminase of any one of claims 35-39, wherein the deaminase has a ratio of adenosine deamination activity to cytidine deamination activity of at least about 0.7:1, 0.8:1, 0.9:1, 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, or 1.5:
1.
41. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:39, wherein the amino acid corresponding to residue 46 of SEQ ID NO:39 is any amino acid except N.
42. A deaminase comprising a mutation at residue N46 in the amino acid sequence of SEQ ID NO: 39 and further comprising at least one mutation at a residue selected from the group consisting of G26, A28, L34, N46, R48, R64, Q71, S73, N96, G105, H154 and A162, or a corresponding mutation in a homologous adenosine deaminase.
43. The deaminase of claim 41 or 42, wherein the deaminase is capable of deaminating cytidine in DNA.
44. The deaminase according to any one of claims 41-43, wherein the deaminase comprises at least one mutation selected from N46I, N46V, N46L and N46C in the amino acid sequence of SEQ ID NO: 39, and further comprises S73P and H154Q mutations, or corresponding mutations in homologous adenosine deaminases.
45. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation at position N46V in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues S73P, G105S and H154Q, or corresponding mutations in homologous adenosine deaminases.
46. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation at position N46L in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues G26R, R48P, S73P, N96H and H154Q, or corresponding mutations in homologous adenosine deaminases.
47. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation at position N46V in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues Q71H, S73P and H154Q, or corresponding mutations in homologous adenosine deaminases.
48. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation at position N46C in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues S73P, H154Q and A162V, or corresponding mutations in homologous adenosine deaminases.
49. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation at position N46I in the amino acid sequence of SEQ ID NO: 39 and further comprises a mutation at residue H154Q, or a corresponding mutation in a homologous adenosine deaminase.
50. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation selected from the group consisting of N46T, N46V, N46C and N46L in the amino acid sequence of SEQ ID NO: 39, and further comprises a mutation at position H154Q, or a corresponding mutation in a homologous adenosine deaminase.
51. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation at N46T in the amino acid sequence of SEQ ID NO: 39, or a corresponding mutation in a homologous adenosine deaminase.
52. The deaminase according to any one of claims 41-43, wherein the deaminase comprises at least one mutation selected from the group consisting of N46V, N46L and N46C in the amino acid sequence of SEQ ID NO: 39, and further comprises one or more mutations selected from the group consisting of residues S73P, S73Y and A162V, or corresponding mutations in homologous adenosine deaminases.
53. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46C mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues S73P and A162V, or corresponding mutations in homologous adenosine deaminases.
54. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46C mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises a mutation at residue S73P, or a corresponding mutation in a homologous adenosine deaminase.
55. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46C mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues S73Y and A162V, or corresponding mutations in homologous adenosine deaminases.
56. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46V mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues Q71H and S73P, or corresponding mutations in homologous adenosine deaminases.
57. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46V mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises a mutation at residue S73P, or a corresponding mutation in a homologous adenosine deaminase.
58. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation at N46L in the amino acid sequence of SEQ ID NO: 39, or a corresponding mutation in a homologous adenosine deaminase.
59. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46V mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises a mutation at residue S73P, or a corresponding mutation in a homologous adenosine deaminase.
60. The deaminase according to any one of claims 41-43, wherein the deaminase comprises a mutation at N46V in the amino acid sequence of SEQ ID NO: 39, or a corresponding mutation in a homologous adenosine deaminase.
61. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46V mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises a mutation at residue R48P, or a corresponding mutation in a homologous adenosine deaminase.
62. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46C mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises a mutation at residue S73P, or a corresponding mutation in a homologous adenosine deaminase.
63. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46L mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues L34M and S73P, or corresponding mutations in homologous adenosine deaminases.
64. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46L mutation in the amino acid sequence of SEQ ID NO: 39 and further comprises a mutation at residue S73P, or a corresponding mutation in a homologous adenosine deaminase.
65. The deaminase of any one of claims 41-43, wherein the deaminase comprises an N46L mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues R48P, R64K and S73P, or corresponding mutations in homologous adenosine deaminases.
66. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO: 39, wherein the deaminase comprises mutations at Q71S and H154Q in the amino acid sequence of SEQ ID NO: 39, or corresponding mutations in homologous adenosine deaminases.
67. A deaminase comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of SEQ ID NO:39, wherein the amino acid corresponding to residue 46 of SEQ ID NO:79 is any amino acid except N.
68. A deaminase comprising a mutation at residue T79 in the amino acid sequence of SEQ ID NO: 39 and further comprising at least one mutation at a residue selected from the group consisting of A28, N46, R48, S73, N96 and G105, or a corresponding mutation in a homologous adenosine deaminase.
69. The deaminase of claim 67 or 68, wherein the deaminase is capable of deaminating cytidine in DNA.
70. The deaminase of any one of claims 67-69, wherein the deaminase comprises at least one mutation selected from N79T or N79P in the amino acid sequence of SEQ ID NO: 39, and further comprises one or more mutations selected from the group consisting of N46L, N46V, N46I, R48, R48P, S73P, S73Y, S73H, N96H and G105S, or corresponding mutations in homologous adenosine deaminases.
71. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues N46L, S73P, N79T and N96H, or corresponding mutations in homologous adenosine deaminases.
72. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues N46L, S73P and N79T, or corresponding mutations in homologous adenosine deaminases.
73. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues R48A, S73P and N79T, or corresponding mutations in homologous adenosine deaminases.
74. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39 and further comprises a mutation at residue N46V, or a corresponding mutation in a homologous adenosine deaminase.
75. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues N46V and S73P, or corresponding mutations in homologous adenosine deaminases.
76. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues A28V, N46L, R48A, S73Y and N96H, or corresponding mutations in homologous adenosine deaminases.
77. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues N46I and S73P, or corresponding mutations in homologous adenosine deaminases.
78. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues N46V and S73P, or corresponding mutations in homologous adenosine deaminases.
79. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79P mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues R48P and S73H, or corresponding mutations in homologous adenosine deaminases.
80. The deaminase of any one of claims 67-69, wherein the deaminase comprises an N79T mutation in the amino acid sequence of SEQ ID NO: 39, and further comprises mutations at residues A28V, N46I, R48A and S73Y, or corresponding mutations in homologous adenosine deaminases.
81. A deaminase comprising mutations at Q71S and H154Q in the amino acid sequence of SEQ ID NO: 39, or corresponding mutations in homologous adenosine deaminases.
82. The deaminase of any one of claims 41-81, wherein the cytidine deamination activity of the deaminase exceeds the cytidine deamination activity of TadA-Dual.
83. The deaminase of any one of claims 41-82, wherein the deaminase has a ratio of adenosine deamination activity to cytidine deamination activity of at least about 0.001:1, 0.005:1, 0.007:1, 0.01:1, 0.05:1, 0.07:1, or 0.1:
1.
84. A cytidine deaminase which has evolved from an adenosine deaminase by continuous evolution and / or discontinuous evolution.
85. The cytidine deaminase of claim 84, wherein the adenosine deaminase comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 99.5% identical to the amino acid sequence of any one of SEQ ID NOs: 34-39, 41-54, 33, 315, 317-323, 326, 354 and 355.
86. A base editor comprising a nucleic acid programmable DNA binding protein (napDNAbp) domain and a TadA-CD domain comprising a deaminase according to any one of claims 1-85.
87. A base editor according to claim 86, wherein the napDNAbp domain is a nickase.
88. A base editor according to claim 86 or 87, wherein the napDNAbp domain is selected from SpCas9n, dCas9, CasX, CasY, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, Cas9-NG, LbCas12a, enAsCas12a, SaCas9, SaCas9-KKH, circularly arranged Cas9, Argonaute (Ago) domain, SmacCas9, Spy-macCas9, SpCas9-VRQR, SpCas9-NRRH, SpaCas9-NRTH, SpCas9-NRCH, eNme2Cas9, eNme2-C Cas9, enCjCas9, SauriCas9, Cas9-NG-VRQR and variants thereof.
89. A base editor according to any one of claims 86-88, wherein the napDNAbp domain comprises an amino acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 74-77, 343-346, 347, 348, 351-353 and 356-358.
90. The base editor of any one of claims 86-88, wherein the napDNAbp domain is a Cas9n or SpCas9-NG domain.
91. A base editor according to any one of claims 86-90, wherein the napDNAbp domain comprises the amino acid sequence shown in SEQ ID NO:77 or 343.
92. The base editor of any one of claims 86-91, wherein the napDNAbp domain is an eNme2-C Cas9 domain.
93. A base editor according to any one of claims 86-92, wherein the napDNAbp domain is an enCjCas9 domain.
94. The base editor of any one of claims 86-93, wherein the napDNAbp domain is a SaCas9 domain.
95. A base editor according to any one of claims 91-94, wherein the napDNAbp domain comprises an amino acid sequence shown in any one of SEQ ID NOs: 347, 348 and 353.
96. A base editor according to any one of claims 86-95, wherein the base editor provides an efficiency of at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75% or 80% conversion of cytosine to thymine when contacted with DNA comprising a target sequence selected from the group consisting of CTT, CTC, CTA, CTG, CCT, CCC, CCA, CCG, CAT, CAC, CAA, CAG, CGT, CGC, CGA, CGG, TCT, TCC, TCG, ACT, ACC, ACA, ACG, GCT, GCC, GCA, GCG, TTC, TAC, TGC, ATC, AAC, AGC, GTC, GAC and GGC.
97. The base editor of any one of claims 86-96, wherein the base editor causes an off-target editing frequency in the range of about 0.1% to about 0.35%.
98. The base editor of any one of claims 86-97, wherein the base editor further comprises one or more UGI domains.
99. The base editor of any one of claims 86-98, wherein the base editor further comprises two UGI domains.
100. The base editor of any one of claims 86-99, wherein the base editor comprises one or more nuclear localization sequences (NLS).
101. The base editor of any one of claims 86-100, wherein the base editor further comprises a bipartite nuclear localization signal (bpNLS).
102. A base editor according to claim 101, wherein the bipartite nuclear localization signal comprises an amino acid sequence selected from the group consisting of: KRTADGSEFEPKKKRKV (SEQ ID NO: 155), KRPAATKKAGQAKKKK (SEQ ID NO: 276), KKTELQTTNAENKTKKL (SEQ ID NO: 277), KRGINDRNFWRGENGRKTR (SEQ ID NO: 278) and RKSGKIAAIVVKRPRK (SEQ ID NO: 279).
103. A base editor according to claim 101 or 102, wherein the bipartite nuclear localization signal comprises the amino acid sequence shown in SEQ IDNO:276 or 155.
104. A base editor according to any one of claims 86-103, wherein the base editor comprises the following structure: NH2-[first nuclear localization sequence]-[TadA-CD domain]-[napDNAbp domain]-[first UGI domain]-[second UGI domain]-[second nuclear localization sequence]-COOH, wherein each instance of "]-[" indicates the presence of an optional linker sequence; optionally, wherein the base editor comprises the following structure: NH2-[first NLS]-[TadA-CD domain]-[SaCas9n]-[UGI domain]-[UGI domain]-[second NLS]-COOH; NH2-[first NLS]-[TadA-CD domain]-[eNme2-C Cas9n]-[UGI domain]-[UGI domain]-[second NLS]-COOH; NH2-[first NLS]-[TadA-CD domain]-[CjCas9n]-[UGI domain]-[UGI domain]-[second NLS]-COOH; or NH2-[first NLS]-[TadA-CD domain]-[SpCas9-NG]-[UGI domain]-[UGI domain]-[second NLS]-COOH.
105. A base editor according to any one of claims 86-104, wherein the TadA-CD domain and the napDNAbp domain are connected via a linker comprising the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 165); the napDNAbp domain and the first UGI domain are connected via a linker comprising the amino acid sequence of SGGSGGSGGS (SEQ ID NO: 166); the first UGI domain and the second UGI domain are connected via a linker comprising the amino acid sequence of SGGSGGSGGS (SEQ ID NO: 166); and / or the second UGI domain and the second nuclear localization sequence are connected via a linker sequence comprising the amino acid sequence of SGGS (SEQ ID NO: 160).
106. A base editor according to any one of claims 86-105, wherein the base editor comprises an amino acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 19-31.
107. A base editor according to any one of claims 86-106, wherein the base editor comprises any one of the amino acid sequences shown in SEQ ID NO:19-31.
108. A complex comprising a base editor according to any one of claims 86-107 and a guide RNA bound to the napDNAbp domain of the base editor.
109. The complex of claim 108, wherein the guide RNA is 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, or 200 nucleotides in length.
110. The complex of claim 108 or 109, wherein the guide RNA is 15 to 80 nucleotides in length and comprises a sequence of at least 10, at least 15, or at least 20 consecutive nucleotides that are complementary to a target sequence.
111. The complex of any one of claims 108-110, wherein the guide RNA comprises a sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 or 40 consecutive nucleotides that are complementary to the target sequence.
112. The complex of claim 110 or 111, wherein the target sequence is a DNA sequence.
113. The complex of any one of claims 110-112, wherein the target sequence is located in the genome of an organism.
114. The complex of claim 113, wherein the organism is a prokaryotic organism.
115. The complex of claim 114, wherein the prokaryotic organism is a bacterium.
116. The complex of claim 113, wherein the organism is a eukaryotic organism.
117. The complex of claim 116, wherein the eukaryotic organism is a plant or a fungus.
118. The complex of claim 116, wherein the eukaryotic organism is a mammal.
119. The complex of claim 118, wherein the mammal is a rodent.
120. The complex of claim 118, wherein the mammal is a human.
121. The complex of any one of claims 110-120, wherein the target sequence is located in the genome of a cell.
122. The complex of claim 121, wherein the cell is a plant cell, a rodent cell, or a human cell.
123. The complex of claim 121 or 122, wherein the cell is a T cell or a hematopoietic stem cell (HSC).
124. A polynucleotide encoding a base editor according to any one of claims 86-107.
125. The polynucleotide of claim 124, wherein the polynucleotide is codon optimized for expression in human cells.
126. The polynucleotide of claim 124 or 125, wherein the polynucleotide is codon optimized for expression in a mammalian cell.
127. A vector comprising the polynucleotide according to any one of claims 124-126.
128. The vector of claim 127, wherein the vector comprises a heterologous promoter driving expression of the polynucleotide.
129. The vector of claim 127 or 128, further comprising a polynucleotide encoding a guide RNA (gRNA).
130. The vector of any one of claims 127-129, wherein the vector is an mRNA construct.
131. The vector of any one of claims 127-130, wherein the vector is a recombinant AAV vector.
132. A vector according to claim 129, wherein the orientation of the polynucleotide encoding the gRNA is reversed relative to the polynucleotide according to any one of claims 124-126.
133. A recombinant adeno-associated virus (rAAV) particle comprising the AAV vector of any one of claims 127-132.
134. A cell comprising a deaminase according to any one of claims 1-85, a base editor according to any one of claims 86-107, a complex according to any one of claims 108-123, a polynucleotide according to any one of claims 124-126, a vector according to any one of claims 127-132, or an rAAV particle according to claim 133.
135. The cell of claim 134, wherein the cell is a T cell.
136. The cell of claim 134, wherein the cell is a stem cell.
137. The cell of claim 134, wherein the cell is a human hematopoietic stem cell (HSC).
138. A cell according to any one of claims 134-137, wherein the cell has been obtained from a subject and contacted ex vivo with a base editor according to any one of claims 86-107, a complex according to any one of claims 108-123, or a vector according to any one of claims 127-132.
139. A pharmaceutical composition comprising a base editor according to any one of claims 86-107, a complex according to any one of claims 108-123, or a vector according to any one of claims 127-132.
140. The pharmaceutical composition of claim 139, further comprising a pharmaceutically acceptable excipient.
141. A method comprising contacting a nucleic acid with a base editor according to any one of claims 86-107 or a complex according to any one of claims 108-123.
142. The method of claim 141, wherein the nucleic acid comprises a target sequence in the genome of a cell.
143. The method of claim 141 or 142, wherein the nucleic acid is DNA.
144. The method of any one of claims 141-143, wherein the nucleic acid is double-stranded DNA.
145. The method of any one of claims 142-144, wherein the target sequence comprises a sequence associated with a disease or disorder.
146. The method of any one of claims 142-145, wherein the target sequence comprises a sequence in the BCL11A enhancer.
147. The method of any one of claims 142-146, wherein the target sequence comprises a sequence in a CCR5 or CXCR4 gene.
148. The method of any one of claims 142-147, wherein the target sequence comprises a point mutation associated with a disease or disorder.
149. The method of any one of claims 141-148, wherein the activity of the base editor or the complex results in correction of the point mutation.
150. The method of any one of claims 142-149, wherein the target sequence comprises a T to C point mutation associated with a disease or condition, and wherein deamination of the mutant C base results in a sequence not associated with the disease or condition.
151. The method of any one of claims 142-150, wherein the target sequence comprises an A to G point mutation associated with a disease or condition, and wherein deamination of a mutant C base complementary to the G base of the A to G point mutation results in a sequence not associated with the disease or condition.
152. The method of claim 150 or 151, wherein deamination of the mutant C results in a change in the amino acid encoded by the mutant codon.
153. The method of any one of claims 150-152, wherein deamination of the mutant C results in a codon encoding a wild-type amino acid.
154. The method of claim 151, wherein deamination of the C base complementary to the G base of the A to G point mutation results in a change in the amino acid encoded by the mutant codon.
155. The method of any one of claims 150-154, wherein the deamination results in the introduction of a stop codon.
156. The method of any one of claims 150-154, wherein the deamination results in the removal of a stop codon.
157. The method of claim 155 or 156, wherein the stop codon comprises the nucleic acid sequence 5'-TAG-3', 5'-TAA-3' or 5'-TGA-3'.
158. The method of any one of claims 150-157, wherein the deamination results in the introduction of a splice site.
159. The method of any one of claims 150-157, wherein the deamination results in removal of a splice site.
160. The method of any one of claims 150-157, wherein the deamination results in the introduction of a mutation in a gene promoter.
161. The method of claim 160, wherein the mutation results in increased transcription of a gene operably linked to the gene promoter.
162. The method of claim 160 or 161, wherein the mutation results in a reduction in transcription of a gene operably linked to the gene promoter.
163. The method of any one of claims 150-162, wherein the deamination results in the introduction of a mutation in a gene repressor.
164. The method of claim 163, wherein the mutation results in increased transcription of a gene operably linked to the gene repressor.
165. The method of claim 163 or 164, wherein the mutation results in a reduction in transcription of a gene operably linked to the gene repressor.
166. The method of any one of claims 142-165, wherein the target sequence encodes a protein, and wherein the point mutation is located in a codon and results in a change in the amino acid encoded by the mutant codon compared to the wild-type codon.
167. The method of any one of claims 141-166, wherein the contacting step is performed in a subject.
168. The method of any one of claims 141-167, wherein the contacting step is performed in vitro or ex vivo.
169. The method of claim 167, wherein the subject has been diagnosed with a disease or condition.
170. The method of claim 169, wherein the disease or condition is HIV / AIDS or sickle cell disease.
171. The method of any one of claims 142-170, wherein the target sequence comprises a DNA sequence 5'-NCN-3', wherein N is A, T, C or G.
172. The method of claim 171, wherein the C in the center of the 5'-NCN-3' sequence is deaminated.
173. The method of claim 171 or 172, wherein the C in the center of the 5'-NCN-3' sequence is changed to T.
174. The method of any one of claims 141-174, wherein the method results in less than 20%, 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1%, 0.5%, 0.2% or 0.1% indel formation.
175. The method of any one of claims 141-174, wherein the method provides a ratio of cytidine deamination activity to adenosine deamination activity of at least about 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 17:1, 19:1, 20:1, 21:1, 23:1, 25:1, 30:1, or greater than 30:
1.
176. The method of any one of claims 141-175, wherein the deaminase has a ratio of cytidine deamination activity to adenosine deamination activity of at least about 10:
1.
177. The method of any one of claims 141-176, wherein the deaminase has a ratio of cytidine deamination activity to adenosine deamination activity of at least about 20:
1.
178. The method of any one of claims 141-177, wherein the method provides an off-target editing frequency of less than 1%, less than 0.75%, less than 0.5%, less than 0.4%, less than 0.35%, less than 0.25%, less than 0.2%, less than 0.15%, or less than 0.1%.
179. The method of any one of claims 141-178, wherein the method provides an off-target editing frequency of about 0.35% or less.
180. The method of any one of claims 141-179, wherein the method results in a ratio of on-target:off-target editing of about 25:1, 50:1, 65:1, 75:1, 80:1, 85:1, 90:1, 95:1, 100:1, 110:1, 125:1, or more than 125:
1.
181. The method of any one of claims 141-180, wherein the method results in a ratio of on-target:off-target editing of about 90:1 or higher in the CXCR4 or CCR5 gene.
182. The method of any one of claims 141-181, wherein the target sequence comprises a target window, wherein the target window comprises a target nucleotide base pair.
183. The method of claim 182, wherein the target window is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides in length.
184. A method comprising administering to a subject a vector according to any one of claims 127-132, a cell according to any one of claims 134-138, or a pharmaceutical composition according to claim 139 or 140.
185. The method of claim 184, wherein the subject is a mammal.
186. The method of claim 184 or 185, wherein the subject is a human.
187. The method of any one of claims 184-186, wherein the administering step comprises ex vivo engineering the cells of any one of claims 134-138 and administering the cells to the subject.
188. A kit comprising a nucleic acid construct comprising: (a) a nucleic acid sequence encoding a base editor according to any one of claims 86-107; (b) a nucleic acid sequence encoding a gRNA; and (c) one or more heterologous promoters driving expression of the sequence of (a) and / or the sequence of (b).
189. The kit of claim 188, further comprising an expression construct encoding a guide RNA backbone, wherein the construct comprises a cloning site positioned at a position that allows a nucleic acid sequence identical or complementary to a target sequence to be cloned into the guide RNA backbone.
190. Use of (a) a base editor according to any one of claims 86-107 and (b) a guide RNA that targets the base editor of (a) to a target C:G nucleotide base pair in a double-stranded DNA molecule in DNA editing.
191. Use of a base editor according to any one of claims 86-107, a complex according to any one of claims 108-123, a cell according to any one of claims 134-138, or a pharmaceutical composition according to claim 130 or 140 as a medicine.
192. Use of the base editor of any one of claims 86-107, the complex of any one of claims 108-123, the cell of any one of claims 134-138, or the pharmaceutical composition of claim 130 or 140 as a drug for treating sickle cell disease.
193. A vector system comprising: (i) selecting a plasmid comprising an isolated nucleic acid encoding an adenosine deaminase, which comprises, in the following order: an adenosine deaminase protein and a sequence encoding the N-terminal portion of a split intein; (ii) a first helper plasmid comprising, in the following order: a sequence encoding a guide RNA operably controlled by a Lac promoter and a sequence encoding an M13 bacteriophage gene III (gIII) peptide operably controlled by a T7 RNA promoter; (iii) a second helper plasmid comprising, in the following order: a sequence encoding the C-terminal portion of the split intein and a sequence encoding a dCas9-UGI fusion; and (iv) a third helper plasmid comprising a non-coding strand and a coding strand, wherein the coding strand comprises an expression construct, and the expression construct comprises, in the following order: a promoter, a ribosome binding site, and a sequence encoding T7 RNA polymerase and a degradation determinant tag, wherein the non-coding strand opposite to the 3' end of the sequence encoding T7 RNA polymerase comprises a CAA sequence.
194. The vector system of claim 193, wherein the split intein is an Npu (N. punctata) intein.
195. The vector system of claim 193 or 194, wherein the adenosine deaminase is TadA-8e.
196. The vector system of any one of claims 193-195, further comprising a mutagenic plasmid.
197. A cell comprising the vector system according to any one of claims 193-196.
198. A method for selecting a cytidine deaminase evolved from an adenosine deaminase using one or more rounds of PACE or PANCE evolution.
199. The method of claim 198, wherein the one or more rounds of PACE or PANCE evolution comprise: Selected phage encoding a mutant TadA8e protein fused to the NpuN intein, a. A first plasmid encoding the NpuC intein fused to dCas9-UGI, b. a second plasmid encoding gene III (gIII) driven by a T7 or proT7 promoter and encoding sgRNA, and c. A third plasmid encoding a T7 RNA polymerase-degradin fusion.
200. The method of claim 199, wherein the T7 RNA polymerase-degradant fusion contains a target sequence at the interface between the T7 RNA polymerase and the degron domains.
201. The method of any one of claims 198-200, wherein the target sequence contains one or more cytosine nucleotides that, when edited to thymine, insert a stop codon between the T7 RNA polymerase and the degron domains of the T7 RNA polymerase-degron fusion.
Citation Information
Patent Citations
Transgenic animals secreting desired proteins into milk
EP0264166A1
CAS9 proteins including ligand-dependent inteins
US10077453B2
Adenosine nucleobase editors and uses thereof
US10113163B2
Nucleobase editors and uses thereof
US10167457B2
Process and device for inerting an aircraft fuel tank
US20020158167A1
Cited By
Saccharopolyspora sourced cytosine deaminase and basic group editing system thereof aiming at chicken genome
CN121294410A