Methods and compositions for improving genome editing specificity and efficiency of cytosine base editors
Patent Information
- Application Number
- JP2025532902
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-06
- Filing Date
- 2023-12-05
- Publication Date
- 2026-02-17
AI Technical Summary
Current cytosine base editors face issues such as bystander editing, off-target effects, and unwanted excision of uracil intermediates, leading to DNA damage and reduced editing efficiency, particularly when delivered via RNP complexes.
The use of single-stranded DNA-binding domains (SSBs) covalently linked to cytosine base editors (CBEs) to narrow the editing window and improve specificity, combined with free-standing uracil glycosylase inhibitors (UGIs) and a sulfonated polysaccharide-based RNP electroporation enhancer to enhance editing efficiency and reduce off-target effects.
This approach significantly reduces bystander editing, increases editing specificity, and enhances efficiency, allowing precise C-to-T conversions without increasing off-target effects, even at difficult-to-edit sites.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 386,187, filed December 6, 2022, the entire contents of which are incorporated herein in their entirety. [Background technology]
[0002] Targeted genome modification is a powerful tool for genetic manipulation of DNA, including the engineering of eukaryotic cells, embryos, and animals. For example, exogenous sequences can be incorporated into targeted genomic locations and / or specific endogenous DNA (e.g., chromosomal) sequences can be deleted, inactivated, or modified. The CRISPR-Cas9 genome editing system has become a widespread method for introducing genetic modifications into diverse cell types and organisms. While simple double-stranded DNA breaks by Cas9 enable gene inactivation through the formation of insertions and deletions, precise editing can be achieved through homology-directed repair (HDR) or base editing. HDR often poses significant challenges in introducing the desired edit; donor DNA can be toxic in some cell types; and the requirement for double-stranded DNA breaks (DSBs) poses the risk of unintended, deleterious repair consequences. In contrast, base editing utilizes a deaminase domain to introduce C-to-T or A-to-G substitutions at target sites without the need for donor DNA. In cytosine base editors (CBEs), cytosine bases are deaminated to uracil, which is "read" as thymine by DNA polymerase, resulting in the introduction of a cytosine-to-thymine substitution. Additionally, base editing utilizes CRISPR effectors such as Cas9 and Cas12a, which have fully or partially inactive DNA cleavage functions, which prevent the formation of DSBs.
[0003] Since its initial publication by Komor et al. (Nature, 2016), cytosine base editing has been widely used in diverse organisms, from bacteria to humans, and is being investigated for the correction of pathogenic mutations in therapeutic contexts. Komor's cytosine base editing system utilizes a catalytically inactive Cas9 (dCas9) containing Asp10Ala and His840Ala mutations that inactivate its nuclease activity while still retaining its ability to bind to DNA in a guide RNA-programmed manner without cleaving the DNA backbone. As previously mentioned, cytosine deamination is catalyzed by cytosine deaminase, resulting in the conversion of cytosine to uracil. Uracil has the base-pairing properties of thymine and, upon replication or repair, creates an AT base pair.
[0004] However, this editing technology still has certain drawbacks that need to be addressed. One drawback is so-called "bystander editing," in which adjacent cytosine residues other than the intended residue in the R-loop are converted to thymine, causing off-target effects. Several efforts to address this issue have previously been attempted with some success. These efforts include the use of a strong peptide linker between the deaminase and Cas9 nickase to narrow the editing window and thus minimize bystander editing (Tan et al., Nat. Commun. 2019) and protein engineering on the deaminase domain to restrict C-to-T editing to certain nucleotide motifs (Gehrke et al., Nat. Biotechnol. 2018; Kim et al., Nat. Biotechnol. 2017) or to reduce the frequency of editing at multiple cytosine residues within the editing window (Jin et al., Molecular Cell 2020). However, each of these approaches has its own limitations.
[0005] Another issue associated with cytosine base editing is the unwanted excision of uracil intermediates during the C-to-T conversion process. Uracil does not naturally occur in DNA. The most common causes of deoxyuridine in DNA are misincorporation in place of deoxythymidine and spontaneous deamination of cytosine, both of which are mutagenic. Therefore, the presence of deoxyuracil triggers a DNA damage response by uracil DNA glycosylase, resulting in the excision of the uracil base. Translesion synthesis at abasic sites can introduce unwanted cytosine-to-adenine or cytosine-to-guanine substitutions. Abasic sites are also prone to DNA strand breaks, which, in the context of a CBE containing Cas9 nickase activity, can result in double-stranded DNA breaks (DSBs), frequently resulting in DNA insertions and deletions. To address these unwanted repair events that can occur during cytosine base editing, the second-generation cytosine base editor (BE2) (Komor et al., Nature, 2016) possesses a phage-derived uracil glycosylase inhibitor (UGI) covalently attached to the C-terminus of Cas9, which has been shown to improve C-to-T conversion rates and overall editing efficiency compared to the first-generation cytosine base editor (BE1) that does not contain a UGI (Komor et al., Nature, 2016). An additional UGI was added to the fourth-generation cytosine base editor (BE4) to enhance efficacy (Komor et al., Sci. Adv., 2017). Currently, nearly all cytosine base editors contain two UGI-like structures linked in tandem to the C-terminus of the Cas effector. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Komor et al., Nature, 2016 [Non-patent document 2] Tan et al., Nat. Commun. 2019 [Non-patent document 3] Gehrke et al., Nat Biotechnol, 2018 [Non-patent document 4] Kim et al., Nat Biotechnol, 2017 [Non-patent document 5] Jin et al., Molecular Cell, 2020 [Non-patent document 6] Komor et al., Sci.Adv. 2017 Summary of the Invention [Problem to be solved by the invention]
[0007] Plasmid delivery of base editors and single guide RNAs (sgRNAs) is currently common practice in base editing. However, plasmid delivery carries the risks of increased off-target effects due to prolonged overexpression of the base editor, toxicity in some cell types due to sensitivity to exogenous DNA, and potential integration of plasmid DNA into the host genome. Ribonucleoprotein (RNP) complex delivery of base editor recombinant proteins and chemically synthesized sgRNAs is thought to greatly reduce or eliminate these risks. However, expression and purification of base editor recombinant proteins from E. coli hosts can be challenging because elevated deaminase levels can be toxic to bacterial cells. To overcome this barrier, some have relied on the use of immortalized human HEK93 cells (Jiang et al., Sci. Adv. 2021). In addition to protein purification challenges, editing rates can be lower following RNP complex delivery compared to plasmid delivery. Therefore, there is a need for base editor proteins that can be produced on a commercial scale that result in a high rate of cytosine to thymine substitutions, deaminate in a relatively narrow window, and allow for the introduction of precise cytosine to thymine substitutions without bystander editing. [Means for solving the problem]
[0008] (Summary of the Invention) Among various aspects of the present invention are compositions and methods that substantially improve cytosine base editing specificity by using an ssDNA-binding domain (SSB) to limit bystander editing within R-loops. Bystander editing converts adjacent cytosine residues within an R-loop other than the intended cytosine to thymine, but this significantly limits the use of cytosine base editors as precise gene editing tools for correcting pathogenic mutations. Without wishing to be bound by theory, we hypothesize that SSBs covalently linked to cytosine base editors (CBEs) may compete with the deaminase domain for binding to specific portions of the single-stranded DNA region formed after Cas9 target binding and R-loop formation, thereby confining deaminase activity to a narrower segment on the single-stranded DNA region. SSBs are ubiquitous and vary widely in size and DNA-binding footprints. For proof-of-concept, we fused T4 phage-derived SSB (SEQ ID NO: 1) or T7 phage-derived SSB (SEQ ID NO: 2) to the N-terminus of human APOBEC3A or APOBEC3B deaminase, which in turn was covalently linked to the N-terminus of SpCas9 nickase.
[0009] In some embodiments, it is contemplated that a fusion protein may contain one or more UGIs at the C-terminus of the fusion protein. In other embodiments, one or more UGIs may be free-standing, i.e., not fused to the fusion protein, but may be separate from the fusion protein. In yet other embodiments, UGI(s) fused to a fusion protein are specifically excluded from the present invention. In still other embodiments, a UGI may be both fused and free-standing. In still other embodiments, the present invention contemplates methods and materials for use with cytosine base editing that exclude one or more of the fused and / or free-standing UGIs.
[0010] To our knowledge, this is the first successful use of the free-standing UGI protein in cytosine base editing applications. Other attempts have been made by Jang et al. (Science Advances (2021) "High purity production and precise editing of DNA base editing ribonucleoproteins" DOI: 10.1126 / sciadv.abg2661) and Wang et al. (Cell Research (2017) "Enhanced base editing by co-expression of free uracil DNA glycosylase inhibitor" DOI: 10.1038 / cr.2017.111). However, in these papers, UGI was overexpressed from a plasmid.
[0011] The recombinant proteins were expressed and purified from E. coli and tested for C-to-T base editing in human immortalized HEK293 cells using chemically synthesized sgRNA via RNP complex delivery. Compared to their non-SSB counterparts, all tested SSB-CBE fusion proteins improved editing specificity by narrowing the editing window by at least one residue. The results demonstrate that this novel strategy can reduce CBE bystander editing and therefore increase editing specificity. Compared to previous methods that rely on deaminase mutagenesis to bias editing toward specific DNA sequence motifs or robust peptide linkers to narrow the editing window, this novel strategy utilizes SSBs with distinct DNA binding footprints and affinities to tailor the size of the editing window to specific needs. Those skilled in the art will be able to identify appropriate SSBs for use with the present invention based on the guidance herein (see, e.g., Guo and Malik et al., Biomolecules, 2022, 12(9), 1187).
[0012] Another aspect of the present invention discloses CBE protein fusion compositions that improve editing efficiency over conventional CBEs when C-to-T editing is performed by delivery of preassembled RNP complexes. As disclosed in the prior art, conventional CBEs contain at least one covalently attached UGI as a means of improving editing efficiency over UGI-free CBEs. This teaching is based on delivery of CBE-encoding plasmid DNA (Komor et al., Nature, 2016; Komor et al., Sci. Adv., 2017). However, we reasoned that a covalently attached UGI at the C-terminus of the CBE could actually interfere with target binding by CRISPR effector nucleases and reduce editing efficiency in the absence of overexpression by plasmid DNA. To test this hypothesis, we constructed paired CBEs with and without a covalently attached UGI and purified the recombinant proteins from E. coli. After transfecting these proteins into human immortalized HEK293 cells in combination with chemically synthesized sgRNA in RNP complexes, we indeed found that CBEs without UGIs were more efficient at converting C to T than CBEs with covalently linked UGIs on all targets tested. Thus, we have uncovered a novel phenomenon that contrasts with the teachings of the prior art and established an improved CBE protein fusion composition that enhances cytosine base editing via RNP complex delivery.
[0013] We further demonstrated that the use of a free-standing recombinant UGI can further improve CBE editing efficiency and specificity, regardless of whether the CBE contains a covalently linked UGI. We engineered a recombinant UGI (SEQ ID NO: 3) containing a bacillus phage-derived UGI, a c-MYC nuclear localization signal (NLS), and an SV40 large T antigen NLS, and purified the protein from E. coli. Cotransfection of recombinant UGI and recombinant CBE substantially increased C-to-T editing efficiency. Furthermore, UGI cotransfection improved editing product purity by reducing C-to-G / A transversions and indel formation. Although UGI efficiency was evident regardless of whether the CBE contained a covalently linked UGI, a UGI-free CBE is preferred because it achieves optimal editing results. The respective enhancing effects of a free-standing UGI and a CBE without a covalently linked UGI appear to be additive to some extent. To our knowledge, the composition and use of this novel recombinant UGI protein as an effective enhancer have not been reported elsewhere.
[0014] Another embodiment of the present invention provides a method for enhancing CBE RNP editing efficiency using a sulfonated polysaccharide-based RNP electroporation enhancer. It is well recognized that RNP complex delivery of gene editing reagents can reduce off-target effects and help eliminate the risk of plasmid DNA-induced cytotoxicity and random genome integration compared to plasmid DNA delivery. However, editing efficiency with RNP complex delivery can be lower than with plasmid DNA delivery due to rapid protein turnover in the absence of overexpression. To address this issue in base editing, we reasoned that the dextran sulfate-based SpCas9 nuclease enhancer (MilliporeSigma Product No. PEXBUF, Burlington, MA) could also be effective with CBE. Indeed, we found that this enhancer can improve C-to-T editing efficiency by several-fold, particularly at difficult-to-edit target sites. Furthermore, this enhancer can be used in combination with CBE with or without a covalently attached UGI. However, it is preferable to use this enhancer in conjunction with a free-standing UGI. This enhancer may also enhance other editing events, such as C to G / A transversions and indel formation, in the absence of a free-standing UGI. This unique combination of an RNP enhancer and a recombination UGI has not been disclosed elsewhere.
[0015] The present invention further discloses a method for extending C-to-T editing at cytosine residues that are inaccessible using conventional methods. These cytosine residues are located at the 5' end, approximately 18 or more nucleotides upstream from the protospacer adjacent motif (PAM). In the context of a conventional 20-nt guide RNA spaser, the cytosine residues are at the end of an R-loop or in a base-paired state, making them inaccessible to CBE. We found that by extending the guide RNA spacer on a synthetic sgRNA from the conventional 18 nt to 19 nt, or from 20 nt to 21 nt, 22 nt, 23 nt, or more, cytosine residues located 18 nt to 23 nt from the PAM can be efficiently converted to thymine residues by CBE. We believe that the use of guide RNA spacers longer than 23 nt may further expand the editing range. We also found that spacer extension significantly shifts the editing window proportionally from the PAM without significantly increasing the size of the editing window. Therefore, this strategy provides a means to extend the CBE editing range on each available PAM, increasing overall genome coverage. It is contemplated that this strategy may also be applied to crRNA (CRISPR RNA) design and tracrRNA (transactivating crRNA). The crRNA directs the Cas effector nuclease to bind to DNA or RNA sequences that are complementary to the "guide" portion of the crRNA and adjacent to the protospacer adjacent motif (PAM) (for DNA targets) or sequences without extensive complementarity to the crRNA repeats (for RNA targets).
[0016] In summary, this disclosure describes novel compositions and methods for improving CBE specificity and efficiency using RNP complex delivery. Additionally, these novelties may also be utilized in the context of mRNA delivery or plasmid DNA delivery. Furthermore, these novelties may be utilized individually or in various combinations without departing from the principles of the present invention.
[0017] While the present invention contemplates the use of cytosine deaminase in the compositions and methods of the present invention, the present invention also contemplates the use of adenosine deaminase in the compositions and methods of the present invention. Similar to the way cytosine base editors convert cytosine (C)-guanine (G) binding pairs to thymidine (T)-adenine (A) binding pairs, adenosine base editors convert AT binding pairs to GC binding pairs. Adenosine deaminases are known to those of skill in the art. See, e.g., Gaudelli et al., 2017, Nature, 551, pp. 464-471; Richter et al., 2020, Nature Biotechnology, 38, pp. 883-891. Those of skill in the art will be able to use adenosine deaminase in the present invention without undue experimentation in light of the teachings herein.
[0018] Furthermore, while the present invention contemplates the use of T4 and / or T7 bacteriophage SSBs, the present invention also contemplates the use of other SSBs. SSBs are known to be found in all kingdoms of life (Antony & Lohman, 2019, Semin Cell Dev Biol, 86:102-111). SSBs have structural diversity, and these structures are known to those of skill in the art (see Antony & Lohman). Furthermore, the physiological mechanisms by which SSBs function are also known to those of skill in the art. For example, the functional structure of SSBs is the oligonucleotide / oligosaccharide-binding (OB) fold, a protein domain that binds to ssDNA and facilitates various protein-protein interactions. One OB fold consists of a five-stranded β-sheet that wraps around to form a closed β-barrel, with an α-helix capped at the end between the third and fourth β-strands. Based on the number of OB-folds, SSBs can be classified into two groups: simple SSBs containing only one OB-fold, and higher-order SSBs containing multiple OB-folds (see Antony & Lohman). SSBs may also or additionally contain K-homology (KH) domains, RNA recognition motifs (RRMs), and Worley domains, the structures and functions of which are known to those skilled in the art (see Guo and Malik, 2022, Biomolecules, 12:1187-1199, "for a detailed description of the OB-folds and the other domains / motifs").
[0019] In one aspect, the present invention contemplates a composition suitable for modifying cytosine residues in a DNA sequence, comprising a single-stranded DNA-binding domain (SSB), a deaminase, a catalytically modified Cas protein, a guide RNA, and optionally one or more free and / or fused uracil glycosylase inhibitors (UGIs), to form a base-editing RNP complex, optionally with a UGI. It is also contemplated that this system works with crRNA- and tracrRNA-based systems.
[0020] In another embodiment, the present invention contemplates that the SSB is directly linked to the deaminase.
[0021] In another embodiment, the present invention contemplates that the SSB is indirectly linked to the deaminase.
[0022] In another aspect, the present invention contemplates that the SSB is not covalently linked to any of the deaminase, catalytically modified Cas protein, guide RNA, or optional UGI.
[0023] In another embodiment, the present invention contemplates that the deaminase is covalently linked to a catalytically engineered Cas protein.
[0024] In another aspect, the present invention contemplates that the one or more UGIs are not covalently linked to any of the deaminase, catalytically modified Cas protein, or guide RNA.
[0025] In another embodiment, the present invention contemplates that one or more SSBs are of viral, prokaryotic, or eukaryotic origin.
[0026] In another aspect, the present invention contemplates that the deaminase is selected from a cytosine deaminase or an adenosine deaminase.
[0027] In another embodiment, the present invention contemplates that the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain.
[0028] In another aspect, the present invention contemplates that the catalytically engineered Cas protein comprises a catalytically inactive RuvC nuclease domain and a catalytically inactive HNH nuclease domain.
[0029] In another embodiment, the present invention contemplates that the catalytically engineered Cas protein is Cas9.
[0030] In another embodiment, the present invention contemplates that the catalytically engineered Cas protein is Cas12.
[0031] As discussed below, other RNA-guided endonucleases and partially or completely inactivated mutants thereof are contemplated.
[0032] In another aspect, the present invention contemplates at least one nucleic acid encoding one or more of the SSB, deaminase, catalytically modified Cas protein, guide RNA, and optionally one or more free and / or fused UGI of the compositions of the invention.
[0033] In another aspect, the present invention contemplates at least one expression vector comprising a nucleic acid encoding at least one or more of the SSB, deaminase, catalytically modified Cas protein, guide RNA, and optionally one or more free and / or fused UGIs of the compositions of the invention.
[0034] In another aspect, the present invention further contemplates a composition suitable for modifying a cytosine residue in a DNA sequence, the composition comprising a deaminase, a catalytically modified Cas protein, a guide RNA, and one or more free UGIs to form a base-editing RNP complex with the free UGIs.
[0035] In another embodiment, the present invention contemplates that the deaminase is covalently linked to a catalytically engineered Cas protein.
[0036] In another aspect, the present invention contemplates that the one or more UGIs are not covalently linked to any of the deaminase, catalytically modified Cas protein, or guide RNA.
[0037] In another aspect, the present invention contemplates that the deaminase is selected from a cytosine deaminase or an adenosine deaminase.
[0038] In another embodiment, the present invention contemplates that the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain.
[0039] In another aspect, the present invention contemplates that the catalytically engineered Cas protein comprises a catalytically inactive RuvC nuclease domain and a catalytically inactive HNH nuclease domain.
[0040] In another embodiment, the present invention contemplates that the catalytically engineered Cas protein is Cas9.
[0041] In another embodiment, the present invention contemplates that the catalytically engineered Cas protein is Cas12.
[0042] In another aspect, the present invention contemplates that the deaminase, catalytically modified Cas protein, and guide RNA are encoded by one or more nucleic acids, and the UGI is a peptide.
[0043] In another aspect, the present invention contemplates one or more expression vectors comprising one or more of the nucleic acids of the invention.
[0044] In another aspect, the present invention further contemplates a method for modifying a DNA sequence, comprising: 1) providing an RNP complex comprising a single-stranded DNA binding domain (SSB), a deaminase, a catalytically modified Cas protein, and a guide RNA, and optionally 2) one or more free uracil glycosylase inhibitors (UGIs); introducing the RNP complex and optionally the free UGIs into a recipient cell, wherein the DNA of the recipient cell has at least one cytosine residue converted to a thymine residue.
[0045] In another embodiment, the present invention contemplates a method in which SSB is directly linked to the deaminase.
[0046] In another embodiment, the present invention contemplates a method wherein the SSB is indirectly linked to the deaminase.
[0047] In another aspect, the present invention contemplates a method in which a deaminase is covalently linked to a catalytically engineered Cas protein.
[0048] In another aspect, the present invention contemplates a method wherein the one or more UGIs are not covalently attached to the deaminase, catalytically modified Cas protein, or guide RNA.
[0049] In another aspect, the present invention contemplates a method wherein one or more SSBs are of viral, prokaryotic, or eukaryotic origin.
[0050] In another aspect, the present invention contemplates a method wherein the deaminase is selected from a cytosine deaminase or an adenosine deaminase.
[0051] In another aspect, the present invention contemplates a method wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain.
[0052] In another aspect, the present invention contemplates a method, wherein the catalytically engineered Cas protein comprises a catalytically inactive RuvC nuclease domain and a catalytically inactive HNH nuclease domain.
[0053] In another aspect, the present invention contemplates a method wherein the catalytically engineered Cas protein is Cas9.
[0054] In another aspect, the present invention contemplates a method wherein the catalytically engineered Cas protein is Cas12.
[0055] In another aspect, the present invention contemplates a method in which the RNP complex and said UGI are introduced into a recipient cell as a recombinant protein.
[0056] In another aspect, the present invention contemplates a method in which the RNP complex and said UGI are introduced into a recipient cell as RNA.
[0057] In another aspect, the present invention contemplates a method wherein the RNP complex and said UGI are introduced into a recipient cell as at least one expression vector.
[0058] In another embodiment, the present invention contemplates a method in which a recombinant protein or nucleic acid is introduced into a recipient cell by electroporation in the presence of a sulfonated polysaccharide-based RNP electroporation enhancer.
[0059] In another aspect, the present invention contemplates a method wherein the guide RNA has a target-complementary sequence of 20-30 nucleotides in length.
[0060] In another aspect, the present invention contemplates a method for modifying a DNA sequence, comprising: 1) providing an RNP complex comprising a deaminase peptide, a catalytically modified Cas protein, and a guide RNA; and 2) one or more free uracil glycosylase inhibitors (UGIs); and introducing the RNP complex and the free UGIs into a recipient cell, wherein the DNA of the recipient cell has at least one cytosine residue converted to a thymine residue.
[0061] In another aspect, the present invention contemplates a method in which a deaminase is covalently linked to a catalytically engineered Cas protein.
[0062] In another aspect, the present invention contemplates a method wherein the one or more UGIs are not covalently attached to the deaminase, catalytically modified Cas protein, or guide RNA.
[0063] In another aspect, the present invention contemplates a method wherein the deaminase is selected from a cytosine deaminase or an adenosine deaminase.
[0064] In another aspect, the present invention contemplates a method wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain.
[0065] In another aspect, the present invention contemplates a method, wherein the catalytically engineered Cas protein comprises a catalytically inactive RuvC nuclease domain and a catalytically inactive HNH nuclease domain.
[0066] In another aspect, the present invention contemplates a method wherein the catalytically engineered Cas protein is Cas9.
[0067] In another aspect, the present invention contemplates a method wherein the catalytically engineered Cas protein is Cas12.
[0068] In another aspect, the present invention contemplates a method in which the RNP complex and UGI are introduced into a recipient cell as recombinant proteins.
[0069] In another embodiment, the present invention contemplates a method in which a recombinant protein is introduced into a recipient cell by electroporation in the presence of a sulfonated polysaccharide-based RNP electroporation enhancer.
[0070] In another aspect, the present invention contemplates a method wherein the guide RNA has a target-complementary sequence of 20-30 nucleotides in length.
[0071] In another aspect, the present invention contemplates a method for introducing UGI into a eukaryotic cell, the method comprising: a) providing at least one recipient cell; and b) at least one UGI covalently linked to at least one NLS, forming a UGI-NLS molecule; and transfecting the UGI-NLS molecule into the recipient cell.
[0072] The present invention further contemplates that the UGI-NLS molecule is a recombinant protein.
[0073] The present invention further contemplates that the UGI-NLS molecule is encoded by a nucleic acid.
[0074] The present invention contemplates compositions comprising at least one UGI covalently linked to at least one NLS to form a UGI-NLS molecule.
[0075] The present invention further contemplates a composition consisting essentially of at least one UGI covalently linked to at least one NLS to form a UGI-NLS molecule.
[0076] The present invention further contemplates that the UGI-NLS molecule is a recombinant protein.
[0077] The present invention further contemplates that the UGI-NLS molecule is encoded by a nucleic acid. [Brief explanation of the drawings]
[0078] [Figure 1A] Figure 1 shows the results of CBE with N-terminal SSB domain fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. (E) shows a graph of C to T substitutions that fall within a given window width. [Figure 1B] Figure 1 shows the results of CBE with N-terminal SSB domain fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. (E) shows a graph of C to T substitutions that fall within a given window width. [Figure 1C] Figure 1 shows the results of CBE with N-terminal SSB domain fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. (E) shows a graph of C to T substitutions that fall within a given window width. [Figure 1D] Figure 1 shows the results of CBE with N-terminal SSB domain fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. (E) shows a graph of C to T substitutions that fall within a given window width. [Figure 1E]Figure 1 shows the results of CBE with N-terminal SSB domain fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. (E) shows a graph of C to T substitutions that fall within a given window width. [Figure 2A] indicates the percentage of sequencing reads containing substitutions at each cytosine residue with and without 2xUGI fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. [Figure 2B] indicates the percentage of sequencing reads containing substitutions at each cytosine residue with and without 2xUGI fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. [Figure 2C] indicates the percentage of sequencing reads containing substitutions at each cytosine residue with and without 2xUGI fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. [Figure 2D] indicates the percentage of sequencing reads containing substitutions at each cytosine residue with and without 2xUGI fusions. (A) Target: EMX1-15, (B) Target: HEKSite2, (C) Target: HBB03, (D) Target: RNF2. [Figure 3A] Figure 1 shows (A) the percentage of sequencing reads containing insertions or deletions plotted for the four targets; (B) the percentage of cytosine substitutions that are C to T; and (C) the percentage of reads containing at least one C to T substitution within the protospacer. [Figure 3B] Figure 1 shows (A) the percentage of sequencing reads containing insertions or deletions plotted for the four targets; (B) the percentage of cytosine substitutions that are C to T; and (C) the percentage of reads containing at least one C to T substitution within the protospacer. [Figure 3C]Figure 1 shows (A) the percentage of sequencing reads containing insertions or deletions plotted for the four targets; (B) the percentage of cytosine substitutions that are C to T; and (C) the percentage of reads containing at least one C to T substitution within the protospacer. [Figure 4A] shows the percentage of reads with substitutions at each cytosine in the protospacer plotted from C to T (gray), C to G (white), and C to A (black). A-D contain data for minimal CBE variants, which do not contain an SSB domain: (A) target EMX1-15, (B) target HEKSite2, (C) target HBB03, (D) target RNF2. F-H contain data for CBE variants with an N-terminal SSB domain: (E) target EMX1-15; (F) target HEKSite2; (G) target HBB03; (H) target RNF2. [Figure 4B] shows the percentage of reads with substitutions at each cytosine in the protospacer plotted from C to T (gray), C to G (white), and C to A (black). A-D contain data for minimal CBE variants, which do not contain an SSB domain: (A) target EMX1-15, (B) target HEKSite2, (C) target HBB03, (D) target RNF2. F-H contain data for CBE variants with an N-terminal SSB domain: (E) target EMX1-15; (F) target HEKSite2; (G) target HBB03; (H) target RNF2. [Figure 4C] shows the percentage of reads with substitutions at each cytosine in the protospacer plotted from C to T (gray), C to G (white), and C to A (black). A-D contain data for minimal CBE variants, which do not contain an SSB domain: (A) target EMX1-15, (B) target HEKSite2, (C) target HBB03, (D) target RNF2. F-H contain data for CBE variants with an N-terminal SSB domain: (E) target EMX1-15; (F) target HEKSite2; (G) target HBB03; (H) target RNF2. [Figure 4D] shows the percentage of reads with substitutions at each cytosine in the protospacer plotted from C to T (gray), C to G (white), and C to A (black). A-D contain data for minimal CBE variants, which do not contain an SSB domain: (A) target EMX1-15, (B) target HEKSite2, (C) target HBB03, (D) target RNF2. F-H contain data for CBE variants with an N-terminal SSB domain: (E) target EMX1-15; (F) target HEKSite2; (G) target HBB03; (H) target RNF2. [Figure 4E] shows the percentage of reads with substitutions at each cytosine in the protospacer plotted from C to T (gray), C to G (white), and C to A (black). A-D contain data for minimal CBE variants, which do not contain an SSB domain: (A) target EMX1-15, (B) target HEKSite2, (C) target HBB03, (D) target RNF2. F-H contain data for CBE variants with an N-terminal SSB domain: (E) target EMX1-15; (F) target HEKSite2; (G) target HBB03; (H) target RNF2. [Figure 4F] shows the percentage of reads with substitutions at each cytosine in the protospacer plotted from C to T (gray), C to G (white), and C to A (black). A-D contain data for minimal CBE variants, which do not contain an SSB domain: (A) target EMX1-15, (B) target HEKSite2, (C) target HBB03, (D) target RNF2. F-H contain data for CBE variants with an N-terminal SSB domain: (E) target EMX1-15; (F) target HEKSite2; (G) target HBB03; (H) target RNF2. [Figure 4G]shows the percentage of reads with substitutions at each cytosine in the protospacer plotted from C to T (gray), C to G (white), and C to A (black). A-D contain data for minimal CBE variants, which do not contain an SSB domain: (A) target EMX1-15, (B) target HEKSite2, (C) target HBB03, (D) target RNF2. F-H contain data for CBE variants with an N-terminal SSB domain: (E) target EMX1-15; (F) target HEKSite2; (G) target HBB03; (H) target RNF2. [Figure 4H] shows the percentage of reads with substitutions at each cytosine in the protospacer plotted from C to T (gray), C to G (white), and C to A (black). A-D contain data for minimal CBE variants, which do not contain an SSB domain: (A) target EMX1-15, (B) target HEKSite2, (C) target HBB03, (D) target RNF2. F-H contain data for CBE variants with an N-terminal SSB domain: (E) target EMX1-15; (F) target HEKSite2; (G) target HBB03; (H) target RNF2. [Figure 5] shows the results of adding dextran sulfate to the co-transfection mixture of each CBE variant with UGI. [Figure 6A] Figure 1 shows the results of increasing the length of the sgRNA. Figures A-C show the percentage of reads with C to T substitutions for three targets and three spacer lengths: (A) target EMX1-15; (B) target RNF2; (C) target HBB03. Figures D-G show the percentage of reads carrying a single-presence editing allele: (D) target EMX1-15 without an SSB; (E) target RNF2 without an SSB; (F) target EMX1-15 with a T4 SSB; (G) target RNF2 with a T4 SSB. [Figure 6B]Figure 1 shows the results of increasing the length of the sgRNA. Figures A-C show the percentage of reads with C to T substitutions for three targets and three spacer lengths: (A) target EMX1-15; (B) target RNF2; (C) target HBB03. Figures D-G show the percentage of reads carrying a single-presence editing allele: (D) target EMX1-15 without an SSB; (E) target RNF2 without an SSB; (F) target EMX1-15 with a T4 SSB; (G) target RNF2 with a T4 SSB. [Figure 6C] Figure 1 shows the results of increasing the length of the sgRNA. Figures A-C show the percentage of reads with C to T substitutions for three targets and three spacer lengths: (A) target EMX1-15; (B) target RNF2; (C) target HBB03. Figures D-G show the percentage of reads carrying a single-presence editing allele: (D) target EMX1-15 without an SSB; (E) target RNF2 without an SSB; (F) target EMX1-15 with a T4 SSB; (G) target RNF2 with a T4 SSB. [Figure 6D] Figure 1 shows the results of increasing the length of the sgRNA. Figures A-C show the percentage of reads with C to T substitutions for three targets and three spacer lengths: (A) target EMX1-15; (B) target RNF2; (C) target HBB03. Figures D-G show the percentage of reads carrying a single-presence editing allele: (D) target EMX1-15 without an SSB; (E) target RNF2 without an SSB; (F) target EMX1-15 with a T4 SSB; (G) target RNF2 with a T4 SSB. [Figure 6E] Figure 1 shows the results of increasing the length of the sgRNA. Figures A-C show the percentage of reads with C to T substitutions for three targets and three spacer lengths: (A) target EMX1-15; (B) target RNF2; (C) target HBB03. Figures D-G show the percentage of reads carrying a single-presence editing allele: (D) target EMX1-15 without an SSB; (E) target RNF2 without an SSB; (F) target EMX1-15 with a T4 SSB; (G) target RNF2 with a T4 SSB. [Figure 6F]Figure 1 shows the results of increasing the length of the sgRNA. Figures A-C show the percentage of reads with C to T substitutions for three targets and three spacer lengths: (A) target EMX1-15; (B) target RNF2; (C) target HBB03. Figures D-G show the percentage of reads carrying a single-presence editing allele: (D) target EMX1-15 without an SSB; (E) target RNF2 without an SSB; (F) target EMX1-15 with a T4 SSB; (G) target RNF2 with a T4 SSB. [Figure 6G] Figure 1 shows the results of increasing the length of the sgRNA. Figures A-C show the percentage of reads with C to T substitutions for three targets and three spacer lengths: (A) target EMX1-15; (B) target RNF2; (C) target HBB03. Figures D-G show the percentage of reads carrying a single-presence editing allele: (D) target EMX1-15 without an SSB; (E) target RNF2 without an SSB; (F) target EMX1-15 with a T4 SSB; (G) target RNF2 with a T4 SSB. [Figure 7A] Figure 1 shows the results of delivery of CBE and UGI delivered as mRNA. (A) Percent GFP-positive cells; (B) Percent C to T conversions; (C) Percent C to G conversions; (D) Percent C to A conversions; (E) Percent indels from human K562 cells after transfection. [Figure 7B] Figure 1 shows the results of delivery of CBE and UGI delivered as mRNA. (A) Percent GFP-positive cells; (B) Percent C to T conversions; (C) Percent C to G conversions; (D) Percent C to A conversions; (E) Percent indels from human K562 cells after transfection. [Figure 7C] Figure 1 shows the results of delivery of CBE and UGI delivered as mRNA. (A) Percent GFP-positive cells; (B) Percent C to T conversions; (C) Percent C to G conversions; (D) Percent C to A conversions; (E) Percent indels from human K562 cells after transfection. [Figure 7D]Figure 1 shows the results of delivery of CBE and UGI delivered as mRNA. (A) Percent GFP-positive cells; (B) Percent C to T conversions; (C) Percent C to G conversions; (D) Percent C to A conversions; (E) Percent indels from human K562 cells after transfection. [Figure 7E] Figure 1 shows the results of delivery of CBE and UGI delivered as mRNA. (A) Percent GFP-positive cells; (B) Percent C to T conversions; (C) Percent C to G conversions; (D) Percent C to A conversions; (E) Percent indels from human K562 cells after transfection. [Figure 8A] Figure 1 shows the effect of using an NLS on co-delivered UGI proteins compared to commercially available UGI without an NLS (NEB), as measured as the percentage of reads with any C to T edit. (A) Targets EMX1-15, (B) Targets HEKSite2. [Figure 8B] Figure 1 shows the effect of using an NLS on co-delivered UGI proteins compared to commercially available UGI without an NLS (NEB), as measured as the percentage of reads with any C to T edit. (A) Targets EMX1-15, (B) Targets HEKSite2. DETAILED DESCRIPTION OF THE INVENTION
[0079] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which this invention belongs. The following references provide those skilled in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker, ed. 1988); The Glossary of Genetics, 5th ed., R. Rieger et al., (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless otherwise specified.
[0080] When introducing elements of the present disclosure or preferred embodiment(s) thereof, the articles "a," "an," "the," and "said" are intended to mean that there are one or more of the elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.
[0081] The transitional terms "comprising," "consisting essentially of," and "consisting of" have the meanings given them in MPEP 2111.03 (Manual of Patent Examining Procedures; U.S. Patent and Trademark Office, 9th Edition, Revision Feb 2023 [R-07.2022]). Any claim using the transitional term "consisting essentially of" is understood to recite only essential elements (i.e., basic and novel features) of the invention, and any other elements recited in a dependent claim are understood to be non-essential to the invention recited in the dependent claim. Similarly, any additional elements beyond the claimed elements described in the prior art reference(s) are excluded from a claim by use of the transitional term "consisting essentially of."
[0082] As used herein, the term "endogenous sequence" refers to a chromosomal sequence that is native to a cell.
[0083] The term "exogenous" as used herein refers to a sequence that is not native to the cell or a chromosomal sequence that is in a chromosomal location that is different from its natural location in the genome of the cell.
[0084] As used herein, a "gene" refers to a DNA region (including exons and introns) that encodes a gene product, as well as all such DNA regions that regulate the production of that gene product, whether or not such regulatory sequences are adjacent to the coding and / or transcribed sequence. Thus, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, border regions, origins of replication, matrix attachment sites, and locus control regions.
[0085] The term "complementary" or "complementarity" refers to the association of double-stranded nucleic acids by base pairing through specific hydrogen bonds. Base pairing may be standard Watson-Crick base pairing (e.g., 5'-AGT C-3' pairs with the complementary sequence 3'-TCA G-5'). Base pairing may be Hoogsteen or reverse Hoogsteen hydrogen bonding. Complementarity is typically measured with respect to the duplex region, thus excluding, for example, overhangs. Complementarity between two strands in a duplex region may be partial, and expressed as a percentage (e.g., 70%) if only a portion of the base pairs are complementary. Non-complementary bases are "mismatches." Complementarity can also be complete (i.e., 100%) if all base pairs in the duplex region are complementary.
[0086] The term "homologous" refers to the degree to which two or more sequences are identical. Two sequences are considered to be homologous if they hybridize to the same sequence under a defined set of conditions. The defined conditions include, but are not limited to, buffer formulation, temperature, and sequence concentration.
[0087] The term "heterologous" refers to an entity that is not endogenous or native to the cell of interest. For example, a heterologous protein is a protein that originates or is originally derived from an external source, such as an exogenously introduced nucleic acid sequence. In some instances, the heterologous protein is not normally produced by the cell of interest.
[0088] The terms "nucleic acid" and "polynucleotide" refer to a deoxyribonucleotide or ribonucleotide polymer in linear or circular conformation and in single- or double-stranded form. For purposes of this disclosure, these terms should not be construed as limiting with respect to the length of the polymer. The terms can encompass known analogues of natural nucleotides as well as nucleotides that have been modified in the base, sugar, and / or phosphate moieties (e.g., phosphorothioate backbones). Generally, analogues of a particular nucleotide have the same base-pairing specificity; i.e., an analogue of A will base pair with T.
[0089] The term "synthetic nucleic acid" refers to a nucleotide sequence that is synthesized in vitro (e.g., in a laboratory and by hand or using a nucleic acid synthesizer) and whose sequence is not found in nature. The sequence may be, for example, DNA or RNA or modifications thereof described below, and may be of any length and any sequence of nucleotides, so long as the sequence does not occur in nature.
[0090] The term "nucleotide" refers to a deoxyribonucleotide or a ribonucleotide. A nucleotide may be a standard nucleotide (i.e., adenosine, guanosine, cytidine, thymidine, and uridine) or a nucleotide analog. A nucleotide analog is a nucleotide with a modified purine or pyrimidine base or a modified ribose moiety. A nucleotide analog may be a naturally occurring nucleotide (e.g., inosine) or a non-naturally occurring nucleotide. Non-limiting examples of modifications to the sugar or base moiety of a nucleotide include the addition (or removal) of acetyl, amino, carboxyl, carboxymethyl, hydroxyl, methyl, phosphoryl, and thiol groups, and the substitution of carbon and nitrogen atoms of the base with other atoms (e.g., 7-deazapurines). Nucleotide analogs also include dideoxynucleotides, 2'-O-methyl nucleotides, locked nucleic acids (LNAs), peptide nucleic acids (PNAs), and morpholinos.
[0091] The terms "polypeptide" and "protein" are used interchangeably to refer to a polymer of amino acid residues.
[0092] The term "inactivated" in the context of the present invention may refer to the deletion of one or more amino acids or the substitution of one or more amino acids in a target protein such that one or more functions of the protein are eliminated or reduced to a level where the activity is less than 75%, 50%, 40%, 30%, 20%, 15%, 10%, 5%, 4%, 3%, 2%, or 1% of the level of activity of the active protein. In the present invention, a Cas-like protein (e.g., Cas9) is catalytically inactivated as a result of two amino acids being substituted with alanine residues, thereby suppressing its ability to cleave double-stranded DNA, as detailed elsewhere herein. Reduced activity may also be referred to as "partially inactivated." The term "catalytically modified" in the context of the present invention refers to a subject protein that has been inactivated or partially inactivated, meaning that one or more of the catalytic activities of the subject protein have been reduced or eliminated. The nuclease activity of the Cas proteins of the present invention may be partially or completely inactivated, rendering the Cas protein a nickase (which cleaves only one DNA strand; discussed in more detail below) or catalytically inactive (which does not cleave any of the DNA strands), respectively.
[0093] In the context of the present invention, the term "directly linked" with respect to proteins and polypeptides means that two proteins are joined (i.e., fused) (e.g., via a peptide bond) to form a contiguous protein or polypeptide without the incorporation of additional amino acid residues between the two joined / fused proteins.
[0094] In the context of the present invention, the term "indirectly linked" means that one or more amino acids are incorporated between the two linked / fused proteins.
[0095] Techniques for determining nucleic acid and amino acid sequence identity are known in the art. Typically, such techniques involve determining the nucleotide sequence of the mRNA for a gene and / or the amino acid sequence encoded thereby and comparing these sequences to a second nucleotide or amino acid sequence. Genomic sequences can also be determined and compared in this manner. Generally, identity refers to exact nucleotide-to-nucleotide or amino acid-to-amino acid matches between two polynucleotide or polypeptide sequences, respectively. Two or more sequences (polynucleotide or amino acid) can be compared by determining their percent identity. The percent identity of two sequences, whether nucleic acid or amino acid sequences, is calculated by dividing the number of exact matches between two aligned sequences by the length of the shorter sequence and multiplying by 100. Approximate alignment for nucleic acid sequences is provided by the local homology algorithm of Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981). This algorithm was developed by Dayhoff, Atlas of Protein Sequences and Structure, MO Dayhoff ed., 5 suppl. 3:353-358, National Biomedical Research Foundation, Washington, DC, USA, and can be applied to amino acid sequences using a normalized score matrix according to Gribskov, Nucl. Acids Res. 14(6):6745-6763 (1986). A representative implementation of this algorithm for determining percent identity of sequences is provided by the Genetics Computer Group (Madison, Wis.) in its "BestFit" practical application. Other suitable programs for calculating percent identity or similarity between sequences are generally known in the art; for example, another alignment program is BLAST, used with default parameters.For example, BLASTN and BLASTP can be used with the following default parameters: genetic code=standard; filter=none; strand=both; cutoff=60; expect=10; Matrix=BLOSUM62; Descriptions=50 sequences; sort by=HIGH SCORE; Databases=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+Swiss protein+Spupdate+PIR. More information about these programs can be found on the GenBank website.
[0096] Because it is believed that various changes can be made in the cells and methods described above without departing from the scope of the invention, it is intended that all matter contained in the above description and in the examples given below be interpreted as illustrative and not in a limiting sense.
[0097] CRISPR / Cas proteins and systems To understand the present invention, it is helpful to understand CRISPR / Cas protein systems generally and in the context of the present invention.
[0098] CRISPR / Cas9-based systems In its basic form, the CRISPR / Cas system introduces a double-stranded break near the binding site of a guide RNA (gRNA). The Cas9 protein and gRNA form a ribonucleoprotein (RNP) complex. This complex can be transfected into recipient cells (e.g., by using extracellular vesicles) or a plasmid or viral vector encoding the complex can be used. The gRNA targets the complex to a desired location in the genome through the formation of an R-loop-triple-stranded nucleic acid structure composed of a DNA:RNA hybrid and a displaced single-stranded DNA, where the Cas9 protein (endonuclease) creates a double-stranded DNA break, allowing the sequence to be modified. Such modifications can take the form of non-homologous end joining (NHEJ), which creates random small insertions or deletions (indels), or homology-directed repair (HDR). NHEJ is useful for generating knockout mutations, while HDR is useful for making desired modifications to target sequences.
[0099] Cytosine base editor As discussed above, traditional CRISPR / Cas9 technology faces challenges in efficiency and specificity. Cytosine base editing has been developed to increase both the efficiency and specificity of CRISPR / Cas9 technology (Komor, 2016). Cytosine base editing allows for the direct, irreversible conversion of one targeted DNA base pair to another, without the need for dsDNA backbone cleavage or a donor template. Rather, the targeted cytosine (C) is converted to uracil (U) by cytidine deaminase tethered to the RNP complex. Uracil has the base-pairing properties of thymine (T). Thus, during DNA replication or repair, the targeted C:G pair is converted to a T:A pair.
[0100] In one embodiment of the present invention, an ssDNA-binding domain (SSB) is used to restrict bystander editing within an R-loop. Bystander editing is where adjacent non-target cytosine residues within the R-loop are converted to thymine. This is important because bystander editing limits the use of cytosine base editors as precise gene editing tools (e.g., correcting pathogenic mutations). The SSB is linked to a cytosine base editor (CBE), for example, by a covalent interaction. Without being limited by theory, it is believed that the SSB may compete with the CBE for binding to the ssDNA region formed after Cas9 target binding and R-loop formation.
[0101] (I) RNA-guided endonuclease RNA-guided endonucleases, such as Cas9, may contain at least one nuclear localization signal (NLS), at least one nuclease domain, and at least one domain that interacts with a guide RNA (gRNA) to direct the endonuclease / deaminase complex of the present invention to a specific cytosine for deamination. Nucleic acids encoding RNA-guided endonucleases, as well as methods for modifying chromosomal sequences in eukaryotic cells or embryos using RNA-guided endonucleases, are also known. RNA-guided endonucleases interact with specific gRNAs, each of which directs the endonuclease to a specific target site, where the deaminase can convert the targeted cytosine to a uracil residue, resulting in the conversion of a CG base pair to an AT base pair. Because specificity is provided by the gRNA, RNA-based endonucleases are essentially universal and can be used with different gRNAs to target different genomic sequences. The methods disclosed herein can be used to target and modify specific chromosomal sequences at targeted locations in the genome of a cell or embryo. Furthermore, targeting is specific and off-target effects are limited.
[0102] Many forms of guide RNA (gRNA) are known in the art. Generally, gRNA is an RNA molecule capable of directing an RNA-binding protein or endonuclease to a specific nucleic acid sequence / target site through base pairing. A gRNA can comprise a single RNA molecule, such as a crRNA, or two RNA molecules, such as a crRNA and a tracrRNA (or trRNA). In some embodiments, a gRNA can further comprise an accessory RNA or DNA molecule. In some embodiments, a crRNA and a tracrRNA can be covalently linked together to form a chimeric guide RNA or a single guide RNA (srRNA). It is also well established in the art that gRNAs can be introduced into target cells in different forms, either alone or in combination with their cognate gRNA-binding proteins or endonucleases. gRNAs can be expressed from DNA vectors introduced into target cells. gRNAs can be synthesized by in vitro transcription or chemical reactions before being introduced into target cells. Those skilled in the art will recognize that chemical synthesis allows for many chemical modifications on gRNAs. For example, certain modifications, such as 2'-O methyl group modifications and phosphorothioate linkage modifications, can be introduced into gRNAs to alter their stability or performance. Non-standard nucleotides, such as DNA and LNA, can also be introduced into gRNAs during chemical synthesis to alter their specificity or performance. Guide RNAs that may have different morphologies or contain different chemical modifications may be used in conjunction with the present disclosure without departing from the spirit of the disclosure.
[0103] The present disclosure provides fusion proteins, which comprise a CRISPR / Cas-like protein or a fragment thereof, where the Cas-like protein is fully or partially catalytically inactivated and retains its ability to bind to DNA but is unable to create a double-stranded break in the target DNA. Each fusion protein is guided by a specific gRNA to a specific chromosomal sequence, where an associated (e.g., tethered or otherwise linked) cytosine deaminase converts the target cytosine to uracil.
[0104] The RNA-guided endonuclease may contain at least one nuclear localization signal, which allows the endonuclease to enter the nucleus of eukaryotic cells and embryos, such as non-human one-cell embryos. The RNA-guided endonuclease also contains at least one nuclease domain and at least one domain that interacts with the gRNA. The RNA-guided endonuclease is directed to a specific nucleic acid sequence (or target site) by the gRNA. The gRNA interacts with the RNA-guided endonuclease and the target site so that the associated cytosine deaminase can convert the target cytosine to uracil after being directed to the target site. Because the gRNA confers specificity to the target cleavage, the RNA-guided endonuclease is universal (provided its ability to cleave DNA is abolished) and can modify different target cytosines when used with different gRNAs. The RNA-guided endonuclease can be a protein, can be encoded by an isolated nucleic acid (i.e., RNA or DNA), can be encoded by a vector containing a nucleic acid encoding the RNA-guided endonuclease, or can be a protein-RNA complex comprising the RNA-guided endonuclease plus a gRNA.
[0105] RNA-guided endonucleases can be derived from clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) systems. CRISPR / Cas systems can be type I, type II, or type III systems, as known to those skilled in the art. Non-limiting examples of suitable CRISPR / Cas proteins include Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3, and the like. (or CasE), Cse4 (or CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966. One of skill in the art can modify any RNA-guided endonuclease to inactivate its catalytic activity for use with a cytosine base editor system.
[0106] In one embodiment, the RNA-guided endonuclease is derived from a type II CRISPR / Cas system. In a particular embodiment, the RNA-guided endonuclease is derived from a Cas9 protein. The Cas9 protein is capable of inhibiting the growth of bacteria such as Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Strepto- sporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, and the like. acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp. sp.), Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis variabilis), Nodularia spumigena, Nostoc sp.The plant may be derived from Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, or Acaryochloris marina.
[0107] Generally, CRISPR / Cas proteins contain at least one RNA recognition domain and / or RNA binding domain. The RNA recognition and / or RNA binding domain interacts with the guide RNA. CRISPR / Cas proteins also typically contain a nuclease domain (i.e., a DNase or RNase domain; however, for the purposes of cytosine base editors, these are disabled), a DNA binding domain, a helicase domain, an RNase domain, a protein-protein interaction domain, a dimerization domain, and other domains.
[0108] The CRISPR / Cas-like protein can be a wild-type CRISPR / Cas protein, a modified CRISPR / Cas protein, or a fragment of a wild-type or modified CRISPR / Cas protein. Modifications to the CRISPR / Cas-like protein can increase nucleic acid binding affinity and / or specificity, alter enzymatic activity (e.g., inactivate its ability to cleave DNA), and / or change another property of the protein. For example, the nuclease (i.e., DNase, RNase) domain of the CRISPR / Cas-like protein can be modified, deleted, or inactivated. Alternatively, CRISPR / Cas-like proteins can be truncated to remove domains that are not essential for the function of the fusion protein. CRISPR / Cas-like proteins can also be truncated or modified to optimize the activity of the effector domain of the fusion protein.
[0109] In some embodiments, the CRISPR / Cas-like protein can be derived from a wild-type Cas9 protein or a fragment thereof. In other embodiments, the CRISPR / Cas-like protein can be derived from a modified Cas9 protein. For example, modifying the amino acid sequence of the Cas9 protein can alter one or more properties of the protein (e.g., nuclease activity, affinity, stability, etc.). Alternatively, domains of the Cas9 protein that are not involved in RNA-guided cleavage can be removed from the protein so that the modified Cas9 protein is smaller than the wild-type Cas9 protein.
[0110] Generally, Cas9 proteins contain at least two nuclease (i.e., DNase) domains. For example, Cas9 proteins can contain a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC and HNH domains act in concert to cleave single strands and create double-strand breaks in DNA (Jinek et al., Science, 337:816-821). In some embodiments, Cas9-derived proteins can be engineered to contain only one functional nuclease domain (either a RuvC-like or HNH-like nuclease domain). For example, Cas9-derived proteins can be engineered to delete one of the nuclease domains or mutated so that one of the nuclease domains is no longer functional (i.e., has no nuclease activity). In some embodiments, where one of the nuclease domains is inactive, the Cas9-derived protein introduces nicks in double-stranded nucleic acids (such proteins are called "nickases") but is unable to cleave double-stranded DNA. For example, an aspartate to alanine (D10A) conversion in the RuvC-like domain converts the Cas9-derived protein into a nickase. Similarly, a histidine to alanine (H840A or H839A) conversion in the HNH domain converts the Cas9-derived protein into a nickase. Each nuclease domain can be modified using well-known methods such as site-directed mutagenesis, PCR-mediated mutagenesis, and complete gene synthesis, as well as other methods known in the art. Modification of both of these domains inactivates the nuclease and nickase activities.
[0111] The RNA-guided endonuclease may contain at least one nuclear localization signal (NLS). Generally, an NLS comprises a stretch of basic amino acids. Nuclear localization signals are known in the art (see, e.g., Lange et al., J. Biol. Chem. 2007, 282:5101-5105). For example, in one embodiment, the NLS can be a monopartite sequence such as PKKKRKV (SEQ ID NO: 30) or PKKKRRV (SEQ ID NO: 31). In another embodiment, the NLS can be a bipartite sequence. In yet another embodiment, the NLS can be KRPAATKKAGQAKKKK (SEQ ID NO: 32). The NLS can be located at the N-terminus, C-terminus, or internal position of the RNA-guided endonuclease.
[0112] In some embodiments, the RNA-guided endonuclease can further comprise at least one cell-penetrating domain. In one embodiment, the cell-penetrating domain can be a cell-penetrating peptide sequence derived from the HIV-1 TAT protein. By way of example, the TAT cell-penetrating sequence can be GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 33). In another embodiment, the cell-penetrating domain can be TLM (PLSSIFSRIGDPPKKKRKV; SEQ ID NO: 34), a cell-penetrating peptide sequence derived from human hepatitis B virus. In yet another embodiment, the cell-penetrating domain can be MPG (GALFLGWLGAAGSTMGAPKKKRKV; SEQ ID NO: 35 or GALFLGFLGAAGSTMGAWSQPKKKRKV; SEQ ID NO: 36). In additional embodiments, the cell-penetrating domain can be Pep-1 (KETWWETWWTEWSQPKKKRKV; SEQ ID NO: 37), VP22, a cell-penetrating peptide derived from herpes simplex virus, or a polyarginine peptide sequence. The cell-penetrating domain can be located at the N-terminus, C-terminus, or internal position of the protein.
[0113] In yet other embodiments, the RNA-guided endonuclease can also include at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In some embodiments, the marker domain can be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed), and the like. Monomeric marker domains include mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, JRed), and orange fluorescent protein (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato), or any other suitable fluorescent protein. In other embodiments, the marker domain can be a purification tag and / or an epitope tag.Representative tags include, but are not limited to, glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, (SEQ ID NO: 44), biotin carboxyl carrier protein (BCCP), and calmodulin.
[0114] In certain embodiments, the RNA-guided endonuclease may be part of a protein-RNA complex that includes a gRNA, which interacts with the RNA-guided endonuclease to direct the endonuclease to a specific target site, where, for example, the 5' end of the guide RNA base pairs with a specific protospacer sequence.
[0115] (II) Fusion Protein Another aspect of the present disclosure provides a fusion protein comprising a CRISPR / Cas-like protein or fragment thereof and, for example, one or more SSBs or one or more UGIs. The CRISPR / Cas-like protein is directed to a target site by a gRNA, where an effector modifies or affects the target nucleic acid sequence. In the present invention, the "effector domain" is a linked deaminase. The fusion protein can further comprise at least one additional domain selected from a nuclear localization signal, a cell penetration domain, or a marker domain.
[0116] (a) CRISPR / Cas-like proteins The fusion protein comprises a CRISPR / Cas-like protein or a fragment thereof. CRISPR / Cas-like proteins are described in detail above in section (I). The CRISPR / Cas-like protein can be located at the N-terminus, C-terminus, or at an internal position of the fusion protein.
[0117] In some embodiments, the CRISPR / Cas-like protein of the fusion protein can be derived from a Cas9 protein. The Cas9-derived protein can be wild-type, modified, or a fragment thereof. In some embodiments of the invention, the Cas9-derived protein is modified to inactivate both functional nuclease domains (RuvC-like or HNH-like nuclease domains). In some embodiments, both the RuvC-like nuclease domain and the HNH-like nuclease domain can be modified or removed such that the Cas9-derived protein is unable to nick or cleave double-stranded nucleic acids. In yet other embodiments, all nuclease domains of the Cas9-derived protein can be modified or removed such that the Cas9-derived protein lacks all nuclease activity. Moreover, the RuvC-like domain or the HNH-like domain can be inactivated independently of each other.
[0118] In any of the above embodiments, any or all of the nuclease domains can be inactivated by one or more deletion, insertion, and / or substitution mutations using well-known methods such as site-directed mutagenesis, PCR-mediated mutagenesis, and complete gene synthesis, as well as other methods known in the art.
[0119] (III) a nucleic acid encoding an RNA-guided endonuclease or a fusion protein Another aspect of the present disclosure provides a nucleic acid encoding either the RNA-guided endonuclease or the fusion protein described in sections (I) and (II) above, respectively. The nucleic acid can be RNA or DNA. In one embodiment, the nucleic acid encoding the RNA-guided endonuclease or the fusion protein is mRNA. The mRNA can be 5'-capped and / or 3'-polyadenylated. In another embodiment, the nucleic acid encoding the RNA-guided endonuclease or the fusion protein is DNA. The DNA can be present in a vector (see below).
[0120] Nucleic acids encoding RNA-guided endonucleases or fusion proteins can be codon-optimized for efficient translation into proteins in cells, plants, or animals of interest. For example, codons can be optimized for expression in humans, mice, rats, hamsters, cows, pigs, cats, dogs, fish, amphibians, plants, yeast, insects, etc. Programs for codon optimization are available as freeware. Commercially available codon optimization programs are also available.
[0121] In some embodiments, the DNA encoding the RNA-guided endonuclease or fusion protein can be operably linked to at least one promoter regulatory sequence. In some iterations, the DNA coding sequence can be operably linked to a promoter regulatory sequence for expression in a eukaryotic cell or animal of interest. The promoter regulatory sequence can be constitutive, regulated, or tissue-specific. Suitable constitutive promoter regulatory sequences include, but are not limited to, the cytomegalovirus immediate early promoter (CMV), simian virus (SV40) promoter, adenovirus major late promoter, Rous sarcoma virus (RSV) promoter, mouse mammary tumor virus (MMTV) promoter, phosphoglycerate kinase (PGK) promoter, elongation factor (EF1) alpha promoter, ubiquitin promoter, actin promoter, tubulin promoter, immunoglobulin promoter, fragments thereof, or any combination of the foregoing. Examples of suitable regulated promoter regulatory sequences include, but are not limited to, promoter regulatory sequences regulated by heat shock, metals, steroids, antibiotics, or alcohol. Non-limiting examples of tissue-specific promoters include the B29 promoter, CD14 promoter, CD43 promoter, CD45 promoter, CD68 promoter, desmin promoter, elastase 1 promoter, endoglin promoter, fibronectin promoter, Flt-1 promoter, GFAP promoter, GPIIb promoter, ICAM-2 promoter, INF-β promoter, Mb promoter, Nphsl promoter, OG-2 promoter, SP-B promoter, SYN1 promoter, and WASP promoter. The promoter sequence can be wild-type or modified for more efficient or effective expression. In one exemplary embodiment, the coding DNA can be operably linked to a CMV promoter for constitutive expression in mammalian cells.
[0122] In certain embodiments, the sequence encoding the RNA-guided endonuclease or fusion protein can be operably linked to a promoter sequence recognized by a phage RNA polymerase for in vitro mRNA synthesis. In such embodiments, the in vitro transcribed RNA can be purified for use in the methods detailed in sections (IV) and (V) below. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence or a variant of a T7, T3, or SP6 promoter sequence. In an exemplary embodiment, DNA encoding the fusion protein is operably linked to a T7 promoter using T7 RNA polymerase for in vitro mRNA synthesis.
[0123] In alternative embodiments, the sequence encoding the RNA-guided endonuclease or fusion protein can be operably linked to a promoter sequence for in vivo expression of the RNA-guided endonuclease or fusion protein in bacterial or eukaryotic cells. In such embodiments, the expressed protein can be purified for use in the methods detailed in sections (IV) and (V) below. Suitable bacterial promoters include, but are not limited to, the T7 promoter, the lac operon promoter, the trp promoter, variants thereof, and combinations thereof. An exemplary bacterial promoter is tac, a hybrid of the trp and lac promoters. Non-limiting examples of suitable eukaryotic promoters are listed above.
[0124] In additional embodiments, the DNA encoding the RNA-guided endonuclease or fusion protein can also be linked to a polyadenylation signal (e.g., SV40 polyA signal, bovine growth hormone (BGH) polyA signal, etc.) and / or at least one transcription termination sequence. Furthermore, the sequence encoding the RNA-guided endonuclease or fusion protein can also be linked to a sequence encoding at least one nuclear localization signal, at least one cell penetration domain, and / or at least one marker domain, which are described in detail in section (I) above.
[0125] In various embodiments, the DNA encoding the RNA-guided endonuclease or fusion protein can be present in a vector. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, etc.). In one embodiment, the DNA encoding the RNA-guided endonuclease or fusion protein is present in a plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, and variants thereof. The vector can include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Additional information can be found in "Current Protocols in Molecular Biology," Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual," Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd ed., 2001.
[0126] In some embodiments, the expression vector containing a sequence encoding an RNA-guided endonuclease or a fusion protein can further contain a sequence encoding a gRNA. The sequence encoding the gRNA is generally operably linked to at least one transcriptional control sequence for expression of the gRNA in a cell or embryo of interest. For example, the DNA encoding the gRNA can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters.
[0127] (IV) Methods for modifying chromosomal sequences using RNA-guided endonucleases Another aspect of the present disclosure includes a method for modifying a chromosomal sequence in a eukaryotic cell or embryo. The method includes introducing into a eukaryotic cell or embryo (i) at least one enzymatically inactivated RNA-guided endonuclease (or a nucleic acid encoding it) containing at least one nuclear localization signal or a nucleic acid encoding at least one RNA-guided endonuclease containing at least one nuclear localization signal, (ii) at least one gRNA or DNA encoding at least one gRNA, and optionally (iii) at least one cytosine deaminase, and (iv) a uracil DNA glycosylase inhibitor (UGI). The method further includes culturing the cell or embryo so that each gRNA directs the enzymatically inactivated RNA-guided endonuclease to a targeted site in the chromosomal sequence, where the cytosine deaminase converts the cytosine to uracil at the targeted site, and the uracil is repaired (i.e., converted) by a DNA repair process such that the chromosomal sequence is modified.
[0128] Thus, as discussed herein, targeted chromosomal sequences can be modified or inactivated. For example, a single base change (SNP) can result in an altered protein product or introduce a "stop" codon into the reading frame of a coding sequence, thereby inactivating or "knocking out" the sequence so that no protein product is made.
[0129] In other embodiments, the method can include introducing two (or more) RNA-guided endonucleases (or encoding nucleic acids) and two gRNAs (or encoding DNAs) into a cell or embryo, where the RNA-guided endonucleases modify two cytosine bases, which can be within a few base pairs, within tens of base pairs, or separated by thousands of base pairs.
[0130] (a) RNA-guided endonuclease The method includes introducing into a cell or embryo at least one RNA-guided endonuclease comprising at least one nuclear localization signal or a nucleic acid encoding at least one RNA-guided endonuclease comprising at least one nuclear localization signal. Such RNA-guided endonucleases and nucleic acids encoding RNA-guided endonucleases are described above in sections (I) and (III), respectively. Such RNA-guided endonucleases may be preloaded with a gRNA.
[0131] In some embodiments, the RNA-guided endonuclease can be introduced into a cell or embryo as an isolated protein. In such embodiments, the RNA-guided endonuclease can further comprise at least one cell-penetrating domain, which facilitates cellular uptake of the protein. In other embodiments, the RNA-guided endonuclease can be introduced into a cell or embryo as an mRNA molecule. In yet other embodiments, the RNA-guided endonuclease can be introduced into a cell or embryo as a DNA molecule. Generally, a DNA sequence encoding the fusion protein is operably linked to a promoter sequence that functions in the cell or embryo of interest. The DNA sequence can be linear, or the DNA sequence can be part of a vector. In yet other embodiments, the fusion protein can be introduced into a cell or embryo as an RNA-protein complex comprising the fusion protein and a gRNA.
[0132] In an alternative embodiment, the DNA encoding the RNA-guided endonuclease may further comprise a sequence encoding a gRNA. Generally, the sequences encoding the RNA-guided endonuclease and the gRNA are each operably linked to an appropriate promoter control sequence that allows expression of the RNA-guided endonuclease and the gRNA, respectively, in a cell or embryo. The DNA sequences encoding the RNA-guided endonuclease and the gRNA may further comprise additional expression control, regulatory, and / or processing sequence(s). The DNA sequences encoding the RNA-guided endonuclease and the gRNA may be linear or may be part of a vector.
[0133] (b) Guide RNA (gRNA) The method also includes introducing into the cell or embryo at least one gRNA or DNA encoding at least one gRNA, wherein the gRNA interacts with an enzymatically inactivated RNA-guided endonuclease to direct the endonuclease to a specific target site in a chromosomal sequence.
[0134] Each gRNA may contain three regions: a first region that is complementary to a target site in the chromosomal sequence, a second internal region that forms a stem-loop structure, and a third region that remains essentially single-stranded. The first region of each gRNA is different so that each gRNA guides the fusion protein to a specific target site. The second and third regions of each gRNA can be the same in all gRNAs.
[0135] The first region of the gRNA, termed the spacer, is complementary to the sequence of the target site in the chromosomal sequence (i.e., the protospacer sequence) so that the first region of the gRNA can base pair with the target site. In various embodiments, the first region of the gRNA can comprise from about 10 nucleotides to more than about 25 nucleotides. For example, the region of base pairing between the first region of the gRNA and the target site in the chromosomal sequence can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more than 25 nucleotides long. In representative embodiments, the first region of the gRNA is about 19, 20, 21, 22, or 23 nucleotides long.
[0136] The gRNA also comprises a second region that forms a secondary structure. In some embodiments, the secondary structure comprises a stem (or hairpin) and a loop. The lengths of the loop and stem can vary. For example, the loop can range from about 3 to about 10 nucleotides in length, and the stem can range from about 6 to about 20 base pairs in length. The stem can include one or more bulges of 1 to about 10 nucleotides in length. Thus, the total length of the second region can range from about 16 to about 60 nucleotides in length. In a representative embodiment, the loop is about 4 nucleotides in length, and the stem comprises about 12 base pairs.
[0137] The gRNA may also include a third region that remains essentially single-stranded. Thus, the third region has no complementarity to any chromosomal sequence in the cell of interest and no complementarity to the rest of the gRNA. The length of the third region can vary. Generally, the third region is greater than about 4 nucleotides in length. For example, the length of the third region can range from about 5 to about 60 nucleotides in length.
[0138] The combined length of the second and third regions (also called the universal or scaffold regions) of the gRNA can range from about 20 to about 120 nucleotides in length. In one embodiment, the combined length of the second and third regions of the gRNA ranges from about 20 to about 100 nucleotides in length.
[0139] In some embodiments, the gRNA comprises a single molecule containing all three regions, known as the sgRNA. In other embodiments, the gRNA can comprise two separate molecules. The first RNA molecule can comprise the first region of the gRNA and half of the "stem" of the second region of the gRNA, known as the crRNA. The second RNA molecule can comprise the other half of the "stem" of the second region of the gRNA and the third region of the gRNA, known as the tracrRNA. Thus, in this embodiment, the first and second RNA molecules each comprise a sequence of nucleotides that is complementary to each other. For example, in one embodiment, the first and second RNA molecules each comprise a sequence (of about 6 to about 20 nucleotides) that base-pairs with the other sequence to form a functional gRNA.
[0140] In some embodiments, the gRNA can be introduced into a cell or embryo as an RNA molecule. The RNA molecule can be transcribed in vitro. Alternatively, the RNA molecule can be chemically synthesized.
[0141] In other embodiments, the gRNA can be introduced into a cell or embryo as a DNA molecule. In such cases, the DNA encoding the gRNA can be operably linked to a promoter regulatory sequence for expression of the gRNA in the cell or embryo of interest. For example, the RNA coding sequence can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6 or H1 promoters. In representative embodiments, the RNA coding sequence is linked to a mouse or human U6 promoter. In other representative embodiments, the RNA coding sequence is linked to a mouse or human H1 promoter.
[0142] The DNA molecule encoding the gRNA can be linear or circular. In some embodiments, the DNA sequence encoding the gRNA can be part of a vector. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors. In a representative embodiment, the DNA encoding the gRNA is present in a plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, and variants thereof. The vector can include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc.
[0143] In embodiments in which both the RNA-guided endonuclease and the gRNA are introduced into the cell as DNA molecules, each can be part of separate molecules (e.g., one vector containing the fusion protein coding sequence and a second vector containing the gRNA coding sequence) or both can be part of the same molecule (e.g., one vector containing the coding (and regulatory) sequences for both the fusion protein and the gRNA).
[0144] (c) Target site The RNA-guided endonuclease in conjunction with the gRNA is directed to a target site in the chromosomal sequence, where the RNA-guided endonuclease (enzymatically inactivated) binds to the chromosomal sequence. The target site has no sequence constraints other than that its sequence be immediately adjacent to a consensus sequence. This consensus sequence is also known as a protospacer adjacent motif (PAM). Examples of PAMs include, but are not limited to, NGG, NGGNG, and NNAGAAW (where N is defined as any nucleotide and W is defined as A or T). As detailed in section (IV)(b) above, the spacer region of the gRNA is complementary to the protospacer of the target sequence. Typically, the spacer region of the gRNA is approximately 19-21 nucleotides in length. Thus, in certain embodiments, the sequence of the target site in the chromosomal sequence is 5'-N. 19~21 -NGG-3'.
[0145] The target site can be in the coding region of a gene, an intron of a gene, a regulatory region of a gene, a non-coding region between genes, etc. The gene can be a protein-coding gene or an RNA-coding gene. The gene can be any gene of interest.
[0146] (d) introduction into a cell or embryo The RNA-targeting endonuclease(s) (or encoding nucleic acid) and gRNA(s) (or encoding DNA) can be introduced into cells or embryos by various means. In some embodiments, the cells or embryos are transfected. Suitable transfection methods include calcium phosphate-mediated transfection, nucleofection (or electroporation), cationic polymer transfection (e.g., DEAE-dextran or polyethyleneimine), viral transduction, virosome transfection, virion transfection, liposome transfection, cationic liposome transfection, immunoliposome transfection, non-liposomal lipid transfection, dendrimer transfection, heat shock transfection, magnetofection, lipofection, gene gun delivery, impale infection, sonoporation, phototransfection, and specialized agent-enhanced uptake of nucleic acids. Transfection methods are well known in the art (see, e.g., "Current Protocols in Molecular Biology," Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual," Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd ed., 2001). In other embodiments, molecules are introduced into cells or embryos by microinjection. Typically, the embryo is a fertilized one-cell stage embryo of the species of interest. For example, molecules can be injected into the pronucleus of a single-cell embryo.
[0147] The RNA-targeting endonuclease(s) (or encoding nucleic acid), gRNA(s) (or DNA encoding the gRNA), and optional uracil glycosylase inhibitor can be introduced into the cell or embryo simultaneously or sequentially. The ratio of RNA-targeting endonuclease(s) (or encoding nucleic acid) to gRNA(s) (or encoding DNA) is generally approximately stoichiometric so that they can form an RNA-protein complex. In one embodiment, the DNA encoding the RNA-targeting endonuclease and the DNA encoding the gRNA are delivered together in a plasmid vector.
[0148] (e) Culturing cells or embryos The method includes maintaining the cell or embryo under suitable conditions such that the gRNA(s) direct the cytosine base editor fusion protein(s) to the targeted site(s) in the chromosomal sequence and the cytosine deaminase introduces at least one cytosine to uracil conversion into the chromosomal sequence such that the chromosomal sequence is modified by substitution of at least one nucleotide.
[0149] In embodiments in which a premature stop codon is introduced, the chromosomal sequence may be inactivated or "knocked out." An inactivated protein-encoding chromosomal sequence will not give rise to the protein encoded by the wild-type chromosomal sequence.
[0150] In embodiments in which chromosomal sequences are altered without knocking out or inactivating the sequence, the encoded protein may be 1) altered to result in a protein with reduced or increased function, or 2) reverted to or approximated wild-type sequence and / or function if a naturally occurring mutation is corrected. Generally, cells are maintained under conditions suitable for cell growth and / or maintenance. Suitable cell culture conditions are well known in the art and are described, for example, in Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. (2005) Nature 435:646-651; and Lombardo et al. (2007) Nat. Biotechnology 25:1298-1306. Those skilled in the art will recognize that methods for culturing cells are known in the art and can vary depending on the cell type. Routine optimization may be used in all cases to determine the best technique for a particular cell type.
[0151] Embryos can be cultured in vitro (e.g., in cell culture). Typically, embryos are cultured at an appropriate temperature and in an appropriate medium with the necessary O2 / CO2 ratio to allow editing to occur. Suitable non-limiting examples of medium include M2, M16, KSOM, BMOC, and HTF medium. Those skilled in the art will recognize that culture conditions can and will vary depending on the species of the embryo. Routine optimization may be used in all cases to determine the best culture conditions for a particular species of embryo. In some cases, cell lines may be derived from in vitro cultured embryos (e.g., embryonic stem cell lines).
[0152] Alternatively, the embryo may be cultured in vivo by transferring the embryo into the uterus of a female host. Generally speaking, the female host is from the same or similar species as the embryo. Preferably, the female host is pseudopregnant. Methods for preparing pseudopregnant female hosts are known in the art. Additionally, methods for transferring embryos into female hosts are known. Culturing the embryo in vivo allows the embryo to develop and result in the birth of an animal derived from the embryo. Such an animal would contain the modified chromosomal sequence in all cells of its body.
[0153] (f) Cell types and embryonic types A variety of eukaryotic cells and embryos are suitable for use in the methods. For example, the cells can be human cells, non-human mammalian cells, non-mammalian vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, unicellular eukaryotes, or prokaryotes. Generally, the embryo is a non-human mammalian embryo. In certain embodiments, the embryo can be a single-cell non-human mammalian embryo. Exemplary mammalian embryos, including single-cell embryos, include, but are not limited to, mouse, rat, hamster, rodent, rabbit, cat, dog, sheep, pig, cow, horse, and non-human primate embryos. In yet other embodiments, the cell can be a stem cell. Suitable stem cells include, but are not limited to, embryonic stem cells, embryonic stem cell-like stem cells, fetal stem cells, adult stem cells, pluripotent stem cells, induced pluripotent stem cells, multipotent stem cells, oligopotent stem cells, unipotent stem cells, and others. In exemplary embodiments, the cell is a mammalian cell.
[0154] Non-limiting examples of suitable mammalian cells include Chinese hamster ovary (CHO) cells, baby hamster kidney (BHK) cells; mouse myeloma NSO cells, mouse embryonic fibroblast 3T3 cells (NIH3T3), mouse B lymphoma A20 cells; mouse melanoma B16 cells; mouse myoblast C2C12 cells; mouse myeloma SP2 / 0 cells; mouse embryonic mesenchymal C3H-10T1 / 2 cells; mouse carcinoma CT26 cells, and mouse prostate DuCuP cells. ;Mouse breast cancer MTP6 cells;Mouse hepatocellular carcinoma Hepa1c1c7 cells;Mouse myeloma J5582 cells;Mouse epithelial MTD-1A cells;Mouse cardiac MyEnd cells;Mouse kidney RenCa cells;Mouse pancreatic RIN-5F cells;Mouse melanoma X64 cells;Mouse lymphoma YAC-1 cells;Rat glioblastoma 9L cells;Rat B lymphoma RBL cells;Rat neuroblastoma B35 cells;Rat hepatocellular carcinoma (HTC);Buffalo rat liver BRL Examples of mammalian cell lines include 3A cells, canine kidney cells (MDCK), canine mammary gland (CMT) cells, rat osteosarcoma D17 cells, rat monocyte / macrophage DH82 cells, monkey kidney SV-40 transformed fibroblast (COS7) cells, monkey kidney CVI-76 cells, African green monkey kidney (VERO-76) cells, human embryonic kidney cells (HEK293, HEK293T), human cervical carcinoma cells (HELA), human lung cells (W138), human liver cells (Hep G2), human U2-OS osteosarcoma cells, human A549 cells, human A-431 cells, and human K562 cells. An extensive list of mammalian cell lines can be found in the American Type Culture Collection (ATCC, Manassas, Va.).
[0155] (V) Methods for using fusion proteins to modify or regulate the expression of chromosomal sequences. Another aspect of the present disclosure includes a method for modifying a chromosomal sequence or modulating expression of a chromosomal sequence in a cell or embryo, the method comprising introducing into a cell or embryo (a) at least one cytosine base editor fusion protein or a nucleic acid encoding at least one fusion protein, where the fusion protein comprises a CRISPR / Cas-like protein or a fragment thereof and an effector domain or a tethered protein encoding a cytosine deaminase, and (b) at least one gRNA or DNA encoding the gRNA, where the gRNA guides the CRISPR / Cas-like protein of the fusion protein to a targeted site in the chromosomal sequence, and the cytosine deaminase of the fusion protein modifies the chromosomal sequence or modulates expression of the chromosomal sequence by modifying a regulatory sequence associated with the chromosomal sequence.
[0156] Fusion proteins comprising a CRISPR / Cas-like protein or fragment thereof and an effector domain are described in detail above in section (II). Generally, fusion proteins further disclosed herein may include at least one nuclear localization signal. Nucleic acids encoding fusion proteins are described above in section (III). In some embodiments, the fusion protein can be introduced into a cell or embryo as an isolated protein (which may further include a cell-penetrating domain). Furthermore, the isolated fusion protein can be part of a protein-RNA complex that includes a gRNA. In other embodiments, the fusion protein can be introduced into a cell or embryo as an RNA molecule (which may be capped and / or polyadenylated). In yet other embodiments, the fusion protein can be introduced into a cell or embryo as a DNA molecule. For example, the fusion protein and gRNA can be introduced into a cell or embryo as separate DNA molecules or as part of the same DNA molecule. Such a DNA molecule can be a plasmid vector.
[0157] (VI) Genetically modified cells and animals The present disclosure encompasses genetically modified cells, non-human embryos, and non-human animals comprising at least one chromosomal sequence modified using an enzymatically inactivated RNA-guided endonuclease-mediated or fusion protein-mediated process, e.g., using the methods described herein. The present disclosure provides cells comprising at least one DNA or RNA molecule or fusion protein encoding an RNA-guided endonuclease or fusion protein targeted to a chromosomal sequence of interest, at least one gRNA, and optionally one or more free-standing or fusion uracil glycosylase inhibitors (UGIs). The present disclosure also provides non-human embryos comprising at least one DNA or RNA molecule encoding an enzymatically inactivated RNA-guided endonuclease or fusion protein targeted to a chromosomal sequence of interest, at least one gRNA, and optionally one or more UGIs.
[0158] The present disclosure provides genetically modified non-human animals, non-human embryos, or animal cells comprising at least one modified chromosomal sequence. The modified chromosomal sequence may be (1) inactivated or (2) modified to have altered expression or produce an altered protein product. The chromosomal sequence is modified using an RNA-guided endonuclease-mediated or fusion protein-mediated process using the methods described herein.
[0159] As discussed, one aspect of the present disclosure provides genetically modified animals in which at least one chromosomal sequence has been modified. In one embodiment, the genetically modified animal comprises at least one inactivated chromosomal sequence. The modified chromosomal sequence may be inactivated so that the sequence is not transcribed and / or a functional protein product is not produced. Thus, a genetically modified animal comprising an inactivated chromosomal sequence may be termed a "knockout." As a result of the mutation, the targeted chromosomal sequence is inactivated and no functional protein is produced. An inactivated chromosomal sequence does not comprise an exogenously introduced sequence. Also included herein are genetically modified animals in which two, three, four, five, six, seven, eight, nine, or ten or more chromosomal sequences have been inactivated.
[0160] In another embodiment, the modified chromosomal sequence can be altered so that it encodes a mutant protein product. For example, a genetically modified animal containing the modified chromosomal sequence can contain targeted point mutation(s) or other modifications such that an altered protein product is produced. In one embodiment, the chromosomal sequence can be altered so that at least one nucleotide is changed and the expressed protein contains one altered amino acid residue (a missense mutation). In another embodiment, the chromosomal sequence can be altered to contain more than one missense mutation such that more than one amino acid is changed. The altered or mutant protein can have an altered property or activity, such as altered substrate specificity, altered enzymatic activity, or altered kinetic rate, compared to the wild-type protein (or compared to the unaltered protein, if correcting a naturally occurring genetic mutation).
[0161] In another embodiment, the genetically modified animal can contain at least one genetic change that can activate a non-expressed gene, termed a "knock-in." The chromosomally modified sequence can encode, for example, a protein similar to an orthologous protein, an endogenous protein, or a combination of both.
[0162] In yet another embodiment, a genetically modified animal can contain at least one modified chromosomal sequence encoding a protein such that the expression pattern of the protein is altered. For example, a regulatory region controlling the expression of the protein, such as a promoter or transcription factor binding site, can be altered so that the protein is overproduced, or the tissue-specific or temporal expression of the protein is altered, or a combination thereof. Alternatively, the expression pattern of the protein can be altered in combination with the use of a conditional knockout system. A non-limiting example of a conditional knockout system is the Cre-lox recombination system. The Cre-lox recombination system contains the enzyme Cre recombinase, a site-specific DNA recombinase that can catalyze the recombination of nucleic acid sequences between specific sites (lox sites) in a nucleic acid molecule. Methods for using this system to produce temporal and tissue-specific expression are known in the art. Generally, a genetically modified animal is generated that has lox sites flanking the chromosomal sequence. The genetically modified animal containing the lox-flanked chromosomal sequence can then be bred with another genetically modified animal that expresses Cre recombinase. Progeny animals are then produced that contain the lox-flanked chromosomal sequences and the Cre recombinase, and the lox-flanked chromosomal sequences recombine to result in a deletion or inversion of the protein-encoding chromosomal sequence. Expression of the Cre recombinase can be transiently and conditionally regulated to result in transient and conditionally regulated recombination of the chromosomal sequence.
[0163] In any of these embodiments, the genetically modified animals disclosed herein can be heterozygous for the modified chromosomal sequence. Alternatively, the genetically modified animals can be homozygous for the modified chromosomal sequence.
[0164] The genetically modified animals disclosed herein can be crossbred to produce animals containing more than one modified chromosomal sequence or to produce animals that are homozygous for one or more modified chromosomal sequences. For example, two animals containing the same modified chromosomal sequence can be crossbred to produce an animal that is homozygous for the modified chromosomal sequence. Alternatively, animals with different modified chromosomal sequences can be crossbred to produce an animal that contains both modified chromosomal sequences.
[0165] In other embodiments, animals containing the modified chromosomal sequence can be interbred to combine the modified chromosomal sequence with other genetic backgrounds, including, by non-limiting example, wild-type genetic backgrounds, genetic backgrounds with deletion mutations, genetic backgrounds with targeted integrations, and genetic backgrounds with non-targeted integrations.
[0166] The term "animal," as used herein, refers to a human or non-human animal. An animal may be an embryo, a child, or an adult. Suitable animals include vertebrates such as mammals, birds, reptiles, amphibians, shellfish, and fish. Examples of suitable mammals include, but are not limited to, rodents, companion animals, livestock, and primates. Non-limiting examples of rodents include mice, rats, hamsters, gerbils, and guinea pigs. Suitable companion animals include, but are not limited to, cats, dogs, rabbits, hedgehogs, and ferrets. Non-limiting examples of livestock include horses, goats, sheep, pigs, cows, llamas, and alpacas. Suitable primates include, but are not limited to, capuchin monkeys, chimpanzees, lemurs, macaques, marmosets, tamarins, spider monkeys, squirrel monkeys, and vernacular monkeys. Non-limiting examples of birds include chickens, turkeys, ducks, and geese. Alternatively, the animal may be an invertebrate, such as an insect, nematode, or the like. Non-limiting examples of insects include fruit flies and mosquitoes. A representative animal is a rat. Non-limiting examples of suitable rat strains include Dahl Salt Sensitive, Fischer 344, Lewis, Long-Evans Hooded, Sprague-Dawley, and Wistar. In one embodiment, the animal is not a genetically modified mouse. In each of the above iterations of animals suitable for the present invention, the animal does not contain exogenously introduced, randomly integrated transposon sequences.
[0167] A further aspect of the present disclosure provides a genetically modified cell or cell line comprising at least one modified chromosomal sequence. The genetically modified cell or cell line can be derived from any of the genetically modified animals disclosed herein. Alternatively, the chromosomal sequence can be modified in the cells described herein above (in the paragraph describing chromosomal sequence modification in animals) using the methods described herein. The present disclosure also encompasses lysates of the cells or cell lines.
[0168] In a preferred embodiment, the cell is a eukaryotic cell. Suitable host cells include fungi or yeasts, such as Pichia, Saccharomyces, or Schizosaccharomyces; insect cells, such as SF9 cells from Spodoptera frugiperda or S2 cells from Drosophila melanogaster; and animal cells, such as mouse, rat, hamster, non-human primate, or human cells. Exemplary cells are mammalian. Mammalian cells can be primary cells. Cells can be of various cell types, such as fibroblasts, myoblasts, T or B cells, macrophages, epithelial cells, etc.
[0169] The cells can be eukaryotic or prokaryotic. For example, cells of bacterial, protist, plant, animal, and fungal origin are contemplated. Animals may include, but are not limited to, insects and mammals.
[0170] When a mammalian cell line is used, the cell line can be any established cell line or primary cell type, or a cell line yet to be described. The cell line can be adherent or non-adherent, or the cell line can be grown under conditions that promote adherent, non-adherent, or organotypic growth using standard techniques known to those skilled in the art. Non-limiting examples of suitable mammalian cells and cell lines are provided in section (IV)(g) herein. In yet other embodiments, the cell can be a stem cell. Non-limiting examples of suitable stem cells are provided in section (IV)(g).
[0171] The present disclosure also provides a genetically modified non-human embryo comprising at least one modified chromosomal sequence. The chromosomal sequence can be modified in an embryo as described herein above (in the paragraph describing chromosomal sequence modification in an animal) using the methods described herein. In one embodiment, the embryo is a non-human fertilized, single-cell stage embryo of the animal species of interest. Representative mammalian embryos, including single-cell embryos, include, but are not limited to, mouse, rat, hamster, rodent, rabbit, cat, dog, sheep, pig, cow, horse, and primate embryos.
[0172] Additionally, as known to those skilled in the art, non-mammalian cells and cell lines can be used, including, but not limited to, plant, insect, and prokaryotic cells and cell lines. Such cells may be newly developed for a specific purpose or may be readily available, such as from ATCC (Bethesda, MD) and other sources known to those skilled in the art. In short, the compositions and methods of the present invention should be useful for base editing in any cell that has a nucleic acid and Cas-based repair pathway, including viruses after infection into a host cell.
[0173] (VII) Kit Yet another aspect of the present invention provides kits for carrying out the methods described above.
[0174] The kits provided herein generally include instructions for carrying out the steps detailed above. The instructions included in the kit may be affixed to packaging materials, included as a package insert, or included as a downloadable file. The instructions are typically written or printed material, but are not limited to such. Any medium capable of storing such instructions and communicating them to an end user is contemplated by the present disclosure. Such media include, but are not limited to, electronic storage media (e.g., magnetic disks, tapes, cartridges, chips), optical media (e.g., CD ROM), and the like. As used herein, the term "instructions" can include the address of an internet site that provides the instructions.
[0175] It is contemplated that various changes can be made to the above processes and kits without departing from the scope of the present invention, and therefore, all material contained in the examples given above and below is intended to be interpreted in an illustrative and not limiting sense. [Example]
[0176] [Example 1] CBEs with N-terminal SSB fusions result in a narrower editing window than CBEs without the SSB domain Cytosine base editor (CBE) proteins were constructed by fusing human APOBEC3A or the C-terminal domain of human APOBEC3A to the amino terminus of the SpCas9 nickase protein and two uracil glycosylase inhibitor (UGI) domains to the C-terminus (SEQ ID NOs: 4 and 5). Novel CBE proteins were constructed by fusing the single-stranded DNA binding protein (SSB) domain from T4 or T7 bacteriophage to the amino terminus of the above CBE (SEQ ID NOs: 6–9). All proteins were expressed and purified from E. coli BL21AI by autoinduction and nickel column chromatography. They were stored at -80°C before use in a buffer containing 10% glycerol, 300 mM KCl, 20 mM HEPES (pH 7.5), and 1 mM DTT. Synthetic single guide RNAs (sgRNAs) targeting four human genomic sites were purchased from MilliporeSigma. The spacer sequences of the sgRNAs are given in Table 1. Each experimental condition was tested in two technical replicates.
[0177] [Table 1]
[0178] Ribonucleoprotein (RNP) complexes were prepared by incubating 15 μg of CBE protein and 200 pmol of sgRNA in a final volume of 10 μL in buffer (20 mM HEPES, 100 mM KCl, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) at room temperature for 15 minutes. The molar ratio of CBE to sgRNA was approximately 1:3. RNPs were kept on ice until transfection. HEK293 cells were obtained from ATCC and grown at 37°C and 5% CO2 in DMEM supplemented with 10% FBS, 2 mM L-glutamine, 1 mM sodium pyruvate, and 0.1 mM non-essential amino acids. Cells were cultured at 1.67 × 10 4 Cell / tissue culture surface area 1cm 2 At the time of transfection, cells were trypsinized to obtain a single cell suspension, washed twice with Hank's balanced salt solution, and plated at 2.5 × 10 cells per 100 μL. 5 Cells were resuspended in Nucleofector Solution V (Lonza, Basel, CH). Nucleofection was performed by mixing 100 μL of the prepared cell suspension with 10 μL of complexed CBE RNP by pipetting up and down six times before transferring to a cuvette for electroporation using program Q-001 on a Nucleofector 2b instrument. Nucleofected cells were immediately transferred to a 6-well plate containing 2 mL of prewarmed medium per well and grown for 3 days before harvesting.
[0179] Genomic DNA was harvested by trypsinizing the transfected cells and resuspending them in 75 μL of QuickExtract reagent (Lucigen). The suspension was incubated at 60°C for 15 minutes and at 95°C for 15 minutes. The genomic region targeted by CBE was amplified by PCR using JumpStart Taq ReadyMix (MilliporeSigma, Burlington, MA) and the following cycle conditions: 94°C / 2 minutes; 25 cycles of 94°C / 30 seconds, 62°C / 30 seconds, 72°C / 45 seconds; and 72°C / 5 minutes. Primers are listed in Table 1. The PCR product underwent a second round of amplification using Illumina index primers and JumpStart Taq ReadyMix and the following conditions: 95°C / 3 minutes; 9 cycles of 95°C / 30 seconds, 55°C / 30 seconds, 72°C / 30 seconds; and 72°C / 5 minutes. Indexed PCR products were purified with Select-a-Size DNA Clean & Concentrator MagBeads (Zymo, Irvine, CA) using 1.2x beads in volume, quantified with PicoGreen (ThermoFisher, Waltham, MA), and pooled according to DNA content. Pools were diluted to 4 nM. Sequencing was performed on an Illumina MiSeq instrument using a 300-cycle kit to obtain single-end reads. FASTQ files for each sample were analyzed using custom analysis scripts.
[0180] The results are presented in Figure 1. Figures 1A-1D plot the percentage of sequence reads containing C to T substitutions at each cytosine residue at each target site. Positions are indicated by their distance from the 5' end of the target sequence. Values are the mean ± standard deviation for two replicates. All transfections were performed on the same day using the same cells. The results demonstrate that absolute and relative editing efficiencies are target specific. The editing profiles show that the editing window for CBE variants with N-terminal SSB fusions is shifted away from the PAM compared to the editing window for CBEs lacking the SSB fusion.
[0181] In Figure 1E, the percentage of C to T substitutions that fall within a window of a given width is plotted against the width of that window for each variant tested. The results show that, regardless of whether the deaminase used is APOBEC3A (open symbols, black line) or APOBEC3B (filled symbols, dashed line), CBE variants with N-terminal SSB fusions (black and white diamonds and black and white triangles) have a narrower editing window than CBE variants without SSB fusions (black and white circles). Variants with N-terminal SSB fusions generate 50% C to T substitutions within a 4-nt window, while variants without SSB require at least 5 nt to contain the same percentage of editing. Furthermore, CBE variants with N-terminal SSB fusions generate 67% C to T substitutions within an approximately 5-nt window, while variants without SSB require a 7-nt window to contain the same percentage of editing.
[0182] [Example 2] CBE RNPs without 2xUGI have higher editing activity than CBE RNPs with C-terminal 2xUGI fusions A minimal cytosine base editor (CBE) protein without a UGI was constructed by fusing the C-terminal domain of human APOBEC3B to the amino terminus of the SpCas9 nickase protein (SEQ ID NO: 10). The minimal CBE was further modified to contain T4 phage SSB at the amino terminus of the CBE (SEQ ID NO: 11). The CBE protein with 2x UGI used in this example was as described in Example 1. The protein was expressed and purified from E. coli as in Example 1. Synthetic single guide RNAs (sgRNAs) targeting four human genomic sites were purchased from MilliporeSigma. The spacer sequences of the sgRNAs are given in Table 1. Each experimental condition was tested in two technical replicates. Ribonucleoprotein (RNP) complexes were prepared by incubating 40 pmol of CBE protein and 120 pmol of sgRNA in a final volume of 10 μL in buffer (20 mM HEPES, 100 mM KCl, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) at room temperature for 15 minutes. RNPs were placed on ice until transfection. Transfection was performed in HEK293 cells as described in Example 1. Genomic DNA was purified, genomic target sites were amplified, libraries were prepared, sequenced, and analyzed as described in Example 1.
[0183] The results are presented in Figures 2A-2D. The percentage of sequencing reads containing substitutions at each cytosine residue in the protospacer is plotted for CBE variants with and without 2xUGI fusions. The results show that CBE variants lacking UGI as part of the fusion protein have higher substitution rates than CBE variants with 2xUGI for all targets tested. The editing rate for CBE variants without UGI was up to 3.7-fold higher than the rate for 2xUGI variants. The rate of C-to-T substitutions was similar for CBE variants with and without 2xUGI fusions, suggesting that the fused UGI domain does not function to restrict C-to-A and C-to-G substitutions as seen in plasmid delivery formats.
[0184] [Example 3] Dose-dependent reduction of indel formation and C-to-A and C-to-G substitutions using co-transfection of uracil glycosylase inhibitor protein and CBE RNP A cytosine base editor (CBE) protein was constructed by fusing the C-terminal domain of human APOBEC3B to the amino terminus of the SpCas9 nickase protein and two uracil glycosylase inhibitor (UGI) domains to the C terminus. This CBE further contained a single-stranded DNA-binding protein (SSB) domain from T4 bacteriophage fused to the N-terminus to form the final structure N-SSB-deaminase-nCas9-C, where N and C represent the amino- and carboxyl-termini of the protein, as described in Example 1. Recombinant UGI (SEQ ID NO: 3) containing bacillus phage UGI, a c-MYC nuclear localization signal (NLS), and an SV40 large T antigen NLS was purified from E. coli. The CBE protein was purified from E. coli as described in Example 1. Synthetic single guide RNAs (sgRNAs) targeting four human genomic sites were purchased from MilliporeSigma. The spacer sequences of the sgRNAs are given in Table 1. Each experimental condition was tested in two technical replicates.
[0185] Ribonucleoprotein (RNP) complexes were prepared by incubating 15 μg of CBE protein, 200 pmol of sgRNA, and 0–15 μg (0–1185 pmol) of UGI in a final volume of 10 μL in buffer (20 mM HEPES, 100 mM KCl, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) at room temperature for 15 minutes. The molar ratio of CBE to sgRNA was approximately 1:3. RNPs were kept on ice until transfection. Transfection was performed in HEK293 cells as described in Example 1. Genomic DNA was purified, genomic target sites amplified, libraries prepared, and sequenced as described in Example 1. FASTQ files for each sample were analyzed using custom analysis scripts and the CRIS.py software package (Connelly & Pruett-Miller, Sci. Rep. 2019).
[0186] The results are presented in Figure 3. In Figure 3A, the percentage of sequencing reads containing insertions or deletions is plotted for four targets edited with CBE RNP. Values are the mean ± standard deviation for two replicates. Co-transfection of the UGI protein with the RNP complex results in a dose-dependent decrease in the indel rate. At the highest dose of UGI, the indel rate is at least 40% lower than the control transfection without UGI (target HBB03) and up to 80% lower than the control (targets EMX1-15 and RNF2).
[0187] In Figure 3B, the percentage of C to T cytosine substitutions is plotted. Values are the mean ± standard deviation for two replicates. Cotransfection of the UGI protein with the RNP complex resulted in a dose-dependent increase in the rate of C to T cytosine substitutions, thus reducing the rate of undesired C to A and C to G substitutions. At the highest dose of UGI, the C to T rate was at least 20% higher than control transfections without UGI (EMX1-15) and up to fourfold higher than the control (HEKSite2).
[0188] In Figure 3C, the percentage of reads containing at least one C to T substitution within the protospacer is plotted. Values are the mean ± standard deviation for two replicates. Co-transfection of the UGI protein with the RNP complex does not, at least minimally, decrease the C to T editing rate. At target HEKSite2, co-transfection of the UGI protein results in a dose-dependent increase in the C to T editing rate, along with an increased proportion of C to T substitutions. This indicates that the decrease in C to A and C to G substitutions, as well as the decrease in indels, is not due to reduced editing activity.
[0189] [Example 4] Addition of dextran sulfate to cotransfection of uracil glycosylase inhibitor protein and CBE RNP enhances C to T substitution rates A cytosine base editor (CBE) protein was constructed by, at a minimum, fusing the C-terminal domain of human APOBEC3B to the amino terminus of the SpCas9 nickase protein. One CBE variant used in this example was constructed by further fusing the single-stranded DNA-binding protein (SSB) domain from T4 bacteriophage to the N-terminus of the above CBE. The CBE and bacillus phage uracil glycosylase inhibitor (UGI) protein were expressed and purified from E. coli as described in Examples 1 and 3, respectively. Synthetic single guide RNAs (sgRNAs) targeting four human genomic sites were purchased from MilliporeSigma. The spacer sequences of the sgRNAs are given in Table 1. Each experimental condition was tested in two technical replicates.
[0190] Ribonucleoprotein (RNP) complexes were prepared by incubating CBE protein, sgRNA, and 0 or 7 μg of UGI in a final volume of 10 μL in buffer (20 mM HEPES, 100 mM KCl, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) at room temperature for 15 minutes. For CBE variants without SSB, 80 pmol of protein and 200 pmol of sgRNA were used, and for CBE variants with SSB, 40 pmol of protein and 120 pmol of sgRNA were used. RNPs were kept on ice until transfection.
[0191] Dextran sulfate sodium salt (product number: D8906) with an average molecular weight greater than 500 kDa was purchased from MilliporeSigma. Dextran sulfate solution was prepared by dissolving the chemical in water at 50 μg / μL and sterilized by filtration through a 0.22 μm filter. The stock solution was diluted with water to prepare a working solution of 1 μg / μL.
[0192] HEK293 cells were obtained from ATCC and grown in DMEM supplemented with 10% FBS, 2 mM L-glutamine, 1 mM sodium pyruvate, and 0.1 mM non-essential amino acids at 37°C and 5% CO. Cells were cultured at 1.67 x 10 4 Cell / tissue culture surface area 1cm 2 Cells were trypsinized to obtain a single cell suspension, washed twice with Hank's balanced salt solution, and plated at approximately 2.5 × 10 cells per 100 μL. 5 Cells were resuspended in Nucleofector solution V (Lonza). Dextran sulfate was added to the cell suspension to a final concentration of 0.5 μg per 100 μL and mixed well by swirling. Nucleofection was performed by mixing 100 μL of the prepared cell suspension with 10 μL of complexed CBE RNP by pipetting up and down six times before transferring to a cuvette for electroporation using program Q-001 on a Nucleofector 2b instrument. Nucleofected cells were immediately transferred to a 6-well plate containing 2 mL of prewarmed medium per well and grown for 3 days before harvesting.
[0193] Genomic DNA was purified, genomic target sites amplified, libraries prepared, and sequenced as described in Example 1. FASTQ files for each sample were analyzed using custom analysis scripts and the BE-Analyzer web-based analysis tool (Hwang et al., BMC Bioinformatics, 2018).
[0194] The results are presented in Table 2 and Figure 4. Table 2 provides the percentage of reads with any C to T substitutions for each CBE variant, each target, and each transfection condition, as determined using the BE-Analyzer. Values are means ± standard deviations. The data show that while UGI and dextran sulfate alone increase the likelihood of target editing, as indicated by the increased percentage of reads containing at least one C to T substitution, the combination of UGI and dextran sulfate results in a higher likelihood of editing than either alone. In many cases, the effect is greater than additive.
[0195] [Table 2]
[0196] In Figures 4A-4H, the percentage of reads with substitutions at each cytosine in the protospacer is plotted: C to T in gray, C to G in white, and C to A in black. (A) & (E) Targets EMX1-15; (B) & (F) Targets HEKSite2; (C) & (G) Targets HBB03; (D) & (H) Targets RNF2. Figures A-D contain data for the minimal CBE variant, which does not contain an SSB domain; Figures E-H contain data for the CBE variant with an N-terminal SSB domain. The data further confirm that the combination of UGI and dextran sulfate results in a higher C to T conversion rate at positions within the editing window than either component alone. Furthermore, the combination of UGI and dextran sulfate results in reduced rates of undesired C to G and C to A substitutions compared to dextran sulfate alone.
[0197] [Example 5] Addition of dextran sulfate to cotransfection of uracil glycosylase inhibitor protein and CBE RNP enhances the reduction of C-to-A and C-to-G substitutions resulting from uracil glycosylase inhibitor Cytosine base editor (CBE) proteins were constructed and purified as in Example 4. Bacillus phage uracil glycosylase inhibitor (UGI) was purified as in Example 3. Synthetic single guide RNAs (sgRNAs) targeting four human genomic sites were purchased from MilliporeSigma. The spacer sequences of the sgRNAs are given in Table 1. Each experimental condition was tested in two technical replicates. Ribonucleoprotein (RNP) complexes and dextran sulfate solutions were prepared as in Example 4. HEK293 cells were cultured and transfected as in Example 4. Genomic DNA was purified, genomic target sites were amplified, libraries were prepared, sequenced, and analyzed as described in Example 1.
[0198] The results are presented in Figure 5. For each CBE variant, RNP was transfected alone, cotransfected with UGI protein, dextran sulfate, or both. The percent reduction in the rate of C-to-A or C-to-G substitutions by UGI is plotted for transfections with (dark gray) and without (light gray) dextran sulfate. The data show that the effect of UGI on reducing C-to-A and C-to-G substitutions is enhanced when dextran sulfate is included in the transfection. When 40 pmol of a CBE variant containing an N-terminal SSB fusion was used, the effect of UGI was enhanced by 5–15%. When 80 pmol of a CBE variant without an SSB fusion was used, the effect of UGI was enhanced by 20–35%.
[0199] [Example 6] Increasing sgRNA length moves the editing window further away from the PAM Cytosine base editor (CBE) proteins were constructed and purified as in Example 4. Synthetic single guide RNAs (sgRNAs) with spacers of 20 nt, 21 nt, and 22 nt in length targeting three human genomic sites were purchased from MilliporeSigma. The spacer sequences of the sgRNAs are given in Table 3. Each experimental condition was tested in two technical replicates. Dextran sulfate solution was prepared as in Example 4.
[0200] [Table 3]
[0201] Ribonucleoprotein (RNP) complexes were prepared by incubating 40 pmol of CBE protein and 120 pmol of sgRNA in a final volume of 10 μL in buffer (20 mM HEPES, 100 mM KCl, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) at room temperature for 15 minutes. Transfection was performed in HEK293 cells as described in Example 4. Genomic DNA was purified, genomic target sites were amplified, libraries were prepared, sequenced, and analyzed as described in Example 3. PCR primers are listed in Table 1.
[0202] The results are presented in Figure 6. Figures 6A-6C plot the percentage of reads with a C to T substitution at each cytosine in the protospacer. In these plots, the protospacer cytosine is displayed as the distance from the PAM (e.g., position -3 is three nucleotides upstream of the PAM). Values are the mean ± standard deviation for two replicates. The data show that as sgRNA length increases, editing increases at positions more distal to the PAM. In the case of target HBB03, this contains a cytosine residue at position -23, which is outside the sgRNA targeting sequence for the longest guide tested. 6A-6C show the percentage of reads with a C to T substitution for three targets and three spacer lengths: (A) target EMX1-15; (B) target RNF2; and (C) target HBB03.
[0203] Figures 6D-6G plot the percentage of reads carrying single-residue edited alleles. Values are the mean ± standard deviation for two replicates. The data show that extending the guide increases the absolute rate of alleles with a single C-to-T edit distal to the PAM and simultaneously decreases the rate of alleles with a single C-to-T edit proximal to the PAM. When the target site contains multiple editable cytosines and the desired edit is distal to the PAM, extending the guide can facilitate obtaining the desired edit. Figures 6D-6G show the percentage of reads carrying single-residue edited alleles: (D) EMX1-15 targeted without SSB; (E) RNF2 targeted without SSB; (F) EMX1-15 targeted with T4 SSB; and (G) RNF2 targeted with T4 SSB.
[0204] [Example 7] In vitro transcribed T4 SSB-containing cytosine base editor and UGI mRNA improves base editing efficiency and accuracy As described in Examples 2 and 3, a plasmid vector encoding an APOBEC3B-derived cytosine base editor (SEQ ID NO: 10; herein designated A3B), a T4 phage SSB-containing APOBEC3B-derived cytosine base editor (SEQ ID NO: 11; herein designated T4 SSB-A3B), and a free UGI (SEQ ID NO: 3) were each constructed with human codon optimization. Each gene was preceded by a T7 RNA polymerase promoter for in vitro RNA transcription using CleanCap (TriLink Biotechnologies, San Diego, CA). Each plasmid was linearized by restriction digestion with Pmel (New England Biolabs, Ipswich, MA) and purified by two rounds of phenol / chloroform extraction. The linearized plasmid DNA was then used for mRNA production using a HiScribe T7 mRNA kit (New England Biolabs, Ipswich, MA) and CleanCap Reagent AG (TriLink Bio). mRNA quality was verified on an Agilent 2100 Bioanalyzer using the NA 6000 Nano kit (Agilent, Santa Clara, CA).
[0205] Human K562 cells harboring the Y93H (T to C mutation in DNA) mutation in GFP integrated into the human EMX1 locus were used for the experiments. The mutation abolished GFP fluorophore formation and therefore inactivated protein fluorescence activity. Cells were harvested from liquid nitrogen and grown for 1 week at 37°C and 5% CO2 in Iscove's modified Dulbecco's medium (Sigma-Aldrich, St. Louis, MO) supplemented with 10% FBS and 2 mM L-glutamine. One day before transfection, cells were diluted to 0.25 × 10 per mL. 6 Cells were plated at approximately 0.5 x 10 per mL at the time of transfection. 6 The cells were washed twice with Hank's balanced salt solution and then diluted to approximately 0.6 x 10 cells per 100 μL. 6Cells were resuspended in Nucleofector solution V (Lonza; Bend, OR). Each transfection sample contained 8 μg of base editor mRNA and 200 pmol of sgRNA with a guide sequence of 5'-CUGAAGGUCACGUACAAGAG-3' (SEQ ID NO: 41). A subset of samples also contained 4 μg of UGI mRNA. RNase-free water was used as a negative control. The molar ratio of UGI mRNA to base editor mRNA was approximately 5:1. Nucleofection was performed by first mixing 100 μL of cells with the transfection sample by gently pipetting up and down without introducing air bubbles before transferring into a cuvette for immediate electroporation on an Amaxa instrument (Lonza) using program T-016. After nucleofection, cells were immediately transferred to a 6-well plate containing 2 mL of prewarmed medium per well and then grown at 37°C and 5% CO2.
[0206] Flow cytometry analysis was performed on a MACSQuant Analyzer (Miltenyi Biotec, San Diego, CA) 5 days after transfection, and data were analyzed using the FlowJo program (BD Biosciences, Franklin Lake, NJ). Genomic DNA was harvested from transfected cells 5 days after transfection using QuickExtract solution (Lucigen, Middleton, WI). The targeted genomic region was PCR amplified with a pair of NGS primers using the JumpStart™ Taq ReadyMix™ for quantitative PCR kit (MilliporeSigma, Burlington, MA) under the following cycling conditions: 98°C / 2 min; 34 cycles of 98°C / 15 s, 62°C / 30 s, and 72°C / 45 s; 72°C / 5 min. The NGS primers were 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGNNNNNNCCTGAAGTTCATCTGCACCACC-3' (SEQ ID NO: 42) (forward) and 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGNNNNNNCACGTTGTGGCTGTTGTAGTTGTA-3' (SEQ ID NO: 43) (reverse). The primary PCR product was then reamplified with Illumina index primers using the JumpStart™ Taq ReadyMix™ for quantitative PCR kit (MilliporeSigma, Burlington, MA) with the following cycling conditions: 95°C / 3 min; 8 cycles of 95°C / 30 sec, 55°C / 30 sec, and 72°C / 30 sec; 72°C / 5 min. Indexed PCR products were purified using the Select-a-Size DNA Clean & Concentrator Kit (Zymo, Irvine, CA) and quantified using PicoGreen (ThermoFisher, Waltham, MA). PCR products were then normalized and pooled to generate NGS libraries. NGS was performed using an Illumina MiSeq instrument and a 2x300bp kit (San Diego, CA).The FASTQ files for each sample were analyzed using a custom analysis program.
[0207] The results are presented in Figures 7A–7E. Figure 7A shows the percentage of cells that became GFP-positive after transfection with two cytosine base editor mRNAs (A3B and T4 SSB-A3B) in the presence or absence of UGI mRNA. The results indicate that adding UGI mRNA during transfection increased the number of GFP-positive cells by approximately 50%. Figures 7B–7D show the percentage of sequencing reads with C to T (Figure 7B), C to G (Figure 7C), or C to A (Figure 7D) substitutions by position from NGS analysis of transfected cells. Figure 7E shows the percentage of sequencing reads with indels from NGS analysis of transfected cells. These results indicate that adding UGI mRNA during transfection increased C to T substitutions while simultaneously reducing C to G or C to A substitutions and indel formation. These results further demonstrate that the T4 SSB-containing base editor mRNA reduced off-target effects on two of the three unintended C nucleotides (C1, C11, and C15) within the protospacer compared with the base editor mRNA without the SSB domain, while these two base editor mRNAs had similar C-to-T conversion efficiencies on the intended C9 nucleotide.
[0208] [Example 8] UGI used with a nuclear localization sequence (NLS) A cytosine base editor (CBE) protein was constructed by fusing the C-terminal domain of human APOBEC3B to the amino terminus of the SpCas9 nickase protein via a rigid peptide linker and two uracil glycosylase inhibitor (UGI) domains to the C-terminus (SEQ ID NO: 12). Recombinant UGI was constructed by fusing the bacillus phage UGI, the c-MYC nuclear localization signal (NLS), and the SV40 large T antigen NLS (SEQ ID NO: 3). Recombinant UGI without the NLS sequence was purchased from NEB (catalog number M0281). All proteins were expressed and purified from E. coli BL21AI by autoinduction and nickel column chromatography and stored at -80°C prior to use in a buffer containing 10% glycerol, 300 mM KCl, 20 mM HEPES (pH 7.5), and 1 mM DTT. Synthetic single guide RNAs (sgRNAs) targeting EMX1-15 and HEKSite2 were purchased from MilliporeSigma. The spacer sequences of the sgRNAs are given in Table 1. Each experimental condition was tested in two technical replicates.
[0209] Ribonucleoprotein (RNP) complexes were prepared by incubating 15 μg of CBE protein, 200 pmol of sgRNA, and 158–1185 pmol of UGI in a final volume of 10 μL in buffer (20 mM HEPES, 100 mM KCl, 0.5 mM DTT, 0.1 mM EDTA, pH 7.5) at room temperature for 15 min. The molar ratio of CBE to sgRNA was approximately 1:3. RNPs were placed on ice until transfection. HEK293 cells were obtained from ATCC and grown at 37°C and 5% CO2 in DMEM supplemented with 10% FBS, 2 mM L-glutamine, 1 mM sodium pyruvate, and 0.1 mM non-essential amino acids. Cells were cultured at 1.67 × 10 cells / ml two days before transfection. 4 Cell / tissue culture surface area 1cm 2 At the time of transfection, cells were trypsinized to obtain a single cell suspension, washed twice with Hank's balanced salt solution, and plated at approximately 2.5 × 10 cells per 100 μL. 5Cells were resuspended in Nucleofector solution V (Lonza). Nucleofection was performed by mixing 100 μL of the prepared cell suspension with 10 μL of complexed CBE RNP by pipetting up and down six times before transferring to a cuvette for electroporation using program Q-001 on a Nucleofector 2b device (Lonza). Nucleofected cells were immediately transferred to a 6-well plate containing 2 mL of prewarmed medium per well and grown for 3 days before harvesting.
[0210] Genomic DNA was harvested by trypsinizing transfected cells and resuspending them in 50 μL of QuickExtract reagent (Lucigen). The suspension was incubated at 60°C for 15 minutes and at 95°C for 15 minutes. The genomic region targeted by CBE was amplified by PCR using JumpStart Taq ReadyMix (MilliporeSigma) and the following cycling conditions: 94°C / 2 minutes; 25 cycles of 94°C / 30 seconds, 62°C / 30 seconds, and 72°C / 45 seconds; and 72°C / 5 minutes. Primers are listed in Table 1. PCR products underwent two rounds of amplification using Illumina index primers and JumpStart Taq ReadyMix and the following cycling conditions: 95°C / 3 minutes; 9 cycles of 95°C / 30 seconds, 55°C / 30 seconds, and 72°C / 30 seconds; and 72°C / 5 minutes. Indexed PCR products were purified with Select-a-Size DNA Clean & Concentrator MagBeads (Zymo, Irvine, CA) using 1.2x beads in volume, quantified with PicoGreen (ThermoFisher, Waltham, MA), and pooled according to DNA content. Pools were diluted to 4 nM. Sequencing was performed on an Illumina MiSeq instrument using a 300-cycle kit to obtain single-end reads. FASTQ files for each sample were analyzed using the CRIS.py software package (Connelly & Pruett-Miller, Sci. Rep. 2019).
[0211] Results are presented in Figures 8A and 8B. For each target: (A) EMX1-1; and (B) HEKSite2; the percentage of reads with any C to T edits is plotted for each UGI condition, as described above and given on the x-axis of the graph. Values are the mean ± standard deviation for two replicates. At both target sites, co-delivery of the UGI protein and NLS increases the rate of C to T editing. In contrast, NEB UGI shows a rate of editing similar to or lower than the control sample without the NLS but without the UGI protein alone.
Claims
1. A composition suitable for modifying cytosine residues in a DNA sequence, the composition comprising a single-stranded DNA binding domain (SSB), a deaminase, a catalytically modified Cas protein, and a guide RNA to form a base-editing RNP complex.
2. A composition further comprising one or more free and / or fused uracil glycosylase inhibitors (UGIs), and forming a base-edited RNP complex with the UGIs.
3. 2. The composition of claim 1, wherein the SSB is directly linked to the deaminase.
4. 2. The composition of claim 1, wherein the SSB is indirectly linked to the deaminase.
5. 2. The composition of claim 1, wherein the SSB is not covalently bound to any of the deaminase, catalytically modified Cas protein, or guide RNA.
6. The composition described in claim 2, wherein the SSB is not covalently bound to the UGI.
7. 2. The composition of claim 1, wherein the deaminase is covalently linked to a catalytically engineered Cas protein.
8. 3. The composition of claim 2, wherein the one or more UGIs are not covalently linked to any of the deaminase, catalytically modified Cas protein, or guide RNA.
9. The composition of claim 1 , wherein the one or more SSBs are of viral origin.
10. The composition of claim 1 , wherein one or more SSBs are derived from a prokaryote.
11. The composition of claim 1 , wherein one or more SSBs are from a eukaryote.
12. The composition of claim 1 , wherein the deaminase is a cytosine deaminase.
13. 2. The composition of claim 1, wherein the deaminase is adenosine deaminase.
14. 2. The composition of claim 1, wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain.
15. 2. The composition of claim 1, wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain and a catalytically inactive HNH nuclease domain.
16. 2. The composition of claim 1, wherein the catalytically engineered Cas protein is Cas9.
17. The composition of claim 1, wherein the catalytically modified Cas protein is Cas12.
18. 18. At least one nucleic acid encoding one or more of the SSB, deaminase, catalytically modified Cas protein, and guide RNA of the composition of any one of claims 1 to 17.
19. At least one nucleic acid encoding one or more of the free and / or fused UGIs of the composition described in any of claims 2, 6 and 8.
20. At least one expression vector comprising a nucleic acid encoding at least one or more of the SSB, deaminase, catalytically modified Cas protein, and guide RNA of the composition of any one of claims 1 to 17.
21. At least one expression vector comprising a nucleic acid encoding at least one of one or more free and / or fused UGIs described in any of claims 2, 6 and 8.
22. The composition of claim 1, further comprising a nuclear localization sequence (NLS).
23. A composition suitable for modifying cytosine residues in a DNA sequence, the composition comprising a deaminase, a catalytically modified Cas protein, a guide RNA, and one or more free UGIs, to form a base-editing RNP complex with the free UGIs.
24. 24. The composition of claim 23, wherein the deaminase is covalently linked to a catalytically engineered Cas protein.
25. 24. The composition of claim 23, wherein the one or more UGIs are not covalently linked to any of the deaminase, catalytically modified Cas protein, or guide RNA.
26. 24. The composition of claim 23, wherein the deaminase is a cytosine deaminase.
27. 24. The composition of claim 23, wherein the deaminase is adenosine deaminase.
28. 24. The composition of claim 23, wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain.
29. 24. The composition of claim 23, wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain and a catalytically inactive HNH nuclease domain.
30. 24. The composition of claim 23, wherein the catalytically engineered Cas protein is Cas9.
31. 24. The composition of claim 23, wherein the catalytically modified Cas protein is Cas12.
32. 24. The composition of claim 23, wherein the deaminase, catalytically modified Cas protein, and guide RNA are encoded by one or more nucleic acids, and the UGI is a peptide or is encoded by one or more nucleic acids.
33. 33. One or more expression vectors comprising one or more of the nucleic acids of claim 32.
34. The composition of any one of claims 23 to 33, further comprising a nuclear localization sequence (NLS).
35. 1. A method for modifying a DNA sequence, comprising: a) 1) providing an RNP complex comprising a single-stranded DNA binding domain (SSB), a deaminase, a catalytically engineered Cas protein, and a guide RNA; and b) introducing the RNP complex into a recipient cell; c) A method wherein the DNA of the recipient cell has at least one cytosine residue converted to a thymine residue.
36. The method of claim 35, further comprising in a) 2) providing one or more free uracil glycosylase inhibitors (UGIs), and further comprising in b) introducing the free UGIs into the recipient cell.
37. 36. The method of claim 35, wherein the SSB is directly linked to the deaminase.
38. 36. The method of claim 35, wherein the SSB is indirectly linked to the deaminase.
39. 36. The method of claim 35, wherein the deaminase is covalently linked to a catalytically engineered Cas protein.
40. 37. The method of claim 36, wherein the one or more UGIs are not covalently linked to the deaminase, catalytically modified Cas protein, or guide RNA.
41. 36. The method of claim 35, wherein one or more SSBs are of viral origin.
42. 36. The method of claim 35, wherein one or more SSBs are of prokaryotic origin.
43. 36. The method of claim 35, wherein one or more SSBs are of eukaryotic origin.
44. 36. The method of claim 35, wherein the deaminase is a cytosine deaminase.
45. 36. The method of claim 35, wherein the deaminase is adenosine deaminase.
46. 36. The method of claim 35, wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain.
47. 36. The method of claim 35, wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain and a catalytically inactive HNH nuclease domain.
48. 36. The method of claim 35, wherein the catalytically engineered Cas protein is Cas9.
49. 36. The method of claim 35, wherein the catalytically modified Cas protein is Cas12.
50. 37. The method of claim 36, wherein the RNP complex and UGI are introduced into the recipient cell as recombinant proteins.
51. 37. The method of claim 36, wherein the RNP complex and UGI are introduced into the recipient cell as RNA.
52. 37. The method of claim 36, wherein the RNP complex and UGI are introduced into the recipient cell as at least one expression vector.
53. 51. The method of claim 50, wherein the recombinant protein is introduced into the recipient cell by electroporation in the presence of a sulfonated polysaccharide-based RNP electroporation enhancer.
54. 36. The method of claim 35, wherein the guide RNA has a target-complementary sequence of 20 to 30 nucleotides in length.
55. 36. The method of claim 35, further comprising a nuclear localization sequence (NLS).
56. 1. A method for modifying a DNA sequence, comprising: a) providing 1) an RNP complex comprising a deaminase peptide, a catalytically engineered Cas protein, and a guide RNA; 2) one or more free uracil glycosylase inhibitors (UGIs); b) introducing the RNP complex and the free UGI into a recipient cell; c) A method wherein the DNA of the recipient cell has at least one cytosine residue converted to a thymine residue.
57. 57. The method of claim 56, wherein the deaminase is covalently linked to a catalytically engineered Cas protein.
58. 57. The method of claim 56, wherein the one or more UGIs are not covalently linked to the deaminase, catalytically modified Cas protein, or guide RNA.
59. 57. The method of claim 56, wherein the deaminase is a cytosine deaminase.
60. 57. The method of claim 56, wherein the deaminase is adenosine deaminase.
61. 57. The method of Claim 56, wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain.
62. 57. The method of Claim 56, wherein the catalytically modified Cas protein comprises a catalytically inactive RuvC nuclease domain and a catalytically inactive HNH nuclease domain.
63. 57. The method of claim 56, wherein the catalytically modified Cas protein is Cas9.
64. 57. The method of claim 56, wherein the catalytically modified Cas protein is Cas12.
65. 57. The method of claim 56, wherein the RNP complex and UGI are introduced into the recipient cell as recombinant proteins.
66. 66. The method of claim 65, wherein the recombinant protein is introduced into the recipient cell by electroporation in the presence of a sulfonated polysaccharide-based RNP electroporation enhancer.
67. 57. The method of claim 56, wherein the guide RNA has a target-complementary sequence of 20 to 30 nucleotides in length.
68. 68. The method of any one of claims 56 to 67, further comprising a nuclear localization sequence (NLS).
69. 1. A method for introducing a UGI into a eukaryotic cell, comprising: providing a) at least one recipient cell and b) at least one UGI covalently linked to at least one NLS, forming a UGI-NLS molecule; b) transfecting a UGI-NLS molecule into a recipient cell.
70. 70. The method of claim 69, wherein the UGI-NLS molecule is a recombinant protein.